CIP 0049 Incentivizing Cold Backups for Super Validators to Enhance Network Resilience
- Hello,Please find the CIP-0049 open for discussion at the PR https://github.com/global-synchronizer-foundation/cips/pull/31In order to move to a vote we are requesting a Super Validator to sponsor the CIP and a second Super Validator to endorse the CIP.Any SV sponsors/endorses or changes needed?Dr. Amanda L. Martin
814-359-6544Director of Program Management - toggle quoted message Show quoted texthow will an SV prove they are actually maintaining backups the way this process would design / prescribe ?will there be some regularly ceremony (quarterly?) where a prod parallel instance is stood up to prove it? what's the plan there?On Mon, Mar 3, 2025 at 11:15 AM DrAmandaLMartin via lists.sync.global <amartin=linuxfoundation.org@...> wrote:Hello,Please find the CIP-0049 open for discussion at the PR https://github.com/global-synchronizer-foundation/cips/pull/31In order to move to a vote we are requesting a Super Validator to sponsor the CIP and a second Super Validator to endorse the CIP.Any SV sponsors/endorses or changes needed?Dr. Amanda L. Martin
814-359-6544Director of Program Management--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message. - toggle quoted message Show quoted text
Could we also define what we mean by cold? Are we talking about fully air gapped cold storage? Same question as Eric – how often?
Is there a forum to discuss this CIP?
Kinga Z. Bosse
COO|MPCH|www.mpch.com
From: cip-discuss@... <cip-discuss@...> on behalf of Eric Saraniecki via lists.sync.global <eric=digitalasset.com@...>
Date: Monday, March 3, 2025 at 11:24 AM
To: cip-discuss@... <cip-discuss@...>
Subject: Re: [cip-discuss] CIP-0049 Incentivizing Cold Backups for Super Validators to Enhance Network Resiliencehow will an SV prove they are actually maintaining backups the way this process would design / prescribe ?
will there be some regularly ceremony (quarterly?) where a prod parallel instance is stood up to prove it? what's the plan there?
On Mon, Mar 3, 2025 at 11:15 AM DrAmandaLMartin via lists.sync.global <amartin=linuxfoundation.org@...> wrote:Hello,
Please find the CIP-0049 open for discussion at the PR https://github.com/global-synchronizer-foundation/cips/pull/31
In order to move to a vote we are requesting a Super Validator to sponsor the CIP and a second Super Validator to endorse the CIP.
Any SV sponsors/endorses or changes needed?
Dr. Amanda L. Martin
814-359-6544Director of Program Management
--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message. - Eric, I’m really open to any ideas here because the size of the network is massiveand it's growing fast.In my mind, the simplest and probably quickest way to do it is to have a SV node operating as a “verifier”,ie, regularly restoring the DB from the backups and catching up with the network.Backing up and restoring the DB should be quick if we use snapshotting (LVM snapshots for example).How much strain will this put on the network if this node needs to sync up one week worth of dataor one month of data, is up for testing. I don’t know how it will affect the network performance.Probably a question that people with deeper knowledge on canton consensus mechanism could answer.More frequent backups mean less strain on the network but more expensive.The people in this list are definitely far more capable than me in assessing the best mechanismand feasibility of each process, so let's kickstart the discussion and see where we land.
Y. - Good point. I believe "Cold" should not be used there since it's confusing.Just backups that are recoverable.
On how often, it has to do with the recovery and verification mechanism as well asI responded to Eric. I believe that a decentralised network, to be considered bulletproof, thisshould not even be a discussion since we have enough nodes for this not to be an issue.For reference, we keep our snapshots hourly for cometBFT and every 4 hours for postgres.Y. - toggle quoted message Show quoted text
A regular process for DR testing like how the Fia does it might be useful. https://www.fia.org/fia/fia-disaster-recovery-exercise.
From: cip-discuss@... <cip-discuss@...> on behalf of Eric Saraniecki via lists.sync.global <eric=digitalasset.com@...>
Date: Monday, March 3, 2025 at 11:24 AM
To: cip-discuss@... <cip-discuss@...>
Subject: Re: [cip-discuss] CIP-0049 Incentivizing Cold Backups for Super Validators to Enhance Network Resiliencehow will an SV prove they are actually maintaining backups the way this process would design / prescribe ? will there be some regularly ceremony (quarterly?) where a prod parallel instance is stood up to prove it? what's the plan there?
ZjQcmQRYFpfptBannerStart
This Message Is From an External Sender
This message came from outside your organization.
ZjQcmQRYFpfptBannerEnd
how will an SV prove they are actually maintaining backups the way this process would design / prescribe ?
will there be some regularly ceremony (quarterly?) where a prod parallel instance is stood up to prove it? what's the plan there?
On Mon, Mar 3, 2025 at 11:15 AM DrAmandaLMartin via lists.sync.global <amartin=linuxfoundation.org@...> wrote:Hello,
Please find the CIP-0049 open for discussion at the PR https://github.com/global-synchronizer-foundation/cips/pull/31
In order to move to a vote we are requesting a Super Validator to sponsor the CIP and a second Super Validator to endorse the CIP.
Any SV sponsors/endorses or changes needed?
Dr. Amanda L. Martin
814-359-6544Director of Program Management
--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message.This message and any attachments are intended only for the use of the addressee and may contain information that is privileged and confidential. If the reader of the message is not the intended recipient or an authorized representative of the intended recipient, you are hereby notified that any dissemination of this communication is strictly prohibited. If you have received this communication in error, please notify us immediately by e-mail and delete the message and any attachments from your system.
- toggle quoted message Show quoted text
While some of the technical details are discussed I wanted to also raise a tokenomics questions. This line from the CIP could use some additional detail.
>> Super Validators that maintain and verifiably prove the integrity of their backups will receive an enhanced reward, equivalent to an approximate bonus of 20–30% (in Canton terms, a +1 weight boost on current rewards).
Specifically, is it a bonus of 20-30% or a bonus of +1 weight boost? If it is in absolute terms, the award for a T3 SV would be 100% and for a T1 SV would be 10%. Are we looking to make the reward consistent in absolute terms or in relative terms? If we were to do something here, it needs to be configured thoughtfully. For instance, could a T4 0.5 weight SV add 200% to its weight by just running a backup?
Regardless of what is being proposed here, as with any CIP that is tokenomics related, the structure needs to be precise and leave no room for interpretation.From: cip-discuss@... <cip-discuss@...> On Behalf Of Prakash Neelakantan via lists.sync.global
Sent: Monday, March 3, 2025 12:54 PM
To: cip-discuss@...
Subject: [ext] Re: [cip-discuss] CIP-0049 Incentivizing Cold Backups for Super Validators to Enhance Network ResilienceA regular process for DR testing like how the Fia does it might be useful. https: //www. fia. org/fia/fia-disaster-recovery-exercise. From: cip-discuss@ lists. sync. global <cip-discuss@ lists. sync. global> on behalf of Eric Saraniecki via lists. sync. global
A regular process for DR testing like how the Fia does it might be useful. https://www.fia.org/fia/fia-disaster-recovery-exercise.
From: cip-discuss@... <cip-discuss@...> on behalf of Eric Saraniecki via lists.sync.global <eric=digitalasset.com@...>
Date: Monday, March 3, 2025 at 11:24 AM
To: cip-discuss@... <cip-discuss@...>
Subject: Re: [cip-discuss] CIP-0049 Incentivizing Cold Backups for Super Validators to Enhance Network Resiliencehow will an SV prove they are actually maintaining backups the way this process would design / prescribe ? will there be some regularly ceremony (quarterly?) where a prod parallel instance is stood up to prove it? what's the plan there?
how will an SV prove they are actually maintaining backups the way this process would design / prescribe ?
will there be some regularly ceremony (quarterly?) where a prod parallel instance is stood up to prove it? what's the plan there?
On Mon, Mar 3, 2025 at 11:15 AM DrAmandaLMartin via lists.sync.global <amartin=linuxfoundation.org@...> wrote:
Hello,
Please find the CIP-0049 open for discussion at the PR https://github.com/global-synchronizer-foundation/cips/pull/31
In order to move to a vote we are requesting a Super Validator to sponsor the CIP and a second Super Validator to endorse the CIP.
Any SV sponsors/endorses or changes needed?
Dr. Amanda L. Martin
814-359-6544Director of Program Management
--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message.This message and any attachments are intended only for the use of the addressee and may contain information that is privileged and confidential. If the reader of the message is not the intended recipient or an authorized representative of the intended recipient, you are hereby notified that any dissemination of this communication is strictly prohibited. If you have received this communication in error, please notify us immediately by e-mail and delete the message and any attachments from your system.
This e-mail and any attachments may contain information that is confidential and proprietary and otherwise protected from disclosure. If you are not the intended recipient of this e-mail, do not read, duplicate or redistribute it by any means. Please immediately delete it and any attachments and notify the sender that you have received it by mistake. Unintended recipients are prohibited from taking action on the basis of information in this e-mail or any attachments. The DRW Companies make no representations that this e-mail or any attachments are free of computer viruses or other defects. - Chris, you are right - this is very ambiguous and I realised that It's not complete.The idea was to have a relative bonus model of 20% with a cap of +1 in the weight, whatever is smaller.No validator should exceed a +1 weight increase by gaining 20% more rewards.
To break it down with an example, with relative bonus of 20%, a SV with a weight of 10 would
be getting the equivalent rewards of a having a weight of 12 which is way to much
for maintaining a backup. So I propose to have a cap of +1 in weight equivalent in termsof the rewards. So they will be getting a 10% bonus instead.
Another solution - that now that I think about it I like more - is to have a simplereward pool of X canton coins, that is split equally between the SV thathold and verify the backups. This simplifies the logistics a lot since thereis no linear increase in rewards and it's irrelevant from the weights, so more fair.
The parameters I propose here are not set. Everyone is runninga different setup in terms of costs. We need to incentivise enough in order to have ALL SV keepbackups but not to the point that the network token economy is affected.
I strongly suggest that we focus on formalising the backup and verification process first.Then, we will have a good idea of what the costs for running this network-wide is.Last step is to formalise the economic incentives.
Y. - Yiannis,This could've been an option however there are some challenges:1) there should be a way to restore data from the network inception till one month before current time without gaps2) catching up could be pretty slow (something like 1/10 of real time)Another option could be picking a few arbitrary (somewhat random) points in time, fetching states at those points from the archive and comparing them to some references. This is also not without challenges however this is probably somewhat close to how the archive would be used when it's needed.Stas
- toggle quoted message Show quoted textHi,I don't think the catchup option works for old backups unfortunately. SVs can only catch up iff they are within the cometbft and sequencer pruning interval of 30 days. So for the main goal here of testing that SVs preserve backups essentially forever, this is not suited. On top of that there are extra complications: Catchup is relatively slow and an SV is not gonna be able to mint rewards while it is catching up.So I think some option where SVs compare data between each other from their backups without actually trying to run a full node against them is more realistic. We'll need to define how exactly that comparison is going to work for cometbft, sequencer, mediator and participant though.On Tue, Mar 4, 2025 at 8:18 AM Stanislav German-Evtushenko via lists.sync.global <stasge=sbisecsol.com@...> wrote:Yiannis,This could've been an option however there are some challenges:1) there should be a way to restore data from the network inception till one month before current time without gaps2) catching up could be pretty slow (something like 1/10 of real time)Another option could be picking a few arbitrary (somewhat random) points in time, fetching states at those points from the archive and comparing them to some references. This is also not without challenges however this is probably somewhat close to how the archive would be used when it's needed.Stas--Moritz KieferArchitect - Canton NetworkDigital Asset, creators of Daml
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message. - Ok so based on the feedback from @stanislav and @moritz we agree that the verification should bein the form or random comparison of data between the backups, instead of full restore and validationof the network state.
I like this approach, so we can probably discuss this via slack or somewhere, come up with asolid and agreed upon formal method and then update the proposal to include it.
What do you think? - > random comparison of data between the backupsThat would require that the backups are taken at the same time, wouldn't it? Otherwise the backups that different SVs have will not be identical. Not sure we can guarantee that level of synchronization for taking backups.
- toggle quoted message Show quoted textYiannis, Proof Group is a fan of your proposal here and will definitely support in running a backup node. I think the increased SV weighting works well to incentivize this.We would be happy to sponsor this CIP with you when appropriate.On Mon, Mar 3, 2025 at 10:53 PM Yiannis Varelas via lists.sync.global <y=fivenorth.io@...> wrote:Chris, you are right - this is very ambiguous and I realised that It's not complete.The idea was to have a relative bonus model of 20% with a cap of +1 in the weight, whatever is smaller.No validator should exceed a +1 weight increase by gaining 20% more rewards.
To break it down with an example, with relative bonus of 20%, a SV with a weight of 10 would
be getting the equivalent rewards of a having a weight of 12 which is way to much
for maintaining a backup. So I propose to have a cap of +1 in weight equivalent in termsof the rewards. So they will be getting a 10% bonus instead.
Another solution - that now that I think about it I like more - is to have a simplereward pool of X canton coins, that is split equally between the SV thathold and verify the backups. This simplifies the logistics a lot since thereis no linear increase in rewards and it's irrelevant from the weights, so more fair.
The parameters I propose here are not set. Everyone is runninga different setup in terms of costs. We need to incentivise enough in order to have ALL SV keepbackups but not to the point that the network token economy is affected.
I strongly suggest that we focus on formalising the backup and verification process first.Then, we will have a good idea of what the costs for running this network-wide is.Last step is to formalise the economic incentives.
Y. - Can you please elaborate on this? My understanding is that we can query specific data timestamps.
thanks. - Hi Chris, this is great news, happy to have proof group as a sponsor!
- toggle quoted message Show quoted text> Can you please elaborate on this? My understanding is that we can query specific data timestamps.From a live system we can definitely query specific data timestamps. But here we're talking about database snapshots kept in cold storage. These are static blobs of data, often in binary format, with no supporting live system to query from.At least my understanding of the proposed CIP is not to keep this data live in the system. We will still prune it from the live system, so it cannot easily be queried by hitting some endpoint. We will only keep historical snapshots of the databases in cold storage.ItaiOn Wed, Mar 5, 2025 at 12:31 AM Yiannis Varelas via lists.sync.global <y=fivenorth.io@...> wrote:Can you please elaborate on this? My understanding is that we can query specific data timestamps.
thanks.--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message. - toggle quoted message Show quoted text
Not to distract the conversation, but wanted to close the loop re: being more precise on the weight changes associated with this.
In short, I’m ok with focusing on the technical implementation details first and re-visiting the incentive construct later. I like the idea of uses caps to address corner cases. I’ll give some thought to the question of weight increase vs. token award pool but my lean is to weight increases for now.
Chris
From: cip-discuss@... <cip-discuss@...> On Behalf Of Itai Segall via lists.sync.global
Sent: Wednesday, March 5, 2025 7:48 AM
To: cip-discuss@...
Subject: [ext] Re: [cip-discuss] CIP-0049 Incentivizing Cold Backups for Super Validators to Enhance Network Resilience> Can you please elaborate on this? My understanding is that we can query specific data timestamps. From a live system we can definitely query specific data timestamps. But here we're talking about database snapshots kept in cold storage.
> Can you please elaborate on this? My understanding is that we can query specific data timestamps.
From a live system we can definitely query specific data timestamps. But here we're talking about database snapshots kept in cold storage. These are static blobs of data, often in binary format, with no supporting live system to query from.
At least my understanding of the proposed CIP is not to keep this data live in the system. We will still prune it from the live system, so it cannot easily be queried by hitting some endpoint. We will only keep historical snapshots of the databases in cold storage.
Itai
On Wed, Mar 5, 2025 at 12:31 AM Yiannis Varelas via lists.sync.global <y=fivenorth.io@...> wrote:
Can you please elaborate on this? My understanding is that we can query specific data timestamps.
thanks.
--
Itai Segall
Director of Engineering, Canton Network / +1 551 208 8835
Digital Asset, creators of Daml
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message.This e-mail and any attachments may contain information that is confidential and proprietary and otherwise protected from disclosure. If you are not the intended recipient of this e-mail, do not read, duplicate or redistribute it by any means. Please immediately delete it and any attachments and notify the sender that you have received it by mistake. Unintended recipients are prohibited from taking action on the basis of information in this e-mail or any attachments. The DRW Companies make no representations that this e-mail or any attachments are free of computer viruses or other defects. - Chris, thanks for following up. I'm OK with caps, just don't know how much it changes the mechanism of minting/distributing tokens.
If it's "minimal invasion" then I probably favour this one as well.
Y. - toggle quoted message Show quoted textHi Itail,I think we need to develop a backup prover. The prover can be run by any SV. Periodically the prover will raise a challenge to ask among its peers for some pieces of data. The prover will need to figure out how to get that data and respond to that challenge.To support different backup solutions, we can implement a plug-in architecture similar to terraform, where terraform provider is a separate program that runs and communicates with the main terraform through rpc call. we can provide a few provider providers such as restic backup, borg backup, ebs snapshot etc.There are 2 kinds of data in our specific setup1. Level DB dataThis is straightforward to easily mount a backup snapshot into disk, and have a simple Go program that reads the key and returns its value. Or restore an EBS snapshot.
CometBFT, currently in our config, uses this https://github.com/syndtr/goleveldb which is fairly simple to write a small program that opens the data (either from snapshot mount through FUSE or an EBS volume) and fetch the key.2. Postgres DataSimilarly as above, with Postgres data, the prover can restore an instance from their snapshot data. Then it can query for the relevant data, and then shutdown the postgres instance. For non-cloud backup solutions like pgbackrest or pgbarman, we can pick one and write a provider to handle restoration, read the data out.For picking a data point out to compare, the prover can try to pick the one that has not changed during the backup window. Example we pick a key or a row that has not changed in 24hours, then regardless which time the backup is taken during the day, it should have that value.What do you think ?On Wed, Mar 5, 2025 at 5:48 AM Itai Segall via lists.sync.global <itai.segall=digitalasset.com@...> wrote:> Can you please elaborate on this? My understanding is that we can query specific data timestamps.From a live system we can definitely query specific data timestamps. But here we're talking about database snapshots kept in cold storage. These are static blobs of data, often in binary format, with no supporting live system to query from.At least my understanding of the proposed CIP is not to keep this data live in the system. We will still prune it from the live system, so it cannot easily be queried by hitting some endpoint. We will only keep historical snapshots of the databases in cold storage.ItaiOn Wed, Mar 5, 2025 at 12:31 AM Yiannis Varelas via lists.sync.global <y=fivenorth.io@...> wrote:Can you please elaborate on this? My understanding is that we can query specific data timestamps.
thanks.--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message. - toggle quoted message Show quoted textThanks Vinh,At a high level, that makes perfect sense and I've been toying around with the same idea in my mind over the last few days.One thing I currently do not have an answer for, which might be a blocker, is how does the challenger know whether the answer is correct. In order for this to be something that minting depends on, you would need somehow to get 2/3 of SVs to attest that they accept the answer.Then there's some interesting challenges around who operates the challengers itself, how do SVs submit their responses without revealing the answer to others, etc. but those can probably be resolved fairly easily.On Fri, Mar 7, 2025 at 8:51 AM Vinh Nguyen via lists.sync.global <v=fivenorth.io@...> wrote:Hi Itail,I think we need to develop a backup prover. The prover can be run by any SV. Periodically the prover will raise a challenge to ask among its peers for some pieces of data. The prover will need to figure out how to get that data and respond to that challenge.To support different backup solutions, we can implement a plug-in architecture similar to terraform, where terraform provider is a separate program that runs and communicates with the main terraform through rpc call. we can provide a few provider providers such as restic backup, borg backup, ebs snapshot etc.There are 2 kinds of data in our specific setup1. Level DB dataThis is straightforward to easily mount a backup snapshot into disk, and have a simple Go program that reads the key and returns its value. Or restore an EBS snapshot.
CometBFT, currently in our config, uses this https://github.com/syndtr/goleveldb which is fairly simple to write a small program that opens the data (either from snapshot mount through FUSE or an EBS volume) and fetch the key.2. Postgres DataSimilarly as above, with Postgres data, the prover can restore an instance from their snapshot data. Then it can query for the relevant data, and then shutdown the postgres instance. For non-cloud backup solutions like pgbackrest or pgbarman, we can pick one and write a provider to handle restoration, read the data out.For picking a data point out to compare, the prover can try to pick the one that has not changed during the backup window. Example we pick a key or a row that has not changed in 24hours, then regardless which time the backup is taken during the day, it should have that value.What do you think ?On Wed, Mar 5, 2025 at 5:48 AM Itai Segall via lists.sync.global <itai.segall=digitalasset.com@...> wrote:> Can you please elaborate on this? My understanding is that we can query specific data timestamps.From a live system we can definitely query specific data timestamps. But here we're talking about database snapshots kept in cold storage. These are static blobs of data, often in binary format, with no supporting live system to query from.At least my understanding of the proposed CIP is not to keep this data live in the system. We will still prune it from the live system, so it cannot easily be queried by hitting some endpoint. We will only keep historical snapshots of the databases in cold storage.ItaiOn Wed, Mar 5, 2025 at 12:31 AM Yiannis Varelas via lists.sync.global <y=fivenorth.io@...> wrote:Can you please elaborate on this? My understanding is that we can query specific data timestamps.
thanks.--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message.--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message. - toggle quoted message Show quoted text
I don't want to make it too complicated, but there are "proof of storage" systems for e.g. filecoin, which I believe solve this problem. I'd be happy to look into this further if there is interest.
On 3/7/25 10:22, Itai Segall via lists.sync.global wrote:
Thanks Vinh,
At a high level, that makes perfect sense and I've been toying around with the same idea in my mind over the last few days.One thing I currently do not have an answer for, which might be a blocker, is how does the challenger know whether the answer is correct. In order for this to be something that minting depends on, you would need somehow to get 2/3 of SVs to attest that they accept the answer.
Then there's some interesting challenges around who operates the challengers itself, how do SVs submit their responses without revealing the answer to others, etc. but those can probably be resolved fairly easily.
On Fri, Mar 7, 2025 at 8:51 AM Vinh Nguyen via lists.sync.global <v=fivenorth.io@...> wrote:
Hi Itail,
I think we need to develop a backup prover. The prover can be run by any SV. Periodically the prover will raise a challenge to ask among its peers for some pieces of data. The prover will need to figure out how to get that data and respond to that challenge.
To support different backup solutions, we can implement a plug-in architecture similar to terraform, where terraform provider is a separate program that runs and communicates with the main terraform through rpc call. we can provide a few provider providers such as restic backup, borg backup, ebs snapshot etc.
There are 2 kinds of data in our specific setup
1. Level DB data
This is straightforward to easily mount a backup snapshot into disk, and have a simple Go program that reads the key and returns its value. Or restore an EBS snapshot.
CometBFT, currently in our config, uses this https://github.com/syndtr/goleveldb which is fairly simple to write a small program that opens the data (either from snapshot mount through FUSE or an EBS volume) and fetch the key.
2. Postgres Data
Similarly as above, with Postgres data, the prover can restore an instance from their snapshot data. Then it can query for the relevant data, and then shutdown the postgres instance. For non-cloud backup solutions like pgbackrest or pgbarman, we can pick one and write a provider to handle restoration, read the data out.
For picking a data point out to compare, the prover can try to pick the one that has not changed during the backup window. Example we pick a key or a row that has not changed in 24hours, then regardless which time the backup is taken during the day, it should have that value.
What do you think ?
On Wed, Mar 5, 2025 at 5:48 AM Itai Segall via lists.sync.global <itai.segall=digitalasset.com@...> wrote:
> Can you please elaborate on this? My understanding is that we can query specific data timestamps.
From a live system we can definitely query specific data timestamps. But here we're talking about database snapshots kept in cold storage. These are static blobs of data, often in binary format, with no supporting live system to query from.At least my understanding of the proposed CIP is not to keep this data live in the system. We will still prune it from the live system, so it cannot easily be queried by hitting some endpoint. We will only keep historical snapshots of the databases in cold storage.
Itai
On Wed, Mar 5, 2025 at 12:31 AM Yiannis Varelas via lists.sync.global <y=fivenorth.io@...> wrote:
Can you please elaborate on this? My understanding is that we can query specific data timestamps.
thanks.
--
Itai Segall
Director of Engineering, Canton Network / +1 551 208 8835Digital Asset, creators of Daml
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message.
--
Itai Segall
Director of Engineering, Canton Network / +1 551 208 8835Digital Asset, creators of Daml
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message. - Good point. We would only need 2/3 of the SVs who are actually storing backups to accept the answer, right?
- Hi all - just bringing this up again since we have now gone through the devnet reset.@Wayne I believe 2/3 is what we need yes.@Ryan this is an excellent idea. Had a look at filecoin and seems that they spin outa business that does exactly that.Might worth looking into this as a complete solution.
- toggle quoted message Show quoted textAs discussed in the last weekly ops meeting, I would like to share a few approaches to bring this up for discussion again.In the new CometBFT we are going to use Postgres as storage instead of LevelDB so we can focus all the effort on proving Postgres data. If we do need to perform that, LevelDB is actually quite easy to prove because they are just key<-> values with limited functionality and can easily restore a snapshot and mount. Backup tooling like `restic snapshot` even allow to mount a snapshot and we can read the leveldb file with any client example https://github.com/liderman/leveldb-cli/blob/master/main.goFor Postgres, the primary challenge is access postgres backup data in a timely manner without restoring it into a real database.
There are however many modern Postgres tools that allow us to export Postgres data as CSV or Parquet file to object storage and allow us to query these directly using SQL interfaces without the need to restore the data to a Postgres instance.
The process is very similar to a full pg_dump, but instead of exporting to CSV/SQL or custom format, we export the data to Parquet files, and the file can be stored on object storage like S3/GCS.
Example tooling: https://github.com/CrunchyData/pg_parquet/ https://bnmoch3.org/notes/2025/postgres-iceberg-sync-duckdb/These parquet files can then be query directly using SQL with tool such as AWS Athena(https://docs.aws.amazon.com/athena/latest/ug/columnar-storage.html), Druio, DuckDB(https://duckdb.org/docs/stable/guides/network_cloud_storage/s3_import.html)We then can write a service that performs SQL queries to verify data consistency between the backup. If we take backup daily, we will need to craft a query that has not been modified in say the last 24 hours, then no matter when a backup was taken, as long as it's in the 24 hours window, the backup is correct and will be consistent across all SV.Example, today is 2025-04-15. We can ask "give me a list of transfers between 2025-04-13 -> 2025-04-14. Then no matter when the backup is taken on 2025-04-15, the data will be consistent.Let me know if this idea is sound and we can expand further into some real http service that can perform this kind of query.On Wed, Mar 19, 2025 at 9:31 AM Yiannis Varelas <y@...> wrote:Hi all - just bringing this up again since we have now gone through the devnet reset.@Wayne I believe 2/3 is what we need yes.@Ryan this is an excellent idea. Had a look at filecoin and seems that they spin outa business that does exactly that.Might worth looking into this as a complete solution. - toggle quoted message Show quoted textThanks Vinh,There are still many questions in my mind that need to be answered. Just to name a few:- How storage-efficient is this format of backup? Sounds like it cannot be very succinct, if it needs to support sql-like querying. So what's the expected cost of maintaining all these backups in this format for eternity?- Does this mean that every SV needs to implement their own export, given that cloud providers do not support an export of this format? CloudSQL supports only pg_dump format for example- How is this verification step planned to be carried out, in practice? Manually? By whom? Automatically? By what service and how?BestItaiOn Tue, Apr 15, 2025 at 2:54 AM Vinh Nguyen via lists.sync.global <v=fivenorth.io@...> wrote:As discussed in the last weekly ops meeting, I would like to share a few approaches to bring this up for discussion again.In the new CometBFT we are going to use Postgres as storage instead of LevelDB so we can focus all the effort on proving Postgres data. If we do need to perform that, LevelDB is actually quite easy to prove because they are just key<-> values with limited functionality and can easily restore a snapshot and mount. Backup tooling like `restic snapshot` even allow to mount a snapshot and we can read the leveldb file with any client example https://github.com/liderman/leveldb-cli/blob/master/main.goFor Postgres, the primary challenge is access postgres backup data in a timely manner without restoring it into a real database.
There are however many modern Postgres tools that allow us to export Postgres data as CSV or Parquet file to object storage and allow us to query these directly using SQL interfaces without the need to restore the data to a Postgres instance.
The process is very similar to a full pg_dump, but instead of exporting to CSV/SQL or custom format, we export the data to Parquet files, and the file can be stored on object storage like S3/GCS.
Example tooling: https://github.com/CrunchyData/pg_parquet/ https://bnmoch3.org/notes/2025/postgres-iceberg-sync-duckdb/These parquet files can then be query directly using SQL with tool such as AWS Athena(https://docs.aws.amazon.com/athena/latest/ug/columnar-storage.html), Druio, DuckDB(https://duckdb.org/docs/stable/guides/network_cloud_storage/s3_import.html)We then can write a service that performs SQL queries to verify data consistency between the backup. If we take backup daily, we will need to craft a query that has not been modified in say the last 24 hours, then no matter when a backup was taken, as long as it's in the 24 hours window, the backup is correct and will be consistent across all SV.Example, today is 2025-04-15. We can ask "give me a list of transfers between 2025-04-13 -> 2025-04-14. Then no matter when the backup is taken on 2025-04-15, the data will be consistent.Let me know if this idea is sound and we can expand further into some real http service that can perform this kind of query.On Wed, Mar 19, 2025 at 9:31 AM Yiannis Varelas <y@...> wrote:Hi all - just bringing this up again since we have now gone through the devnet reset.@Wayne I believe 2/3 is what we need yes.@Ryan this is an excellent idea. Had a look at filecoin and seems that they spin outa business that does exactly that.Might worth looking into this as a complete solution.--
This message, and any attachments, is for the intended recipient(s) only, may contain information that is privileged, confidential and/or proprietary and subject to important terms and conditions available at http://www.digitalasset.com/emaildisclaimer.html. If you are not the intended recipient, please delete this message.