Zero knowledge proof
By decomposing the calculation and building zero-knowledge proofs through recursive SNARK and PCD methods, the problem of computational complexity and excessive space consumption when generating zero-knowledge proofs in the prior art is solved, and efficient and scalable proof generation is achieved.
Patent Information
- Application Number
- CN202380068409.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-23
- Filing Date
- 2023-08-23
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to efficiently generate zero-knowledge proofs for proving that Merkel leaves meet predefined criteria, especially when processing larger data, the prover algorithm consumes too much time and space.
The recursive SNARK and proof-carrying data (PCD) method are used to decompose the entire calculation into distributed computing. Each node only executes the given subroutines and constructs the final SNARK through the PCD scheme to verify the recursive nature of the incoming proof.
The zero-knowledge proof of scalable and incremental calculations is implemented, which can effectively prove that any large hash image is known, and the simplicity of the proof generation ensures that the size of the proof is constant or only logarithmic to the proof size.
Smart Images

Figure CN119948803A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for generating a zero-knowledge proof for proving that a leaf of a Merkle tree satisfies a predefined criterion, and a computer system for implementing the method. Background Art
[0002] Blockchain refers to a distributed data structure in which a copy of the blockchain is maintained at each of a plurality of nodes in a distributed peer-to-peer (P2P) network (hereinafter referred to as a "blockchain network") and is widely disclosed. The blockchain includes a series of data blocks, each of which includes one or more transactions. Except for the so-called "coinbase transaction", each transaction points to a previous transaction in a sequence, which can span one or more blocks back to one or more coinbase transactions. The coinbase transaction will be discussed further below. Transactions submitted to the blockchain network are included in new blocks. The process of creating a new block is generally referred to as "mining", which involves each of a plurality of nodes competing to perform "proof of work", i.e., solving a cryptographic puzzle based on a representation of a defined set of ordered and verified valid pending transactions waiting to be included in a new block of the blockchain. It should be noted that the blockchain can be pruned at some nodes, and the publication of blocks can be achieved by publishing only the block header.
[0003] Transactions in a blockchain can be used for one or more of the following purposes: transferring digital assets (i.e., a certain number of digital tokens); sorting a set of entries in a virtualized ledger or registry; receiving and processing time-stamped entries; and / or sorting index pointers by time. Blockchains can also be used to implement hierarchical additional functions on blockchains. For example, a blockchain protocol may allow additional user data or data indexes to be stored in a transaction. There is no pre-specified limit on the maximum data capacity that can be stored in a single transaction, so increasingly complex data can be incorporated. For example, this can be used to store electronic documents, audio, or video data in a blockchain.
[0004] The nodes of the blockchain network (commonly referred to as "miners") perform a distributed transaction registration and verification process, which will be described in more detail later. In summary, in this process, nodes verify transactions and insert them into block templates, which attempt to identify a valid proof-of-work solution for the block template. Once a valid solution is found, the new block is propagated to other nodes of the network, so that each node can record the new block on the blockchain. In order to record a transaction in the blockchain, a user (e.g., a blockchain client application) sends the transaction to one of the nodes in the network for propagation. The nodes receiving the transaction can compete to find a proof-of-work solution that incorporates the verified valid transaction into the new block. Each node is configured to execute the same node protocol, which will include one or more conditions for confirming that the transaction is valid. Invalid transactions will not be propagated or incorporated into the block. Assuming that the transaction has been verified to be valid and thus accepted on the blockchain, the transaction (including any user data) will therefore be registered and indexed as an immutable public record on each node in the blockchain network.
[0005] The node that successfully solves the proof-of-work puzzle to create the latest block is usually rewarded with a new transaction called a "coinbase transaction" that distributes the amount of digital assets, i.e. the number of tokens. The detection and rejection of invalid transactions is performed by the actions of competing nodes, which act as agents of the network and are incentivized to report and prevent improper behavior. The widespread publication of information allows users to continuously audit the performance of nodes. Publishing only block headers allows participants to ensure the continued integrity of the blockchain.
[0006] In the "output-based" model (sometimes referred to as the UTXO-based model), the data structure of a given transaction includes one or more inputs and one or more outputs. Any spendable output includes an element that specifies the amount of a digital asset, which can be derived from the ongoing sequence of transactions. Spendable outputs are sometimes called UTXOs ("unspent transaction outputs"). Outputs may also include a locking script that specifies future redemption conditions for the output. A locking script is a predicate that defines the conditions necessary to verify and transfer a digital token or asset. Each input of a transaction (except for a coinbase transaction) includes a pointer (i.e., a reference) to such an output in a previous transaction, and may also include an unlocking script for unlocking the locking script pointing to the output. Thus, consider a pair of transactions, referred to as a first transaction and a second transaction (or "target" transaction). The first transaction includes at least one output that specifies the amount of a digital asset, and includes a locking script that defines one or more conditions for unlocking the output. The second (target) transaction includes at least one input and an unlocking script, the at least one input including a pointer to the output of the first transaction; the unlocking script is used to unlock the output of the first transaction.
[0007] In such a model, when the second (target) transaction is sent to the blockchain network to be propagated and recorded in the blockchain, one of the validity conditions applied at each node will be that the unlocking script satisfies all of the one or more conditions defined in the locking script of the first transaction. Another condition will be that the output of the first transaction has not been redeemed by another earlier valid transaction. Any node that finds the target transaction invalid according to any of these conditions will not propagate the transaction (as a valid transaction, but may register an invalid transaction) nor include the transaction in a new block to be recorded in the blockchain.
[0008] Another transaction model is the account-based model. In this case, each transaction is defined not by reference to the UTXO of the previous transaction in the past transaction sequence, but by reference to the absolute account balance. The current state of all accounts is stored individually by the node in the blockchain and is constantly updated. Summary of the invention
[0009] Known succinct zero-knowledge arguments of knowledge (SNARKs) are used to learn a hash preimage or a Merkle tree statement, e.g. to prove knowledge of an authentication path consistent with a Merkle root, usually proving knowledge of a witness taken from a fixed domain, or of varying sizes but bounded by a small constant. This is because of the holistic approach taken, i.e. expressing the entire computation as a single circuit, and then proving the satisfiability of that circuit in a single computation. In fact, the larger the witness, the larger the circuit, and the more time / space the prover algorithm consumes.
[0010] This paper presents new methods for generating zero-knowledge proofs that differ from known monolithic methods. Instead, recursive SNARKs are used, or more specifically, proof carrying data (PCD). PCD is a primitive for proving that a distributed computation (whose transcript can be described by a graph) was evaluated correctly. Each node appends to its output an easily verifiable proof that (i) its inputs, outputs, and local data satisfy a given predicate Π(z in ,z loc ,z out )=1, and (ii) the validity of the proof attached to the input data. Due to the recursive nature of proof generation (which verifies the incoming proof), the verifier only needs to verify the proof produced by the last (sink) node that computes the transcript.
[0011] The focus is on the following two aspects:
[0012] a) Interpret the entire computation as a "distributed" computation. Thus, the (potentially large) computation is split into a series of small subroutines (resulting in a manageable circuit). Each node executes only a given subroutine instantiation, and in particular, if the subroutine is a step in a loop, there will be as many nodes as the number of loop iterations.
[0013] b) Use the existing PCD scheme to express compliant computational transcripts and construct the final SNARK (which calls the PCD algorithm internally).
[0014] Any PCD scheme can be used for step (b). In Section 8.1, the choice of curves when using pair-based preprocessing PCD is discussed.
[0015] According to one aspect disclosed herein, there is provided a computer-implemented method for generating a zero-knowledge proof, the zero-knowledge proof being used to prove that each of a plurality of data blocks corresponding to a Merkle tree satisfies a predefined criterion, wherein the Merkle tree comprises a plurality of leaf hash values and a plurality of internal hash values, wherein the plurality of internal hash values are arranged in layers, wherein a plurality of leaf nodes are mapped to the plurality of leaf hash values, and wherein a plurality of internal nodes are mapped to the plurality of internal hash values, wherein the method comprises: executing the plurality of leaf nodes, wherein each leaf node is configured to: receive a corresponding data block and a corresponding data block proof, wherein the corresponding data block proof is used to prove that the data block satisfies the predefined criterion; verify the corresponding data block proof; based on the corresponding data block to calculate a data block hash; and, output the data block hash; execute the multiple internal nodes, wherein each of the multiple internal nodes is configured to: receive a corresponding hash value from each of two previous nodes of the Merkle tree; calculate an output hash value based on the received corresponding hash value; and, output the output hash value; wherein the received corresponding hash value for a first layer of the multiple internal nodes is a corresponding data block hash received from a corresponding leaf node, and wherein the received corresponding hash value for each other layer of the multiple internal nodes is a corresponding output hash value received from a corresponding internal node; wherein the output hash value calculated by a final internal node in a final layer of the Merkle tree is a Merkle root corresponding to the Merkle tree.
[0016] The present disclosure provides a succinct zero-knowledge argument of knowledge (SNARK) for hash-based statements. Proof generation is scalable and can be computed in an incremental manner. For example, to prove knowledge of an arbitrarily large SHA256 preimage (e.g., a 1GB or even larger preimage), the prover's memory requirements can be the same as those for proving knowledge of a 512-bit preimage.
[0017] In general, the prover's runtime scales well with the size of the private input (witnesses, which can be, for example, a large preimage or many leaves of a Merkle tree). This means that there are no strict requirements on the prover's hardware (RAM).
[0018] Furthermore, proof generation can be paused and resumed at a later stage, not necessarily by the same prover. Specifically, proof generation can be distributed across multiple nodes that know only a portion of the private input. This is possible due to the incremental nature of the SNARKs provided in this paper.
[0019] The succinct property of SNARKs also guarantees that the size of the proof is constant regardless of the size of the witness (or only logarithmically in the size of the witness). BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To facilitate an understanding of the embodiments of the present disclosure and to show how such embodiments may be implemented, reference will now be made, by way of example only, to the accompanying drawings, in which:
[0021] Figure 1 is a schematic block diagram of a system for implementing a blockchain;
[0022] Figure 2 Some examples of transactions that may be recorded in a blockchain are schematically shown;
[0023] Figure 3 Schematically shows the calculated transcript of the function f(x,y)∶=(2(x+y),3(x+y)) with bounded noise;
[0024] Figure 4 SHA2 transcripts are schematically shown;
[0025] Figure 5 The input relationship of the SHA2 node is schematically shown;
[0026] Figure 6 An exemplary method for proving knowledge of a preimage using a zero-knowledge proof is provided;
[0027] Figure 7 Schematically illustrating a Merkle tree used to generate a zero-knowledge proof to prove that each leaf of the Merkle tree satisfies a criterion;
[0028] Figure 8 An exemplary method for proving that each leaf of a Merkle tree satisfies a criterion is provided;
[0029] Fig. 9 Schematic illustration of multiple predicates for efficient and scalable zero-knowledge proofs Transcripts of
[0030] Fig.10 An exemplary method for purchasing data using scalable zero-knowledge proofs is shown. DETAILED DESCRIPTION
[0031] 1. Exemplary System Overview
[0032] Figure 1 An exemplary system 100 for implementing a blockchain 150 is shown. The system 100 may include a packet-switched network 101, typically a wide area internet such as the Internet. The packet-switched network 101 includes a plurality of blockchain nodes 104, which may be arranged to form a peer-to-peer (P2P) network 106 within the packet-switched network 101. Although not shown, the blockchain nodes 104 may be arranged as a nearly complete graph. Thus, each blockchain node 104 is highly connected to other blockchain nodes 104.
[0033] Each blockchain node 104 includes a computer device of a peer, and different nodes 104 belong to different peers. Each blockchain node 104 includes a processing device, which includes one or more processors, such as one or more central processing units (CPUs), accelerator processors, special processors and / or field programmable gate arrays (FPGAs), and other devices, such as application-specific integrated circuits (ASICs). Each node also includes a memory, that is, a computer-readable memory in the form of a non-transitory computer-readable medium. The memory may include one or more memory units, which use one or more memory media, such as magnetic media such as hard disks, electronic media such as solid-state drives (SSDs), flash memory, or electrically erasable programmable read-only memories (EEPROMs), and / or optical media such as optical disk drives.
[0034] The blockchain 150 includes a series of data blocks 151, wherein a respective copy of the blockchain 150 is maintained at each of the plurality of blockchain nodes 104 in the distributed or blockchain network 106. As described above, maintaining a copy of the blockchain 150 does not necessarily mean storing the blockchain 150 in its entirety. Instead, the blockchain 150 can be pruned as long as each blockchain node 150 stores a block header (discussed below) for each block 151. Each block 151 in the blockchain includes one or more transactions 152, wherein a transaction in this context refers to a data structure. The nature of the data structure will depend on the type of transaction protocol used as part of the transaction model or plan. A given blockchain uses a particular transaction protocol throughout. In a common transaction protocol, the data structure of each transaction 152 includes at least one input and at least one output. Each output specifies an amount of a digital asset represented as a property, an example of which is an output that is cryptographically locked to a user 103 (requiring the user's signature or other solution to unlock, thereby redeeming or spending). Each input points to the output of a previous transaction 152, thereby linking these transactions.
[0035] Each block 151 also includes a block pointer 155, which points to a previously created block 151 in the blockchain to define the order of blocks 151. Each transaction 152 (except for the coinbase transaction) includes a pointer to a previous transaction to define the order of the transaction sequence (Note: the sequence of transactions 152 can branch). The blockchain of blocks 151 is traced back to the genesis block (Gb) 153, which is the first block in the blockchain. One or more original transactions 152 earlier in the blockchain 150 point to the genesis block 153, not to the previous transaction.
[0036] Each blockchain node 104 is configured to forward transactions 152 to other blockchain nodes 104, so that transactions 152 are propagated throughout the network 106. Each blockchain node 104 is configured to create blocks 151 and store corresponding copies of the same blockchain 150 in its corresponding memory. Each blockchain node 104 also maintains an ordered set (or "pool") 154 of transactions 152 waiting to be incorporated into block 151. Ordered pools 154 are often referred to as "memory pools". In this article, the term is not intended to be limited to any particular blockchain, protocol, or model. The term refers to a set of ordered transactions that a node 104 has accepted as valid, and for which the node 104 is forced to not accept any other transaction that attempts to spend the same output.
[0037] In a given current transaction 152j, an input (or each input) includes a pointer that references an output of a previous transaction 152i in the transaction sequence, specifying that the output is to be redeemed or "spent" in the current transaction 152j. Spending or redeeming does not necessarily mean transferring a financial asset, although this is certainly a common application. More generally, spending can be described as consuming an output, or allocating it to one or more outputs in another subsequent transaction. In general, a previous transaction can be any transaction in an ordered set 154 or any block 151. Although a previous transaction 152i will need to exist and be verified to be valid in order to ensure that the current transaction is valid, it is not necessary for a previous transaction 152i to exist when the current transaction 152j is created or even sent to the network 106. Therefore, in this article, "previous" refers to the predecessor in a logical sequence linked by a pointer, and not necessarily the creation time or sending time in a time sequence, and therefore, does not necessarily exclude the situation where transactions 152i, 152j are created or sent out of order (see the discussion of isolated transactions below). A previous transaction 152i can also be called a predecessor transaction or a predecessor transaction.
[0038] The inputs of the current transaction 152j also include input authorizations, such as the signature of the user 103a to whom the outputs of the previous transaction 152i are locked. In turn, the outputs of the current transaction 152j may be cryptographically locked to the new user or entity 103b. Thus, the current transaction 152j may transfer the amounts defined in the inputs of the previous transaction 152i to the new user or entity 103b defined in the outputs of the current transaction 152j. In some cases, a transaction 152 may have multiple outputs to split the input amounts among multiple users or entities (one of which may be the original user or entity 103a for changes). In some cases, a transaction may also have multiple inputs to aggregate the amounts from multiple outputs of one or more previous transactions and reallocate them to one or more outputs of the current transaction.
[0039] According to an output-based transaction protocol, such as Bitcoin, when a party 103, such as an individual user or an organization, wishes to issue a new transaction 152j (either by an automated program employed by the party or manually), the issuing party sends the new transaction from its computer terminal 102 to a recipient. The issuing party or recipient will ultimately send the transaction to one or more blockchain nodes 104 of the network 106 (now typically a server or data center, but in principle it can also be other user terminals). It is also not excluded that the party 103 issuing the new transaction 152j can send the transaction directly to one or more blockchain nodes 104, and in some examples, the transaction may not be sent to the recipient. The blockchain node 104 receiving the transaction checks whether the transaction is valid according to the blockchain node protocol applied at each blockchain node 104. The blockchain node protocol typically requires the blockchain node 104 to check whether the cryptographic signature in the new transaction 152j matches the expected signature, which depends on the previous transaction 152i in the ordered sequence of transactions 152. In such an output-based transaction protocol, this may include checking whether the cryptographic signature or other authorization of the party 103 included in the input of the new transaction 152j matches the condition defined in the output of the previous transaction 152i that the new transaction spends (or "allocates"), where the condition typically includes at least checking whether the cryptographic signature or other authorization in the input of the new transaction 152j unlocks the output of the previous transaction 152i to which the input of the new transaction is linked. The condition may be defined at least in part by a script included in the output of the previous transaction 152i. Alternatively, this may be determined solely by the blockchain node protocol, or may be determined by a combination thereof. In either case, if the new transaction 152j is valid, the blockchain node 104 forwards it to one or more other blockchain nodes 104 in the blockchain network 106. These other blockchain nodes 104 apply the same test according to the same blockchain node protocol, and therefore forward the new transaction 152j to one or more other nodes 104, and so on. In this way, the new transaction is propagated throughout the network of blockchain nodes 104.
[0040] In the output-based model, the definition of whether a given output (e.g., UTXO) is allocated (or "spent") is whether it is validly redeemed by the input of another subsequent transaction 152j according to the blockchain node protocol. Another condition for a transaction to be valid is that the output of the previous transaction 152i that it attempts to redeem has not been redeemed by another transaction. Similarly, if invalid, transaction 152j will not be propagated (unless marked as invalid and propagated for reminder) or recorded in the blockchain 150. This prevents double spending, that is, the transaction processor allocates the output of the same transaction more than once. On the other hand, the account-based model prevents double spending by maintaining account balances. Because there is also a defined transaction order, the account balance has a single defined state at all times.
[0041] In addition to verifying that transactions are valid, blockchain nodes 104 compete to be the first node to create a block of transactions in a process generally referred to as mining, which is supported by "proof of work". At a blockchain node 104, a new transaction is added to an ordered pool 154 of valid transactions that have not yet appeared in a block 151 recorded on the blockchain 150. The blockchain nodes then compete to assemble a new valid transaction block 151 of transactions 152 in the ordered transaction set 154 by attempting to solve a cryptographic puzzle. Typically, this involves searching for a "nonce" value such that when the nonce is juxtaposed with a representation of the ordered pool 154 of pending transactions and hashed, the output of the hash value satisfies a predetermined condition. For example, the predetermined condition may be that the output of the hash value has a certain predefined number of leading zeros. Note that this is only one specific type of proof of work puzzle, and other types are not excluded. A property of a hash function is that it has an unpredictable output relative to its input. Therefore, this search can only be performed by brute force, consuming a large amount of processing resources at each blockchain node 104 that attempts to solve the puzzle.
[0042] The first blockchain node 104 that solves the puzzle announces the puzzle solution on the network 106, providing the solution as proof, which can then be easily checked by other blockchain nodes 104 in the network (once a solution to the hash value is given, it is directly possible to check whether the solution satisfies the output of the hash value). The first blockchain node 104 propagates a block to other nodes that accept the block to reach a threshold consensus, thereby enforcing the protocol rules. The ordered set of transactions 154 is then recorded by each blockchain node 104 as a new block 151 in the blockchain 150. A block pointer 155 is also assigned to the new block 151n pointing to the previously created block 151n-1 in the blockchain. The large amount of work required to create a proof-of-work solution (e.g., in the form of a hash) signals the intention of the first node 104 to follow the blockchain protocol. These rules include not accepting a transaction as valid if it spends or allocates the same output as a previously verified valid transaction, otherwise it is called a double spend. Once created, the block 151 cannot be modified because it is identified and maintained at each blockchain node 104 in the blockchain network 106. The block pointers 155 also impose an order on the blocks 151. Because transactions 152 are recorded in ordered blocks at each blockchain node 104 in the network 106, an immutable public ledger of transactions is provided.
[0043] It should be noted that the different blockchain nodes 104 competing to solve a puzzle at any given time may do so based on different snapshots of the pool 154 of transactions that have not yet been published at any given time, depending on when they began searching for a solution or the order in which they received the transactions. The person solving the corresponding puzzle first defines the transactions 152 included in the new block 151n and their order, and updates the current pool of unpublished transactions 154. The blockchain nodes 104 then continue to compete to create blocks from the newly defined ordered pool of unpublished transactions 154, and so on. In addition, there is a protocol to resolve any "forks" that may occur, where two blockchain nodes 104 solve puzzles within a short time of each other, thereby propagating conflicting views of the blockchain between the nodes 104. In short, the fork direction that is the longest becomes the final blockchain 150. It should be noted that this does not affect users or agents of the network, as the same transaction will appear in both forks.
[0044] According to the Bitcoin blockchain (and most other blockchains), a node that successfully constructs a new block 104 is granted the ability to newly allocate an additional, accepted amount of digital assets in a new special type of transaction that allocates an additional limited amount of digital assets (as opposed to an inter-agent or inter-user transaction, which transfers a certain amount of digital assets from one agent or user to another). This special type of transaction is often called a "coinbase transaction", but may also be called a "starting transaction" or "generating transaction". It usually forms the first transaction of a new block 151n. The proof of work signals the intention of the node that constructs the new block to follow the protocol rules, thereby allowing the particular transaction to be redeemed later. The blockchain protocol rules may require a maturation period, such as 100 blocks, before the special transaction can be redeemed. Typically, a regular (non-generating) transaction 152 will also specify an additional transaction fee in one of its outputs to further reward the blockchain node 104 that created the block 151n in which the transaction was published. This fee is often called a "transaction fee" and is discussed below.
[0045] Due to the resources involved in transaction verification and publication, typically at least each blockchain node 104 takes the form of a server comprising one or more physical server units, or even an entire data center. However, in principle, any given blockchain node 104 may take the form of a user terminal or a group of user terminals networked together.
[0046] The memory of each blockchain node 104 stores software configured to run on the processing device of the blockchain node 104 to perform its corresponding role according to the blockchain node protocol and process transactions 152. It should be understood that any action attributed to the blockchain node 104 herein can be performed by software running on the processing device of the corresponding computer device. The node software can be implemented in one or more applications at the application layer or a lower layer such as an operating system layer or a protocol layer, or any combination of these layers.
[0047] Computer devices 102 of each of the parties 103 acting as consuming users are also connected to the network 101. These users can interact with the blockchain network 106 but do not participate in verifying transactions or constructing blocks. Some of these users or agents 103 can act as senders and receivers in transactions. Other users can interact with the blockchain 150 without acting as senders or receivers. For example, some parties can act as storage entities that store a copy of the blockchain 150 (e.g., having obtained a copy of the blockchain from the blockchain node 104).
[0048] Some or all of the parties 103 may be connected as part of a different network, such as a network overlaid on the blockchain network 106. Users of the blockchain network (often referred to as "clients") may be referred to as being part of a system that includes the blockchain network 106; however, these users are not blockchain nodes 104 because they do not perform the roles required of blockchain nodes. Instead, each party 103 may interact with the blockchain network 106, thereby utilizing the blockchain 150 by connecting to (i.e., communicating with) the blockchain node 106. For illustrative purposes, two parties 103 and their corresponding devices 102 are shown: a first party 103a and its corresponding computer device 102a, and a second party 103b and its corresponding computer device 102b. It should be understood that more such parties 103 and their corresponding computer devices 102 may exist and participate in the system 100, but for convenience, they are not illustrated. Each party 103 may be an individual or an organization. For illustrative purposes only, in this document, the first party 103a is referred to as Alice and the second party 103b is referred to as Bob, but it should be understood that this is not limited to Alice or Bob, and any reference to Alice or Bob in this document can be replaced with "first party" and "second party" respectively.
[0049] The computer device 102 of each party 103 includes a corresponding processing device, which includes one or more processors, such as one or more CPUs, graphics processing units (GPUs), other accelerator processors, specific application processors and / or FPGAs. The computer device 102 of each party 103 also includes a memory, that is, a computer-readable memory in the form of a non-temporary computer-readable medium. The memory may include one or more memory units, which use one or more memory media, such as magnetic media such as hard disks, electronic media such as SSDs, flash memory or EEPROMs, and / or optical media such as optical disk drives. The memory on the computer device 102 of each party 103 stores software, which includes a corresponding instance of at least one client application 105 that is set to run on the processing device. It should be understood that any action attributed to a given party 103 in this article can be executed by software running on the processing device of the corresponding computer device 102. The computer device 102 of each party 103 includes at least one user terminal, such as a desktop or laptop computer, a tablet computer, a smart phone, or a wearable device such as a smart watch. The computer device 102 of a given party 103 may also include one or more other network resources, such as cloud computing resources accessed through a user terminal.
[0050] The client application 105 may initially be provided to the computer device 102 of any given party 103 via a suitable computer-readable storage medium downloaded from a server, for example, or via a removable storage device such as a removable SSD, flash memory key, removable EEPROM, removable disk drive, floppy disk or tape, an optical disk such as a CD or DVD ROM, or a removable optical drive, etc.
[0051] The client application 105 includes at least a "wallet" function. This has two main functions. One of the functions is to enable the corresponding party 103 to create, authorize (e.g., sign) and send transactions 152 to one or more Bitcoin nodes 104, which are then propagated in the network of blockchain nodes 104 and included in the blockchain 150. Another function is to report to the corresponding party the amount of digital assets it currently owns. In an output-based system, this second function includes collating the amounts defined in the outputs of various transactions 152 belonging to the relevant parties scattered in the blockchain 150.
[0052] Note: While various client functions may be described as being integrated into a given client application 105, this is not necessarily limiting, and rather, any client function described herein may be implemented in a suite consisting of two or more different applications, such as interfacing via an API or one application as a plug-in to another application. More generally, the client functions may be implemented at the application layer or at a lower layer such as an operating system, or any combination of these layers. The following description will be based on the client application 105, but it should be understood that this is not limiting.
[0053] An instance of a client application or software 105 on each computer device 102 is operably coupled to at least one of the blockchain nodes 104 of the network 106. This can enable a wallet function of the client 105 to send transactions 152 to the network 106. The client 105 can also contact the blockchain node 104 to query the blockchain 150 for any transactions to which the corresponding party 103 is a recipient (or indeed to check other parties' transactions in the blockchain 150, because in an embodiment, the blockchain 150 is a public facility that provides transaction trust to some extent through its public visibility). The wallet function on each computer device 102 is configured to formulate and send transactions 152 according to a transaction protocol. As described above, each blockchain node 104 runs software that is configured to verify transactions 152 according to the blockchain node protocol and forward transactions 152 for propagation in the blockchain network 106. The transaction protocol and the node protocol correspond to each other, and a given transaction protocol and a given node protocol together implement a given transaction model. The same transaction protocol is used for all transactions 152 in the blockchain 150. All nodes 104 in the network 106 use the same node protocol.
[0054] When a given party 103 (say Alice) wishes to send a new transaction 152j to be included in the blockchain 150, she will formulate the new transaction according to the relevant transaction protocol (using the wallet function in her client application 105). She will then send the transaction 152 from the client application 105 to one or more blockchain nodes 104 to which she is connected. For example, this may be the blockchain node 104 that Alice's computer 102 is best connected to. When any given blockchain node 104 receives the new transaction 152j, it will process it according to the blockchain node protocol and its corresponding role. This includes first checking whether the newly received transaction 152j meets certain conditions for becoming "valid", specific examples of which will be discussed in detail later. In some transaction protocols, the validity conditions may be configurable on a per-transaction basis through a script included in the transaction 152. Alternatively, the conditions may simply be a built-in function of the node protocol, or defined by a combination of script and node protocol.
[0055] If the newly received transaction 152j passes the validity test (i.e., under the “valid” condition), any blockchain node 104 that receives the transaction 152j will add the new verified valid transaction 152 to the ordered transaction set 154 maintained at the blockchain node 104. Further, any blockchain node 104 that receives the transaction 152j will then propagate the verified valid transaction 152 to one or more other blockchain nodes 104 in the network 106. Since each blockchain node 104 applies the same protocol, it is assumed that the transaction 152j is valid, which means that the transaction will soon be propagated throughout the network 106.
[0056] Once in the ordered pool 154 of pending transactions maintained at a given blockchain node 104, that blockchain node 104 will begin racing to solve the proof-of-work puzzle on the latest version of its respective pool 154 containing the new transaction 152 (keep in mind that other blockchain nodes 104 can attempt to solve the puzzle based on different transaction pools 154. However, whoever solves the puzzle first will define the set of transactions included in the latest block 151. Ultimately, the blockchain node 104 will solve the puzzle for a portion of the ordered pool 154, which includes Alice's transaction 152j). Once the pool 154 including the new transaction 152j completes the proof-of-work, it will immutably become part of one of the blocks 151 in the blockchain 150. Each transaction 152 includes a pointer to an earlier transaction, so the order of transactions is also immutably recorded.
[0057] Different blockchain nodes 104 may initially receive different instances of a given transaction and therefore have conflicting views about which instance is “valid” before one instance is published in a new block 151, at which point all blockchain nodes 104 agree that the published instance is the only valid instance. If a blockchain node 104 accepts one instance as valid and then discovers that a second instance has been recorded in the blockchain 150, the blockchain node 104 must accept this and will discard (i.e., treat as invalid) the instance it originally accepted (i.e., the instance that has not yet been published in block 151).
[0058] Another type of transaction protocol operated by some blockchain networks as part of the account-based transaction model can be called an "account-based" protocol. In the account-based case, each transaction is defined not by reference to the UTXO of a previous transaction in the past sequence of transactions, but by reference to an absolute account balance. The current state of all accounts is stored individually in the blockchain by the nodes of the network and is continuously updated. In such systems, transactions are ordered using an account's running transaction record (also called a "position"). This value is signed by the sender as part of their cryptographic signature and hashed as part of the transaction reference calculation. In addition, optional data fields can also be signed in the transaction. For example, a data field can point to a previous transaction if it contains the ID of the previous transaction.
[0059] 2. UTXO-based model
[0060] Figure 2 An exemplary transaction protocol is shown. This is an example of a UTXO-based protocol. A transaction 152 ("Tx" for short) is the basic data structure of the blockchain 150 (each block 151 includes one or more transactions 152). The following will be described with reference to an output-based or "UTXO"-based protocol. However, this is not limited to all possible embodiments. It should be noted that although the exemplary UTXO-based protocol is described with reference to Bitcoin, it can also be implemented on other exemplary blockchain networks.
[0061] In the UTXO-based model, each transaction ("Tx") 152 includes a data structure that includes one or more inputs 202 and one or more outputs 203. Each output 203 may include an unspent transaction output (UTXO), which can be used as a source of input 202 for another new transaction (if the UTXO has not been redeemed). The UTXO includes a value that specifies the amount of a digital asset. This represents a set of tokens on a distributed ledger. The UTXO may also include a transaction ID of its source transaction and other information. The transaction data structure may also include a header 201, which may include an indicator of the size of the input field 202 and the output field 203. The header 201 may also include an ID for the transaction. In an embodiment, the transaction ID is a hash value of the transaction data (excluding the transaction ID itself) and is stored in the header 201 of the original transaction 152 submitted to the node 104.
[0062] Let's say Alice 103a wishes to create a transaction 152j that transfers the relevant amount of digital assets to Bob 103b. Figure 2 In , Alice's new transaction 152j is labeled "Tx1". This new transaction takes the amount of digital assets locked to Alice in output 203 of the previous transaction 152i in the sequence and transfers at least a portion of such amount to Bob. Figure 2, the previous transaction 152i is labeled "Tx0". Tx0 and Tx1 are just arbitrary labels that do not necessarily mean that Tx0 refers to the first transaction in blockchain 151 and Tx1 refers to a subsequent transaction in pool 154. Tx1 can point to any previous (i.e., preceding) transaction that still has unspent output 203 locked to Alice.
[0063] When Alice creates her new transaction Tx1, or at least when she sends it to the network 106, the previous transaction Tx0 may already be valid and included in block 151 of the blockchain 150. The transaction may have been included in one of the blocks 151 at this time, or it may still be waiting in the ordered set 154, in which case it will be included in the new block 151 soon. Alternatively, Tx0 and Tx1 may be created and sent to the network 106 together; or, if the node protocol allows buffering of "orphan" transactions, Tx0 may even be sent after Tx1. The terms "previous" and "successor" as used herein in the context of transaction sequences refer to the order of transactions in the sequence defined by the transaction pointers specified in the transactions (which transaction points to which other transaction, etc.). They may equally be replaced by "predecessors" and "successors", "predecessors" and "descendants", or "parents" and "children", etc. This does not necessarily refer to the order in which they are created, sent to the network 106, or arrive at any given blockchain node 104. However, subsequent transactions (descendant transactions or "child transactions") pointing to a previous transaction (predecessor transaction or "parent transaction") will not be valid unless the parent transaction is valid. A child transaction that arrives at the blockchain node 104 before the parent transaction is considered an orphan transaction. Depending on the node protocol and / or node behavior, it may be discarded or buffered for a period of time to wait for the parent transaction.
[0064] One of the one or more outputs 203 of the previous transaction Tx0 includes a specific UTXO, labeled UTXO0. Each UTXO includes a value specifying the amount of digital assets represented by the UTXO and a locking script that defines the conditions that the unlocking script in the input 202 of the subsequent transaction must satisfy in order for the subsequent transaction to be valid and thus successfully redeem the UTXO. Typically, the locking script locks the amount to a specific party (the beneficiary of the transaction for that amount). That is, the locking script defines the unlocking condition, which typically includes the following conditions: the unlocking script in the input of the subsequent transaction includes the cryptographic signature of the party to which the previous transaction was locked.
[0065] The locking script (also known as scriptPubKey) is a piece of code written in a domain-specific language recognized by the node protocol. A specific example of such a language is called a "Script" (with a capital S), which can be used by the blockchain network. The locking script specifies the information required to spend the transaction output 203, such as the requirement for Alice's signature. The unlocking script appears in the output of the transaction. The unlocking script (also known as scriptSig) is a piece of code written in a domain-specific language that provides the information required to meet the locking script criteria. For example, it can contain Bob's signature. The unlocking script appears in the input 202 of the transaction.
[0066] Thus in the example shown, UTXO0 in output 203 of Tx0 includes a locking script [Checksig P A ], the locking script requires Alice’s signature Sig P A , in order to redeem UTXO0 (strictly speaking, to make subsequent transactions that attempt to redeem UTXO0 valid). A ] contains Alice's public key P from her public-private key pair A The input 202 of Tx1 includes a pointer to Tx1 (e.g., by its transaction ID (TxID0), which in the embodiment is the hash value of the entire transaction Tx0). The input 202 of Tx1 includes an index identifying UTXO0 in Tx0 to identify it among any other possible outputs of Tx0. The input 202 of Tx1 further includes an unlocking script <Sig P A >, the unlocking script includes Alice's cryptographic signature, which is created by Alice by applying the private key of her key pair to a predetermined portion of data (sometimes called a "message" in cryptography). The data (or "message") that Alice needs to sign to provide a valid signature can be defined by the locking script, the node protocol, or a combination thereof.
[0067] When a new transaction Tx1 arrives at a blockchain node 104, the node applies the node protocol. This includes running the locking script and the unlocking script together to check whether the unlocking script satisfies the conditions defined in the locking script (where the conditions may include one or more criteria). In an embodiment, this involves juxtaposing two scripts:
[0068] <Sig PA> <pa>||[Checksig PA]
[0069] Where "||" means concatenation, "<…>" means putting data on the stack, and "[…]" means a function consisting of locked scripts (in this case, a stack-based language). Similarly, scripts can be run one after another using a common stack instead of concatenating the scripts. In either case, when run together, the scripts use Alice's public key P A (included in the locking script of the output of Tx0) to verify that the unlocking script in the input of Tx1 contains Alice's signature when signing the expected portion of the data. The expected partial data itself (the "message") also needs to be included in order to perform this verification. In an embodiment, the signed data includes the entire Tx1 (so there is no need to include a separate element to plaintext specify the signed partial data, as it is already present).
[0070] Those skilled in the art will be familiar with the details of authentication via public-private cryptography. Basically, if Alice has signed a message using her private key cryptographically, then given Alice's public key and the message in plain text, other entities such as node 104 can authenticate that the message must have been signed by Alice. Signing typically involves hashing the message, signing the hash value, and marking this to the message as a signature, thereby enabling any holder of the public key to authenticate the signature. Therefore, it should be noted that in embodiments, any reference herein to signing a particular data fragment or transaction portion, etc., may mean signing the hash value of that data fragment or transaction portion.
[0071] If the unlocking script in Tx1 satisfies one or more conditions specified in the locking script of Tx0 (so, in the example shown, if Alice's signature is provided and authenticated in Tx1), then the blockchain node 104 considers Tx1 to be valid. This means that the blockchain node 104 will add Tx1 to the pending transaction ordered pool 154. The blockchain node 104 will also forward the transaction Tx1 to one or more other blockchain nodes 104 in the network 106 so that it will propagate throughout the network 106. Once Tx1 is valid and included in the blockchain 150, this will define UTXO0 from Tx0 as spent. It should be noted that Tx1 is only valid if it spends unspent transaction output 203. If it attempts to spend an output that has already been spent by another transaction 152, then Tx1 will be invalid even if all other conditions are met. Therefore, the blockchain node 104 also needs to check whether the UTXO referenced in the previous transaction Tx0 has been spent (that is, whether it has formed a valid input for another valid transaction). This is one of the reasons why it is important for the blockchain 150 to impose a defined order on transactions 152. In practice, a given blockchain node 104 may maintain a separate database marking the UTXO 203 of a transaction 152 as having been spent, but ultimately defining whether a UTXO has been spent depends on whether it has formed a valid input to another valid transaction in the blockchain 150.
[0072] If the total amount specified in all outputs 203 of a given transaction 152 is greater than the total amount pointed to by all its inputs 202, this is another basis for failure in most transaction models. Therefore, such a transaction will not be propagated or included in block 151.
[0073] Note that in the UTXO-based transaction model, a given UTXO needs to be spent as a whole. You cannot "leave behind" a portion of the amount defined in a UTXO as spent while spending another portion. However, the amount of a UTXO can be split between multiple outputs of subsequent transactions. For example, the amount defined in UTXO0 of Tx0 can be split between multiple UTXOs in Tx1. Therefore, if Alice does not want to give all of the amount defined in UTXO0 to Bob, she can use the remaining portion to make change for herself in the second output of Tx1, or to pay another party.
[0074] In practice, Alice will also typically need to include a fee for the Bitcoin node 104 that successfully includes Alice's transaction 104 in block 151. If Alice does not include such a fee, Tx0 may be rejected by the blockchain node 104, and therefore, although technically valid, may not be propagated and included in the blockchain 150 (if the blockchain node 104 does not want to accept transaction 152, the node protocol does not force the blockchain node 104 to accept it). In some protocols, the transaction fee does not require its own separate output 203 (i.e., no separate UTXO is required). Instead, any difference between the total amount pointed to by input 202 and the total amount specified by output 203 of a given transaction 152 will be automatically provided to the blockchain node 104 that issued the transaction. For example, assume that a pointer to UTXO0 is the only input to Tx1, and Tx1 has only one output UTXO1. If the amount of digital assets specified in UTXO0 is greater than the amount specified in UTXO1, the difference may be allocated (or spent) by the node 104 that wins the proof-of-work race to create the block containing UTXO1. Alternatively or additionally, this does not necessarily preclude the possibility of explicitly specifying a transaction fee in one of the UTXOs 203 in its own transaction 152.
[0075] Alice and Bob's digital assets consist of the UTXO locked to them in any transaction 152 anywhere in the blockchain 150. Therefore, in general, the assets of a given party 103 are scattered among the UTXOs of various transactions 152 throughout the blockchain 150. No single number defining the total balance of a given party 103 is stored anywhere in the blockchain 150. The role of the wallet function of the client application 105 is to collate the various UTXO values locked to the respective parties and not yet spent in other subsequent transactions. To achieve this, it can query a copy of the blockchain 150 stored at any one Bitcoin node 104.
[0076] It should be noted that script code is often represented schematically (i.e., using an imprecise language). For example, an opcode (opcode) may be used to represent a particular function. "OP_..." refers to a particular opcode of the scripting language. For example, OP_RETURN is a scripting language opcode that, when preceded by OP_FALSE at the beginning of a locking script, creates an unspendable output of a transaction that can store data within the transaction, thereby immutably recording the data in the blockchain 150. For example, the data may include a file to be stored in the blockchain.
[0077] Typically, the inputs to a transaction contain a digital signature corresponding to the public key PA. In an embodiment, this is based on ECDSA using the elliptic curve secp256k1. The digital signature signs a specific piece of data. In an embodiment, for a given transaction, the signature will sign part of the transaction input and part or all of the transaction output. Signing a specific part of the output depends on the SIGHASH flag. The SIGHASH flag is typically a 4-byte code included at the end of the signature that selects the output to be signed (and is therefore fixed at the time of signing).
[0078] The locking script is sometimes referred to as "scriptPubKey", which means that it usually includes the public key of the party to which the corresponding transaction is locked. The unlocking script is sometimes referred to as "scriptSig", which means that it usually provides the corresponding signature. However, more generally speaking, in all applications of blockchain 150, the conditions for UTXO redemption do not necessarily include verification of the signature. More generally speaking, the scripting language can be used to define any one or more conditions. Therefore, the more general terms "locking script" and "unlocking script" may be preferred.
[0079] 3. Side Channels
[0080] like Figure 1 As shown, the client application on each of Alice and Bob's computer devices 102a, 120b can include additional communication functionality. This additional functionality enables Alice 103a to establish a separate side channel 107 with Bob 103b (at the instigation of any party or third party). The side channel 107 enables data to be exchanged away from the blockchain network. Such communications are sometimes referred to as "off-chain" communications. For example, this can be used to exchange transactions 152 between Alice and Bob without registering the transaction (yet) on the blockchain network 106 or publishing it on the chain 150 until one of the parties chooses to broadcast it to the network 106. Sharing transactions in this way is sometimes referred to as sharing "transaction templates". The transaction template may lack one or more inputs and / or outputs required to form a complete transaction. Alternatively or additionally, the side channel 107 can be used to exchange any other transaction-related data, such as keys, negotiation amounts or terms, data content, etc.
[0081] The side channel 107 may be established via the same packet switching network 101 as the blockchain network 106. Alternatively or additionally, the side channel 301 may be established via a different network such as a mobile cellular network or a local area network such as a wireless local area network, or even via a direct wired or wireless link between Alice's and Bob's devices 102a, 102b. In general, a side channel 107 referred to anywhere in this document may include any one or more links via one or more networking technologies or communication media that are used to exchange data "off-chain", i.e., off-chain. The link bundle or collection as a whole may be referred to as a side channel 107 in the case of multiple links. Therefore, it should be noted that if Alice and Bob are said to exchange certain information or data, etc., via a side channel 107, this does not necessarily mean that all of this data must be sent over exactly the same link or even the same type of network.
[0082] 4. SHA2 hash
[0083] SHA2 is a family of cryptographic hash algorithms that converts an l-bit message M∈{0,1} l As input and produces a d-bit summary H∈{0,1} d The length of M can be within a certain upper limit. The digest is of fixed length. "SHAd" is used to denote a cryptographic hash function of the SHA2 family that outputs a digest of size d.
[0084] SHAd is performed in two steps. First, the l-bit message M is divided into N blocks of fixed size m. To do this, padding is required, where:
[0085] M (1) ||…||M (N-1) ||M (N) :=pad(M)
[0086] Filling is defined as follows: Let k be the smallest integer such that l+1+k≡(ml max ) mod m. Append 1 to the end of M, followed by k zeros. Then, append l corresponding to the binary expression of l max The padding results in at most one extra block. If l bits of message M fit into B blocks (each block is m bits), then after padding there are at most B+1 blocks. max Additional blocks are added only if
[0087] The second step of SHAd iteratively applies a compression function after inputting a block of the message and a previously compressed value:
[0088] CF m,d :{0,1} m ×{0,1} d →{0,1} d
[0089] The first compressed value is the initialization vector IV, and it is set to a constant d-bit array specific to each SHAd function. In summary, the SHAd(M) algorithm is as follows:
[0090] 1M (1) ||…||M (N-1) ||M (N) :=pad(M)
[0091] 2. Set H (0) :=IV
[0092] 3. For i = 1 to N, calculate H i :=CF m,d (H (i-1) ,M (i) )
[0093] 4. Output H: =H (N)
[0094] The following table provides the parameters for the SHA256 and SHA512 functions.
[0095]
[0096] 5. Proof System
[0097] 5.1 zkSNARK
[0098] Let P(x;w)=b∈{0,1} be an efficiently computable binary program that takes a bit string x (instance) as public input and another bit string w (witness) as private input, and outputs a decision bit b. If b=1, then P accepts.
[0099] Related NP relations Given by the instance / witness pairs that make program P accept:
[0100]
[0101] A pre-proccessing succinct non-interactive argument system of knowledge (SNARK) for correctly executing a program P is an algorithmic triple SNARK: = (Gen, Prove, Verify) such that:
[0102] Gen(λ,P)→(pk,vk): After inputting the security parameter λ and the description of the program P, it will output a pair of proving key and verifying key.
[0103] Prove(pk,x,w)→π: After inputting the proving key, public input x, and private input w, it outputs proof π.
[0104] Verify(vk,x,π)→b∈{0,1}: After inputting the verification key, the public input x, and the proof π, it either accepts or rejects the proof.
[0105] Completeness, (knowledge) soundness, and zero-knowledge. A SNARK is complete if the verifier always accepts a proof π generated by the prover SNARK.prove on an input pair (x,w) of public input / private input that makes the program P accept. It is sound if, for all public inputs x that do not have private input w that makes P accept, the verifier is very likely to reject any proof π of x. Furthermore, a proof is said to have knowledge soundness if the witness can also be computed (extracted) efficiently from a valid proof π and the randomness that the (possibly cheating) prover used to generate π (up to some negligible error - the knowledge error). A proof is zero-knowledge if the proof π does not reveal any information about w.
[0106] Succinct proofs are "short" proofs. This means that the proof is logarithmically related to the size of the private input w. More specifically, the size of the proof is ploy(λ)polylog(|w|), where λ is a security parameter. A system has succinct verification (also called fully succinct) if, in addition to the short proof, the verifier runtime is also "fast". That is, the size of the public input x is logarithmically related to the size of the private input w. Therefore, a system is fully succinct if the runtime takes poly(λ)polylog((|x|+|w|) steps.
[0107] 5.2 Data carrying proof
[0108] The Proof-Carrying Data (PCD) scheme provides a means for proving the integrity or correctness of dynamic computations distributed across mutually untrusted nodes. This scheme differs from the multi-party computation protocol in two main ways: the number of nodes is not fixed; and there is no need to worry about the privacy of the computation. The latter ensures that PCD is lighter (no node communication overhead is incurred).
[0109] 5.2.1 Multi-predicate transcripts
[0110] Dynamically computed transcripts It is modeled as a directed acyclic graph G = (V, E) that originates from some source nodes and ends at an output (destination) node. The edges (u, v) ∈ E have data attached to them. Each node v ∈ V performs some operation involving the incoming data. Outgoing data out and (possibly) local data z loc The computation at node v must satisfy a predicate Π. That is,
[0111] Definition. A computed transcript is a tuple So that:
[0112] G = (V, E) is a directed acyclic graph
[0113] · is the node label. (The compliance predicate that the node complies with.)
[0114] ·LOC:V→{0,1} * is another node label. (Local data.)
[0115] PAYLOAD:E→{0,1} * are edge labels. (data flowing into and out of a node)
[0116] Messages and Outputs. For an edge (u,v)∈E, a message z attached to that edge has two parts: its type z.type := (TYPE(u)) is the type of the parent node, and its payload z.payload := PAYLOAD((u,v)) is the actual data. Outputs of the transcript is the set of messages attached to an edge (v,w), where w is the output (sink) node.
[0117] Transcript and output compliance. Let the vector of compliance predicates be If the transcript If Compliance:
[0118] i. Let s∈V. TYPE(s)=0 if and only if s is a source node.
[0119] ii. For all non-source nodes v∈V, let i:=TYPE(v), let is the incoming message of v, let Is the outgoing message, and set If it is local data, (Thus, a node must conform to the predicate given by its type.)
[0120] If there exists a transcript with compliance then message z has compliance such that
[0121] Figure 3 is a schematic representation of a computational transcript of the function f(x,y):=(2(x + y),3(x + y)) with bounded noise. The computational transcript includes two source nodes 302, two output nodes 306, and intermediate nodes 304. All non-source nodes 304, 306 enforce different compliance predicates on their inputs and outputs. The "+" nodes (intermediate nodes 304) are allowed to introduce a bounded noise addend ‖e‖2 < B as local data for their computations.
[0122] 5.2.2 Preprocessing PCD
[0123] Syntax. The preprocessing PCD scheme is an algorithmic triple where the generator takes a compliance predicate as input and outputs a pair of proof key / verification key (pk pcd , vk pcd ). For each non-source node, the prover takes its incoming data z in along with the proof (which attests to the compliance of the parent nodes (assuming they are non-source nodes)), local data z loc and outgoing data z out as input and produces a proof The verifier takes the outgoing data and the proof as input and accepts or rejects. Typically, PCD is built according to a succinct non-interactive argument of knowledge (SNARK) that can be recursive. Existing schemes applicable to recursion are provided and compared in Section 8.1.
[0124] Security (knowledge soundness). If the set of proofs of the outgoing data is acceptable, then it is guaranteed that there exists a computational transcript (and its use to compute efficiently) whose output (sink) node has the outgoing data (i.e., ), and all nodes (all the way back to the source nodes) have incoming data / local data / outgoing data that has compliance. Thus, the output set of proofs attests to the compliance of the entire computational transcript.
[0125] 6. Scalable SNARKs for Hash-Based Statements
[0126] 6.1 Knowing Arbitrarily Large SHA2 Preimages
[0127] SNARKs are defined for the following NP relations:
[0128]
[0129] Thus, it is proved that given a public digest H and an l-bit preimage M (private input) of length l, we know that max and the initialization vector IV that is implicitly used in the evaluation of the SHA2 function.
[0130] 6.1.1 Calculating transcripts
[0131] SHA2 evaluation can be viewed as a dynamically computed transcript. The i-th node receives the current iteration counter i-1 and the current state H (i-1) As input and output update i, H (i) , also called the next iteration counter and the next state, where the next state is obtained by using the i-th message block M (i) The first node receives the initialization vector as input, and the last node uses the message length e to fill the last block M (N) And output H (N) .
[0132] Figure 4 A SHA2 transcript 400 is schematically shown with a source node 302 , an initialization node (init node) 402 , intermediate nodes 304 , and a digest (output) node 306 .
[0133] The message M is divided into a series of message blocks M (i) , so that:
[0134] M∶=(M (1) ||…||M (N) )
[0135] Where, if necessary, the padding block M is defined by the following equation ′(N) :
[0136] M′ (N) ∶=pad(M (N) )
[0137] The summary H is defined as:
[0138] H=SHAd(M)
[0139] The message M may be referred to herein as the preimage, and the message block M (i) It can be called the original image block.
[0140] 6.1.2 Node Compliance
[0141] Transcript 400 includes a series of nodes: source node 302 (type 0), initialization node 402 (type 1), intermediate state node 304 (type 2), and summary node 306 (type 3). Transcript 400 provides a method.
[0142] For each non-source type node 402, 304, 306, the compression function standard is enforced to verify that the predefined compression function has been correctly calculated. Each of these nodes 402, 304, 306 sets the current state H (i-1) As input and apply the compression function to calculate the next state H (i) They also take as input the current iteration counter i-1 and increment it to calculate the next iteration counter i.
[0143] Each of these nodes 402, 304, 306 also performs a compression function evaluation check to verify whether the compression function has been correctly evaluated. This check uses the next message block M (i) The node 402 , 304 , 306 generates a proof to verify the correct evaluation of the compression function at the node 402 , 304 , 306 .
[0144] In this way, each of nodes 402, 304, 306 performs a single iteration of the compression function. This allows the proof to be generated iteratively, which can both reduce the computational requirements of the prover, thereby improving the efficiency of the process, and can generate proofs of arbitrarily large messages. The output of the final node 306 is the digest H and the preimage proof π preim , in addition to proving that the final node 306 executed correctly, it also proves that all previous nodes 402, 304 have executed the compression function correctly and therefore prove that the message M is known.
[0145] In addition to verifying the compression function, the first (initialization) node 402 and the final (digest) node 306 also perform additional verification.
[0146] The first node 402 performs an initialization check to check whether the received initialization vector IV (received as part of its input) is correct. The received initialization vector can be referred to as the current state of the first node 402, i.e., H (0) =IV. The received initialization vector IV is compared with the predefined initialization vector, and if the results are equal, it is determined that the received initialization vector is correct. The predefined initialization vector may be hard-coded into the first node 302. The first node 302 may also verify whether the current iteration counter received at the first node 302 has a first iteration count value, i.e., i in =0. In some embodiments, the first iteration count value may be 1.
[0147] The last node 306 performs a padding check. If the pre-image fits into N m-bit blocks, then the last node 306 consistently uses l to pad the last block M (N) , so the padding length k+1 makes:
[0148] l+k+1=ml max +(Nb)·m
[0149] Where m is the block length, and N-1 is the input iteration counter of the final node 306. Here, b is 0 or 1. If b=1, no extra blocks are used for padding, i.e., the bit length l of the message M is equal to the maximum bit length l max If the above equation is satisfied, the filling condition can be considered to be met.
[0150] To perform padding checking, final node 306 may receive l, b, and / or k as input and check whether the received values satisfy the following equation:
[0151] l+k+1=ml max +(Nb)·m
[0152] If the required padding does not fit in the final message block, an additional padding block is required. If the message length l is max This can happen if the message length is a multiple of 0 or if the padding bits for the message length do not fit into the last message block.
[0153] If additional padding blocks are needed, the final node 306 also receives the padding blocks, also called the padding preimage portion. The length of the padding blocks is such that when concatenated with the message, the total length is equal to the maximum bit length. The final node 306 executes the compression function again, this time taking as input the state and iteration counter that the final node 306 has already calculated. The compression function check uses the padding blocks to check whether the compression function has been correctly evaluated. The proof generated by the final block 306 confirms that both instances of the compression function have been correctly evaluated.
[0154] Conversely, if no additional padding blocks are needed, the final node 306 outputs the state and proof generated relative to the last message block of the message. That is, the final node 306 does not need to perform the compression function a second time.
[0155] The final node 306 verifies the final l of M max The bit is the binary representation of whether l. The definition of M depends on whether padding has been added, as explained below. That is, if no final block is added, the slice of checked bits is located at , and if the final block is added, the checked bits are located in Somewhere in .
[0156] That is, the correct value of k must be provided as input to the final node 306, whereas the value of M' can be set to any value if no padding is added. This is because in the absence of padding, M' is not used for padding enforcement nor for the second compression function evaluation.
[0157] It should be noted that SHA2 always adds some extra bits (called padding) to the end of the message. The difference is in which block the padding check is enforced. is the last block containing message bits, then (i) the padding fits entirely within the block Or (ii) an additional padding block M' is required. In case (i), In case (ii), padding is enforced in M′.
[0158] More generally, the final node 306 receives the final message block and the message length and generates the correct padding for the message based on the message length - Φ described below. pad Step 2 of . This may involve generating additional padding blocks containing padding information. The final node 306 then applies the compression function to the last message block and also to the padding blocks if there are additional padding blocks. The final node 306 generates a proof that verifies the correctness of the padding and the application of the compression function. The proof is based on the received additional input bits b that indicate whether an additional block containing only pure padding information has been generated and passed through the compression function.
[0159] An intermediate state node 304 can only receive input from another intermediate state node 304 or from the first node 402. A digest node 306 can receive input from all types of nodes 302, 304, 402. These input relationships of a SHA2 node are as follows: Figure 5 shown.
[0160] The predicate is defined as To capture these enforcements. Internal toolsΦ IV ,Φ eval ,Φ pad As explained below.
[0161] Initialize the compliance of node 402 (type 1).
[0162] Π init (z in ,z loc ,z out ):
[0163] 1. z in .payload is parsed into counter and status
[0164] 2. z loc Parsing into message blocks
[0165] 3. z out .payload is parsed into counter and status
[0166] 4. Verify z in .type=0.
[0167] 5. Verification Accept or not. / / See Error! Reference source not found.
[0168] 6. Verification Do you accept?
[0169] 7. If all three checks are accepted, output "accept". Otherwise, output "reject".
[0170] Compliance of intermediate state node 304 (type 2).
[0171] Π update (z in ,z loc ,z out ):
[0172] 1. z in .payload is parsed into counter and status
[0173] 2. z loc Parsing into message blocks
[0174] 3. z out .payload is parsed into counter and status
[0175] 4. Verify z in .type∈{1,2}.
[0176] 5. Verification Do you accept?
[0177] 6. If both checks are acceptable, output "accept". Otherwise, output "reject".
[0178] Compliance of summary node 306 (type 3).
[0179] Π digest (z in ,z loc ,z out ):
[0180] 1. z in .payload is parsed into counter and status
[0181] 2. z loc Parses into the message block, the extra padding block (if any), the extra hash state (only used if there is an extra padding block), a counter, the padding length, and an indication of whether the extra block was added position.
[0182] 3. z out .payload resolves to the last state and message length
[0183] 4. Verify z in .type∈{0,1,2}.
[0184] 5. If z in .type=0, then check Accept or not. / / See Error! Reference source not found.
[0185] 6. If b=1, check: Do you accept?
[0186] 7. Otherwise, (b=0, so add extra blocks), check:
[0187] a. Do you accept?
[0188] b. Do you accept?
[0189] 8. Verification Do you accept?
[0190] If all checks are accepted, then output "accept". Otherwise, output "reject".
[0191] The following table shows Tools used internally. Block length m, digest length d, maximum message length l max and the initialization vector IV are hardcoded in the description.
[0192]
[0193] 6.1.3 SNARK
[0194] set up is the PCD scheme to prove the message Compliance. The proof system used to prove knowledge of the SHA2 preimage is the algorithmic triple SNARK defined below preim :=(Gen preim ,Prove preim ,Verify preim ).
[0195] Gen preim (λ,CF SHA2 )→(pk,vk). It combines the security parameter λ and the description of the compression function of SHA2 CF SHA2 As input, and output the proof key pk (including CF SHA2 ) and verification key vk (including CF SHA2 A concise summary of
[0196] Prove preim (pk,(H,M))→π SHA2 It will prove the key pk and the Takes as input and outputs a succinct proof π SHA2 .step:
[0197] 1. Divide the l-bit message M into N blocks M (i) (Each block is m bits). Let k+1 be M (N) Filling length.
[0198] 2. Set H (0) :=IV
[0199] 3. Set up Pi (0) :=⊥ / / Empty proof.
[0200] 4. For i=1 to N, perform the following operations:
[0201] a. Calculate H (i) :=CF m,d (H (i-1) ,M (i) ) / / Assume H (N) =H.
[0202] b. Set input data, local data and output data:
[0203] i. Set z in .payload := (i - 1, H (i-1) )
[0204] ii. If i < N, then set z loc := M (i) and z out .payload = (i, H (i) ) / / Type 1 or type 2 node (non - summary)
[0205] iii. Otherwise, if i = N, then set z loc := (M (i) , M′, H, i,, k, b) and z out .payload = (H (i) , l)
[0206] c. Set the node type:
[0207] i. If i = 1 & N ≥ 2, then set z in .type = 0 and z out .type = 1 / / Initialize node
[0208] ii. If i = 2 & N ≥ 3, then set z in .type = 1 and z out .type = 2: / / First intermediate state node
[0209] iii. If N > i > 2 & N ≥ 3, then set z in .type = 2 and z out .type = 2: / / Remaining intermediate state nodes
[0210] iv. If i = N & N ≥ 3, then set z in .type = 2 and z out .type = 3 / / Summary node (with input from intermediate state nodes)
[0211] v. If i = 2 & N = 2, then set z in .type = 1 and z out .type = 3: / / Summary node (with input from source node)
[0212] vi. If i = 1 & N = 1, then set z in .type = 0 and z out .type = 3 / / Summary node (with input from source node)
[0213] vii. Interpret pk as pk SHA2PCD And calculate
[0214]
[0215] 5. Output π SHA2 : =π (N) .
[0216] It is more efficient for the prover to keep the data corresponding to the current iteration in memory and delete the data of the old iterations. In this way, the output proof π can be calculated in an incremental manner SHA2 .
[0217] Verify preim (H,π SHA2 ,vk)→{"accept","reject"}. It will verify the key vk, the digest H and the proof π SHA2 As input and accept or reject the proof. Acceptance signals that H was correctly computed using SHA2 from the preimage M (which was not provided to the verifier). Steps:
[0218] 1. Parse vk to vk SH A 2PCD .
[0219] 2. Set z out .type:=2 (summary node) and z out .payload:=H.
[0220] 3. Run If accepted, output "accept". Otherwise, output "reject".
[0221] Figure 6 An exemplary method for proving that prover 602 knows preimage M without revealing the preimage to verifier 604 is shown.
[0222] In step 1, the verifier 604 executes Gen preim , to generate a proof key pk and a verification key vk based on a compression function. The compression function used is known to both the prover 602 and the verifier 604. In step 2, the verifier 604 provides the proof key pk to the prover 602, or otherwise makes the proof key available to the prover.
[0223] In step 3, the prover 602 generates a series of original image blocks M (i) If necessary, the prover 602 may also generate padding blocks in this step.
[0224] In step 4, the prover 602 iterates the compression function and generates a corresponding proof for each iteration to prove that the compression function has been correctly executed. As described above, each iteration is executed as a node 402, 304, 306 of the transcript 400. In step 5, the output proof of the final node 306 is set as the original image proof.
[0225] In step 6, the prover 602 provides the verifier with the pre-image proof and the next state generated by the final node 306 (ie, the message digest H). The verifier 604 performs the Verify preim , to verify that the proof is valid for the digest and therefore verify that the prover 602 knows the message M.
[0226] Although shown as a single entity, it should be understood that prover 602 may include multiple computing devices, each of which includes a processor. Each of these computing devices of prover 602 may be configured to execute one or more nodes of transcript 400. The output of each node may be sent to the processor of prover 602 for input to the next node.
[0227] 6.2 Patterns in SHA2 Preimages
[0228] The method described above can be modified to prove that a pattern exists in the preimage of a given digest d. For example, one can prove the statement "the first and last bits of the preimage of d are equal to 1". In general, the method described below provides a method for proving any bit pattern in the preimage M of a digest d and verifying a brute force execution knowing only d but not M.
[0229] 6.2.1 Mode
[0230] Start by defining a brief description of how the pattern is viewed and how it is computed (summary).
[0231] model Represented as two l-bit vectors C:=(C1,…,C l )), the first vector P is the pattern, and the second vector C is the check bit. If every bit C i =1 when M i =P i , then the l-bit string M is consistent with the pattern. The vector P can take any value for the unchecked bits, for example, all unchecked bits of P are set to 0. In other words, the check bit vector C defines which bits of the message are checked, that is, compared with the pattern vector P.
[0232] Definition 2 (Pattern and Summary). Let k be a security parameter, and let l, N, m be integers such that Nm ≥ l. Let Hash: {0,1} 2m+k →{0,1} k Is a collision-resistant hash function. l Mode is a pair of l-bit vectors
[0233] set up is an Nm-bit array that is divided into m-bit blocks, which are obtained by Get, that is, if C j =1, the i-th j Bit is set to p j , otherwise, set it to 0. Similarly, set is the result of dividing C into m-bit blocks and padding the last block with 0 if necessary. Let S (0) is any k-bit string. The summary is a k-bit string So that:
[0234]
[0235] In this article, it may be referred to as a bit pattern array, which includes a series of bit pattern array blocks. In this article, it may be referred to as a check bit array, which includes a series of check bit array blocks.
[0236] Note. As with SHA2, the construction of the summary S follows the Merkle-Damgard structure. Therefore, it can be based on the i-th block P (i) ,C (i) and the (i-1)th intermediate state S (i-1) To calculate the i-th intermediate state S of the summary (i) That is, S (i) :=Hash(P (i) ,C (i) ,S (i-1) ).
[0237] 6.2.3 Statements
[0238] The method presented in this paper proves the following statements:
[0239] "set up Then H is the same as the mode A digest of the consistent message M".
[0240] More formally, this paper provides a proof system to prove instances of the following relation
[0241]
[0242] The length e of the pre-image M is given by the pattern vector The length is given.
[0243] 6.2.4 Calculating transcripts and compliance
[0244] Replace patterns with summaries to justify variable-length statements. The underlying PCD scheme defined in this paper is suitable for use with summaries of patterns rather than with full descriptions. It should be noted that in order to enforce the mode on the ith message block of the SHA2 preimage Just need to know The i-th vector P (i) , C (i) However, in order to ensure that P (i) 、N (i) Is the original public mode part of The previous vector of is enforced for the previous message block (rather than for some other mode).
[0245] to this end, All i-1 previous vectors of can be passed as input to the compliance predicate and output for subsequent iterations. However, this approach will require as many compliance predicates as N (each of which takes an input of a different length), and preimages of different lengths will require different numbers of compliance predicates. The latter means that it is not possible to use the same SNARK pattern Scheme to prove patterns on any two preimages.
[0246] To solve this problem, the i-1 intermediate states of the summary are passed, and the correct generation of the next summary state is enforced in order to check the consistency of the previous pattern block. (i) All have fixed length k, so a single compliance predicate is sufficient.
[0247] The actual construction. The definition of the transcript used to verify the pattern is very similar to the transcript used for the SHA2 preimage (see Section 6.1.2). The difference is that the constraint predicate And the internal logic of edge and node data. Specifically:
[0248] In addition to receiving the ith message block, the ith node (iteration) also receives and The ith m-bit block of is used as local data. Therefore, Furthermore, the side message (outgoing data) now contains the i-th intermediate state of the digest in addition to the SHA2 digest and the iteration counter. in :=(i in ,H in ,S in ), and z out :=(i out ,H out ,S out ).
[0249] · For the predicates in Section 6.1.2 The internal subroutine Φ in eval The call to the subroutine Φ defined below is replaced by pattern It should be noted that the subroutine Φ eval As a subroutine Φ pattern Execute one of the steps in the process.
[0250] The following table provides the predicates used to ensure schema consistency in Used internally as a subroutine. Calculate according to Definition 1.
[0251]
[0252] That is, each node 304, 306, 402 is configured to perform the pattern check as well as the compression function check set forth above.
[0253] Each node receives as additional input a block of bit pattern arrays and a block of check bit arrays corresponding to that block. Each node generates the next summary value S out , and by based on the current summary value S in , bit pattern array and check bit array generate a hash to verify whether the next summary value is evaluated correctly.
[0254] Nodes 304, 306, 402 also use the corresponding check bit array blocks and bit pattern array blocks to check the pattern of the message blocks they are processing. (i) , the i-th message block M (i) and the i-th bit pattern array block P (i) Make a comparison.
[0255] Optimize Φ pattern The size of . The pattern summary can be calculated using a zero-knowledge friendly hash function Hash. This will make the compliance predicate The size and The size of SHA2 is closely related. For example, the Pedersen hash has an R1CS of 2753 constraints. Poseidon has 316 constraints. The downside is that the cryptanalysis of these new constructions is not as well studied as that of SHA2's compression functions.
[0256] 6.2.5 SNARK
[0257] set up is the PCD scheme to prove the message Compliance. Proof system SNARK pattern :=(Gen pattern ,Prove pattern ,Verify pattern ) is defined similarly to the preimage proof system described previously.
[0258] Verifier 604 generates the proof key and verification key as set forth above. The proof key is then provided to prover 602 for use in proving knowledge of the preimage M. Prover 602 also computes the corresponding state of the pattern profile S at each compression function iteration.
[0259] The final node 306 of the transcript provides outputs including the final state H (N) (Abstract H), Final Summary S (N) (Summary(P)) and proof of pattern π pattern The output is provided to a verifier 604, which verifies the pattern proof based on the output. In the method, And the verifier 604 runs To verify z out Given proof of π out .
[0260] 6.3 Merkle Tree Statements
[0261] set up is an NP language. The following method provides a method for proving the following statement:
[0262] "Let H be a byte array, and let e be a non-zero positive integer, then H is the root of a Merkle tree of depth e, whose leaves are ”.
[0263] Therefore, statements about all leaves can be proved.
[0264] Proving a Merkle tree statement by sending leaves is not efficient: first, a tree needs to be built to verify a given root; second, a cardinality relation needs to be To verify 2 e For example, for a tree storing 1 million leaves (each leaf is 1MB), at least 1TB needs to be sent (considering only the leaf data, not the proofs) (which is inefficient) and may not be achievable. The situation is similar for smaller trees storing larger datasets.
[0265] More formally, given the relationship between leaves Provide a concise proof system for the following relation:
[0266]
[0267] Variable length statements. The depth e of the tree is not specified by the relation but is part of the instance. Therefore, contains Merkle trees of arbitrary depth, which in turn implies the proof system SNARK described below merkle It can prove the cardinality relationship Any number of instances of .
[0268] 6.3.1 Guiding from leaf relations to the Merkle tree
[0269] From the leaves Initially, all leaves are cardinality relations , and consider the input The transcript produced by the GetRoot circuit is computed above. The source (leaf) node takes the leaf as input and hashes it. The other (non-leaf) nodes take the two digests (from the two child nodes) as input and hash them. As NP is adopted, it allows SNARKs, so the circuit can be modified as follows. The source node receives data L and a valid proof π b As input, this valid proof proves the statement This means that The complexity of the SNARK prover depends only on the complexity of the base SNARK verifier and the hash.
[0270] Note 1. The method provided in this paper pre-computes the leaf proof π b Another possibility is to assume that the cardinality relation has a predicate vector and Enhancement to adapt the PCD scheme of circuit GetRoot. However, the resulting prover may be more complex.
[0271] Note 2. Use for cardinality relationship A preprocessing (concise) verifier for π that takes a verification key vk as input to verify the proof π b . Therefore, a SNARK for the following relation is provided:
[0272]
[0273] Authentication key exist is hard-coded in the description. The knowledge rationality of the SNARK verifier means that if but (Very likely). This is simply because it is very likely to extract a witness from a valid proof.
[0274] 6.3.2 Merkle Tree Calculation Transcript
[0275] Let Hash:{0,1} 2k →{0,1} k is a cryptographic hash function. The compliance predicate vector is defined as
[0276] Figure 7 An exemplary Merkle tree 700 described herein is provided. The Merkle tree 700 includes four leaves 702, each leaf defining leaf data (also referred to herein as a data block) L i And the corresponding data block proof π i The Merkle tree 700 includes four leaf hash values, to which each of the four leaf nodes 704 is respectively mapped, wherein each leaf node 704 has an associated leaf 702 and is configured to receive block data and data block proofs from the associated leaf 702. The Merkle tree 700 also includes an internal hash to which the internal node 706 is mapped. The nodes 704, 706 mapped to the Merkle tree 700 are arranged in layers, and the nodes of each layer receive as input the output generated by a pair of nodes 702, 704 of the previous layer.
[0277] Data Block Proof π i Confirm data block L i Satisfies a predefined criterion. For example, the criterion may be that the data block matches a predefined pattern, as described in Section 6.2, where the data proof verifies that the data block matches the pattern.
[0278] Each leaf node in the leaf node 704 receives a corresponding data block L i And its associated data proves π i Each leaf node 704 verifies the received proof π i Is it valid? And hash the received data block to generate a data block hash H i :=Hash(L i ).
[0279] The internal nodes 706 in the first layer of internal nodes 706 each receive a data block hash generated by two of the leaf nodes 704. These internal nodes 706 generate a hash in the data block hash, referred to herein as an output hash, which is then provided to the internal nodes 706 in the next layer of internal nodes 706.
[0280] This process is repeated, with each internal node 706 receiving two hash values generated by the internal node 706 of the previous layer, until the final internal node 706a arranged in the last layer of nodes 704, 706 mapped to the Merkle tree 700 generates its hash value, which is the Merkle root of the Merkle tree 700.
[0281] Each of nodes 704, 706 may also compute a proof.
[0282] Each leaf node 704 receives the received data block L i The associated data proves π i , and generates a leaf node certificate that verifies that the node outputs the hash of the input data and that the node has successfully verified the input data certificate. That is, each leaf node certificate verifies that (1) the input leaf certificate is valid (the verification algorithm outputs a 1 on the certificate), and (2) the output block hash is the hash of the input data. Therefore, the leaf node certificate verifies that the data block L i is a leaf of the Merkle tree 700 and verifies that the data block itself meets the predefined criteria.
[0283] Each internal node 706 of the first layer receives a corresponding leaf node certificate and a leaf hash. These internal nodes 706 generate a certificate based on the two received leaf node certificates, referred to herein as an output certificate. Each output certificate confirms that (1) the input certificates are valid, and (2) the output hash is the hash of the two input hashes.
[0284] In a similar manner to generating the output hash, each leaf node 706 of each subsequent layer receives as input two output proofs corresponding to the received hashes generated by the internal node 706 of the previous layer. Each internal node 706 generates an output proof based on the two received proofs. In this way, each output proof confirms the existence of the block data value and the previous hash, and confirms that the leaf data block meets the criteria.
[0285] The output proof generated by the final node 706a confirms that the output hash is the root of the Merkle tree, and the leaves of the Merkle tree meet the criteria. The output proof may be referred to herein as the Merkle tree proof of the Merkle tree 700.
[0286] Nodes 704, 706 may be executed by the same computing device. Alternatively, one or more of these nodes may be executed by different computing devices. In this embodiment, the output hash and proof are sent between computing devices to generate a Merkle tree proof and a Merkle root. The code defining the Merkle tree may be divided into multiple parts, each part defining one of the nodes 704, 706 of the Merkle tree 700, and each computing device stores and executes one or more parts of the code corresponding to the node 704, 706 being executed by the computing device.
[0287] Leaf node 704 (type 1). Leaf node 704 converts data L∈{0,1} 2k And data proves π b (For the statement ) as input z in , calculate the leaf hash H:=Hash(L) and output z out :=(H,0). If L has m<2k bits, then it is padded to the right with 2k-m zeros before hashing. All of these checks are in the predicate Π leaf Specifically, the SNARK verifier’s circuit enforces the cardinality relation to achieve π b The validity of the circuit is verified (the verification key is hard-coded in the circuit), and the correctness of H is enforced by the circuit against the Hash.
[0288] Internal node 706 (Type 2). Internal node 706 takes two inputs and where e≥1 represents the depth of the internal node 706 in the Merkle tree 700, and H (l) ,H (r) ∈{0,1} k . Calculate H: = Hash (H (l) ||H (r) ) and output z out :=(H,e). If e=1, the input comes from two leaf nodes 702. Otherwise, the input comes from the internal nodes 704 of the previous layer. All these checks are in the predicate Π inner Encoding is performed in .
[0289] Hash the leaf data with a size > 2k. The domain of Hash is fixed to 2k. i If the size of > k is large, it can be double-hashed. Therefore, Hash var Set to a cryptographic hash for which knowledge of the preimage can be incrementally proven (e.g., SHA2, with a SNARK from Section 6.1.3). Input Proof π i Prove the following statement "Given a public Existence L i ,w i , so that and Observe the size of leaf data |L i |≠|L j |Not necessarily the same.
[0290] Choice of hash function. As in the case of proving patterns in SHA2 preimages, a zk-friendly hash function such as Pedersen hash or Poseidon can be used in the Merkle tree construction. It should be understood that any hash function can be used.
[0291] 6.3.3 SNARK Proof System
[0292] set up is a PCD scheme that proves that the output (H, e) has Compliance, and set (G b ,P b ,V b ) is a basic SNARK verifier. Relation The SNARK proof system is a triple SNARK merkle :=(Gen merkle ,Prove merkle ,Verify merkle ).
[0293]
[0294] 1. As a basic Generate a key.
[0295] 2. For the Merkle tree Generate a key. / / Predicate In it, the basic verification key vk b Hard coded.
[0296] 3. Output pk: =(pk pcd ,pk b ), vk:=vk pcd .
[0297]
[0298] 1. Analyze pk: =(pk pcd ,pk b ).
[0299] 2. For i = 1 to 2 e , calculate π b,i :=P b (pk b ,L i ,w i ) / / Offline prover.
[0300] 3. Calculate the output proof of the leaf node. For i = 1 to 2 e , perform the following operations: / / Total 2 e leaf nodes.
[0301] a. Let L i , π b,i are the i-th leaf data and leaf relationship respectively Set the input z of the leaf node in, i .payload:=(L i ,π b,i ) and z in,i .type=0 (source node).
[0302] b. Set Setting up the output node And z out, i .type=0 (leaf node).
[0303] c. Calculation output proof
[0304] 4. Calculate the output proof of the internal nodes. For d=1,…,e, repeat the process / / from layer d-1 to layer d.
[0305] a. Input / proof pair As input 2 e-(d-1) .
[0306] b. Take 2 e-d Node output payload / / It should be noted that
[0307] c. For k = 1 to 2 e-d , perform the following operations:
[0308] i. Set the input data to If d=1, the input node is of type 1 (leaf node). Otherwise, the type is 2 (internal node).
[0309] ii. Set the input proof to
[0310] iii. Settings
[0311] iv. Computational Output Proof
[0312] d. Output
[0313] Verify merkle (vk,(H,e),π merkle )→{"accept","reject"}. It will verify the key vk, the digest and the tree depth (H,e) and the proof π merkle as input. Accepting means that H is the root of a Merkle tree of depth e, whose leaves are Instance of .
[0314] step:
[0315] 1. Interpret vk as pk pcd .
[0316] 2. Set z out .type:=2 (internal node) and z out .payload:=(H,e).
[0317] 3. Run If accepted, output "accept". Otherwise, output "reject".
[0318] Figure 8 The proof for each data block L is shown i Example methods that meet the criteria. Figure 8 In the example, the standard is a predefined pattern.
[0319] In step 1, the verifier 604 generates a proof key pk for the criteria that the leaf data must satisfy b and verification key vk b The verifier also generates the proof key pk of the Merkle tree 700 pcd and verification key pk pcd In step 2, the two proof keys pk b , pk pcd The two attestation keys are sent to the attestor 602, or otherwise made available to the attestor.
[0320] In step 3, the prover 602 generates data proofs for each of the data blocks. To generate these proofs, the prover 602 writes the corresponding check bit array block C i Each data block L defined i The bits of the corresponding pattern bit array block P i If the bits match, then the block of data satisfies the pattern criteria and a proof can therefore be generated.
[0321] In step 4, the prover 602 traverses the Merkle tree 700. That is, the prover 602 executes the leaf nodes 704 and the internal nodes 706 to generate the Merkle root and the Merkle tree proof in step 5, as generated by the final node 706a.
[0322] In step 6, the prover 602 sends both the Merkle tree proof π merkle and the Merkle root H to the verifier 604. In step 7, the verifier 604 uses the Merkle root and the Merkle root verification key (generated in step 1) to verify the received Merkle tree proof. In this way, the verifier 604 is convinced that the data blocks used by the prover 602 to generate the Merkle root and the Merkle tree proof satisfy the pattern criteria.
[0323] It should be understood that the criteria that the data block L i must satisfy can be any criteria that can generate a zero - knowledge proof.
[0324] 6.3.4 Proof Aggregation and Universal Trees
[0325] The design described above has two important features.
[0326] Aggregate proof. Two proofs π merkle , v′ merkle for (H, e), (H′, e′) (where e = e′, i.e., the Merkle trees have the same depth) can be combined and a proof π″ for (H″ := Hash(H, H′), e + 1) can be produced by a single call to the PCD prover merkle . It should be noted that H″ is the root of the Merkle tree whose 2 e+1 leaf nodes are in a base relationship If the tree depths are different, e.g., e < e′, then 2 e′-e dummy leaves can be used to duplicate the smaller tree to generate an augmented tree with root H replicated and depth e′, and then the two proofs can be combined. Additionally, the correct augmentation of the smaller tree can be proven incrementally.
[0327] Prove any base relationship. This relationship hard - codes the verification key of a specific relationship as part of its description. A general SNARK for the base relationship can be used to decouple the description from In a general SNARK, there is a common procedure specialize that takes a circuit - independent (general) verification key vk and produces a circuit - specific verification key Therefore, the generic vk can be hard-coded in the circuit and the correct specialization to This allows the circuit specific authentication key is considered as part of the instance. In other words, through a single SNARK proof system, it can be proved that the leaves of the Merkle tree are in any NP language. The general tree relationship is as follows:
[0328]
[0329] Therefore, when the circuit changes, there is no need to change the authentication key.
[0330] 6.4 Possible modifications
[0331] Patterns in intermediate hash states. The ideas in Section 6.2 can be used to prove that the intermediate state H (i) The (internal) compliance predicate Φ of the i-th node midstatePattern It is possible to enforce the outgoing intermediate state H (i) With mode Specifically, it can be proved that a given d-bit string H mid is the ith intermediate state of a given summary H.
[0332] Prove that a keyword search or string does not appear in the preimage. It can be shown that a given short string S of at most m bits appears in some of the SHA2 message blocks (or not). The idea of the compliance predicate is to right-shift the string S by 1 bit, loop mp times, and check if it matches the corresponding p-bit slice of the message block. For example, this can be used to prove that a transaction with an identifier TxID of unknown size is a P2PKH transaction matching a pattern of a P2PKH script (4 bytes), or to prove that it does not contain embedded data indicating a 2-byte string "OP_FALSE OP_RETURN" in the serialization of this transaction.
[0333] Merkle tree proofs of variable size. In addition, consider statements of the form "given a public (H, e, L, i), I know a certification path ap proving that L is the i-th leaf of a Merkle tree with root H and depth e. Furthermore, I know a witness w such that L is Similar to the Merkle tree statement in Section 6.4, the private certification path is variable in size. This can be used for zk-rollup, where accounts are leaves of the Merkle tree and account transfers imply proving knowledge of the Merkle tree proof. Variable-sized Merkle tree proofs allow zk-rollup to handle batches of varying sizes (i.e., the batch size is independent of the instantiation of the underlying SNARK system).
[0334] The relationship between leaves depends on their position in the tree. Therefore, the i-th leaf and the j-th leaf are and Instances of (not necessarily in the same relationship).
[0335] 7. Application
[0336] Some exemplary applications for the above zero-knowledge proof system are provided. It should be understood that these examples are non-limiting. The above proof system is particularly useful for applications in which larger data is encrypted. In known methods, proving that the data is correctly encrypted requires multiple iterations, which is inefficient in terms of time and computation, and may not even be achievable for certain data sizes.
[0337] The above method solves this problem by hashing the data and proving that the prover knows the preimage of the hash.
[0338] 7.1 Scalable Zero-Knowledge Contingent Payments
[0339] Maxwell's contingent payment scheme. The zero-knowledge contingent payment scheme (ZKCP) known in the art is carried out in two steps:
[0340] (1) Buyer Alice specifies the requirements for the data she wants to purchase, e.g., Φ(data, public) = 1.
[0341] (2) The seller Bob sends the (symmetric) ciphertext ct and digest d along with the zkSNARK, proving that the data encrypted by the ciphertext meets the buyer’s requirements and that the symmetric key used for encryption is the preimage of the transmitted digest.
[0342] Once the buyer verifies the zkSNARK, he uses digest d to establish a hashed time lock (HTLC) transaction on the BSV blockchain for the agreed amount. When the seller redeems the funds, he also reveals the symmetric key (the preimage of the digest) and the buyer can decrypt the purchased data.
[0343] Examples of real-world use cases with larger datasets include:
[0344] Movies in HD (or lossless) format;
[0345] Complex proprietary software.
[0346] The requirement in both cases is that the SHA256 digest equals some known bit string h * Therefore, if SHA256(data)=h * , then Φ(data,h * )=1.
[0347] Source of inefficiency. The problem with this approach is that if the data is large (as in the above example), it is expensive to prove in a zero-knowledge manner that the encryption circuit as a whole was correctly evaluated. Encrypting just 1MB of data using a 128-bit block cipher (e.g., AES-CTR) in counter mode requires 65536 iterations of the block cipher.
[0348] Solution. The encryption of the data is proved in an incremental manner. Since the prover adopts an incremental approach, it can handle data of arbitrary size in a scalable manner. In more detail, the data is encrypted using a one-time-pad (OTP) encryption scheme. The key required for this OTP is as long as the data. In order to avoid using an overly large key to redeem the HTLC transaction, a key stretching step can be introduced. Therefore, the data is encrypted using the output key material okm, which is an extension of the short (e.g., 128-bit or 256-bit) input key material ikm using a key derivation function (HKDF).
[0349] HKDF is well known in the art and will not be described in detail herein. In short, HKDF consists of two steps. In the first step, a fixed-length pseudo-random key prk is extracted from the input key material ikm. This step can be performed by HMAC ext Node 906 implements. In the second step, the fixed-length pseudo-random key is expanded into several additional pseudo-random keys H i This step can be performed by multiple HMAC exp Node 908 implements the output key material okm including these additional pseudo-random keys H i .
[0350] What is placed on the chain is the hash of the (short) ikm. That is:
[0351]
[0352] okm:=HKDF(ikm);
[0353] d:=SHA2(ikm).
[0354] Fig. 9 Multi-predicates for efficient and scalable ZKCP are shown has The source nodes are represented by white circles, and the output nodes are represented by black circles. The data is data:=(pt1,…,pt N ), the generated ciphertext is ct:=(ct1,…,ct N ).pt i and ct i is a block of h bits, where h is the range of the underlying hash function used in HKDF, and
[0355] The improvements provided by this method are in two aspects.
[0356] 1) Use recursive zkSNARKs to incrementally prove that the data is correctly encrypted. This means that even when dealing with larger data, the prover's hardware requirements can be very limited. More specifically, the correct hashing of the input key material ikm and the XOR operation of the data and the output key material okm are incrementally proved. The transcript to be proved is Fig. 9 The transcript distinguishes four types of nodes 906, 908, 910, 912. The key extension sub-transcript 902 corresponds to the calculation of HKDF and is performed in two types of HMAC nodes. The difference between these nodes is the size of their input. That is, HMAC ext Node 906 corresponds to the "extraction" step of HKDF, while HMAC exp Node 908 corresponds to the loop of the "Extend" step. Both the key extension 902 and the XORed transcript 904 are dependent on the data length, thus taking advantage of the incremental nature of the scheme.
[0357] 2) In order to further speed up the proof time (by circuit ∏ HMAC′ ,∏ HMAC A zero-knowledge-friendly hash function (e.g., Pedersen or Poseidon) can be used in the HKDF calculation (at each HMAC node / iteration 906, 908). This speeds up the proof time compared to proving the compliance of the transcript produced by AES-CTR, etc.
[0358] Reduce the number of output proofs. The number of proofs generated by the PCD prover is as many as the sink (output) nodes that compute the transcript. Fig. 9 In the example, the N+1 output nodes can be collapsed into two nodes as shown below. i are all considered the i-th leaf of the Merkle tree, and then using the Merkle tree prover in Section 6.3, it can be proved that the root was generated correctly and the leaves are in the correct form. The verifier will receive the root node proof π merkle 、SHA2 node proof π preim And the ciphertext ct:=(ct1,…,ct N ). To verify the well-formedness of ct, the verifier regenerates the Merkle root and verifies π on it. merkle This compression also applies when incremental proof data is encrypted using a block cipher.
[0359] Fig.10 Exemplary methods for the above applications are provided. Fig.10 In the example of , the data requester 1004 acts as the verifier 604 and the data provider 1002 acts as the certifier 602 .
[0360] In step 1, the data requester 1004 requests data from the data provider 1002. The requested data can be any large data, such as an HD movie file or a complex computer program. The data requester 1004 also provides the data provider 1002 with a proof key pk for both the preimage SNARK of Section 6.1 and the Merkle tree SNARK of Section 6.3. In some embodiments, a trusted third party provides the proof key pk to the data requester 1004. The trusted third party provides the verification key vk corresponding to the proof key pk to the data provider 1002, and may also provide the proof key pk to the data provider 1002. In this way, a malicious data requester 1004 cannot obtain information about the data without purchasing the data by simply checking the zk proof (which is generated using the wrong proof key provided by the data requester 1004 and does not retain zero knowledge).
[0361] The data provider 1002 selects input key material ikm, uses HKDF to derive output key material okm, and uses the output key material to generate ciphertext cti for the requested data (step 2). The input key material may be referred to herein as a data encryption key.
[0362] It should be understood that the data provider 1002 can derive the output key material okm before receiving the data request. The data provider 1002 can also derive the ciphertext before the data request, so that the data provider 1002 stores the ciphertext in memory in association with the data for retrieval when receiving the request for the data. The private information required to generate the proof can also be stored in association with it.
[0363] The data provider 1002 also computes a hash of the input key material ikm to compute a digest d, also referred to herein as a keyed hash (step 3). As described above, the data provider 1002 may derive a digest and store the digest in memory prior to receiving a data request.
[0364] The data provider 1002 generates a certificate based on the certification key pk, which verifies both the pre-image and the ciphertext. In this way, it can be ensured that the ciphertext has been generated using the pre-image of the SHA2 digest as the symmetric key, so that the certificate guarantees that the pre-image of the ciphertext and the digest are consistent. For example, the certificate may include the pre-image certificate π preim And the Merkle tree proves π tree , the preimage proof is used to prove in a zero-knowledge manner that the input key material ikm is the preimage of the digest d, and the Merkle tree proof is used to prove that the ciphertext is generated correctly (step 4).
[0365] In step 5, the data provider 1002 provides the ciphertext corresponding to the requested data, digest, and proof to the data requester 1004, or otherwise makes the ciphertext available to the data requester.
[0366] In step 6, the data requester 1004 verifies the digest and ciphertext using the received certification and verification keys.
[0367] If the data requester 1004 is confident that the received ciphertext and digest meet the requirements, the data requester 1004 generates a funding transaction in step 7. The funding transaction provides payment in UTXO for exchanging data. The UTXO is locked to a key corresponding to the data provider 1002. The funding transaction can be an HTLC transaction and can be generated using the digest. In step 8, the data requester 1004 makes the funds available for storage to the blockchain 150.
[0368] In step 9, in order to provide the input key material to the data requester 1004, the data provider 1002 generates a key transaction. The unlocking script of the key transaction unlocks the UTXO of the funding transaction and includes the input key material ikm so that when run with the locking script of the funding transaction, the input material key is verified as the preimage of the digest. In this way, the data provider 1002 provides the key required to decrypt the ciphertext when receiving the funding for the data. In step 10, the key transaction is stored to the blockchain 150.
[0369] The data requester 1004 retrieves input key material from the blockchain 150 in step 11, and uses the input key material to decrypt the ciphertext in step 12 to obtain the requested data.
[0370] 7.2 Fair and Private Digital Market
[0371] Without a trusted third party (TTP), atomic swaps between buyers and sellers cannot guarantee fairness and privacy at the same time. Zero-knowledge contingent payments (ZKCP) use blockchain as a TTP to enable such fair and private transactions. However, these exchanges occur between two parties, which may not be practical. Intermediaries (digital marketplaces) can arrange for the two parties to connect and charge a fee.
[0372] Digital Marketplace. The following digital marketplace designs are available.
[0373] 1. The seller encrypts its data in two layers.
[0374]
[0375] 2. In addition, the seller will also generate a SNARK proof π (sellerID) , which proves that the above outer ciphertext is correctly generated. Therefore, specifically, the proof ensures that (i) ct outer is correctly encrypted (specifically, this means knowing the external encryption key k used outer ), (ii) the Φ compliance of the inner encrypted data satisfies the given predicate Φ, and (iii) the outer ciphertext also encrypts the inner encryption key k inner The hash of .
[0376] 3. The market maintains its database as a Merkle tree whose leaves contain the Uses the scheme from the following section: Error! Reference source not found. It generates a proof for the Merkle root π merkle , which proves the validity of all leaves.
[0377] 4. Buyer acquires tree and verifies roots once and for all.
[0378] 5. The buyer later expresses that he wants to buy N items from seller sellerID. He contacts and informs the seller of his intention to purchase the data items.
[0379] 6. The seller sends the external key to the buyer via a private channel (Maybe more than one).
[0380] 7. The buyer decrypts the outer layer of each received ciphertext, thereby obtaining N inner ciphertexts and the hashed inner key.
[0381]
[0382] It should be noted that the buyer implicitly verifies the N encrypted data items by verifying the (single) proof of the Merkle root in step 4.
[0383] 8. The buyer and seller use the blockchain to perform a fair and private atomic swap. (Maxwell ZKCP protocol). Therefore:
[0384] a. The buyer uses digest d to set up the HTLC contract. (In BSV, this can be done in two transactions.)
[0385] b. The seller embeds the internal key k in the unlocking script inner Redeem the funds as the preimage of d.
[0386] c. The buyer reads the blockchain to retrieve k inner and decrypt compliance data.
[0387] Digital Market Federation. Several digital markets can join together. One entity, the data aggregator will aggregate the Merkle root proofs of all markets, as explained in Section 6.3. Sellers and buyers only need to verify this single master root and upload / download data from different locations.
[0388] 7.3 Partial Blockchain Edit
[0389] The mechanism for proving that a transaction was edited correctly can use SNARKs to prove that a public pattern appears in the preimage (transaction) of a given TxID (SHA256 digest). However, this proof scheme is not scalable: to show that the pattern is distributed over each of the 512-bit blocks of a transaction, m proofs would need to be generated, where m is the number of blocks. For a 1MB transaction, this means verifying 16384 proofs.
[0390] In contrast, the SNARK scheme of Section 6.2 can be used to generate a single proof, regardless of the size of the transaction. The incremental computational nature of the disclosed SNARKs also means that for extremely large transactions (e.g., 1GB of data), the prover can pause proof generation and resume it later.
[0391] 7.4 Efficient Merkle-Processed Transactions
[0392] The identifier of a transaction, TxID′, can be generated by ordering the fields as leaves of a Merkle tree and setting TxID′ as the root. Such a data structure allows the inclusion of fields without revealing the entire transaction which is proven by sending the Merkle tree to the verifier.
[0393] Scalability also becomes an issue when proving the consistency of a Merkle-processed identifier TxID′ and a standard identifier TxID that appears on-chain in a zero-knowledge manner. There are at least as many leaves in a transaction as there are inputs and outputs. Since the number of I / Os in each transaction is different, circuit-specific SNARKs (the most efficient) cannot be used, and therefore a general SNARK must be used. Furthermore, proving the consistency of identifiers for transactions with a large number of I / Os is very time- and space-consuming, and may exceed practical limits.
[0394] Using the SNARKs proposed in Section 6.3 (Merkle Tree Statements), the consistency of both types of identifiers can be proven in a scalable manner. Regardless of the number of I / Os per transaction, a circuit-specific proof system (such as Groth16) can be chosen as needed. The input to each leaf in the tree is proved to be a correct SHA2 hash. Here, the scalable scheme in Section 6.1 can also be utilized when, for example, processing the leaf hash of the locking script field of a larger OP_RETURN data block.
[0395] 8. Further considerations
[0396] 8.1 Comparison of Recursive SNARKs
[0397] The following metrics are used to classify existing preprocessed SNARKs with succinct verifiers.
[0398] Circuit specific: Proof keys / verification keys cannot be reused for different circuits (NP relation). If the keys can be reused, the scheme is general.
[0399] Argument size: small, medium, large. (Smaller is better.)
[0400] Prover runtime: fast, medium, slow.
[0401] · Establishment: Trusted establishment, updatable establishment, transparent establishment.
[0402] o Trustworthy : The party that generates the attestation and verification keys or structured reference strings (SRS) possesses sensitive data that would invalidate the scheme if it were disclosed publicly (particularly to the prover). Trust establishment must be performed in a controlled environment.
[0403] o Updatable : Anyone can update the SRS. This limits the risk of a trusted establishment violating the plausibility, since having just one honest updater is sufficient to maintain plausibility (of the proofs generated after the update occurs).
[0404] o transparent : An untrusted party can generate a proof key and a verification key or SRS.
[0405] Post-quantum security: Is the solution secure in the presence of post-quantum computers?
[0406] Circuit Specific Argument size Prover Runtime Establish Post-quantum security Groth16 yes Small fast Trustworthy no*** GM17 yes Small fast* Trustworthy no**** Marlin no big slow Updatable no Plonk no middle Moderate Updatable no*** Sonic no big slow Updatable no*** Fractal no big slow transparent yes**
[0407] (*)GM17 verification consists of 6 pairings, which will result in a more expensive recursive prover than Groth16, where verification consists of 4 pairings (no precomputation required).
[0408] (**) Security in ROM (non-standard model).
[0409] (***) Safety in the general / algebraic group model (not so good).
[0410] (****)GM17 has the ability to simulate extraction and its safety guarantee is better than Groth16.
[0411] 8.2 PCD from Pairing-Based SNARKs
[0412] Recursive proof compositions or proof-carrying data can be constructed from basic SNARKs with a succinct verifier (an algorithm whose runtime is sublinear in the size of the circuit). Succinct verification is not possible without preprocessing: the verifier must read the circuit whose correct evaluation is being verified at some point, either at preprocessing time or later when given an instance of the relation. Preprocessing (i.e., offline verifier) enables the generation of a short (sublinear) description of the circuit, i.e., the verification key. Such a key is provided to the online verifier along with the public inputs to the circuit.
[0413] Note. Other ways of constructing PCDs, e.g., via compact accumulators, or for describing circuits much smaller than the actual computation, are not considered here.
[0414] 8.2.1 Circuits for Compliance Predicates
[0415] Design of transcript Compliance predicate Each node is verified by the following template circuit C i The satisfiability of i In addition to the verification predicate Π i Is it in the node data z loc ,z out In addition to the above, the circuit also asserts that there is a valid input proof To confirm compliance.
[0416]
[0417] Ensure proper input compliance. How do you ensure that the input complies with the correct predicate? Ensure this by:
[0418] i. Each compliance predicate Π i declares the input type it accepts. Therefore, it only accepts in,i (For a subset, ) will only accept i′: = z in,j Type of input.
[0419] ii. Template circuit C i First check the type of output data z out .t ype Is it equal to the compliance predicate Π i type. That is, z out .type=i. This means that the input z with a valid proof in,j Satisfy circuit C i′ , so that z in,j .type=i′ (because the proof is valid).
[0420] Putting these two together, we can see that the input only satisfies the predicates Π i , so that i′∈T in,i , where T in,i is the current node predicate Π i The set of allowed input types specified in .
[0421] Actual circuit. Low-level details have been avoided for clarity, and many optimizations have been made. In reality, the inputs and logic of the circuits are slightly different. It is important to make each circuit C i The size of C is independent of the number of predicates n. i The internal Merkle tree proves that C i The definition clearly requires moving the verification key to the private input and passing its hash as the public input.
[0422] 8.2.2 Proving the Satisfiability of Circuits
[0423] SNARKs on Elliptic Curve Cycles. For each circuit C outlined above i , consider two preprocessing SNARK schemes (G i,α ,P i,α ,V i,α )、(G i,β ,P i,β ,V i,β ), these schemes are instantiated on elliptic curve cycles. The first scheme (G i,α ,P i,α ,V i,α )prove Satisfiability of arithmetic circuits that lie on the elliptic curve The second solution (G i,β ,P i,β ,V i,β )prove Satisfiability of arithmetic circuits that lie on the elliptic curve Note the recurring pattern: the base domain of the first curve coincides with the scalar domain of the second curve, and vice versa. and
[0424] Two-step proof generation. The first solution (G i,α ,P i,α ,V i,α ) Proof / Verification Circuit C i The circuit is satisfiable as Arithmetic circuit. In order to i Provide input, require proof of input The input proof verifies the node’s input Assume that z in,j Conform to the predicate Π i′ , where i′:=z in, j .type. First prover P i′,α To generate V i′,α Proof of verification π α However, π α Cannot be used directly in C i The input π in,j , because the verifier V i′,α The circuit is Arithmetic circuit (V i′,α Processing points of the first curve So it is in the base domain on, and in Arithmetic circuit simulation The arithmetic is expensive.) To solve this problem, generate proof π β , which proves that π α (Proof of Proof) Validity. More precisely, the "conversion" circuit is constructed as:
[0425]
[0426] The circuit is Arithmetic circuit (because the first verifier V i′,α In the base domain ), and use the second prover P i′,β To generate Proof of the satisfiability of π β Provided to C i The input proof π in,j is the conversion proof π j,β , and embedded as C i The verifier of the subcircuit (in step 3) is V i,β . This is now well defined, since V i,β It can be expressed as Arithmetic circuits.
[0427] 8.2.3 PCD Scheme
[0428] Generator To generate the proof key and verification key: is a compliance predicate. The PCD generator converts the compliance circuit (C1,…,C n ) and its corresponding conversion circuit As input. It uses the following SNARK scheme to generate the proving key / verifying key: (pk i,α ,vk i,α )←G α,i (C i ,λ) and It outputs the proof key pk pcd :=((pk 1,α ,vk 1,α ,…,pk n,α ,vk n,α ),(pk 1,β ,vk 1,β ,…,pk n,β ,vk n,β )),
[0429] and verification key
[0430] vk pcd :=(vk 1,β ,…,vk n,β ).
[0431] Prover To prove that a node satisfies the predicate Π i : It receives node data (input Local loc and output message z out ), input proof and the corresponding authentication key (to verify the input proof) as input. It uses pk i,α As the proof key to generate circuit C i Proof of the satisfiability of π α Then, it will prove that π α "Convert" to π β Therefore, it uses pk i,β Prove as a proof key satisfiability of . It outputs π out : =π β .
[0432] Validator In order to verify the out Conform to the predicate Π i : It receives output data z out And prove π out As input, it uses the verification key vk i,β To calculate b: =V i,β (z out ,π out ). If b accepts, it outputs "accept". Otherwise, it outputs "reject".
[0433] 8.3 Elliptic Curve Pairing
[0434] 8.3.1 Curve family to choose - stick with the MNT curve
[0435] PCD implemented via SNARKs on pairing-friendly elliptic curves can be instantiated on a finite number of curves. The following impossibility result can be proven:
[0436] The Barreto-Naehrig (BN) curve does not have cycles like the elliptic curve.
[0437] ·It can only loop on prime-order curves.
[0438] MNT curves have only cycles of length 2 or 4. The embedding degree must alternate between 4 and 6.
[0439] From the above, we can draw the following conclusion: the only feasible cycle is the MNT4-MNT6 series.
[0440] 8.3.2 Trading safety for efficiency
[0441] If the discrete logarithm problem is in the target group (extended domain If it is easily solved in a subgroup of , then this problem can be solved in any source group. Here, p is the prime order of the basis of the source curve, and k is the embedding degree. k The smaller p is, the easier it is to find discrete pairs in the target population. Conversely, the larger p or k is, the less efficient the pairing is (preferably, with small p and large k).
[0442] For pairing-friendly applications, curves with small embedding degree k or prime number p are desired, but for security reasons, the embedding degree or prime number cannot be too small.
[0443] 8.3.3 Security of MNT Curve
[0444] As of July 2022, to achieve a conservative 128-bit security level in pairing-friendly elliptic curves, the extension field must be 5534 bits to resist Other choices are possible, as summarized in the table below. As mentioned above, the MNT curve only has an embedding degree of 4 or 6. The security must be that of a curve with a smaller embedding degree (4).
[0445] The following table provides three MNT cycles and their corresponding safety levels.
[0446]
[0447] 9. Further comments
[0448] Other variations or uses of the disclosed technology may become apparent to those skilled in the art once given the disclosure herein.The scope of the present disclosure is not limited by the described embodiments but only by the appended claims.
[0449] For example, some of the embodiments above have been described in terms of the Bitcoin network 106, the Bitcoin blockchain 150, and the Bitcoin node 104. However, it should be understood that the Bitcoin blockchain is a specific example of a blockchain 150, and the above description can generally be applied to any blockchain. That is, the present invention is in no way limited to the Bitcoin blockchain. More generally, any references to the Bitcoin network 106, the Bitcoin blockchain 150, and the Bitcoin node 104 above can be replaced with reference to the blockchain network 106, the blockchain 150, and the blockchain node 104, respectively. Blockchains, blockchain networks, and / or blockchain nodes can share some or all of the characteristics of the Bitcoin blockchain 150, the Bitcoin network 106, and the Bitcoin node 104 described above.
[0450] In a preferred embodiment of the present invention, the blockchain network 106 is a Bitcoin network, and the Bitcoin node 104 performs at least all of the functions described in creating, publishing, propagating and storing blocks 151 of the blockchain 150. It is not excluded that there may be other network entities (or network elements) that perform only one or some of these functions but not all of them. That is, a network entity may perform the functions of propagating and / or storing blocks without creating and publishing blocks (remember that these entities are not considered to be nodes of the preferred Bitcoin network 106).
[0451] In other embodiments of the present invention, the blockchain network 106 may not be a Bitcoin network. In these embodiments, it is not excluded that the node can perform at least one or part of the functions of creating, publishing, propagating and storing blocks 151 of the blockchain 150, but not all of the functions. For example, on these other blockchain networks, "node" may be used to refer to a network entity that is configured to create and publish blocks 151 but does not store and / or propagate these blocks 151 to other nodes.
[0452] Even more generally, any reference above to the term "Bitcoin node" 104 may be replaced with the term "network entity" or "network element," where such entity / element is configured to perform some or all of the roles in creating, publishing, propagating, and storing blocks. The functionality of such a network entity / element may be implemented in hardware in the same manner as described above with reference to blockchain node 104.
[0453] Some embodiments have been described in terms of a blockchain network that implements a proof-of-work consensus mechanism to secure the underlying blockchain. However, proof-of-work is only one type of consensus mechanism, and in general embodiments any type of suitable consensus mechanism may be used, such as proof-of-stake, delegated proof-of-stake, proof-of-capacity, or proof-of-elapsed-time. As a specific example, proof-of-stake uses a randomization process to determine which blockchain node 104 has a chance to produce the next block 151. The selected node is often referred to as a verifier. A blockchain node may lock its token for a period of time in order to have a chance to become a verifier. Typically, the node that has locked the largest stake for the longest time is most likely to become the next verifier.
[0454] It should be understood that the above embodiments are described by way of example only. In more general terms, a method, apparatus or program may be provided according to any one or more of the following statements.
[0455] Statement 1. A computer-implemented method for generating a zero-knowledge proof, the zero-knowledge proof being used to prove that each of a plurality of data blocks corresponding to a Merkle tree satisfies a predefined criterion, wherein the Merkle tree comprises a plurality of leaf hash values and a plurality of internal hash values, wherein the plurality of internal hash values are arranged in a hierarchical manner, wherein a plurality of leaf nodes are mapped to the plurality of leaf hash values, and wherein a plurality of internal nodes are mapped to the plurality of internal hash values, wherein the method comprises: executing the plurality of leaf nodes, wherein each leaf node is configured to: receive a corresponding data block and a corresponding data block proof, wherein the corresponding data block proof is used to prove that the data block satisfies the predefined criterion; verify the corresponding data block proof; calculate a data block based on the corresponding data block; block hash; and, outputting the data block hash; executing the plurality of internal nodes, wherein each of the plurality of internal nodes is configured to: receive a corresponding hash value from each of two previous nodes of the Merkle tree; calculate an output hash value based on the received corresponding hash value; and, outputting the output hash value; wherein the received corresponding hash value of a first layer of the plurality of internal nodes is a corresponding data block hash received from a corresponding leaf node, and the received corresponding hash value of each other layer of the plurality of internal nodes is a corresponding output hash value received from a corresponding internal node; wherein the output hash value calculated by a final internal node in a final layer of the Merkle tree is a Merkle root corresponding to the Merkle tree.
[0456] Statement 2. The method according to statement 1, wherein each leaf node is further configured to: generate a leaf node certificate based on the data block hash and the corresponding data block certificate; and output the leaf node certificate.
[0457] Statement 3. A method according to statement 2, wherein each internal node is further configured to: receive a corresponding certificate from each of the two previous nodes; and generate an output certificate based on the received corresponding certificates and the output hash value; wherein the received corresponding certificates of the first layer of the multiple internal nodes are corresponding leaf node certificates received from corresponding leaf nodes, and wherein the received corresponding certificates of each other layer of the multiple internal nodes are corresponding output certificates received from corresponding internal nodes; wherein the output certificate generated by the final internal node in the final layer is a Merkle tree certificate corresponding to the Merkle tree, wherein the Merkle tree certificate confirms that: the output hash calculated by the final internal node is the Merkle root corresponding to the Merkle tree having a data block that meets the predefined criteria.
[0458] Statement 4. The method of statement 3, wherein the method comprises: receiving a Merkle tree proof key, wherein each leaf node proof and output proof are generated based on the Merkle tree proof key; and enabling the Merkle proof to be used by a verification entity, wherein the verification entity has access to a Merkle tree verification key corresponding to the Merkle tree proof key.
[0459] Statement 5. A method according to any of the preceding statements, wherein the method further comprises: calculating the corresponding data block proof for each data block, the corresponding data block proof being used to prove that the data block meets the predefined criteria.
[0460] Statement 6. A method according to any of the preceding statements, wherein the method includes: receiving a data block certification key, wherein the data block certification key corresponds to the predefined standard, and wherein each data block certification is generated based on the data block certification key.
[0461] Statement 7. A method according to any of the preceding statements, wherein the predefined standard enforces a predefined pattern, wherein the step of calculating the corresponding data block proof for each data block comprises: obtaining a corresponding pattern bit array block and a corresponding check bit array block, wherein the pattern bit array block describes a portion of the predefined pattern corresponding to the corresponding data block, wherein the check bit array block defines the bits of the corresponding data block to be checked; determining that the bits of the corresponding data block defined by the corresponding check bit array are equal to the corresponding bits of the corresponding pattern bit array block; and generating the data block proof based on the corresponding pattern bit array block and the corresponding check bit array block.
[0462] Statement 8. A computer-implemented method for providing data to a data requesting entity, wherein the method comprises: generating a Merkle proof according to statement 3, wherein each of the multiple data blocks is an encrypted portion of the data, which is encrypted based on a data encryption key, wherein a predefined criterion is that the encrypted portion has been encrypted based on the data encryption key, wherein the data encryption key is a symmetric key; making the Merkle tree proof and the Merkle root available to the data requesting entity; making the data blocks available to the data requesting entity; obtaining a data encryption key request from the data requesting entity; and, in response to the data encryption key request, making the data encryption key available to the data requesting entity.
[0463] Statement 9. The method according to statement 8, wherein the method further comprises: generating a pre-image proof associated with the data encryption key, wherein the pre-image proof is a zero-knowledge proof for proving knowledge of the data encryption key; and making the pre-image proof and a hash of the data encryption key available to the data requesting entity.
[0464] Statement 10. A method according to statement 8 or 9, wherein the data encryption key request is provided in a funding blockchain transaction, wherein the method further comprises: generating the blockchain transaction for providing the data encryption key by providing a private key and the data encryption key in a first unlocking script of the blockchain transaction, wherein when executed together with a first locking script of the funding blockchain transaction, the first unlocking script is configured to: unlock an unspent transaction output corresponding to the first locking script based on the private key; and, make the data encryption key available to the data requesting entity; and, make the blockchain transaction available to one or more nodes of a blockchain network.
[0465] Statement 11. A method according to any one of statements 8 to 10, wherein each encrypted portion of the data is derived from a corresponding portion of the data, wherein the method further comprises: deriving output key material from the data encryption key, wherein the output key material comprises multiple output key material portions; and encrypting each corresponding portion of the data using the corresponding output key material portion to generate the multiple data blocks.
[0466] Statement 12. The method of statement 11, wherein the output key material is derived from the data encryption key using a hash-based key derivation function.
[0467] Statement 13. A computer-implemented method for decrypting encrypted data, wherein the encrypted data corresponds to request data encrypted based on a data encryption key, wherein the data encryption key is a symmetric key, wherein the method comprises: obtaining a Merkle tree proof and a Merkle root according to statement 3, wherein each of the plurality of data blocks is an encrypted portion of the request data, which is encrypted based on a data encryption key, wherein a predefined criterion is that the encrypted portion has been encrypted based on the data encryption key; verifying the Merkle tree proof based on the Merkle root; obtaining the encrypted data; requesting the data encryption key in response to verifying the Merkle proof; obtaining the data encryption key; and, decrypting the encrypted data based on the data encryption key.
[0468] Statement 14. The method of statement 13, wherein a request for the data encryption key is provided in a funding blockchain transaction, wherein the method further comprises: generating the funding blockchain transaction; and making the funding blockchain transaction available to one or more nodes of a blockchain network.
[0469] Statement 15. A method according to statement 13 or 14, wherein the method further includes: receiving a pre-image proof associated with the data encryption key, wherein the pre-image proof is a zero-knowledge proof for proving knowledge of the data encryption key; receiving a hash of the data encryption key; and, verifying the pre-image proof based on the hash of the data encryption key; wherein the data encryption key is requested in response to verifying the pre-image proof.
[0470] Statement 16. A method according to statements 14 and 15, wherein the funding blockchain includes a first locking script, the first locking script includes the hash of the data encryption key, wherein when executed together with a first unlocking script of a blockchain transaction including the data encryption key, the first locking script is configured to: generate a hash of the data encryption key of the first unlocking script; and verify whether the generated hash is equal to the hash of the data encryption key of the first locking script.
[0471] Statement 17. A computer system, the computer system comprising: at least one computing device, the at least one computing device comprising: a memory, the memory comprising one or more memory units; and a processing device, the processing device comprising one or more processing units, wherein the memory stores one or more code portions configured to be run on the processing device, wherein the code defines a Merkle tree, the Merkle tree is used to generate a zero-knowledge proof, the zero-knowledge proof is used to prove that each of a plurality of data blocks corresponding to the Merkle tree satisfies a predefined standard, wherein the Merkle tree comprises a plurality of leaf nodes and a plurality of internal nodes, wherein the plurality of internal nodes are arranged in layers, wherein each of the one or more code portions defines a leaf node in the plurality of leaf nodes or an internal node in the plurality of internal nodes, wherein the processing unit is configured to execute the one or more code portions, wherein: a portion of a leaf node in the plurality of leaf nodes is defined by When the processing device is executed, the processing device performs the following operations: receiving a corresponding data block and a corresponding data block certificate; verifying the corresponding data block certificate; calculating a data block hash based on the corresponding data block; and, outputting the data block hash; and, a part for defining an internal node among the multiple internal nodes, when executed by the processing device, causes the processing device to perform the following operations: receiving a corresponding hash value from each of two previous nodes of the Merkle tree; calculating an output hash value based on the received corresponding hash value; and, outputting the output hash value; wherein the received corresponding hash value of a first layer of the multiple internal nodes is a corresponding data block hash received from a corresponding leaf node, and the received corresponding hash value of each other layer of the multiple internal nodes is a corresponding output hash value received from a corresponding internal node; wherein the output hash value calculated by a final internal node in a final layer of the Merkle tree is a Merkle root corresponding to the Merkle tree.
[0472] Statement 18. A computer system according to statement 17, wherein the computer system includes a second computing device, the second computing device executes the code portion that defines a leaf node among the multiple leaf nodes, wherein the computing device executes the code portion that defines an internal node among the multiple internal nodes, and wherein the processing device of the computing device is further configured to: receive the data block hash from the second computing device.
[0473] Statement 19. A computer system according to statement 17 or 18, wherein the system also includes a requesting computing device, the requesting computing device includes a memory and a processing device, wherein the memory stores code configured to run on the processing device, and the code is configured to execute a method according to any one of statements 13 to 16 when running on the processing device.
[0474] Statement 20. A computer program embodied on a computer readable storage device and configured to, when executed on one or more processors, perform the method of any one of statements 1 to 16.< / pa>
Claims
1. A computer-implemented method for generating a zero-knowledge proof for proving that each of a plurality of data blocks corresponding to a Merkle tree satisfies a predefined criterion, wherein the Merkle tree comprises a plurality of leaf hash values and a plurality of internal hash values, wherein the plurality of internal hash values are arranged in a hierarchy, wherein a plurality of leaf nodes are mapped to the plurality of leaf hash values, and wherein a plurality of internal nodes are mapped to the plurality of internal hash values, wherein the method comprises: Executing the plurality of leaf nodes, wherein each leaf node is configured to: receiving a corresponding data block and a corresponding data block certificate, wherein the corresponding data block certificate is used to prove that the data block meets the predefined criteria; Verifying the corresponding data block proof; calculating a data block hash based on the corresponding data block; and Output the data block hash; Executing the plurality of internal nodes, wherein each internal node of the plurality of internal nodes is configured to: From each of two previous nodes of the Merkle tree, receiving a corresponding hash value; calculating an output hash value based on the received corresponding hash values; and Outputting the output hash value; wherein the received respective hash values of a first layer of the plurality of internal nodes are respective data block hashes received from respective leaf nodes, and wherein the received respective hash values of each other layer of the plurality of internal nodes are respective output hash values received from respective internal nodes; Wherein the output hash value calculated by the final internal node in the final layer of the Merkle tree is the Merkle root corresponding to the Merkle tree.
2. The method according to claim 1, wherein each leaf node is further configured to: generating a leaf node certificate based on the data block hash and the corresponding data block certificate; and Output the leaf node proof.
3. The method according to claim 2, wherein each internal node is further configured to: receiving a respective certification from each of the two predecessor nodes; and generating an output proof based on the received corresponding proof and the output hash value; wherein the received respective proofs of the first layer of the plurality of internal nodes are respective leaf node proofs received from respective leaf nodes, and wherein the received respective proofs of each other layer of the plurality of internal nodes are respective output proofs received from respective internal nodes; wherein the output proof generated by the final internal node in the final layer is a Merkle tree proof corresponding to the Merkle tree, wherein the Merkle tree proof verifies that the output hash calculated by the final internal node is the Merkle root corresponding to the Merkle tree having a data block that meets the predefined criteria.
4. The method according to claim 3, wherein the method comprises: Receiving a Merkle tree proof key, wherein each leaf node proof and output proof are generated based on the Merkle tree proof key; as well as The Merkle proof is made available to a verifying entity, wherein the verifying entity has access to a Merkle tree verification key corresponding to the Merkle tree proof key.
5. The method according to any one of the preceding claims, wherein the method further comprises: The corresponding data block proof is calculated for each data block, and the corresponding data block proof is used to prove that the data block meets the predefined standard.
6. A method according to any preceding claim, wherein the method comprises: A chunk certification key is received, wherein the chunk certification key corresponds to the predefined standard, wherein each chunk certification is generated based on the chunk certification key.
7. The method according to any of the preceding claims, wherein the predefined standard enforces a predefined pattern, wherein the step of calculating the corresponding data block proof for each data block comprises: Obtaining a corresponding pattern bit array block and a corresponding check bit array block, wherein the pattern bit array block describes a portion of the predefined pattern corresponding to the corresponding data block, and wherein the check bit array block defines bits to be checked of the corresponding data block; determining that the bits of the corresponding data block defined by the corresponding check bit array are equal to corresponding bits of the corresponding pattern bit array block; as well as The data block certificate is generated based on the corresponding pattern bit array block and the corresponding check bit array block.
8. A computer-implemented method for providing data to a data requesting entity, the method comprising: generating a Merkle proof according to claim 3, wherein each of the plurality of data blocks is an encrypted portion of the data encrypted based on a data encryption key, wherein the predefined criterion is that the encrypted portion has been encrypted based on the data encryption key, wherein the data encryption key is a symmetric key; making the Merkle tree proof and the Merkle root available to the data requesting entity; making the data block available to the data requesting entity; Obtaining a data encryption key request from the data requesting entity; as well as In response to the data encryption key request, the data encryption key is made available to the data requesting entity.
9. The method according to claim 8, wherein the method further comprises: generating a preimage proof associated with the data encryption key, wherein the preimage proof is a zero-knowledge proof for proving knowledge of the data encryption key; as well as The pre-image certificate and a hash of the data encryption key are made available to the data requesting entity.
10. The method of claim 8 or 9, wherein the data encryption key request is provided in a funding blockchain transaction, wherein the method further comprises: The blockchain transaction for providing the data encryption key is generated by providing a private key and the data encryption key in a first unlocking script of the blockchain transaction, wherein when executed together with the first locking script of the funding blockchain transaction, the first unlocking script is configured to: unlocking an unspent transaction output corresponding to the first locking script based on the private key; and making the data encryption key available to the data requesting entity; as well as The blockchain transaction is made available to one or more nodes of the blockchain network.
11. A method according to any one of claims 8 to 10, wherein each encrypted portion of the data is derived from a corresponding portion of the data, wherein the method further comprises: deriving output key material from the data encryption key, wherein the output key material comprises a plurality of output key material portions; as well as Each respective portion of the data is encrypted using a corresponding portion of output key material to generate the plurality of data blocks.
12. The method of claim 11, wherein the output key material is derived from the data encryption key using a hash-based key derivation function.
13. A computer-implemented method for decrypting encrypted data, wherein the encrypted data corresponds to request data encrypted based on a data encryption key, wherein the data encryption key is a symmetric key, wherein the method comprises: Obtaining a Merkle tree proof and a Merkle root according to claim 3, wherein each of the plurality of data blocks is an encrypted portion of the request data encrypted based on a data encryption key, wherein the predefined criterion is that the encrypted portion has been encrypted based on the data encryption key; Verifying the Merkle tree proof based on the Merkle root; Obtaining the encrypted data; In response to verifying the Merkle proof, requesting the data encryption key; Obtaining the data encryption key; as well as The encrypted data is decrypted based on the data encryption key.
14. The method of claim 13, wherein the request for the data encryption key is provided in a funding blockchain transaction, wherein the method further comprises: generating said funds blockchain transaction; And making the funding blockchain transaction available to one or more nodes of the blockchain network.
15. The method according to claim 13 or 14, wherein the method further comprises: receiving a preimage proof associated with the data encryption key, wherein the preimage proof is a zero-knowledge proof for proving knowledge of the data encryption key; receiving a hash of the data encryption key; as well as verifying the preimage proof based on the hash of the data encryption key; Wherein the data encryption key is requested in response to verifying the pre-image proof.
16. The method of claims 14 and 15, wherein the funding blockchain comprises a first locking script, the first locking script comprising the hash of the data encryption key, wherein when executed together with a first unlocking script of a blockchain transaction comprising the data encryption key, the first locking script is configured to: generating a hash of the data encryption key for the first unlocking script; and Verifying whether the generated hash is equal to the hash of the data encryption key of the first locking script.
17. A computer system, comprising: At least one computing device, the at least one computing device comprising a memory and a processing device, the memory comprising one or more memory units, the processing device comprising one or more processing units, wherein the memory stores one or more code portions configured to be run on the processing device, wherein the code defines a Merkle tree, the Merkle tree is used to generate a zero-knowledge proof for proving that each of a plurality of data blocks corresponding to the Merkle tree satisfies a predefined criterion, wherein the Merkle tree comprises a plurality of leaf nodes and a plurality of internal nodes, wherein the plurality of internal nodes are arranged in layers, wherein each of the one or more code portions defines one of the leaf nodes or one of the internal nodes, wherein the processing unit is configured to execute the one or more code portions, wherein: The part of defining one of the plurality of leaf nodes, when executed by the processing device, causes the processing device to perform the following operations: Receive the corresponding data block and the corresponding data block certificate; Verifying the corresponding data block proof; calculating a data block hash based on the corresponding data block; and Outputting the data block hash; and The part of defining an internal node of the plurality of internal nodes, when executed by the processing device, causes the processing device to perform the following operations: From each of two previous nodes of the Merkle tree, receiving a corresponding hash value; calculating an output hash value based on the received corresponding hash values; and Outputting the output hash value; wherein the received respective hash values of a first layer of the plurality of internal nodes are respective data block hashes received from respective leaf nodes, and wherein the received respective hash values of each other layer of the plurality of internal nodes are respective output hash values received from respective internal nodes; Wherein the output hash value calculated by the final internal node in the final layer of the Merkle tree is the Merkle root corresponding to the Merkle tree.
18. A computer system according to claim 17, wherein the computer system includes a second computing device, the second computing device executes the code portion defining a leaf node among the multiple leaf nodes, wherein the computing device executes the code portion defining an internal node among the multiple internal nodes, and wherein the processing device of the computing device is further configured to: receive the data block hash from the second computing device.
19. A computer system according to claim 17 or 18, wherein the system further comprises a requesting computing device, the requesting computing device comprising a memory and a processing device, wherein the memory stores a code configured to be run on the processing device, the code being configured to execute the method according to any one of claims 13 to 16 when run on the processing device.
20. A computer program embodied on a computer readable storage device and configured to, when run on one or more processors, perform the method of any one of claims 1 to 16.