Verifiable Index Construction and Verification Method Based on Data Value

By building an efficient verified Merkel IR tree index structure based on data value, the blockchain system's query efficiency and result set credibility problems are solved, efficient and reliable data retrieval and result verification are achieved, and the query performance and data storage flexibility of the blockchain system are improved.

CN114911867BActive Publication Date: 2025-07-18BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210408956.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-07-18
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

Blockchain systems have problems such as low query efficiency, limited query functions, difficult to adjust read and write performance, and lack of flexibility in data storage in terms of query processing. The credibility of search results is difficult to guarantee, especially in the efficient retrieval of manufacturing big data and the verification of result set integrity.

Method used

Build an efficient verified Merkel IR tree index structure based on data value, and through the collaborative work of storage service providers and blockchain, store data and generate verifiable result sets. Clients perform the correctness and integrity of results verification, and use EVMIRT index to improve query efficiency and ensure the reliability of result sets.

Benefits of technology

It effectively reduces the maintenance consumption of Merkel IR trees on the blockchain, improves the retrieval efficiency of target keywords within the data range, and supports the reliability verification of the search result set by the query client to ensure the correctness and integrity of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114911867B_ABST
    Figure CN114911867B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing and verifying a verifiable index based on data value, belonging to the technical field of blockchain data retrieval. The method of the present invention includes an efficient verifiable Merkle IR tree index structure and construction method based on data value, an efficient verifiable top-k retrieval algorithm for data value, and a reliability verification algorithm for the retrieved result set. The present invention can effectively reduce the maintenance cost of maintaining the Merkle IR tree structure on the blockchain, improve the efficiency of top-k queries for data containing target keywords within the query data range on the blockchain, and support the reliability verification of the retrieved result set by the query client, enabling users to verify the correctness and integrity of the retrieved data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of blockchain data retrieval, and particularly relates to a method for constructing and verifying a verifiable index based on data value. Background Art

[0002] Due to its characteristics such as decentralization and immutability, blockchain technology is considered an effective solution for secure data storage and retrieval. However, current research on blockchain underlying data storage and query technologies is relatively scarce. Blockchain systems have problems in query processing, such as low query efficiency, limited query functions, difficult adjustment of read and write performance, and lack of flexibility in data storage, which restricts the development of blockchain system applications. Taking manufacturing big data as an example, to retrieve the order quantity, amount, and transactions containing certain keywords simultaneously, one can only retrieve by traversing all transactions, resulting in low efficiency. At the same time, manufacturing data has characteristics such as a huge scale, high dimensionality, and multi-source heterogeneity. Directly storing manufacturing data in the blockchain is not realistically feasible. To solve this problem, a common method is to combine blockchain with off-chain storage technology to form a hybrid storage blockchain. The original data is sent to an off-chain storage service provider for management, while the encrypted hash of the data is stored on the chain to maintain data integrity. However, the query results returned by the service provider may be tampered with or incomplete. Therefore, it is necessary to consider the credibility of the retrieval results, that is, to ensure the correctness and integrity of the retrieval result set simultaneously. The problem of trustworthy verification of the retrieval result set has been studied more in the field of cloud storage, but there is little research in the field of blockchain data. Therefore, it is necessary to provide a reliability verification function for the retrieval result set while constructing a data index with high retrieval efficiency on the blockchain. The existing Merkle IR tree scheme that simply combines the Merkle hash tree and the IR tree can be used to solve the problem of trustworthy verification of high-dimensional data, but it has disadvantages such as high maintenance costs and low retrieval efficiency on the blockchain. Summary of the Invention

[0003] In view of this, the present invention provides a method for constructing and verifying a verifiable index based on data value, which can achieve high retrieval efficiency, good reliability, and low maintenance costs for a hybrid storage blockchain.

[0004] The technical solution of the present invention is implemented as follows:

[0005] A method for constructing and verifying a verifiable index based on data value, comprising the following steps:

[0006] Step 1: The storage service provider and the blockchain construct an Efficient Verifiable Merkle IR-tree (EVMIRT) index based on data value; the storage service provider stores all EVMIRT tree nodes, while the blockchain stores all data and the EVMIRT root hash in the insertion order;

[0007] Step 2: The blockchain sends the stored root hash to the query client; the storage service provider retrieves the eligible data according to the query request sent by the query client and puts it into the Verifiable Result Set (VRS), generates a verification object (VO) for correctness and integrity verification during the retrieval process, and sends the verifiable result set and the verification structure to the query client together;

[0008] Step 3: The query client reconstructs the EVMIRT according to the verifiable result set and the verification structure, compares the root hash obtained from the reconstructed EVMIRT with the root hash returned by the blockchain. If the two root hash results are the same, the correctness verification of the query result passes; the query client checks that each query result object actually exists in the verification structure and that their scores are less than the scores of other entries returned in the verification structure, then the integrity verification of the query result passes.

[0009] Further, Step 1 is specifically as follows:

[0010] Step A1: The storage service provider stores the data;

[0011] Step A2: Select the data insertion area I according to the data value i ;

[0012] Step A3: If the partition P i in the selected area I max reaches the set maximum size, go to Step A4, otherwise go to Step A5;

[0013] Step A4: Merge all IR trees in the partition P max to the upper-level partition P max-1 , if the upper-level partition is full, merge it to the higher-level partition, repeat this operation until a partition is not full, and empty the merged partition;

[0014] Step A5: Insert the data into the partition P max ;

[0015] The specific steps of Step A5 include:

[0016] Step A51: Select the inserted MIR tree (Merkle Inverted Rectangle tree);

[0017] Step A52: Select the inserted leaf node N;

[0018] Step A53: Insert the data into the leaf node N;

[0019] Step A54: If the data in node N exceeds the node capacity, it needs to be split; otherwise, no split is required. When splitting, redistribute the entries in node N and split it into two new nodes {O, P}; if N is the root node, initialize a new node M and add the two new nodes {O, P} as the children of M, and propagate the node split upward if necessary;

[0020] Step A55: Update the MBR (Minimum Bounding Rectangle), inverted file, and node hash value from N to the root node upward;

[0021] Step A6: While calculating the hash value, the storage service provider updates and stores the content of the above - related nodes, and the hash value of the root node of the EVMIR tree constructed by blockchain storage.

[0022] Further, the hash value calculation formula in step A55 is as follows:

[0023] If node N is a leaf node, H(N)=H(O1|…|O i |IF); where O1,...,O i represent the entries in leaf node N, and IF represents the associated inverted file of leaf node N;

[0024] If node N is an intermediate node, H(N)=H(R1|H(N1)|...|R i |H(N i )|IF), R1,...,R i represent the entries in the node, H(N1),...,H(N i ) represent the hash values of the child nodes of intermediate node N, and IF represents the associated inverted file of intermediate node N.

[0025] Further, step two specifically includes:

[0026] Step B1: The query client sends a data query request including the retrieval range, retrieval quantity, and retrieval keyword to the storage service provider;

[0027] Step B2: The blockchain returns all the root hash values stored to the query client;

[0028] Step B3: Starting from indexing all root nodes by the storage service provider in the EVMIRT, calculate the degree of relevance between the nodes and the query request as a score, and store it in the priority queue; and initialize the verification structure to node entries, child node hash values, and inverted indexes.

[0029] Step B4: Select the entry with the minimum score from the priority queue. If the entry is a data object, go to Step B5; otherwise, go to Step B6.

[0030] Step B5: Put the data object into the verifiable result set VRS.

[0031] Step B6: If the entry is a leaf node, traverse the node entries to calculate the score and store it in the priority queue, and update the verification structure; otherwise, traverse its child nodes, calculate the score and store it in the priority queue, and update the verification structure.

[0032] Step B7: Repeat Steps B4 - B6 until the number of retrieval results in the verifiable result set reaches the number required by the user.

[0033] Step B8: Return the verifiable result set VRS and the verification structure to the query client.

[0034] Furthermore, the score calculation formula in Steps B3 and B6 is

[0035]

[0036] where α ∈ (0, 1) is a parameter used to balance spatial proximity and text relevance; the Euclidean distance between the object O and the query object Q is D ε (Q.loc, O.loc); P(Q.keywords|O.keywords) represents the probability that the document contains the query keywords, maxD represents the maximum distance between the object O and the query object Q, and maxP represents the maximum value in the probability P.

[0037] Furthermore, the principle of updating the verification structure VO in Step B6 is specifically: if it is a leaf node, replace the relevant MBR and hash value of the node in VO with the data object and inverted index in the node; otherwise, replace it with the entry in the node, its child node hash value, and inverted index.

[0038] Beneficial effects:

[0039] The verifiable index construction and verification method based on data value proposed by the present invention can effectively reduce the maintenance consumption of maintaining the Merkle IR tree structure on the blockchain, improve the efficiency of top-k queries for data containing target keywords within the query data range on the blockchain, and support the reliability verification of the retrieval result set by the query client, enabling users to verify the correctness and integrity of the retrieved data. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is the architecture diagram of the hybrid storage blockchain system of the present invention.

[0041] Figure 2 It is the overall structure diagram of the efficient verifiable Merkle IR tree EVMIRT of the present invention.

[0042] Figure 3 It is the detailed structure diagram of the sub-region in the efficient verifiable Merkle IR tree of the present invention.

[0043] Figure 4 It is the index structure diagram of the MIR tree of the present invention.

[0044] Figure 5 It is the flow chart of top-k verifiable retrieval of the present invention.

[0045] Figure 6 It is the flow chart of reliability verification of the retrieval result set of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following are specific embodiments of the present invention with reference to the attached drawings and detailed descriptions.

[0047] The verifiable index construction and verification method based on data value provided by the present invention has a system architecture as Figure 1 shown. The system includes four parts: data owners, off-chain storage service providers, blockchain platforms with smart contract functions, and query clients. The blockchain platform and the off-chain storage service provider together form a hybrid storage blockchain. The data owner sends data key-value pairs in the format of <loc, {w j}, v> to the off-chain storage service provider for storage, and sends the hash of the data query key and value in the format of <loc, {w j}, h(v)> to the blockchain for storage. The off-chain storage service provider and the blockchain maintain the same authentication data structure (ADS). When querying, the client sends a query request Q to the off-chain storage service provider; the associated storage service provider returns the query result R and the generated verification structure VO SP , and the blockchain side returns the verification structure VO chain ; the client combines the query result and the verification structure for verification.

[0048] In the present invention, the verifiable Merkle IR tree EVMIRT is defined as a two-layer index structure (as Figure 2 shown): the first layer is the regional index, and the second layer is the internal index within the region. For the regional index, the overall interval is divided into multiple sub-regions I1, I2, I 3... ; as Figure 3 shown, the internal index within the region contains multiple partitions P1, P2, P3,..., P max , and each partition contains multiple MIR trees. All internal indexes within the region share partition P0; the MIR tree structure is as Figure 4 shown, where each entry in the leaf node corresponds to a piece of data O i , and each entry R i in the internal node corresponds to the minimum bounding rectangle (MBR) of all rectangles in the entries of its child nodes N i . H(N) is a hash value calculated by hashing the node entry, the hash values of its child nodes, and the binary concatenation of the inverted file. For the leaf node, H(N) = H(O1|...|O i |IF), where O1,..., O i represent the entries in the node, and IF represents the inverted file. For the intermediate node, H(N) = H(R1|H(N1)|...|R i |H(N i )|IF), R1,..., R i represent the entries in the node, H(N1),..., H(N i ) represent the hash values of its child nodes, and IF represents the inverted file associated with the node.

[0049] The process of generating the index of the efficient verifiable Merkle IR tree EVMIRT is as follows:

[0050] 1. Determine the sub-regions for indexing according to the value data range;

[0051] 2. Insert the data into the largest partition in the region. If the partition is full, merge the partition to the upper-level partition and then insert it into the new largest partition;

[0052] 3. Calculate the hash value of the inserted node as the Hash field and pass this update upward, and calculate the update of its ancestor nodes and the Hash field;

[0053] 4. Repeat this process iteratively until a verifiable index EVMIRT based on the data value is generated. The top-k verifiable retrieval method for querying data containing the target keyword within the data range (as Figure 5 shown) is as follows:

[0054] 1. The user sends a data value top-k retrieval request q;

[0055] 2. The blockchain platform returns the stored root hash value;

[0056] 3. The storage service provider starts from each root node of the EVMIRT tree, pushes the root node into the priority queue, and initializes the verification structure VO;

[0057] 4. Obtain the item with the smallest score from the priority queue, determine whether the item is data. If it is, add the retrieved data to the verifiable result set VRS. If not, push the child nodes of the node into the priority queue;

[0058] 5. Update the verification structure VO. If the accessed item is a leaf node, replace the relevant content of VO with the content of the leaf node; if the accessed item is a non-leaf node, replace it with the information of its child nodes;

[0059] 6. Repeat steps 3-5 until k retrieval results are added to the verifiable result set VRS.

[0060] Method for verifying the reliability of the retrieval result set (as Figure 6 shown). The specific steps are as follows: After the client sends a retrieval request, it obtains the result set <VRS, VO SP , VO chain >. First, use the blockchain consensus protocol to verify the latest block, that is, verify the reliability of VO chain . Then, reconstruct the EVMIRT tree according to VO SP to calculate the root hash value, and then compare it with the relevant tree root node Hash in VO chain to verify the correctness of the result. Then, call VerifyExist(VRS, VO SP ) to verify that the verifiable result set actually exists in the verification structure VO and the score is less than the scores of other returned entries, that is, verify the integrity of the retrieval results.

[0061] In summary, the above is only a preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing and verifying a verifiable index based on data value, characterized in that, It includes the following steps: Step 1: The storage service provider and the blockchain construct an EVMIRT index based on data value; the storage service provider stores all EVMIRT tree nodes, while the blockchain stores all data and the root hash of the EVMIRT tree in the insertion order. Specifically, Step 1 is as follows: Step A1: The storage service provider stores data; Step A2: Select the data insertion area I according to the data value i ; Step A3: If the partition P in the selected area I i reaches the set maximum size, proceed to Step A4; otherwise, proceed to Step A5. max ​ Step A4: Merge partition P max all IR trees in it to the upper-level partition P max-1 , if the upper-level partition is full, merge it to the upper-upper-level partition, repeat this operation until a partition is not full, and empty the merged partition; Step A5: Insert the data into partition P max ; Specifically, Step A5 includes: Step A51: Select the inserted MIR tree; Step A52: Select the inserted leaf node N; Step A53: Insert the data into the leaf node N; Step A54: If the data in node N exceeds the node capacity, splitting is required, otherwise it is not. When splitting, redistribute the entries in node N and split it into two new nodes {O, P}; if N is the root node, initialize a new node M and add the two new nodes {O, P} as the children of M, and propagate node splitting upward if necessary; Step A55: Update the MBR, inverted file, and node hash values from N to the root node upward; Step A6: The storage service provider updates the storage of the above - related node content while calculating the hash value, and the blockchain stores the root node hash value of the constructed EVMIR tree; Step 2: The blockchain sends the stored root hash to the query client; the storage service provider retrieves the data that meets the conditions according to the query request sent by the query client and puts it into the verifiable result set VRS, generates a verification structure VO for correctness and integrity verification during the retrieval process, and sends the verifiable result set and the verification structure to the query client together; Step 3: The query client reconstructs the EVMIRT according to the verifiable result set and the verification structure, compares the root hash obtained from the reconstructed EVMIRT with the root hash returned by the blockchain. If the two root hash results are the same, the correctness verification of the query result passes; the query client checks that each query result object actually exists in the verification structure and that their scores are less than the scores of other entries returned in the verification structure, then the integrity verification of the query result passes.

2. The verifiable index construction and verification method based on data value according to claim 1, characterized in that The hash value calculation formula in Step A55 is: If node N is a leaf node, H(N) = H(O1|…|O i |IF); where O1,...,O i represent the entries in leaf node N, and IF represents the associated inverted file of leaf node N; If node N is an intermediate node, H(N) = H(R1|H(N1)|...|R i |H(N i )|IF), R1,..., R i represent the entries in the node, H(N1),..., H(N i ) represent the hash values of the child nodes of the intermediate node N, and IF represents the inverted file associated with the intermediate node N.

3. The verifiable index construction and verification method based on data value according to claim 1, characterized in that, Specifically, Step 2 includes: Step B1: The query client sends a data query request including the retrieval range, retrieval quantity, and retrieval keyword to the storage service provider; Step B2: The blockchain returns all the root hash values it stores to the query client; Step B3: Starting from all root nodes indexed by the EVMIRT by the storage service provider, calculate the relevance degree between the nodes and the query request as the score and store it in the priority queue; and initialize the verification structure as node entries, child node hash values, and inverted indexes; Step B4: Select the entry with the smallest score from the priority queue. If the entry is a data object, go to Step B5, otherwise go to Step B6; Step B5: Put the data object into the verifiable result set VRS; Step B6: If the entry is a leaf node, traverse the node entries to calculate the scores and store them in the priority queue, and update the verification structure; otherwise traverse its child nodes, calculate the scores and store them in the priority queue and update the verification structure; Step B7: Repeat Steps B4 - B6 until the number of retrieval results in the verifiable result set reaches the required number of retrievals by the user; Step B8: Return the verifiable result set VRS and the verification structure to the query client.

4. The verifiable index construction and verification method based on data value according to claim 3, characterized in that The score calculation formula in Steps B3 and B6 is Among them, α ∈ (0, 1) is a parameter used to balance spatial proximity and text relevance; the Euclidean distance between object O and query object Q is D ε (Q.loc, O.loc); P(Q.keywords|O.keywords) represents the probability that the document contains the query keywords, maxD represents the maximum distance between object O and query object Q, and maxP represents the maximum value in probability P.

5. The verifiable index construction and verification method based on data value according to claim 3, characterized in that, The principle for updating the verification structure VO in Step B6 is specifically as follows: If it is a leaf node, replace the relevant MBR and hash value of this node in VO with the data object and inverted index in the node; otherwise, replace it with the entry in the node, the hash values of its child nodes, and the inverted index.

Citation Information

Patent Citations

  • Efficient data retrieval method based on improved block storage structure

    CN112800065A

  • A management system of country risk indicators and their items using Block Chain for proving their sources

    KR102032780B1