Novel block chain-based verifiable vector similarity search method
By introducing a compact vector Merkel tree (CVMTree) structure on the blockchain, the problem of poor vector similarity search performance in the prior art is solved, efficient vector similarity search and verification are realized, and the proof generation and verification process is optimized.
Patent Information
- Application Number
- CN202510269809.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-22
AI Technical Summary
The existing blockchain verified vector similarity search scheme lacks efficient vector data arrangement in the data structure, resulting in poor authentication similarity query performance.
Using a compact vector Merkel tree (CVMTree) structure, the HNSW index graph of blockchain data is mapped into a multi-layer tree structure, a compact vector Merkel tree is generated, and a node summary and verification path are designed to optimize the proof generation and verification process.
It significantly reduces the proof size and verification time, improves query efficiency, reduces calculation and storage overhead, and ensures the accuracy and completeness of query results.
Smart Images

Figure CN120353969A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of blockchain and digital asset authentication, and relates to a new blockchain-based verifiable vector similarity search method. Background Art
[0002] Narrow blockchain is a chain data structure formed by combining data blocks in a sequential connection manner according to the time sequence, and is a distributed ledger that is guaranteed to be tamper-proof and unforgeable by cryptographic means. Generalized blockchain technology is a new distributed infrastructure and computing paradigm that uses a block chain data structure to verify and store data, uses a distributed node consensus algorithm to generate and update data, uses cryptographic means to ensure the security of data transmission and access, and uses smart contracts composed of automated script codes to program and operate data.
[0003] In a blockchain-based verifiable query framework, users are usually limited by computing resources and difficult to run a complete blockchain node, so they cannot directly analyze the data on the chain to obtain query results. To address this challenge, users outsource query requests to a third-party service provider (SP) with the necessary resources and capabilities. However, this approach introduces the risk that the SP may act maliciously or return incorrect results. To ensure the credibility of the query, the adoption of an authenticated data structure (such as a Merkle tree) becomes crucial. The verifiable query process is as follows: The service provider (SP) maintains an up-to-date index structure based on the current blockchain state. When receiving a query request from the client, the SP uses this index to retrieve the query result (referred to as result R). In addition, the SP also provides additional data used in the search process to form a verification object (VO). The VO includes the intermediate data involved in the query process and the corresponding cryptographic evidence (such as Merkle proofs). Users can use this evidence to verify the correctness of the query result. In this article, "proof" and "VO" are regarded as the same concept. After receiving the result R and the accompanying proof from the SP, the user first verifies whether the Merkle hash generated from the proof matches the hash stored on the chain. If the root hashes are consistent, it proves that the data used to calculate the result has not been tampered with and is consistent with the chain state. Subsequently, the user needs to verify whether the result R is indeed derived from the data blocks included in the Merkle proof. This step ensures the integrity and accuracy of the internal data of the application and prevents any tampering behavior. This verifiable query framework ensures that users do not need to fully trust the service provider (SP). By cross-referencing the Merkle proof and leveraging the immutability of the blockchain, users can independently verify the integrity of the result. This hybrid model cleverly combines the blockchain with the index mechanism hosted by the SP, achieving a balance between query performance and security guarantee, making it applicable to various scenarios that require efficient data query and retrieval of large datasets.
[0004] With the increasing popularity of blockchain technology, verifiable vector similarity search has become increasingly important. An authenticated data structure (ADS) can provide proofs to verify the correctness of similarity query results obtained from third-party service providers. However, existing ADS solutions often lack efficient vector data arrangements in the data structure, resulting in poor performance in authenticated similarity queries. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a novel blockchain-based verifiable vector similarity search method. The present invention proposes AVSSProof (Verifiable Vector Similarity Search System), a system that realizes efficient authenticated vector similarity search in a hybrid storage blockchain. A novel Compact Vector Merkle Tree (CVMTree) structure is designed in the system to efficiently manage vector data.
[0006] The present invention establishes a data structure of CVMTree and designs algorithms for the construction, proof generation, and verification of this data structure. Based on this innovative data structure, the proof size and the verification construction time of the proof are greatly reduced during the verifiable vector similarity search process.
[0007] The present invention elaborates on the proof generation process and verification process based on the data structure of CVMTree, and conducts an innovative design around the process of forming a proof for this special data structure of CVMtree.
[0008] The technical solution of the present invention is as follows:
[0009] A novel blockchain-based verifiable vector similarity search method, the steps of which include:
[0010] 1) Map the HNSW index graph of blockchain data to a multi-layer tree structure to generate a Compact Vector Merkle Tree CVMTree;
[0011] 2) Generate a node digest for each node in the Compact Vector Merkle Tree CVMTree, where the nodes include leaf nodes and internal nodes; among them, based on the vector information of the leaf node, the list of neighbor IDs neighborsId of the leaf node in each layer of the CVMTree, and the layer number lc where the leaf node is located, generate the digest h of the corresponding leaf node leaf ; based on the neighbor relationship neighborsId_internal of the internal node in the CVMTree, the highest layer number lc_internal of the internal node in the CVMTree, and the digest of each child node of the internal node, generate the digest h of the corresponding internal node internal ;
[0012] Then, based on the summaries of the leaf nodes and the internal nodes, the root node of the CVMTree is generated;
[0013] 3) For the user's query vector q, traverse each layer of the HNSW index graph; during the search process of each layer, record the visited nodes in visitedNodes as evidence for the user verification result, store the M nodes with the highest similarity to the query vector q in each layer in the set W as candidates, and after traversing each layer of the HNSW index graph, select the K nodes with the highest similarity to the query vector q from the set W as the query result R and return it to the user;
[0014] 4) Create a treePaths proof based on the nodes in visitedNodes, and the treePaths proof records the verification paths from each node in visitedNodes to the root node of the CVMTree and returns it to the user.
[0015] Furthermore, starting from the bottom layer of the CVMTree based on treePaths, generate the verification paths from each node in visitedNodes to the root node of the CVMTree layer by layer upwards to obtain the treePaths proof.
[0016] Furthermore, the method for verifying the correctness of the query result R is as follows:
[0017] 1) Verify the data integrity: First, reconstruct the root node of the CVMTree according to the treePaths proof and compare it with the root proof VOchain stored on the blockchain; if the reconstructed root node of the CVMTree is consistent with the root proof VOchain, it proves that the data in the query result R is complete and has not been tampered with;
[0018] 2) Local search and result comparison: Perform a local search through the function SearchGraph based on the nodes in visitedNodes and the treePaths proof to obtain the local query result R'; compare the local query result R' with the received query result R. If R = R', the verification passes; otherwise, the query result R is regarded as invalid and the verification fails.
[0019] Further, the method for maintaining the CVMTree is as follows: when a node v needs to be inserted into the HNSW index graph, first assign a level lc to the node v according to the generated random seed seed; then insert the node v into the lc-th layer of the HNSW index graph and add the node v to the lc-th layer of the CVMTree; wherein, when inserting the node v into the lc-th layer of the HNSW index graph, perform a search neighbor operation on each layer of the HNSW index graph, identify the neighbor closest to the node v, establish a two-way adjacency between the node v and its closest neighbor, and update the neighbor information and parent-child relationship of the node corresponding to the node v in the CVMTree.
[0020] Further, the method for mapping the HNSW index graph of blockchain data to a multi-layer tree structure is as follows: according to the level result of the HNSW index graph, if a node appears in the i-th layer of the HNSW graph index, then map a corresponding node in the i-th layer of the CVMTree to generate a compact vector Merkle tree CVMTree.
[0021] A server, comprising a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the above method.
[0022] A computer-readable storage medium, having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the above method.
[0023] The advantages of the present invention are as follows:
[0024] The solution of the present invention based on ADS has been widely studied in various application fields, including range queries, keyword searches, and subgraph matching. In the ADS-based solution, optimizing the proof size and processing time is crucial. A smaller proof size reduces the network bandwidth required for transmission, while efficient verification and proof generation time are essential for enhancing the user experience. Existing work, such as ANNproof, focuses on reducing the height of the Merkle tree by adopting a sharded Merkle tree method; this method effectively reduces the cost by dividing the tree structure into multiple subtrees, thereby reducing the size of the generated proof. In addition, the reduced tree height minimizes the computational workload required to verify the hash chain; however, this method introduces a new challenge: the need to store multiple Merkle roots in the smart contract, which results in significant storage overhead. Moreover, ANNproof ignores an important aspect: the inherent neighbor relationship of high-dimensional vector data, which may further optimize the proof structure. To address these limitations, the present invention proposes AVSSProof, which introduces a novel authentication data structure called the Compact Vector Merkle Tree (CVMTree), which is specifically designed to optimize proof generation and verification in vector similarity search tasks. The present invention incorporates the insights of the HNSW graph index structure into the design of CVMTree. Therefore, during the similarity search process, vectors that may be relevant or accessed together are strategically placed close to each other in CVMTree to achieve "compact arrangement". As a result, the number of hashes required to construct the proof is reduced, thereby improving the proof size and computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of the method of the present invention.
[0026] Figure 2 It is a schematic diagram of CVMTree construction.
[0027] Figure 3 It is a comparison diagram of the Merkle tree and the Compact Vector Merkle Tree;
[0028] (a) Merkle tree structure, (b) Compact Vector Merkle Tree structure. DETAILED DESCRIPTION OF THE INVENTION
[0029] The present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0030] 1. CVMTree Structure
[0031] Different from the traditional tree structure, the node arrangement of the traditional tree structure does not consider the corresponding HNSW graph index (refer to Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs), and the architecture of CVMTree is closely integrated with the hierarchical structure of the HNSW graph. The present invention maps the hierarchical structure of the HNSW graph into a multi-layer tree structure to generate CVMTree, as Figure 2 shown.
[0032] Specifically, if a node appears in the i-th layer of the HNSW graph index, there is a corresponding node in the i-th layer of CVMTree. Then the present invention defines the summary of the leaf nodes and the summary of the internal nodes in CVMTree.
[0033] Definition 1 (Leaf node summary): Let id represent the unique identifier of the node in CVMTree, and v id represent the graph node in the lc-th layer of the CVMTree HNSW index (there is also a corresponding node vector in the lc-th layer of the HNSW index). The summary of the leaf node v id is calculated based on the vector v id , the list neighborsId of the neighbor IDs of the graph node v id in each layer of HNSW, and the layer number lc. lc is the level where the HNSW graph node is located, that is, the height of the tree. Formally, this is defined as:
[0034] h leaf = h(v id |neighborsId| lc)
[0035] where h is a cryptographic hash function.
[0036] Definition 2 (Internal node summary): Similarly, for the internal nodes in CVMTree, the summary h internal is defined based on the summary values of its child nodes. v id is the vector information of the corresponding HNSW graph node of this internal node (that is, the vector data after feature extraction of data such as text and pictures in the blockchain), neighborsId_internal is the neighbor relationship of the corresponding HNSW graph node of this internal node, and lc_internal is the highest layer number of the corresponding HNSW graph node of this internal node. Let h childj represent the summary of the j-th child node of this node. The summary h_internal of the internal node is defined as:
[0037] hinternal = h(v id ∣neighborsId_internal|h child1 |h child2 |…|h childj |lc_internal)
[0038] By including the hashes of the neighbors of the vector and all child nodes in each layer, the internal node summary preserves the integrity of the entire subtree below it. Based on the summaries of the leaf nodes and internal nodes, the root node of the final CVMTree can be generated, obtaining the root node hash, which plays a role in the final proof verification process.
[0039] 2. CVMTree Construction Process
[0040]
[0041] In this section, the present invention elaborates in detail on the construction process of CVMTree, focusing on exploring its seamless integration mechanism with the insertion process of HNSW (Hierarchical Navigable Small World). The insertion operation of CVMTree is synchronized with the insertion operation of HNSW, that is, while maintaining the HNSW hierarchical graph index, the node insertion of CVMTree is completed. Specifically, when a node is inserted at level 1 of the HNSW graph, the corresponding graph node v n will be placed at height l of CVMTree.
[0042] Definition of Node Relationships in CVMTree
[0043] After determining the height of the nodes in CVMTree, it is necessary to clarify the parent-child relationships between the nodes. Each intermediate node in CVMTree must be assigned at least one parent node. If a node is inserted at level 1 of the HNSW graph index, the HNSW insertion algorithm will first establish its neighbor relationships, and then search for a parent node among the neighbor nodes at level l + 1 for v n . If no suitable neighbor exists, a depth-first search will be performed at a higher level until a suitable parent node is found. The nearest parent node is selected according to the search depth. If multiple nodes are found at the same search depth, they will all be designated as parent nodes. If a node a is inserted at level 1 of the HNSW graph index, a corresponding node v is also inserted in CVMTree n and at height l of CVMTree? The HNSW insertion algorithm will first establish the neighbor relationships for node a, and then select a neighbor node c of node a at level l + 1, and insert this adjacent node c as the parent node of v n into CVMTree.
[0044] Instead, the algorithm also searches among the neighbors of the node, at level l-1, and selects these nodes as child nodes. If no suitable child nodes exist, the node will have no child nodes. By aligning the structure design of the CVMTree with the hierarchical layout of the HNSW graph, the present invention ensures that verification paths often overlap and reuse shared nodes during traversal. This alignment enables more extensive path sharing across levels, significantly reducing the proof size and the computational and communication overhead required to verify query results.
[0045] A notable feature of the CVMTree is that it allows multiple parent nodes to point to the same child node. This design is inspired by the traversal process in the HNSW algorithm, in which nodes at lower levels are likely to be accessed by multiple upper-level nodes. Therefore, in repeated similarity search operations, the target node can be efficiently located within the same subtree. Additionally, the number of child nodes for each CVMTree parent node is not fixed but dynamic. The connectivity depends on the number of nodes in the next level of the HNSW graph that are most likely to be accessed during traversal, thus ensuring the flexibility and optimization of the search process.
[0046] Specific process of node insertion
[0047] When inserting a node, as described in Algorithm 1, the process first utilizes the "GetLevel" function designed in the original HNSW, which assigns a level lc to the node using a random seed seed. Then, ep is selected as the insertion entry point at level lc. The node is inserted into the HNSW graph in each layer from lc to the bottom, and the new node is also added to the CVMTree with a height of lc. When inserting a node, the algorithm performs a "search neighbors" operation in each layer of the HNSW graph to identify the neighbors closest to the node according to the Euclidean distance, as described in line 8. The "efConstruction" parameter specifies the number of neighbor nodes to be found, and the resulting neighbors are stored in the neighbor list of the new node. Line 10 describes the process of establishing a two-way adjacency relationship for the newly inserted node in the HNSW graph, following the same method in the HNSW graph construction algorithm. Once the neighbors of the new node are determined, the algorithm uses "update neighbors" to update the neighboursId information in the CVMTree and calls the "update parent" function to update the parent-child relationship of the node, ensuring that these changes are accurately reflected in the CVMTree, as shown in lines 11-12.
[0048] In summary, the algorithm seamlessly integrates the insertion of node vectors into the HNSW graph with the construction of the CVMTree. By leveraging hierarchical proximity and path sharing, the design of the present invention significantly enhances the performance of the system, making it highly suitable for large-scale vector search scenarios with cryptographic guarantees. This integration not only optimizes search efficiency but also provides higher practicality and scalability for the verifiable query framework by reducing the proof size and verification overhead.
[0049] 3、VO Generation and Results Verification
[0050]
[0051] In a verifiable query framework, the core responsibility of the service provider (SP) is to generate accurate similarity search results and their corresponding proofs. If the SP can provide correct data that enables the user to successfully reproduce the search results, this will serve as strong evidence of the correctness of the query results. Specifically, the SP is responsible for generating the similarity search results R and the proof VOsp := {visitedNodes, treePaths}, and the detailed process is shown in Algorithm 2. Among them, visitedNodes records all the nodes visited by the SP during the similarity search, while treePaths is similar to the Merkle path and records the verification path for each node to ensure the integrity of the node information.
[0052] Two-stage process of proof generation
[0053] The proof generation process is divided into two stages, corresponding to similarity search and the construction of verification paths respectively. In the first stage, the algorithm performs a similarity search on the query vector q by traversing each layer of the hierarchical HNSW graph from top to bottom. When reaching the search boundary of each layer, the algorithm determines the entry point ep and continues the search in the next layer starting from ep. During the search process of each layer, the SP records the visited nodes in visitedNodes, and these nodes will serve as evidence for the user to verify the results. In addition, the algorithm stores the M nodes with the highest similarity found in each layer as candidates in the set W, and selects the K elements closest to the query vector q from W. In the second stage, the "MHT-VO-Generate" function creates treePaths proofs based on the nodes in visitedNodes, enabling users to utilize the VOchain to verify the authenticity and integrity of the nodes. The proof generation starts from the bottom layer of the CVMTree and proceeds layer by layer upward. The treePaths record the verification paths from each node in visitedNodes to the root node of the CVMTree. In each level, the nodes in visitedNodes are included in the proof, and then their parent nodes are also included to generate proofs at higher levels. Due to the compact structure design of the CVMTree, the parent nodes are likely to be part of the search trajectories captured in visitedNodes as well, which significantly reduces the additional computational or storage overhead during the proof generation process. This design advantage will be further verified in the experimental section, and the present invention will provide detailed evidence to prove its efficiency. This two-stage method significantly improves efficiency while ensuring accuracy. The proof generation process strictly follows the original similarity search strategy of the HNSW algorithm, thus ensuring the accuracy of the search results. In addition, due to the high correspondence between the structure of the CVMTree and the HNSW graph, the size of the generated treePaths proofs is significantly reduced, thereby reducing the costs of proof generation and verification. This design not only optimizes the utilization of computing resources but also provides users with an efficient and reliable verification mechanism, making it suitable for large-scale vector search scenarios.
[0054] By combining similarity search with verification path generation, the present invention proposes an efficient and verifiable query framework. This framework not only ensures the accuracy of query results but also significantly reduces the computational and storage overhead by optimizing the proof generation process. The experimental section will further verify its performance, providing strong support for large-scale vector search scenarios.
[0055]
[0056] Next, the present invention details how users utilize VOsp and VOchain to verify the correctness of the query result R according to Algorithm 3. This verification process is divided into two key steps, aiming to ensure the integrity of the data and the accuracy of the query result.
[0057] Step 1: Verify the integrity of the data
[0058] Users first need to verify whether the data provided by the service provider (SP) is accurate. Specifically, users reconstruct the root node of the CVMTree from treePaths and compare it with the root proof VOchain stored on the blockchain. This verification step ensures that the information of each node recorded in visitedNodes in VOsp has not been tampered with by the SP. If the reconstructed root node is consistent with VOchain, it proves that the data provided by the SP is complete and has not been tampered with.
[0059] Step 2: Local search and result comparison
[0060] After confirming the data integrity, the user performs a local search based on visitedNodes and treePaths through the function "SearchGraph". This function simulates the k-nearest neighbor search process performed by the SP and generates a local query result R'. If the SP honestly records all the visited nodes in visitedNodes during the operation, the user can accurately replicate the traversal of the search boundary of each layer of the HNSW graph and obtain the correct entry point for the next layer, ultimately generating the same search result as the SP. Finally, the user compares the locally computed result R' with the received result R. If R = R', the user can confirm that the SP has provided the correct query result. Otherwise, if the results do not match, the SP's result is considered invalid and the verification fails.
[0061] This two-step verification process not only ensures the integrity and correctness of the data but also avoids the need for users to download the entire dataset. By utilizing the CVMTree structure for secure proof generation and efficient verification, the method of the present invention significantly reduces the computational and storage overhead of the client while maintaining strong cryptographic guarantees. This design makes the verification process both efficient and reliable, especially suitable for large-scale vector search scenarios.
[0062] Vector similarity query based on naive Merkle Tree
[0063] First, a baseline ADS is introduced, called the basic vector Merkle tree (BVMTree). In BVMTree, all leaf nodes store actual vector data, while internal nodes are specifically used to save the hash values generated from their child nodes. During the construction of the tree, each inserted vector is assigned a unique identifier. AsFigure 3 As shown, for example, the leaf node d1 stores vector information and related structured data from the HNSW graph, such as its neighbor list. At the same time, the internal node h9 stores the hash value H(h1||h2).
[0064] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention shall be defined by the scope defined in the claims.
Claims
1. A novel blockchain-based verifiable vector similarity search method, the steps of which include: 1) Map the HNSW index graph of blockchain data into a multi-layer tree structure to generate a Compact Vector Merkle Tree (CVMTree); 2) Generate a node digest for each node in the Compact Vector Merkle Tree CVMTree, where the nodes include leaf nodes and internal nodes; among them, generate the digest h of the corresponding leaf node based on the vector information of the leaf node, the list neighborsId of the neighbor IDs of the leaf node in each layer of the CVMTree, and the layer number lc where the leaf node is located leaf ; Generate the digest h of the corresponding internal node based on the neighbor relationship neighborsId_internal of the internal node in the CVMTree, the highest layer number lc_internal of the internal node in the CVMTree, and the digests of each child node of the internal node internal ; Then generate the root node of the CVMTree based on the summaries of the leaf nodes and the internal nodes; 3) For the query vector q of the user, traverse each layer of the HNSW index graph; during the search process of each layer, record the visited nodes in visitedNodes as evidence for the user verification result, store the M nodes with the highest similarity to the query vector q in each layer in the set W as candidates, and after traversing each layer of the HNSW index graph, select the K nodes with the highest similarity to the query vector q from the set W as the query result R and return it to the user; 4) Create a treePaths proof based on the nodes in visitedNodes, and the treePaths proof records the verification paths from each node in visitedNodes to the root node of the CVMTree and returns it to the user.
2. The method according to claim 1, wherein Based on the treePaths, start from the bottom layer of the CVMTree and generate the verification paths from each node in visitedNodes to the root node of the CVMTree layer by layer upward to obtain the treePaths proof.
3. The method according to claim 1 or 2, characterized in that, The method for verifying the correctness of the query result R is as follows: 1) Verify the data integrity: First, reconstruct the root node of the CVMTree according to the treePaths proof and compare it with the root proof VOchain stored on the blockchain; if the reconstructed root node of the CVMTree is consistent with the root proof VOchain, it proves that the data in the query result R is complete and has not been tampered with; 2) Local search and result comparison: Perform a local search through the function SearchGraph based on the nodes in visitedNodes and the treePaths proof to obtain the local query result R'; compare the local query result R' with the received query result R. If R = R', the verification passes; otherwise, the query result R is regarded as invalid and the verification fails.
4. The method according to claim 1, characterized in that The method for maintaining the CVMTree is as follows: When a node v needs to be inserted into the HNSW index graph, first assign a level lc to the node v according to the generated random seed seed; then insert the node v into the lc-th layer of the HNSW index graph and add the node v to the lc-th layer of the CVMTree; among them, when inserting the node v into the lc-th layer of the HNSW index graph, perform a search neighbor operation in each layer of the HNSW index graph to identify the nearest neighbor of the node v, establish a two-way adjacency between the node v and its nearest neighbor, and update the neighbor information and parent-child relationship of the corresponding node of the node v in the CVMTree.
5. The method according to claim 1 or 2, characterized in that, The method of mapping the HNSW index graph of blockchain data to a multi-layer tree structure is as follows: according to the hierarchical result of the HNSW index graph, if a node appears in the i-th layer of the HNSW graph index, then a corresponding node is mapped in the i-th layer of the CVMTree to generate a compact vector Merkle tree CVMTree.
6. A server, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing any one of the methods according to claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements any one of the methods according to claims 1 to 5.