Vector Database Retrieval Acceleration Method and System Based on Quantum Grover-Merkle Tree Algorithm
By using the heterogeneous architecture of the quantum Grover-Merkle Tree algorithm in a vector database, combining classical and quantum computing, the problem of low retrieval efficiency of vector databases is solved, and efficient retrieval acceleration and data integrity verification is achieved.
Patent Information
- Application Number
- CN202411385373.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-09-30
AI Technical Summary
In the prior art, the search process of vector databases consumes more calculation time, resulting in lower efficiency.
The vector database search acceleration method based on the quantum Grover-Merkle Tree algorithm is used to construct the heterogeneous architecture of the quantum vector database and search using a combination of classical and quantum computing.
It significantly accelerates the search process of vector database, improves the search efficiency, and verifies the integrity of the results through quantum resistance hash index, ensuring the reliability of the search results and the immutability of the data.
Smart Images

Figure CN119293138B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database technology, and in particular, to a method and system for accelerating vector database retrieval based on the quantum Grover-Merkle Tree algorithm. Background Art
[0002] A vector database is a database system specifically designed for storing, indexing, querying, and computing vector data. With the rapid development of big data and artificial intelligence technologies, the application of vector data is becoming increasingly widespread in multiple fields, such as image recognition, natural language processing, and recommendation systems. Due to its unique design and high processing capabilities, the vector database has become an ideal choice for processing such data.
[0003] The vector database adopts a vectorized query execution engine, which can process multiple data at once, significantly reducing the computational complexity and improving the processing speed. Compared with traditional relational databases, the vector database has significant advantages in processing large-scale vector data. The vector database uses high-dimensional indexing technologies, such as KD-Tree and LSH, to divide the vector space into multiple hyperplanes and establish an index table for quickly locating and retrieving high-dimensional vector data. This technology enables the vector database to support efficient similarity queries and range queries.
[0004] However, in the retrieval process of the vector database in the prior art, it consumes a relatively large amount of computing time, resulting in low efficiency. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method for accelerating vector database retrieval based on the quantum Grover-Merkle Tree algorithm to eliminate or improve one or more defects existing in the prior art.
[0006] One aspect of the present invention provides a method for accelerating vector database retrieval based on the quantum Grover-Merkle Tree algorithm. The steps of the method include:
[0007] Obtain database text data and convert it into candidate vectors;
[0008] Construct a heterogeneous architecture of a quantum vector database based on the candidate vectors. The heterogeneous architecture of the quantum vector database includes a first hash index tree structure and a second hash index tree structure. Classical computing is based on the first hash index tree structure, and quantum computing is based on the second hash index tree structure;
[0009] In the step where the heterogeneous architecture of the quantum vector database includes the first hash index tree structure and the second hash index tree structure, the hash values of the candidate vectors are generated by a classical KD tree or BallTree, and the first hash index tree structure is constructed; the hash values of the candidate vectors are generated by a quantum-resistant hash function, and the second hash index tree structure is constructed;
[0010] Obtain the query text data, convert the query text data into a query vector correspondingly, and obtain the state of the current quantum computer;
[0011] If the quantum computer is in the standby mode, start the quantum computer to flexibly switch between classical and quantum computing; and calculate the state vector similarity QSS or cosine similarity based on the query vector and the candidate vectors in the second hash index tree structure and perform the Grover retrieval algorithm through quantum computing to retrieve, and output the database text data corresponding to the matching candidate vectors;
[0012] If the quantum computer is not in the standby mode, calculate the cosine similarity based on the query vector and the candidate vectors in the first hash index tree structure and perform classical retrieval, and output the database text data corresponding to the matching candidate vectors.
[0013] Adopting the above solution, this solution first constructs the text data in the database into vector data, and then constructs two tree-shaped data based on the vector data. If the quantum computer is in the standby mode, the quantum computing method can be adopted. The query vector and each candidate vector calculate the similarity by using the Grover retrieval algorithm, and determine whether they match through the similarity. If they match, the text corresponding to the candidate vector is output.
[0014] In some embodiments of the present invention, in the step of constructing the heterogeneous architecture of the quantum vector database based on the candidate vectors, the candidate vectors are used as the root nodes of the first hash index tree structure or the second hash index tree structure, and the candidate vectors are divided multiple times by using a quantum-resistant hash function to construct a quantum-resistant hash Merkle Tree structure.
[0015] In some embodiments of the present invention, in the step of dividing the candidate vectors multiple times by using a quantum-resistant hash function to construct a quantum-resistant hash Merkle Tree structure, in the calculation process of each division, the quantum random number provided by the quantum random number generator is combined with the original data and then hashed to obtain the hash value.
[0016] In some embodiments of the present invention, in the step of calculating the state vector similarity QSS or cosine similarity between the query vector and the candidate vectors in the second hash index tree structure and performing Grover retrieval algorithm retrieval through quantum computing to output the database text data corresponding to the matching candidate vectors, Grover retrieval algorithm retrieval is performed on the query vector and each candidate vector, the state vector similarity QSS or cosine similarity between the query vector and each candidate vector is calculated, the candidate vector matching the query vector is determined based on the similarity, and the integrity of the result is verified through quantum-resistant hash indexing.
[0017] In some embodiments of the present invention, in the step of verifying the integrity of the result through quantum-resistant hash indexing, it is determined layer by layer from any leaf node in the second hash index tree structure to the root node, and it is confirmed whether the hash value of the final root node matches the overall hash value of the Merkle Tree.
[0018] In some embodiments of the present invention, in the step of calculating the cosine similarity between the query vector and the candidate vectors in the first hash index tree structure and performing classical retrieval to output the database text data corresponding to the matching candidate vectors, retrieval is performed based on the candidate vectors in the first hash index tree structure through classical computing, the cosine similarity between the query vector and each candidate vector is calculated, and the most matching candidate vector is determined according to the similarity value.
[0019] In some embodiments of the present invention, in the step of calculating the cosine similarity between the query vector and each candidate vector, the cosine similarity is calculated based on the following formula:
[0020]
[0021] where cos represents the cosine similarity, u represents the query vector, v represents the candidate vector, <u, v> represents the inner product of the query vector and the candidate vector, ‖u‖ represents the norm of the query vector, and ‖v‖ represents the norm of the candidate vector.
[0022] In the specific implementation process, ‖u‖ and ‖v‖ are respectively the norms of vectors u and v, which are defined as:
[0023]
[0024] The calculation method of ‖v‖ is the same.
[0025] In some embodiments of the present invention, during the quantum computing process, the corresponding index of the candidate vector is generated through a quantum circuit, and the efficiency of index retrieval is improved through the Grover retrieval algorithm. Finally, the matching candidate vector and the corresponding database text data are found through the index.
[0026] In some embodiments of the present invention, in the step of constructing a quantum-resistant hash Merkle Tree structure by dividing candidate vectors multiple times using a quantum-resistant hash function, a hash index is constructed through all hash values. When retrieving using the Grover retrieval algorithm, a quantum circuit is used to determine the hash value of each candidate vector, and the integrity of the data is verified based on the path of the hash index, thereby ensuring the reliability of the retrieval result and the immutability of the data.
[0027] The present invention also provides a vector database retrieval acceleration system based on the quantum Grover-Merkle Tree algorithm. The system includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.
[0028] The third aspect of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps implemented by the foregoing vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm are implemented.
[0029] The additional advantages, objectives, and features of the present invention will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following text, or may be learned from the practice of the present invention. The objectives and other advantages of the present invention can be pointed out and obtained specifically in the specification and the accompanying drawings.
[0030] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention.
[0032] Figure 1 It is a schematic diagram of an embodiment of the vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm of the present invention;
[0033] Figure 2 It is a schematic diagram of the processing steps for verifying the data integrity of the Merkle Tree;
[0034] Figure 3 It is a schematic diagram of the processing flow of the Grover retrieval algorithm;
[0035] Figure 4 It is a schematic diagram of a quantum circuit diagram;
[0036] Figure 5 is Figure 4 a schematic diagram of the Oracle circuit in
[0037] Figure 6 a schematic diagram of the overall framework of this solution. Specific implementation manners
[0038] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the implementation manners and the drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but are not used to limit the present invention.
[0039] Herein, it also needs to be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0040] As Figure 1 and 6 shown, the present invention proposes a method for accelerating vector database retrieval based on the quantum Grover-Merkle Tree algorithm. The steps of this method include:
[0041] Step S100, obtaining database text data and converting it into candidate vectors;
[0042] Specifically, obtaining database text data for matching query text data, and correspondingly converting the database text data into candidate vectors;
[0043] In the specific implementation process, the database text data is used to construct a vector database. In the step of obtaining database text data for matching query text data, it can be obtaining database text data when initially constructing the vector database, or it can be the step of adding candidate vectors corresponding to the database text data to the vector database when expanding the vector database.
[0044] In the specific implementation process, in the step of correspondingly converting the database text data into candidate vectors, the text can be converted into vectors in the ways of bag-of-words model, TF-IDF, Word2Vec and BERT.
[0045] Specifically, the process is as follows:
[0046] 1. Data preparation: Prepare a set of text data or data in a relational database as the dataset to be processed.
[0047] 2. Vectorization:
[0048] Use the TF-IDF (Term Frequency-Inverse Document Frequency) method to convert text data into a numerical vector representation.
[0049] TF-IDF calculation formula: TF-IDF(t, d) = TF(t, d) × IDF(t);
[0050] Among them, TF(t, d) represents the term frequency of text or data t in document d in a relational database, and IDF(t) is the inverse document frequency of term t, which is used to measure the universality of this term in all documents.
[0051] Step S200, construct a heterogeneous architecture of the quantum vector database based on candidate vectors. The heterogeneous architecture of the quantum vector database includes a first hash index tree structure and a second hash index tree structure. Classical computing is based on the first hash index tree structure, and quantum computing is based on the second hash index tree structure;
[0052] In the step where the heterogeneous architecture of the quantum vector database includes a first hash index tree structure and a second hash index tree structure, generate the hash value of the candidate vector through a classical KD tree or BallTree, and construct the first hash index tree structure; generate the hash value of the candidate vector through a quantum-resistant hash function, and construct the second hash index tree structure;
[0053] In the specific implementation process, the heterogeneous architecture of the quantum vector database adopted in this solution can adopt different query methods in different situations. On the one hand, it can ensure the efficiency of retrieval queries, and on the other hand, by setting multiple processing schemes, it ensures the feasibility of processing.
[0054] Step S300, obtain the query text data, convert the query text data into a query vector correspondingly, and obtain the state of the current quantum computer;
[0055] In the specific implementation process, the query text data input by the user side will be converted into a query vector, and the conversion method is the same as that of the candidate vector.
[0056] Step S410, if the quantum computer is in the standby mode, then start the quantum computer, and perform state vector similarity QSS or cosine similarity calculation based on the query vector and the candidate vectors in the second hash index tree structure, and execute the Grover retrieval algorithm through quantum computing to retrieve, and output the database text data corresponding to the matching candidate vectors;
[0057] In the specific implementation process, the quantum computing is performed using a quantum computer.
[0058] Step S420: If the quantum computer is not in the standby mode, calculate the cosine similarity between the query vector and the candidate vectors in the first hash index tree structure and perform classical retrieval, and output the database text data corresponding to the matching candidate vectors.
[0059] This solution provides a feasible path for quantum computing and can improve the efficiency of query processing through quantum computing.
[0060] In the specific implementation process, the classical calculation is performed by a computer.
[0061] Adopting the above solution, this solution first constructs the text data in the database into vector data, and then constructs two tree-shaped data based on the vector data. If the quantum computer is in the standby mode, the quantum computing method can be adopted. The query vector and each candidate vector calculate the similarity using the Grover retrieval algorithm, and determine whether they match through the similarity. If they match, the text corresponding to the candidate vector is output.
[0062] In the specific implementation process, the Grover retrieval algorithm is a quantum algorithm used to find a specific item in an unsorted database and has the advantage of square acceleration.
[0063] Time complexity: The classical algorithm requires O(N) time, and the Grover retrieval algorithm only requires O(√N) time.
[0064] The flowchart of the Grover retrieval algorithm is as Figure 3 shown. The main steps of the Grover retrieval algorithm:
[0065] 1. Initialize the superposition state: Initialize the quantum bits into a uniform superposition state of all possible states.
[0066] 2. Apply Oracle: Mark the target state, usually by applying a phase inversion to the target state.
[0067] 3. Apply the diffuser: Enhance the probability amplitude of the target state and reduce the probability of other states.
[0068] 4. Iteration: Repeat the application of Oracle and the diffuser several times until the probability of the target state approaches 1.
[0069] 5. Measurement: Measure the quantum state to obtain the target state.
[0070] Process and mathematical formula:
[0071] 1. Initialize the superposition state: Initialize all possible states into a uniform superposition state:
[0072]
[0073] 2. Oracle Operation: Construct an Oracle, mark the target state, and flip its phase:
[0074]
[0075] 3. Grover Diffusion Operation: Perform the Grover diffusion operation to amplify the probability amplitude of the target state:
[0076] D = 2|ψ><ψ| - I
[0077] 4. Complete Grover Iteration: Repeat the processes in 2 and 3
[0078] 5. Apply k times of Grover Iteration:
[0079] |ψ final 〉 = G k |ψ>
[0080] where the selection of k maximizes the search probability, generally:
[0081]
[0082] The flowchart of the Grover search algorithm is as Figure 3 shown, Figure 3 "Whether it matches the target vector" in it is to calculate whether the value of the cosine similarity between the query vector and the candidate vector matches. In this case, Quantum State Similarity is not used to compare the cosine similarity between the query vector and the candidate vector temporarily because the number of qubits of the current quantum computer is limited and not sufficient to support the calculation of Quantum State Similarity for big data. Here, an interface is reserved in this solution. Once the error rate of the quantum computer decreases and the number of qubits is sufficient, Quantum State Similarity will be used to replace cosine_similarity. Because the calculation speed of Quantum State Similarity is higher than that of the classical cosine_similarity.
[0083] In some embodiments of the present invention, in the step of constructing a heterogeneous architecture of a quantum vector database based on candidate vectors, the candidate vectors are used as the root nodes of the first hash index tree structure or the second hash index tree structure, and a quantum-resistant hash function is used to partition the candidate vectors multiple times to construct a quantum-resistant hash Merkle Tree structure.
[0084] In the specific implementation process, for any non-leaf node, the hash value of the parent node is composed of the hash values of its two child nodes: Hparent = SHA3-256(Hleft + Hright). The finally generated root hash value (MerkleRoot) represents the integrity summary of the entire data set.
[0085] In the specific implementation process, the quantum-resistant hash Merkle Tree structure is a tree-like data structure mainly used to verify data integrity and consistency.
[0086] In some embodiments of the present invention, in the step of constructing the quantum-resistant hash Merkle Tree structure by dividing the candidate vectors multiple times using the quantum-resistant hash function, in the calculation process of each division, the quantum random number provided by the quantum random number generator is combined with the original data and then hashed to obtain the hash value.
[0087] Implementation process of the quantum-resistant hash function:
[0088] With the development of quantum computing technology, traditional hash algorithms may face certain security threats. First, to enhance the quantum computing resistance of the hash function, the quantum random number provided by the quantum random number generator (QRNG) is introduced into the hash calculation. The unpredictability of the quantum random number can significantly improve the security of data encryption and hash operations. Then, a hash function with quantum randomness is used: SHA3-256 is a widely used hash function, and its security performs well in the classical computer environment.
[0089] Process description:
[0090] 1. Generation of quantum random numbers: First, a random bit sequence is generated through the quantum random number generator. This bit sequence is generated using the superposition and uncertainty of quantum states to ensure its unpredictability and irreproducibility. Function name: generate_quantum_random_bits
[0091] Formula: Qrand = QRNG(n);
[0092] Where Qrand represents the generated quantum random bit sequence, and n is the length of the bit sequence.
[0093] Combination of quantum random numbers and data: The generated quantum random bit sequence is combined with the original data to form a new input data. In this way, the randomness and unpredictability of the hash calculation are increased.
[0094] Formula: Dcombined = D + Qrand;
[0095] Among them, D represents the original data, Qrand is the quantum random bit sequence, and Dcombined is the combined data.
[0096] 2. Calculate the quantum-resistant hash value: Input the input data Dcombined combined with quantum random numbers into the SHA-3 hash function to generate the final quantum-resistant hash value.
[0097] Function name: quantum_resistant_SHA3-256_with_qrng
[0098] Formula: H(Dcombined) = SHA3-256(Dcombined);
[0099] Among them, H(Dcombined) represents the hash value generated after combining quantum random numbers.
[0100] In the specific implementation process, in addition to the Grover Merkle tree, the Grover hash index can also be other Grover quantum-resistant hash indexes such as Grover Kyber, Grover Dilithium, and Grover Falcon.
[0101] In the specific implementation process, in addition to constructing the quantum-resistant hash with the Merkle tree used in this case as an example, it can also be constructed with other post-quantum algorithms to build the hash index, such as post-quantum algorithms like CRYSTALS-Kyber, CRYSTALS-Dilithium, Falcon, and SPHINCS+ to implement the quantum-resistant hash index.
[0102] In the specific implementation process, the hash function with quantum randomness: SHA-3 is a widely used hash function, and its security performs well in the classical computer environment. However, with the development of quantum computing technology, traditional hash algorithms may face certain security threats. To enhance the quantum computing resistance of the hash function, the quantum randomness provided by the quantum random number generator (QRNG) can be introduced into the hash calculation. The unpredictability of quantum random numbers can significantly improve the security of data encryption and hash operations.
[0103] First, generate a random bit sequence through the quantum random number generator. This bit sequence is generated using the superposition and uncertainty of quantum states to ensure its unpredictability and irreproducibility;
[0104] The generated quantum random bit sequence is combined with the original data to form a new input data. In this way, the randomness and unpredictability of the hash calculation are increased;
[0105] Input the input data combined with quantum random numbers into the SHA-3 hash function to generate the final quantum-resistant hash value.
[0106] Adopting the above scheme, the technical effects include:
[0107] Improved security: Due to the unpredictability of quantum random numbers, the input data of the hash function becomes more complex and difficult to predict. This significantly enhances the security of the hash value. Even if an attacker understands the working mechanism of the hash function, it is difficult to reproduce or crack the generated hash value.
[0108] Resistance to quantum computing attacks: In the era of quantum computing, traditional hash functions may face attack threats. However, the introduction of quantum random numbers makes the generation process of hash values have uncertainty and randomness, making it difficult for quantum computers to effectively attack them, thus enhancing the quantum computing resistance of the hash function.
[0109] Enhanced irreversibility: The hash function itself has irreversibility, but after introducing quantum random numbers, this irreversibility is further enhanced. The participation of quantum random numbers makes each hash calculation depend on unpredictable quantum states, making cracking and reverse calculation more difficult.
[0110] By introducing quantum random numbers into the SHA-3 hash function, we have implemented a quantum random hash function. This function utilizes the unpredictability of quantum computing to improve the security and collision resistance of the hash value. This method has important application prospects in data encryption and authentication scenarios that require high security. Through these enhanced measures, the quantum random hash function can still maintain a high level of security in the face of quantum computing attacks.
[0111] In the specific implementation process, in the step of dividing the data in the upper-level node into two parts, the data in the upper-level node is evenly divided into two parts, and for each part of the data, a quantum-resistant hash function is used for calculation.
[0112] In some embodiments of the present invention, in the step of calculating the state vector similarity QSS or cosine similarity between the query vector and the candidate vectors in the second hash index tree structure and performing Grover retrieval algorithm retrieval through quantum computing to output the database text data corresponding to the matching candidate vectors, the Grover retrieval algorithm is executed on the query vector and each candidate vector, the state vector similarity QSS or cosine similarity between the query vector and each candidate vector is calculated, the candidate vector matching the query vector is determined based on the similarity, and the integrity of the result is verified through the quantum-resistant hash index.
[0113] Using the above scheme, the Grover search algorithm is a quantum algorithm mainly used for fast searching in an unsorted database, with significant quantum acceleration effect.
[0114] As Figure 2 shown, in the specific implementation process, the implementation of the Grover search algorithm involves the design and execution of a quantum circuit, and finally obtains the search result through quantum measurement. Tools such as matplotlib can be used to visually display the structure and operation steps of the quantum circuit. In this step, this scheme uses a visualization tool to display the structure of the quantum circuit in order to better understand and analyze the working principle of the quantum algorithm.
[0115] Specifically, first, calculate the number of incoming query vectors. This step is to determine how many qubits are needed to represent these query vectors; use binary logarithm to calculate the number of qubits needed and round up. This step determines the number of qubits required in the quantum circuit; create a quantum circuit that contains the calculated qubits. This circuit will be used to execute the Grover search algorithm; apply the Hadamard gate to all qubits to put them in a superposition state, which means that each qubit is simultaneously in the |0> and |1> states, ready for the search operation; construct an Oracle circuit for marking the target state. By checking the similarity between each query vector and the candidate vector to determine the marked state, the Oracle circuit marks the query vector that best matches the candidate vector. This loops through all query vectors and checks if each query vector is close to the candidate vector (implemented by the np.allclose function). If there is a match, apply an operation to the qubits to mark this state; use the Oracle circuit to create the Grover operation. This step constructs a special quantum gate for enhancing the most likely result, i.e., maximizing its amplitude; apply the Grover operation to the entire quantum circuit. This step performs the reflection and expansion operations on the quantum state, aiming to increase the measurement probability of the most likely result; add measurement operations to all qubits. This "collapses" the quantum state into a classical bit state in order to obtain the measurement result; run the quantum circuit using a quantum simulator and execute it multiple times to obtain statistical results. By simulating the operation of a quantum computer, the final measurement result is obtained; find the result that appears most frequently from the measurement results. This is the most likely query vector found by the Grover algorithm; finally, return the index of the most likely query vector, indicating the query vector closest to the candidate vector.
[0116] With the above solution, this solution combines the Merkle Tree with the Grover retrieval algorithm. The Merkle Tree is a hash-based binary tree structure that can efficiently verify the integrity of large-scale data. Traditionally, querying the Merkle Tree requires traversing multiple nodes, and as the amount of data increases, the query overhead also increases. By applying the Grover retrieval algorithm to the leaf nodes of the Merkle Tree, the process of finding these hash indexes can be accelerated. This solution combines the computational advantages of quantum algorithms and the security of the Merkle Tree, and can improve the efficiency of querying and verification while maintaining data integrity. However, its practical application is still restricted by the development of quantum computing resources and hardware, and further optimization and expansion are needed in the future.
[0117] In some embodiments of the present invention, in the step of verifying the integrity of the result through the quantum-resistant hash index, starting from any leaf node of the second hash index tree structure, determine layer by layer upwards to the root node, and confirm whether the hash value of the final root node matches the overall hash value of the Merkle Tree.
[0118] In the specific implementation process, use the verify_proof function to recompute the hash value of each parent node layer by layer by combining the target_hash with the hash values of the sibling nodes in the path.
[0119] Finally, compare the computed root node hash value with the root node hash value stored in the Merkle Tree.
[0120] 3. Verify data integrity:
[0121] The verification process is to combine the hash values of sibling nodes level by level until the root hash is computed:
[0122] Hcurrent = SHA3-256(Hdata + Hsibling). The finally computed root hash should match the root hash value of the Merkle Tree.
[0123] As Figure 2 shown, in the specific implementation process, the Merkle Tree ensures the integrity of data through its structure. The hash value of each node depends on the hash values of its child nodes, which means that if the data of any leaf node changes, the impact will be passed all the way to the root node.
[0124] For each non-leaf node, its hash value is generated by combining the hash values of the left and right child nodes, that is, if any one leaf node changes, it will cause the hash values of all its parent nodes to change, and ultimately cause the root hash value to be different.
[0125] In the specific implementation process, during the data integrity verification process, a hash path (Proof of Inclusion) is constructed to confirm whether specific data belongs to a given Merkle Tree. The verification process calculates the hash path of the given data and gradually combines and calculates the value that matches the root hash, thereby verifying the data integrity; during the verification process, the hash values of sibling nodes are combined level by level until the root hash is calculated.
[0126] Adopting the above solution, through the Merkle Tree, this solution can efficiently verify whether a certain data block is in the set without having to check the entire data set. Verifying a data block only requires checking the hash values on the path from the data block to the root, based on the principle of Proof of Inclusion: given the hash value of the target data and the hash values of all sibling nodes from the target node to the root node, it can be verified whether the target data is in the Merkle Tree.
[0127] In the specific implementation process, the Merkle Tree achieves quantum resistance in the following ways:
[0128] 1. Quantum-resistant hash function: Ensure that the hash value is difficult to be cracked by a quantum computer;
[0129] 2. Structural anti-tampering: Any data modification will cause the hash value of the entire tree to change, preventing data tampering;
[0130] 3. Efficient verification: Quickly verify the inclusion of data to ensure data integrity.
[0131] In summary, through the use of quantum-resistant hash functions and its structural characteristics, the Merkle Tree achieves security in the era of quantum computing and can ensure the integrity and immutability of data under the threats that may be brought by quantum computers.
[0132] In some embodiments of the present invention, the second hash index tree structure in this solution can also adopt Faiss, Annoy, LSH, and KD trees;
[0133] 1. Faiss: Performs excellently in processing large-scale data, especially when using optimized indexes. The query time generally grows logarithmically with the data scale, but may approach linearity in very large scales.
[0134] 2. Annoy: Suitable for approximate nearest neighbor search of high-dimensional data. As the number of trees increases, both the accuracy and time will change accordingly. The performance is relatively stable.
[0135] 3. LSH: It accelerates the search for high-dimensional data through hash functions. The query time depends on the design of the hash function and the parameter ρ, and it performs well under large-scale data.
[0136] 4. KD-tree: It performs excellently in low-dimensional data, but its efficiency drops significantly in high-dimensional data, especially when the data scale is very large.
[0137] The Merkle Tree is mainly used to verify the integrity of data. The query time grows logarithmically with the depth of the tree, and the growth rate is slow, so the query effect is good.
[0138] The comparison of the above algorithms is shown in Table 1 below:
[0139] Table 1
[0140] Algorithm 10,000 Tokens 100,000 Tokens 1,000,000 Tokens 10,000,000 Tokens Faiss ~0.001-0.01s ~0.01-0.1s ~0.1-1.0s ~1.0-10.0s Annoy ~0.001-0.01s ~0.01-0.05s ~0.05-0.5s ~0.5-5.0s LSH ~0.01-0.05s ~0.05-0.1s ~0.1-0.5s ~0.5-2.0s KD Tree ~0.001-0.01s ~0.01-0.1s ~0.1-1.0s ~1.0-10.0s Merkle Tree ~0.001-0.01s ~0.01-0.05s ~0.05-0.1s ~0.1-0.5s
[0141] As can be seen from Table 1, the query time of the quantum-resistant hash function index - Merkle Tree is very stable under different data volumes, and as the data volume increases, the growth rate of its query time is relatively small. That is, it is not only stable in query but also fast in query speed, and is suitable for large-scale vector databases; whether the data volume is 10,000 or 10,000,000 tokens, the query time of the Merkle Tree is between ~0.001 - 0.5 seconds. This stability stems from the logarithmic growth characteristic of the Merkle Tree. Even when the data volume increases, the query time complexity still remains at O(logn), and the Merkle Tree maintains a low growth rate, that is, the Merkle Tree will not increase the query time due to the increase in the query data volume. It is suitable for large-scale vector databases. The design of the Merkle Tree is suitable for the verification of data integrity and anti-tampering. It can not only provide efficient verification operations but also be applicable to vector retrieval.
[0142] Comparison with other algorithms:
[0143] 1. Faiss and KD-tree: The query times are similar when the data volume is small, but as the data volume increases, the query time may increase nearly linearly.
[0144] 2. Annoy and LSH: These two algorithms perform well in medium and small-scale data, but in large-scale data, the query time increases significantly, while the Merkle Tree maintains a low growth rate, that is, the Merkle Tree will not increase the query time due to the increase in the query data volume. It is suitable for large-scale vector databases.
[0145] Applicable scenarios:
[0146] 1. Data integrity verification: The original design intention of the Merkle Tree is for data integrity and anti-tampering verification. In this scenario, it can not only provide efficient verification operations but also be applicable to vector retrieval.
[0147] Combined with other indexing algorithms: The Merkle Tree can also be used in combination with other indexing algorithms. For example, Faiss or Annoy can be used for vector retrieval, and then the Merkle Tree is used to verify the integrity of the retrieval results.
[0148] In some embodiments of the present invention, in the step of calculating the cosine similarity between the query vector and the candidate vectors in the first hash index tree structure and performing classical retrieval, and outputting the database text data corresponding to the matching candidate vectors, classical calculation is used to retrieve based on the candidate vectors in the first hash index tree structure, calculate the cosine similarity between the query vector and each candidate vector, and determine the most matching candidate vector according to the similarity value.
[0149] In some embodiments of the present invention, in the step of calculating the cosine similarity between the query vector and each candidate vector, the cosine similarity is calculated based on the following formula:
[0150]
[0151] where cos represents the cosine similarity, u represents the query vector, v represents the candidate vector, <u, v> represents the inner product of the query vector and the candidate vector, ‖u‖ represents the norm of the query vector, and ‖v‖ represents the norm of the candidate vector.
[0152] In some embodiments of the present invention, during the quantum computing process, the corresponding index of the candidate vector is generated through a quantum circuit, and the efficiency of index retrieval is improved through the Grover retrieval algorithm. Finally, the matching candidate vector and the corresponding database text data are found through the index.
[0153] As Figure 4 and 5 shown, in some embodiments of the present invention, in the step of constructing a quantum-resistant hash Merkle Tree structure by dividing the candidate vectors multiple times using a quantum-resistant hash function, a hash index is constructed through all the hash values. When retrieving using the Grover retrieval algorithm, a quantum circuit is used to determine the hash value of each candidate vector, and the integrity of the data is verified based on the path of the hash index, thereby ensuring the reliability of the retrieval results and the non-tamperability of the data.
[0154] Circuit analysis:
[0155] Hadamard gate (H gate):
[0156] The Hadamard gate on q0, q1, q2 places all qubits in a uniform superposition state. The first step of Grover's algorithm is to create a uniform superposition of all possible states.
[0157] Oracle: (The dark gray part of the circuit)
[0158] Figure 4 The role of the Oracle in the middle is to flip the target state, that is, it applies a phase inversion to the target state. In this figure, the oracle is represented as a single gate and labeled Q. Usually this step includes a series of X gates and multi-controlled X gates.
[0159] Grover Diffuser (the light gray circuit):
[0160] In the figure, the gray gate part on the right represents the Grover Diffuser, and the role of the diffuser is to amplify the amplitude probability of the target state. This usually includes:
[0161] Applying Hadamard gates
[0162] Applying X gates
[0163] Applying a controlled multi-controlled Z gate to n - 1 qubits (usually implemented through a controlled-X gate, MCX gate)
[0164] Applying X gates again
[0165] Finally applying Hadamard gates again
[0166] Measurement:
[0167] The last part of the circuit is the measurement gate, which projects the measurement result of the quantum state onto a classical bit.
[0168] Oracle: The purple gate block in the provided circuit diagram is labeled Q, representing the oracle part of Grover's algorithm. This part implements the marking of the target state by differentiating it through a phase inversion of the target state.
[0169] Grover Diffuser: The part before the measurement gate in the circuit diagram most likely represents the Grover Diffuser, whose purpose is to amplify the amplitude probability of the target state marked by the oracle before.
[0170] Specifically, the expansion of the internal Oracle circuit (the dark gray part of the circuit) is as follows Figure 5 as shown;
[0171] The internal operations of the quantum circuit are as follows:
[0172] 1. Convert the index of the candidate vector into binary format to determine the qubits to be operated on.
[0173] 2. For each qubit, if the corresponding binary digit is 0, apply the X gate (NOT gate) to flip the qubit state.
[0174] 3. Use the multi-control Toffoli gate, which flips the target qubit only when all control qubits are 1.
[0175] 4. Apply the X gate again to restore the quantum state.
[0176] After the final measurement, the binary form corresponding to the index is obtained. The correspondence between the index and the binary representation (i.e., "coding matching index" in the flowchart)
[0177] Index 0 -> Binary "000"
[0178] Index 1 -> Binary "001"
[0179] Index 2 -> Binary "010"
[0180] Index 3 -> Binary "011"
[0181] Index 4 -> Binary "100"
[0182] Index 5 -> Binary "101"
[0183] Index 6 -> Binary "110"
[0184] Index...... -> Binary "......"
[0185] Overall framework diagram of the StateVector of the quantum vector database:
[0186] Quantum inner product similarity formula:
[0187] For two quantum states ∣ψ1> and ∣ψ2>, their similarity calculation can be expressed as:
[0188]
[0189] Where:
[0190] <ψ1∣ψ2> is the inner product of the two quantum states;
[0191] ||ψ1|| and ||ψ2|| are the norms (or lengths) of their respective quantum states.
[0192] It is very similar in form to the cosine similarity formula:
[0193]
[0194] Formula example
[0195] The formula for the quantum inner product can be written in the following form to represent similarity, which can also be referred to as: Quantum State Vector Similarity (QSS):
[0196]
[0197] Formula Explanation
[0198] <ψ1∣ψ2> represents the inner product of the quantum states ∣ψ1> and ∣ψ2>;
[0199] ||ψ1|| represents the norm of the quantum state ∣ψ1>, and the calculation method is
[0200] ||ψ2|| represents the norm of the quantum state ∣ψ2>, and the calculation method is
[0201] The beneficial effects of this solution include:
[0202] 1. The quantum Grover-Merkle Tree algorithm accelerates the search speed of the vector database. It can not only determine the position of the target hash value faster, reducing the traversal steps required in the classical Merkle Tree query process, but also is more suitable for large-scale vector databases.
[0203] 2. The construction of the quantum-resistant hash index not only reduces conflict handling and shortens the query time, but more importantly, it can encrypt the data entering the vector database itself. It is a quantum-resistant hash function that kills multiple birds with one stone.
[0204] 3. The hybrid system architecture design not only needs to seamlessly switch or run the computing resources of classical and quantum computing simultaneously. At the same time, it also needs to meet the algorithm scheduling system, which dynamically allocates tasks to classical or quantum computing nodes according to the nature of the tasks and the availability of computing resources.
[0205] 4. The quantum random hash function enhances irreversibility, making cracking and reverse calculation more difficult.
[0206] The embodiment of the present invention also provides a vector database retrieval acceleration system based on the quantum Grover-Merkle Tree algorithm. The system includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.
[0207] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps implemented by the foregoing method for accelerating vector database retrieval based on the quantum Grover-Merkle Tree algorithm are realized. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0208] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to execute the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0209] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0210] In the present invention, the features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0211] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A vector database retrieval acceleration method based on quantum Grover-Merkle Tree algorithm, characterized in that: The steps of the method include: Obtain database text data and convert it into candidate vectors; Building a heterogeneous architecture of a quantum vector database based on the candidate vectors, the heterogeneous architecture of the quantum vector database comprising a first hash index tree structure and a second hash index tree structure, classical computing is based on the first hash index tree structure, and quantum computing is based on the second hash index tree structure; In the step that the heterogeneous architecture of the quantum vector database includes a first hash index tree structure and a second hash index tree structure, a hash value of a candidate vector is generated by a classic KD tree or a BallTree to construct the first hash index tree structure; a hash value of the candidate vector is generated by a quantum resistant hash function to construct the second hash index tree structure; Obtain query text data, convert the query text data into a query vector, and obtain the current state of the quantum computer; If the quantum computer is in standby mode, the quantum computer is started, and a state vector similarity QSS or cosine similarity calculation is performed based on the query vector and the candidate vector in the second hash index tree structure, and a Grover retrieval algorithm is performed through quantum computing to output the database text data corresponding to the matching candidate vector; If the quantum computer is not in the standby mode, cosine similarity calculation is performed based on the query vector and the candidate vectors in the first hash index tree structure, and a classical search is performed to output the database text data corresponding to the matching candidate vectors.
2. The vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm according to claim 1 is characterized in that: In the step of constructing a heterogeneous architecture of a quantum vector database based on a candidate vector, the candidate vector is used as a root node of a first hash index tree structure or a second hash index tree structure, and a quantum-resistant hash function is used to divide the candidate vector multiple times to construct a quantum-resistant hash Merkle Tree structure.
3. The vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm according to claim 2 is characterized in that: In the step of using a quantum-resistant hash function to divide the candidate vector multiple times and constructing a quantum-resistant hash Merkle Tree structure, in the calculation process of each division, the quantum random number provided by the quantum random number generator is combined with the original data and then hashed to obtain a hash value.
4. The vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm according to claim 1 is characterized in that: In the step of performing state vector similarity QSS or cosine similarity calculation based on the query vector and the candidate vectors in the second hash index tree structure and performing Grover retrieval algorithm retrieval through quantum computing, and outputting database text data corresponding to the matching candidate vector, Grover retrieval algorithm retrieval is performed on the query vector and each candidate vector, the state vector similarity QSS or cosine similarity calculation between the query vector and each candidate vector is calculated, the candidate vector matching the query vector is determined based on the similarity, and the integrity of the result is verified through the quantum resistant hash index.
5. The vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm according to claim 4 is characterized in that: In the step of verifying the integrity of the result through the quantum-resistant hash index, a determination is made layer by layer from any leaf node in the second hash index tree structure to the root node to confirm whether the hash value of the final root node matches the overall hash value of the Merkle Tree.
6. The vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm according to claim 1, characterized in that: In the step of performing cosine similarity calculation based on the query vector and the candidate vectors in the first hash index tree structure and executing classical retrieval, and outputting the database text data corresponding to the matching candidate vector, retrieval is performed based on the candidate vectors in the first hash index tree structure through classical calculation, the cosine similarity between the query vector and each candidate vector is calculated, and the best matching candidate vector is determined based on the similarity value.
7. The vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm according to claim 6 is characterized in that: In the step of calculating the cosine similarity between the query vector and each candidate vector, the cosine similarity is calculated based on the following formula: Among them, cos represents cosine similarity, u represents query vector, and v represents candidate vector.<u,v> represents the calculation of the inner product of the query vector and the candidate vector, ‖u‖ represents the norm of the query vector, and ‖v‖ represents the norm of the candidate vector.
8. The vector database retrieval acceleration method based on quantum Grover-Merkle Tree algorithm according to claim 1, characterized in that: During the quantum computing process, the corresponding index of the candidate vector is generated through the quantum circuit, and the efficiency of index retrieval is improved through the Grover retrieval algorithm. Finally, the matching candidate vector and the corresponding database text data are found through the index.
9. The vector database retrieval acceleration method based on the quantum Grover-Merkle Tree algorithm according to claim 2, characterized in that: In the step of using quantum-resistant hash functions to divide candidate vectors multiple times and constructing a quantum-resistant hash Merkle Tree structure, a hash index is constructed using all hash values. When searching using the Grover retrieval algorithm, a quantum circuit is used to determine the hash value of each candidate vector, and the integrity of the data is verified based on the path of the hash index, thereby ensuring the reliability of the retrieval results and the immutability of the data.
10. A vector database retrieval acceleration system based on quantum Grover-Merkle Tree algorithm, characterized in that: The system includes a computer device, which includes a processor and a memory. The memory stores computer instructions. The processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Block chain storage data query method and device
CN113157735A
Super-computing-oriented quantum search simulation method and system
CN116227615A