A Vector Database Retrieval Method and System Accelerated by Semantic Hashing
By constructing a phased hash representation set by time partitioning of historical semantic vector samples, and combining dynamic clustering and hash code drift trajectory model, adaptive calibration hash code is generated, which solves the problem of hash code distribution drift in dynamic incremental scenarios, realizes semantic consistency retrieval and efficient index construction, and improves retrieval accuracy and speed.
Patent Information
- Application Number
- CN202510472581.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the dynamic incremental scenarios of the prior art, traditional static semantic hashing models cannot adapt to data distribution drift, resulting in the hash code gradually deviating from the original clustering center, and the semantic similarity measurement distortion and retrieval recall rate are reduced.
By partitioning historical semantic vector samples by time, a phased hash representation set is constructed, and a dynamic cluster matching and historical hash code drift trajectory model is combined to generate an adaptive calibration hash code, and an accelerated index structure is constructed for similar data retrieval.
It effectively solves the hash code distribution drift problem in dynamic incremental scenarios, ensures the consistency of semantic mapping, improves the retrieval recall and accuracy, and improves the search speed and index construction efficiency.
Smart Images

Figure CN119990144B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data retrieval, and more specifically, to a method and system for retrieving a vector database accelerated by semantic hashing. Background Art
[0002] In the context of big data, vector databases are widely used for similarity retrieval of multi-modal data such as text, images, and speech. Existing technologies mainly use semantic embedding combined with hash coding to map high-dimensional data into discrete hash codes for fast retrieval. Existing technologies generally adopt a static coding strategy, which cannot fully adapt to the distribution drift problem caused by continuously increasing data, resulting in the hash codes of new data gradually deviating from the original clustering center, thereby causing semantic similarity measurement distortion and a decrease in retrieval recall rate. Therefore, how to maintain the consistency of the hash code distribution and improve the retrieval accuracy in a dynamic incremental scenario has become a key problem to be solved.
[0003] To solve the above problems, a technical solution is provided. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method and system for retrieving a vector database accelerated by semantic hashing to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for retrieving a vector database accelerated by semantic hashing includes the following steps:
[0007] Obtain a plurality of historical semantic vector samples, and divide them into a plurality of time partition groups based on the original data time period to which each semantic vector sample belongs;
[0008] Perform an embedding model fitting operation on each time partition group respectively, extract the corresponding semantic embedding feature set, and generate an initial hash code space for each feature set to form a plurality of stage hash representation sets;
[0009] For the current newly added data, perform a dynamic clustering matching operation according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set to obtain a target hash center candidate set for the current newly added data;
[0010] Based on the target hash center candidate set, combine the historical hash code drift trajectory model to generate an adaptive calibration hash code for the current newly added data;
[0011] Write the adaptive calibration hash code of the currently newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to achieve semantic consistency retrospective query of historical data.
[0012] In a preferred embodiment, obtain multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs. Specifically:
[0013] Set a fixed time interval length according to the timestamp generated by the original data corresponding to each historical semantic vector sample;
[0014] According to the fixed time interval length, sequentially divide all historical semantic vector samples into multiple consecutive and non-overlapping time partition groups; the number of historical semantic vector samples in each time partition group is greater than a preset number threshold.
[0015] In a preferred embodiment, perform an embedding model fitting operation on each time partition group respectively, extract the corresponding semantic embedding feature set, and generate an initial hash code space for each feature set to form multiple phased hash representation sets. Specifically:
[0016] Input all historical semantic vector samples in each time partition group into a pre-trained semantic embedding model, and perform forward inference calculation to obtain the high-dimensional semantic embedding feature vector corresponding to each historical semantic vector sample;
[0017] Perform a hash function mapping operation on all the high-dimensional semantic embedding feature vectors obtained in each time partition group in sequence to obtain the corresponding binary initial hash code representation result;
[0018] Store all the binary initial hash code representation results obtained from each time partition group respectively to form the phased hash representation set corresponding to each time partition group.
[0019] In a preferred embodiment, for the currently newly added data, perform a dynamic clustering matching operation according to the time period to which the currently newly added data belongs and the mapping position of the corresponding semantic similarity vector in each phased hash representation set to obtain the target hash center candidate set of the currently newly added data. Specifically:
[0020] Determine the latest time partition group to which it belongs based on the timestamp generated by the original data corresponding to the currently newly added data;
[0021] In the phased hash representation set corresponding to the latest time partition group, by calculating the cosine similarity between the semantic embedding feature of the currently newly added data and the historical semantic vector sample, select the hash code corresponding to the historical semantic vector sample with a cosine similarity greater than the similarity threshold as the clustering matching candidate;
[0022] Cluster the hash codes in the clustering match candidates according to the clustering algorithm to determine the target hash center candidate set corresponding to the currently newly added data.
[0023] In a preferred embodiment, based on the target hash center candidate set, combined with the historical hash code drift trajectory model, generate the adaptive calibration hash code of the currently newly added data, specifically:
[0024] Construct the historical hash code drift trajectory model according to the temporal variation law of the hash centers within each historical time partition group;
[0025] Determine the candidate hash code with the smallest distance as the adaptive calibration hash code of the currently newly added data through the Euclidean distance between each candidate hash code in the target hash center candidate set and the hash code predicted by the historical hash code drift trajectory model at the current moment.
[0026] In a preferred embodiment, write the adaptive calibration hash code of the currently newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to achieve semantic consistency retrospective query of historical data, specifically:
[0027] In the vector database, use the adaptive calibration hash code of the currently newly added data as the index primary key to establish the mapping relationship between the adaptive calibration hash code and the original semantic vector features of the currently newly added data;
[0028] When the retrieval request arrives, generate the adaptive calibration hash code corresponding to the retrieval request, and use the hash index structure in the vector database to locate the corresponding hash bucket, and return the historical semantic vector data that matches the adaptive calibration hash code of the retrieval request to complete the fast retrieval of similar data.
[0029] In a preferred embodiment, when the retrieval request arrives, generate the adaptive calibration hash code corresponding to the retrieval request, and use the hash index structure in the vector database to locate the corresponding hash bucket, and return the historical semantic vector data that matches the adaptive calibration hash code of the retrieval request to complete the fast retrieval of similar data, specifically:
[0030] Input the retrieval request into the pre-trained semantic embedding model to generate the corresponding high-dimensional semantic embedding feature vector;
[0031] Input the high-dimensional semantic embedding feature vector representation corresponding to the retrieval request into the historical hash code drift trajectory model to calculate the drift vector corresponding to the retrieval request time;
[0032] Perform vector superposition operation on the drift vector and the high-dimensional semantic embedding feature vector corresponding to the retrieval request to generate the adaptive calibration hash code corresponding to the retrieval request;
[0033] Query the hash index structure in the vector database with the adaptive calibration hash code, locate the hash bucket with the same hash code, and extract the historical semantic vector data that matches the adaptive calibration hash code corresponding to the retrieval request from the hash bucket to complete the fast retrieval of similar data.
[0034] On the other hand, the present invention provides a vector database retrieval system with semantic hashing acceleration, including a time partition division module, an embedding model fitting module, a dynamic clustering matching module, a hash code calibration module, and an index construction and retrieval module;
[0035] Time partition division module: Obtain multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs;
[0036] Embedding model fitting module: Perform embedding model fitting operations on each time partition group respectively, extract the corresponding semantic embedding feature sets, and generate an initial hash code space for each feature set to form multiple phased hash representation sets;
[0037] Dynamic clustering matching module: For the currently newly added data, perform dynamic clustering matching operations according to the time period to which the currently newly added data belongs and the mapping positions of the corresponding semantic similarity vectors in each phased hash representation set to obtain the target hash center candidate set of the currently newly added data;
[0038] Hash code calibration module: Based on the target hash center candidate set, combine the historical hash code drift trajectory model to generate the adaptive calibration hash code of the currently newly added data;
[0039] Index construction and retrieval module: Write the adaptive calibration hash code of the currently newly added data into the vector database to construct an acceleration index structure, and perform similar data retrieval based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data.
[0040] Technical effects and advantages of a vector database retrieval method and system with semantic hashing acceleration according to the present invention:
[0041] By partitioning historical semantic vector samples by time, constructing phased hash representation sets, combining dynamic clustering matching and the historical hash code drift trajectory model, the generation of adaptive calibration hash codes is realized. Effectively solves the problem of hash code distribution drift between old and new data in the traditional static semantic hashing model in the dynamic incremental scenario, ensures the consistency of semantic mapping, thereby improving the retrieval recall rate and query accuracy, and enhancing the retrieval speed and index construction efficiency. When facing large-scale and continuously updated data, it can maintain high stability and scalability, providing reliable technical support for the fast and accurate retrieval of massive data. Brief Description of the Drawings
[0042] Figure 1 Schematic diagram of a semantic hashing accelerated vector database retrieval method of the present invention;
[0043] Figure 2 Schematic diagram of the structure of a semantic hashing accelerated vector database retrieval system of the present invention. Detailed implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Embodiment 1
[0046] Figure 1 A semantic hashing accelerated vector database retrieval method of the present invention is given, which includes the following steps:
[0047] Obtain a plurality of historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs;
[0048] Perform an embedding model fitting operation on each time partition group respectively, extract the corresponding semantic embedding feature set, and generate an initial hash code space for each feature set to form a plurality of stage hash representation sets;
[0049] For the currently newly added data, perform a dynamic clustering matching operation according to the time period to which the currently newly added data belongs and the mapping positions of the corresponding semantic similarity vectors in each stage hash representation set to obtain a target hash center candidate set of the currently newly added data;
[0050] Based on the target hash center candidate set, combine with the historical hash code drift trajectory model to generate an adaptive calibration hash code for the currently newly added data;
[0051] Write the adaptive calibration hash code of the currently newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to realize semantic consistency backtracking query of historical data.
[0052] Specifically, obtaining a plurality of historical semantic vector samples and dividing them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs includes:
[0053] Set a fixed time interval length according to the timestamp generated by the original data corresponding to each historical semantic vector sample;
[0054] Specifically, a generation timestamp is extracted from the original data of each historical semantic vector sample, and a fixed time span is determined. For example, the fixed time interval is 24 hours, 1 week, or 1 month. Exemplarily, if the timestamp format of the historical semantic vector sample is "year / month / day hour:minute:second", "24 hours" can be set as the fixed time interval length, that is, the data within every 24 hours is grouped into one set.
[0055] According to the fixed time interval length, all historical semantic vector samples are sequentially divided into multiple consecutive and non-overlapping time partition groups; the number of historical semantic vector samples within each time partition group is greater than a preset number threshold;
[0056] Specifically, using the determined fixed time interval length, all historical semantic vector samples are sorted in chronological order and divided into several consecutive and non-overlapping time partition groups. It is required that the number of historical semantic vector samples in each time partition group be greater than the preset number threshold to ensure sufficient data volume for processing, facilitating model training or statistical analysis.
[0057] For example, if the fixed time interval length is 24 hours and the preset number threshold is 100, then only when the number of historical semantic vector samples within 24 hours is greater than 100, this time period is valid; otherwise, adjacent time periods are merged or ignored.
[0058] Specifically, an embedding model fitting operation is performed on each time partition group respectively, a corresponding semantic embedding feature set is extracted, and an initial hash code space is generated for each feature set, forming multiple phased hash representation sets, including:
[0059] All historical semantic vector samples in each time partition group are input into a pre-trained semantic embedding model, and forward inference calculation is performed to obtain the high-dimensional semantic embedding feature vector corresponding to each historical semantic vector sample;
[0060] Specifically, for each set of historical semantic vector samples grouped by time, each historical semantic vector sample therein is input into a pre-trained semantic embedding model, and forward calculation is performed using the semantic embedding model to obtain a high-dimensional semantic embedding feature vector. The high-dimensional semantic embedding feature vector accurately reflects the semantic information of the input data. For example, after a semantic embedding model inputs a piece of text, it outputs a 512-dimensional vector. Each historical semantic vector sample goes through this process to generate the corresponding high-dimensional semantic embedding feature vector.
[0061] All the high-dimensional semantic embedding feature vectors obtained for each time partition group are sequentially subjected to a hash function mapping operation to obtain the corresponding binary initial hash code representation result;
[0062] Specifically, apply a hash function to each high-dimensional semantic embedding feature vector to map the continuous real-valued vector into a binary code of a fixed length.
[0063] For example, set the hash function as: ;
[0064] where, represents the binary output obtained by calculating the th hash function on the input high-dimensional semantic embedding feature vector ; represents the input high-dimensional semantic embedding feature vector, which is used to represent the semantic information of the input data (such as historical semantic vector samples); represents the weight vector or hyperplane direction vector associated with the th hash function, which is used to perform an inner product calculation with the input high-dimensional semantic embedding feature vector to define a hyperplane so that the input high-dimensional semantic embedding feature vector can be binarized after projection; represents the bias constant in the th hash function, which is used to adjust the position of the hyperplane so that the mapping result better conforms to a specific data distribution.
[0065] Perform calculations on each high-dimensional semantic embedding feature vector to obtain a binary code of length 64, which is the initial hash code representation result.
[0066] Store all the binary initial hash code representation results obtained for each time partition group separately to form a corresponding phased hash representation set for each time partition group;
[0067] Specifically, save all the binary hash code sets obtained within each time partition group separately to form independent hash representation sets that correspond one-to-one with the time partition groups, which is convenient for retrieval and matching according to time information.
[0068] For example, all the binary codes generated by the time partition group on March 15 are stored as set A, and those on March 16 are stored as set B. Set A and set B respectively form their corresponding phased hash representation sets.
[0069] Specifically, for the current newly added data, perform a dynamic clustering and matching operation based on the time period to which the current newly added data belongs and the mapping positions of the corresponding semantic similarity vectors in each phased hash representation set to obtain a target hash center candidate set for the current newly added data, including:
[0070] Determine the latest time partition group to which it belongs based on the timestamp of the original data corresponding to the current newly added data;
[0071] Specifically, by obtaining the generation timestamp of the currently newly added data, it is determined which pre-divided time partition group it should be classified into.
[0072] For example, if the generation time of the currently newly added data is 10:30 on March 17, 2023, it is classified into the time partition group of "March 17, 2023".
[0073] In the set of stage hash representations corresponding to the latest time partition group, by calculating the cosine similarity between the semantic embedding features of the currently newly added data and the historical semantic vector samples, the hash code corresponding to the historical semantic vector sample with a cosine similarity greater than the similarity threshold is selected as a clustering matching candidate.
[0074] Specifically, first, using the same pre-trained semantic embedding model as the historical semantic vector samples, the currently newly added data is converted into a high-dimensional semantic embedding feature vector, and then the cosine similarity is calculated with the high-dimensional vectors of each historical semantic vector sample in the latest time partition group. The calculation formula is:
[0075] ; where represents the cosine value between two high-dimensional semantic embedding feature vectors, and its value range is from -1 to 1, which is used to measure the direction similarity degree of the two high-dimensional semantic embedding feature vectors; represents the high-dimensional semantic embedding feature vector generated by the currently newly added data through the pre-trained semantic embedding model; represents the high-dimensional semantic embedding feature vector generated by the historical semantic vector sample through the pre-trained semantic embedding model.
[0076] Exemplarily, if the cosine value is 0.85 and the set similarity threshold is 0.8, it is considered that the historical sample is semantically similar to the currently newly added data, and the corresponding binary hash code is selected as a clustering matching candidate.
[0077] According to the clustering algorithm, cluster the hash codes in the clustering matching candidates to determine the target hash center candidate set corresponding to the currently newly added data;
[0078] Specifically, apply the clustering algorithm (such as clustering center calculation or density clustering) to the selected candidate hash code set, group the candidate hash codes according to similarity, and the central value of each cluster in the grouping result is regarded as a candidate target hash center.
[0079] Exemplarily, assume the candidate hash codes are {1010, 1011, 1001, 0110}, and after being clustered by the clustering algorithm into two clusters, the center of the first cluster is 1010, and the center of the second cluster is 0110, then the target hash center candidate set is {1010, 0110}.
[0080] Specifically, based on the target hash center candidate set and combined with the historical hash code drift trajectory model, an adaptive calibration hash code for the currently newly added data is generated, including:
[0081] Construct a historical hash code drift trajectory model according to the temporal variation law of hash centers within each historical time partition group;
[0082] Specifically, analyze the change trend of the hash center values calculated in different historical time partition groups over time, and construct a mathematical model to predict the hash center position at the current time point.
[0083] For example, if the hash centers are respectively in several consecutive days, a linear regression can be used to fit the change trend: ; where represents the predicted hash center at the current moment, which is a representative code in the hash code space and is used to correct the hash code generated by the currently newly added data; represents the hash center at the reference time point and serves as the initial value of the historical hash code drift trajectory model; is the slope, indicating the rate of change of the hash center over time, that is, the value by which the hash center changes per unit time; represents the time corresponding to the current prediction of the hash center; represents the reference time point.
[0084] Determine the candidate hash code with the minimum distance as the adaptive calibration hash code for the currently newly added data through the Euclidean distance between each candidate hash code in the target hash center candidate set and the hash code at the current moment predicted by the historical hash code drift trajectory model;
[0085] Specifically, for each candidate hash code in the target hash center candidate set calculate the Euclidean distance from the predicted hash center at the current moment
[0086] ; where represents the Euclidean distance between the candidate hash code and the predicted hash center ; represents the number of bits or vector dimension of the hash code; represents the value of the candidate hash code at the -th bit (usually a binary value 0 or 1); represents the value of the predicted hash center at the -th bit.
[0087] After calculating all candidate codes, select the candidate code with the smallest Euclidean distance as the adaptive calibration hash code.
[0088] Specifically, write the adaptive calibration hash code of the currently newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data, including:
[0089] In the vector database, use the adaptive calibration hash code of the currently newly added data as the index primary key to establish a mapping relationship between the adaptive calibration hash code and the original semantic vector features of the currently newly added data;
[0090] Specifically, the adaptive calibration hash code obtained by the currently newly added data is represented in a fixed format (such as a binary string) and is used as an index key in the database. At the same time, the original semantic vector features of the currently newly added data (that is, the high-dimensional semantic embedding feature vector obtained through the semantic embedding model) are associated and stored with the adaptive calibration hash code, forming a one-to-one or many-to-one mapping relationship.
[0091] For example, the adaptive hash code generated by the new data is "10101010", and its semantic embedding vector is (0.45, 0.67,..., 0.12); record the key-value pair in the database: "10101010" to (0.45, 0.67,..., 0.12).
[0092] When a retrieval request arrives, generate the adaptive calibration hash code corresponding to the retrieval request, and use the hash index structure in the vector database to locate the corresponding hash bucket, and return the historical semantic vector data that matches the adaptive calibration hash code of the retrieval request to complete the fast retrieval of similar data;
[0093] Specifically, after receiving the retrieval request, first process the retrieval input data through a pre-trained semantic embedding model, etc., to generate an adaptive calibration hash code in the same storage format as in the database. Then, through the pre-constructed hash index structure in the database (such as a data structure based on a hash table), use this hash code to quickly locate the hash bucket storing the same (or approximate) hash code. Finally, retrieve the historical semantic vector data that matches it from this hash bucket.
[0094] Exemplarily, if the adaptive calibration hash code generated by the retrieval request is "10101010", all records with the key "10101010" in the corresponding hash bucket are found through the index in the database, and the semantic vector data in these records is returned.
[0095] Specifically, when a retrieval request arrives, an adaptive calibration hash code corresponding to the retrieval request is generated, and the corresponding hash bucket is located in the vector database using the hash index structure, and historical semantic vector data matching the adaptive calibration hash code of the retrieval request is returned to complete the fast retrieval of similar data, including:
[0096] Input the retrieval request into a pre-trained semantic embedding model to generate a corresponding high-dimensional semantic embedding feature vector;
[0097] Specifically, when a retrieval request is received, first input the data included in the retrieval request (such as query text, picture, or other form of data) into the same trained semantic embedding model, and use the forward inference process of the semantic embedding model to calculate a high-dimensional semantic embedding feature vector, which represents the semantic content of the retrieval request. For example, when the user inputs "query example text", the semantic embedding model outputs a semantic embedding feature vector , whose dimension is the same as that of the historical data vector.
[0098] Input the high-dimensional semantic embedding feature vector representation corresponding to the retrieval request into the historical hash code drift trajectory model to calculate the drift vector corresponding to the time of the retrieval request;
[0099] Specifically, using the constructed historical hash code drift trajectory model, combine the time information of the current retrieval request with its high-dimensional semantic embedding feature to calculate the drift vector, which is used to reflect the change trend of the hash code at the current time: ; where represents the drift vector; is the slope, which represents the rate of change of the hash center over time, that is, the value of the change of the hash center per unit time; represents the time corresponding to the current hash center to be predicted; represents the reference time point.
[0100] Perform a vector superposition operation on the drift vector and the high-dimensional semantic embedding feature vector corresponding to the retrieval request to generate an adaptive calibration hash code corresponding to the retrieval request;
[0101] Specifically, perform a vector addition operation on the high-dimensional semantic embedding feature vector and the drift vector to obtain a corrected vector. Then apply the hash function to the corrected vector to convert it into a binary form, which is the adaptive calibration hash code corresponding to the retrieval request.
[0102] Query the adaptive calibration hash code in the hash index structure in the vector database, locate the hash bucket with the same hash code, and extract the historical semantic vector data matching the adaptive calibration hash code corresponding to the retrieval request from the hash bucket to complete the fast retrieval of similar data;
[0103] Specifically, the generated adaptive calibration hash code corresponding to the retrieval request is used as the query key, and the hash bucket storing the data with the same or similar hash codes is located through the pre-established hash index structure (such as the organization method based on the hash table) in the vector database. All historical data records matching the query hash code are extracted from this hash bucket, and the rapid retrieval is completed.
[0104] Embodiment 2
[0105] The difference between Embodiment 2 and Embodiment 1 of the present invention is that this embodiment introduces a semantic hashing accelerated vector database retrieval system.
[0106] Figure 2 The structural schematic diagram of a semantic hashing accelerated vector database retrieval system of the present invention is given. A semantic hashing accelerated vector database retrieval system includes a time partition division module, an embedding model fitting module, a dynamic clustering matching module, a hash code calibration module, and an index construction and retrieval module;
[0107] Time partition division module: Obtain multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs;
[0108] Embedding model fitting module: Perform embedding model fitting operations on each time partition group respectively, extract the corresponding semantic embedding feature sets, and generate an initial hash code space for each feature set to form multiple phased hash representation sets;
[0109] Dynamic clustering matching module: For the currently newly added data, perform dynamic clustering matching operations according to the time period to which the currently newly added data belongs and the mapping positions of the corresponding semantic similarity vectors in each phased hash representation set to obtain the target hash center candidate set of the currently newly added data;
[0110] Hash code calibration module: Based on the target hash center candidate set, combine the historical hash code drift trajectory model to generate the adaptive calibration hash code of the currently newly added data;
[0111] Index construction and retrieval module: Write the adaptive calibration hash code of the currently newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to realize the semantic consistency backtracking query of historical data.
[0112] All the above formulas are dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0113] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0114] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0115] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0116] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0117] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module. It may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0118] In addition, in each embodiment of the present application, each functional module may be integrated into a processing module, or each module may exist physically alone, or two or more modules may be integrated into one module.
[0119] If the described function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0120] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
[0121] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for retrieving a vector database accelerated by semantic hashing, characterized in that The steps are as follows: Obtain multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Perform an embedding model fitting operation on each time partition group respectively, extract the corresponding semantic embedding feature set, and generate an initial hash code space for each feature set to form multiple phased hash representation sets; For the current newly added data, perform a dynamic clustering matching operation according to the time period to which the current newly added data belongs and the mapping positions of the corresponding semantic similarity vectors in each phased hash representation set to obtain a target hash center candidate set for the current newly added data; Based on the target hash center candidate set, combine the historical hash code drift trajectory model to generate an adaptive calibration hash code for the current newly added data; Write the adaptive calibration hash code of the current newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to achieve semantic consistency retrospective query of historical data.
2. The vector database retrieval method with semantic hashing acceleration according to claim 1, wherein Obtain multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs. Specifically: Set a fixed time interval length according to the timestamp generated by the original data corresponding to each historical semantic vector sample; According to the fixed time interval length, sequentially divide all historical semantic vector samples into multiple continuous and non-overlapping time partition groups; the number of historical semantic vector samples in each time partition group is greater than a preset quantity threshold.
3. The semantic hashing accelerated vector database retrieval method according to claim 2, characterized in that Perform an embedding model fitting operation on each time partition group respectively, extract the corresponding semantic embedding feature set, and generate an initial hash code space for each feature set to form multiple phased hash representation sets. Specifically: Input all historical semantic vector samples in each time partition group into a pre-trained semantic embedding model, and perform forward inference calculation to obtain the high-dimensional semantic embedding feature vector corresponding to each historical semantic vector sample; Perform a hash function mapping operation on all the high-dimensional semantic embedding feature vectors obtained in each time partition group in sequence to obtain the corresponding binary initial hash code representation result; Store all the binary initial hash code representation results obtained from each time partition group respectively to form a phased hash representation set corresponding to each time partition group.
4. A method for retrieving a vector database with semantic hashing acceleration according to claim 3, characterized in that For the current newly added data, perform a dynamic clustering matching operation according to the time period to which the current newly added data belongs and the mapping positions of the corresponding semantic similarity vectors in each phased hash representation set to obtain a target hash center candidate set for the current newly added data. Specifically: Determine the latest time partition group to which it belongs based on the timestamp generated by the original data corresponding to the current newly added data; In the phased hash representation set corresponding to the latest time partition group, calculate the cosine similarity between the semantic embedding feature of the current newly added data and the historical semantic vector samples, and select the hash code corresponding to the historical semantic vector sample with a cosine similarity greater than the similarity threshold as the clustering matching candidate; Perform clustering processing on the hash codes in the clustering matching candidates according to the clustering algorithm to determine the target hash center candidate set corresponding to the current newly added data.
5. A method for retrieving a vector database with semantic hashing acceleration according to claim 4, characterized in that, Based on the target hash center candidate set, combined with the historical hash code drift trajectory model, an adaptive calibration hash code for the currently newly added data is generated, specifically as follows: According to the temporal variation law of the hash centers within each historical time partition group, construct a historical hash code drift trajectory model; Determine the candidate hash code with the smallest distance as the adaptive calibration hash code for the currently newly added data through the Euclidean distance between each candidate hash code in the target hash center candidate set and the hash code predicted by the historical hash code drift trajectory model at the current moment.
6. The method for retrieving a vector database with semantic hashing acceleration according to claim 5, wherein Write the adaptive calibration hash code of the currently newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to achieve semantic consistency retrospective query of historical data, specifically as follows: In the vector database, use the adaptive calibration hash code of the currently newly added data as the index primary key to establish a mapping relationship between the adaptive calibration hash code and the original semantic vector features of the currently newly added data; When a retrieval request arrives, generate an adaptive calibration hash code corresponding to the retrieval request, and use the hash index structure in the vector database to locate the corresponding hash bucket, and return the historical semantic vector data that matches the adaptive calibration hash code of the retrieval request to complete the fast retrieval of similar data.
7. A method for retrieving a vector database with semantic hashing acceleration according to claim 6, characterized in that, When a retrieval request arrives, generate an adaptive calibration hash code corresponding to the retrieval request, and use the hash index structure in the vector database to locate the corresponding hash bucket, and return the historical semantic vector data that matches the adaptive calibration hash code of the retrieval request to complete the fast retrieval of similar data, specifically as follows: Input the retrieval request into the pre-trained semantic embedding model to generate a corresponding high-dimensional semantic embedding feature vector; Input the high-dimensional semantic embedding feature vector representation corresponding to the retrieval request into the historical hash code drift trajectory model to calculate the drift vector corresponding to the retrieval request time; Perform a vector superposition operation on the drift vector and the high-dimensional semantic embedding feature vector corresponding to the retrieval request to generate an adaptive calibration hash code corresponding to the retrieval request; Query the adaptive calibration hash code in the hash index structure in the vector database to locate the hash bucket with the same hash code, and extract the historical semantic vector data that matches the adaptive calibration hash code corresponding to the retrieval request from the hash bucket to complete the fast retrieval of similar data.
8. A semantic hashing-accelerated vector database retrieval system for implementing a semantic hashing-accelerated vector database retrieval method according to any one of claims 1-7, characterized in that It includes a time partition division module, an embedding model fitting module, a dynamic clustering matching module, a hash code calibration module, and an index construction and retrieval module; Time partition division module: Obtain multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Embedding model fitting module: Perform embedding model fitting operations on each time partition group respectively, extract the corresponding semantic embedding feature sets, and generate an initial hash code space for each feature set to form multiple phased hash representation sets; Dynamic clustering matching module: For the currently newly added data, perform dynamic clustering matching operations according to the time period to which the currently newly added data belongs and the mapping positions of the corresponding semantic similarity vectors in each phased hash representation set to obtain the target hash center candidate set of the currently newly added data; Hash Code Calibration Module: Based on the target hash center candidate set, combined with the historical hash code drift trajectory model, generate an adaptive calibration hash code for the currently newly added data; Index Construction and Retrieval Module: Write the adaptive calibration hash code of the currently newly added data into the vector database to construct an accelerated index structure, and perform similar data retrieval based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data.
Citation Information
Patent Citations
Distributed index method based on LSH (Locality Sensitive Hashing)
CN103744934A
Rapid cross-modal retrieval method and system for incremental data carrying new categories
CN113326289A