Semantic hash accelerated vector database retrieval method and system

By time partitioning and phased hash representation sets of historical semantic vector samples, combining dynamic cluster matching and historical hash code drift trajectory model, adaptive calibration hash codes are generated, which solves the problem of hash code distribution drift in dynamic incremental scenarios, and improves retrieval accuracy and recall rate.

CN119990144AActive Publication Date: 2025-05-13SHANGHAI JUXIAN NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510472581.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art cannot effectively maintain the consistency of hash code distribution in dynamic incremental scenarios, resulting in a decrease in semantic similarity measurement distortion and retrieval recall.

Method used

By obtaining historical semantic vector samples and partitioning by time, a phased hash representation set is constructed, combining dynamic cluster matching and historical hash code drift trajectory model, an adaptive calibration hash code is generated to ensure the consistency of semantic mapping.

Benefits of technology

It effectively solves the problem of hash code distribution drift, improves the search recall rate and query accuracy, improves the search speed and index construction efficiency, and maintains stability and scalability in large-scale and continuous updates of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990144A_ABST
    Figure CN119990144A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic hash accelerated vector database retrieval method and system, and particularly relates to the technical field of data retrieval. The method is used for solving the problem that a traditional static semantic hash model cannot adapt to new and old data distribution drift in a dynamic increment scene. The method comprises the following steps of: dividing a historical semantic vector sample into a plurality of time partition groups according to an original data timestamp, and performing embedded model fitting on each time partition group to generate an initial hash code space and form a stage hash representation set; performing dynamic clustering matching on the current newly-added data according to the time period to which the current newly-added data belong and the mapping position of the historical similar sample in the hash representation set of each stage to obtain a target hash center candidate set; and further combining a historical hash code drift trajectory model to generate an adaptive calibration hash code, writing the adaptive calibration hash code into a vector database to construct an accelerated index structure, and realizing rapid retrieval and semantic consistency backtracking query of similar data through the adaptive calibration hash code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data retrieval technology, and more specifically, to a semantic hash-accelerated vector database retrieval method and system. Background Art

[0002] In the context of big data, vector databases are widely used for similarity retrieval of multimodal data such as text, images, and speech. The existing technology mainly uses semantic embedding combined with hash coding to map high-dimensional data into discrete hash codes to achieve fast retrieval. The existing technology generally adopts a static coding strategy, which cannot fully adapt to the distribution drift problem caused by the continuous incremental data, causing the hash code of the newly added data to gradually deviate from the original cluster center, thereby causing the semantic similarity measurement to be distorted and the retrieval recall rate to decrease. Therefore, how to maintain the consistency of hash code distribution and improve retrieval accuracy in dynamic incremental scenarios has become a key problem that needs to be solved.

[0003] In order to solve the above problems, a technical solution is now provided. Summary of the invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a vector database retrieval method and system accelerated by semantic hashing to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions: A vector database retrieval method accelerated by semantic hashing comprises the following steps: Acquire multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Perform embedding model fitting operation on each time partition group, extract corresponding semantic embedding feature set, and generate initial hash code space for each feature set to form multiple staged hash representation sets; For the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic clustering matching operation is performed to obtain the target hash center candidate set of the current newly added data; Based on the target hash center candidate set and the historical hash code drift trajectory model, an adaptive calibration hash code for the current newly added data is generated; The adaptive calibration hash code of the current newly added data is written into the vector database to build an accelerated index structure, and similar data is retrieved based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data.

[0006] In a preferred embodiment, a plurality of historical semantic vector samples are obtained, and are divided into a plurality of time partition groups based on the original data time period to which each semantic vector sample belongs, specifically: According to the timestamp of each historical semantic vector sample corresponding to the original data generation, a fixed time interval length is set; According to the length of the fixed time interval, all historical semantic vector samples are divided into a plurality of continuous and non-overlapping time partition groups in sequence; the number of historical semantic vector samples in each time partition group is greater than a preset number threshold.

[0007] In a preferred embodiment, an embedding model fitting operation is performed on each time partition group, the corresponding semantic embedding feature set is extracted, and an initial hash code space is generated for each feature set to form multiple staged hash representation sets, specifically: Input the pre-trained semantic embedding model into all historical semantic vector samples in each time partition group, and perform forward reasoning calculation to obtain the high-dimensional semantic embedding feature vector corresponding to each historical semantic vector sample; All high-dimensional semantic embedding feature vectors obtained from each time partition group are sequentially subjected to hash function mapping operations to obtain the corresponding binary initial hash code representation results; All binary initial hash code representation results obtained for each time partition group are stored separately to form a phased hash representation set corresponding to each time partition group.

[0008] In a preferred embodiment, for the current newly added data, a dynamic cluster matching operation is performed according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set to obtain the target hash center candidate set of the current newly added data, specifically: Determine the latest time partition group to which the newly added data belongs based on the timestamp generated by the original data; In the phased hash representation set corresponding to the latest time partition group, by calculating the cosine similarity between the semantic embedding features of the current newly added data and the historical semantic vector samples, the hash codes corresponding to the historical semantic vector samples whose cosine similarity is greater than the similarity threshold are selected as cluster matching candidates; The hash codes in the cluster matching candidates are clustered according to the clustering algorithm to determine the target hash center candidate set corresponding to the current newly added data.

[0009] In a preferred embodiment, based on the target hash center candidate set and in combination with the historical hash code drift trajectory model, an adaptive calibration hash code for the current newly added data is generated, specifically: According to the temporal change law of the hash center in each historical time partition group, a historical hash code drift trajectory model is constructed; The candidate hash code with the smallest distance is determined as the adaptive calibration hash code for the current newly added data through the Euclidean distance between each candidate hash code in the target hash center candidate set and the hash code at the current moment predicted by the historical hash code drift trajectory model.

[0010] In a preferred embodiment, the adaptive calibration hash code of the current newly added data is written into the vector database to construct an accelerated index structure, and similar data is retrieved based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data, specifically: In the vector database, the adaptive calibration hash code of the current newly added data is used as the index primary key, and a mapping relationship between the adaptive calibration hash code and the original semantic vector features of the current newly added data is established; When a retrieval request arrives, an adaptive calibrated hash code corresponding to the retrieval request is generated, and the corresponding hash bucket is located in the vector database using the hash index structure, and the historical semantic vector data matching the adaptive calibrated hash code of the retrieval request is returned to complete the rapid retrieval of similar data.

[0011] In a preferred embodiment, when a search request arrives, an adaptive calibration hash code corresponding to the search request is generated, and the corresponding hash bucket is located in the vector database using the hash index structure, and the historical semantic vector data matching the adaptive calibration hash code of the search request is returned to complete the rapid search of similar data, specifically: Input the search request into the pre-trained semantic embedding model to generate the corresponding high-dimensional semantic embedding feature vector; The high-dimensional semantic embedding feature vector representation corresponding to the retrieval request is input into the historical hash code drift trajectory model, and the drift vector corresponding to the retrieval request time is calculated; Performing vector superposition operation on the drift vector and the high-dimensional semantic embedding feature vector corresponding to the retrieval request to generate an adaptive calibration hash code corresponding to the retrieval request; The adaptive calibration hash code is queried in the hash index structure in the vector database to locate the hash bucket with the same hash code, and the historical semantic vector data matching the adaptive calibration hash code corresponding to the retrieval request is extracted from the hash bucket to complete the rapid retrieval of similar data.

[0012] On the other hand, the present invention provides a semantic hash accelerated vector database retrieval system, including a time partition division module, an embedding model fitting module, a dynamic cluster matching module, a hash code calibration module, and an index construction and retrieval module; Time partition division module: obtains multiple historical semantic vector samples and divides them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Embedding model fitting module: performs embedding model fitting operations on each time partition group, extracts the corresponding semantic embedding feature set, and generates an initial hash code space for each feature set to form multiple staged hash representation sets; Dynamic cluster matching module: For the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic cluster matching operation is performed to obtain the target hash center candidate set of the current newly added data; Hash code calibration module: Based on the target hash center candidate set and the historical hash code drift trajectory model, it generates an adaptive calibration hash code for the current newly added data; Index construction and retrieval module: write the adaptive calibrated hash code of the current newly added data into the vector database to build an accelerated index structure, and retrieve similar data based on the adaptive calibrated hash code to achieve semantic consistency backtracking query of historical data.

[0013] The technical effects and advantages of the semantic hash accelerated vector database retrieval method and system of the present invention are as follows: By partitioning historical semantic vector samples by time, constructing a phased hash representation set, and combining dynamic cluster matching with the historical hash code drift trajectory model, the generation of adaptive calibrated hash codes is achieved. It effectively solves the problem of hash code distribution drift between new and old data in the traditional static semantic hash model under dynamic incremental scenarios, ensures the consistency of semantic mapping, thereby improving retrieval recall and query accuracy, and improving retrieval speed and index building efficiency. When facing large-scale, continuously updated data, it can maintain high stability and scalability, providing reliable technical support for the fast and accurate retrieval of massive data. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A schematic diagram of a vector database retrieval method accelerated by semantic hashing according to the present invention; Figure 2 A schematic diagram of the structure of a vector database retrieval system accelerated by semantic hashing according to the present invention. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0016] Example 1

[0017] Figure 1The present invention provides a vector database retrieval method accelerated by semantic hashing, which comprises the following steps: Acquire multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Perform embedding model fitting operation on each time partition group, extract corresponding semantic embedding feature set, and generate initial hash code space for each feature set to form multiple staged hash representation sets; For the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic clustering matching operation is performed to obtain the target hash center candidate set of the current newly added data; Based on the target hash center candidate set and the historical hash code drift trajectory model, an adaptive calibration hash code for the current newly added data is generated; The adaptive calibration hash code of the current newly added data is written into the vector database to build an accelerated index structure, and similar data is retrieved based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data.

[0018] Specifically, multiple historical semantic vector samples are obtained, and based on the original data time period to which each semantic vector sample belongs, they are divided into multiple time partition groups, including: According to the timestamp of each historical semantic vector sample corresponding to the original data generation, a fixed time interval length is set; Specifically, a timestamp is extracted from the original data of each historical semantic vector sample to determine a fixed time span. For example, the fixed time interval is 24 hours, 1 week or 1 month. For example, if the timestamp format of the historical semantic vector sample is "year / month / day hour: minute: second", "24 hours" can be set as the fixed time interval length, that is, the data within every 24 hours is grouped together.

[0019] According to the length of the fixed time interval, all historical semantic vector samples are divided into a plurality of continuous and non-overlapping time partition groups in sequence; the number of historical semantic vector samples in each time partition group is greater than a preset number threshold; Specifically, all historical semantic vector samples are sorted in chronological order using a fixed time interval and divided into several continuous and non-overlapping time partition groups. The number of historical semantic vector samples in each time partition group is required to be greater than a preset threshold to ensure sufficient processing data volume for model training or statistical analysis.

[0020] For example, if the fixed time interval is 24 hours and the preset number threshold is 100, the time period is valid only when the number of historical semantic vector samples within 24 hours is greater than 100; otherwise, adjacent time periods are merged or ignored.

[0021] Specifically, an embedding model fitting operation is performed on each time partition group, the corresponding semantic embedding feature set is extracted, and an initial hash code space is generated for each feature set to form multiple staged hash representation sets, including: Input the pre-trained semantic embedding model into all historical semantic vector samples in each time partition group, and perform forward reasoning calculation to obtain the high-dimensional semantic embedding feature vector corresponding to each historical semantic vector sample; Specifically, for each set of historical semantic vector samples grouped by time, each historical semantic vector sample is input into a pre-trained semantic embedding model, and the semantic embedding model is used for forward calculation to obtain a high-dimensional semantic embedding feature vector. The high-dimensional semantic embedding feature vector accurately reflects the semantic information of the input data. For example, after a piece of text is input into the semantic embedding model, a 512-dimensional vector is output. Each historical semantic vector sample undergoes this process to generate a corresponding high-dimensional semantic embedding feature vector.

[0022] All high-dimensional semantic embedding feature vectors obtained from each time partition group are sequentially subjected to hash function mapping operations to obtain the corresponding binary initial hash code representation results; Specifically, a hash function is applied to each high-dimensional semantic embedding feature vector to map continuous real number vectors into fixed-length binary codes.

[0023] For example, setting the hash function to: ; in, Indicates The hash function is used to embed the input high-dimensional semantic feature vector The binary output obtained after calculation; A high-dimensional semantic embedding feature vector representing the input, which is used to represent the semantic information of the input data (such as historical semantic vector samples); Indicates The weight vector or hyperplane direction vector associated with the hash function is used to embed the input high-dimensional semantic feature vector Perform inner product calculation and define a hyperplane so that the input high-dimensional semantic embedding feature vector can be binarized after projection; Indicates The bias constant in the hash function is used to adjust the position of the hyperplane so that the mapping result is more consistent with the specific data distribution.

[0024] Each high-dimensional semantic embedding feature vector is calculated to obtain a binary code with a length of 64, which is the initial hash code representation result.

[0025] All binary initial hash code representation results obtained for each time partition group are stored separately to form a phased hash representation set corresponding to each time partition group; Specifically, all binary hash code sets obtained in each time partition group are saved separately to form an independent hash representation set corresponding to the time partition group one by one, so as to facilitate retrieval and matching according to time information.

[0026] For example, all binary codes generated by the time partition group on March 15 are stored as set A, and those on March 16 are stored as set B. Set A and set B constitute their respective stage hash representation sets.

[0027] Specifically, for the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic clustering matching operation is performed to obtain the target hash center candidate set of the current newly added data, including: Determine the latest time partition group to which the newly added data belongs based on the timestamp generated by the original data; Specifically, by obtaining the generation timestamp of the current newly added data, it is determined into which pre-divided time partition group the newly added data should be included.

[0028] For example, if the current new data generation time is 10:30 on March 17, 2023, it will be included in the time partition group of "March 17, 2023".

[0029] In the phased hash representation set corresponding to the latest time partition group, by calculating the cosine similarity between the semantic embedding features of the current newly added data and the historical semantic vector samples, the hash codes corresponding to the historical semantic vector samples whose cosine similarity is greater than the similarity threshold are selected as cluster matching candidates; Specifically, the current newly added data is first converted into a high-dimensional semantic embedding feature vector using the same pre-trained semantic embedding model as the historical semantic vector sample, and then the cosine similarity is calculated with the high-dimensional vector of each historical semantic vector sample in the latest time partition group. The calculation formula is: ;in, Represents the cosine value between two high-dimensional semantic embedding feature vectors, with a value range of -1 to 1, which is used to measure the directional similarity of two high-dimensional semantic embedding feature vectors; Represents the high-dimensional semantic embedding feature vector generated by the pre-trained semantic embedding model for the current newly added data; Represents the high-dimensional semantic embedding feature vector generated by the pre-trained semantic embedding model of the historical semantic vector sample.

[0030] Exemplarily, if the cosine value is 0.85 and the set similarity threshold is 0.8, the historical sample is considered to be semantically similar to the current new data, and its corresponding binary hash code is selected as the cluster matching candidate.

[0031] Cluster the hash codes in the cluster matching candidates according to the clustering algorithm to determine the target hash center candidate set corresponding to the current newly added data; Specifically, a clustering algorithm (such as cluster center calculation or density clustering) is applied to the screened candidate hash code set, the candidate hash codes are grouped according to similarity, and the center value of each cluster in the grouping result is regarded as a candidate target hash center.

[0032] Exemplarily, assuming that the candidate hash codes are {1010, 1011, 1001, 0110}, they are divided into two clusters through the clustering algorithm, the center of the first cluster is 1010, and the center of the second cluster is 0110, then the target hash center candidate set is {1010, 0110}.

[0033] Specifically, based on the target hash center candidate set and combined with the historical hash code drift trajectory model, an adaptive calibration hash code for the current newly added data is generated, including: According to the temporal change law of the hash center in each historical time partition group, a historical hash code drift trajectory model is constructed; Specifically, the changing trend of the hash center values ​​calculated in different historical time partition groups over time is analyzed, and a mathematical model is constructed to predict the hash center position at the current time point.

[0034] For example, if the hash centers are , linear regression can be used to fit the changing trend: ;in, Indicates the hash center of the current moment of the prediction, which is a representative code in the hash code space and is used to correct the hash code generated by the current newly added data; Indicates at the reference time point The hash center of is used as the initial value of the historical hash code drift trajectory model; is the slope, which indicates the rate at which the hash center changes over time, that is, the value of the hash center change per unit time; Indicates the time when the hash center needs to be predicted; Indicates a reference time point.

[0035] The candidate hash code with the smallest distance is determined as the adaptive calibration hash code for the current newly added data through the Euclidean distance between each candidate hash code in the target hash center candidate set and the hash code at the current moment predicted by the historical hash code drift trajectory model; Specifically, for each candidate hash code in the target hash center candidate set The hash center of the current moment with the prediction Calculate the Euclidean distance, the calculation formula is: ;in, Represents a candidate hash code With predicted hash center The Euclidean distance between The number of bits or vector dimensions representing the hash code; Represents a candidate hash code In the The numeric value of the bit (usually the binary value 0 or 1); Represents the predicted hash center In the The value of the bit.

[0036] After calculating all candidate codes, the candidate code with the smallest Euclidean distance is selected as the adaptive calibration hash code.

[0037] Specifically, the adaptive calibration hash code of the current newly added data is written into the vector database to build an accelerated index structure, and similar data is retrieved based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data, including: In the vector database, the adaptive calibration hash code of the current newly added data is used as the index primary key, and a mapping relationship between the adaptive calibration hash code and the original semantic vector features of the current newly added data is established; Specifically, the adaptive calibration hash code obtained for the current newly added data is represented in a fixed format (e.g., a binary string) and is used as an index key in the database. At the same time, the original semantic vector features of the current newly added data (i.e., the high-dimensional semantic embedding feature vector obtained by the semantic embedding model) are associated with the adaptive calibration hash code and stored to form a one-to-one or many-to-one mapping relationship.

[0038] For example, the adaptive hash code generated for new data is “10101010”, and its semantic embedding vector is (0.45, 0.67, ..., 0.12); the key-value pair is recorded in the database: “10101010” to (0.45, 0.67, ..., 0.12).

[0039] When a search request arrives, an adaptive calibrated hash code corresponding to the search request is generated, and the corresponding hash bucket is located in the vector database using the hash index structure, and the historical semantic vector data matching the adaptive calibrated hash code of the search request is returned to complete the rapid search of similar data; Specifically, after receiving a search request, the search input data is first processed by a pre-trained semantic embedding model to generate an adaptive calibrated hash code that is consistent with the storage format in the database. Then, the hash code is used to quickly locate the hash bucket that stores the same (or similar) hash code through the pre-built hash index structure in the database (e.g., a data structure based on a hash table). Finally, the historical semantic vector data that matches it is retrieved from the hash bucket.

[0040] Exemplarily, if the adaptive calibration hash code generated by the retrieval request is "10101010", all records with the key "10101010" in the corresponding hash bucket are found in the database through the index, and the semantic vector data in these records are returned.

[0041] Specifically, when a search request arrives, an adaptive calibration hash code corresponding to the search request is generated, and the corresponding hash bucket is located in the vector database using the hash index structure, and the historical semantic vector data matching the adaptive calibration hash code of the search request is returned to complete the rapid search of similar data, including: Input the search request into the pre-trained semantic embedding model to generate the corresponding high-dimensional semantic embedding feature vector; Specifically, when a search request is received, the data contained in the search request (such as query text, images, or other forms of data) is first input into the semantic embedding model that has also been trained, and a high-dimensional semantic embedding feature vector is calculated using the forward reasoning process of the semantic embedding model to represent the semantic content of the search request. For example, if a user inputs "query sample text", the semantic embedding model outputs a semantic embedding feature vector , the dimension is consistent with the historical data vector.

[0042] The high-dimensional semantic embedding feature vector representation corresponding to the retrieval request is input into the historical hash code drift trajectory model, and the drift vector corresponding to the retrieval request time is calculated; Specifically, the constructed historical hash code drift trajectory model is used to combine the time information of the current retrieval request with its high-dimensional semantic embedding features to calculate the drift vector, which is used to reflect the change trend of the hash code at the current time: ;in, represents the drift vector; is the slope, which indicates the rate at which the hash center changes over time, that is, the value of the hash center change per unit time; Indicates the time when the hash center needs to be predicted; Indicates a reference time point.

[0043] Performing vector superposition operation on the drift vector and the high-dimensional semantic embedding feature vector corresponding to the retrieval request to generate an adaptive calibration hash code corresponding to the retrieval request; Specifically, high-dimensional semantics are embedded into feature vectors With drift vector Perform vector addition to obtain a corrected vector, and then convert the corrected vector into a binary form using a hash function, which is the adaptive calibration hash code corresponding to the retrieval request.

[0044] The adaptive calibration hash code is queried in the hash index structure in the vector database to locate the hash bucket with the same hash code, and the historical semantic vector data matching the adaptive calibration hash code corresponding to the retrieval request is extracted from the hash bucket to complete the rapid retrieval of similar data; Specifically, the adaptive calibration hash code corresponding to the generated search request is used as the query key, and the hash bucket storing the same or similar hash code data is located through the hash index structure (e.g., an organization based on a hash table) pre-established in the vector database. All historical data records matching the query hash code are extracted from the hash bucket, thus completing the fast search.

[0045] Example 2

[0046] The difference between Example 2 of the present invention and Example 1 is that this example introduces a vector database retrieval system accelerated by semantic hashing.

[0047] Figure 2 A structural schematic diagram of a semantic hash accelerated vector database retrieval system of the present invention is given, a semantic hash accelerated vector database retrieval system, including a time partition division module, an embedded model fitting module, a dynamic cluster matching module, a hash code calibration module and an index construction and retrieval module; Time partition division module: obtains multiple historical semantic vector samples and divides them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Embedding model fitting module: performs embedding model fitting operations on each time partition group, extracts the corresponding semantic embedding feature set, and generates an initial hash code space for each feature set to form multiple staged hash representation sets; Dynamic cluster matching module: For the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic cluster matching operation is performed to obtain the target hash center candidate set of the current newly added data; Hash code calibration module: Based on the target hash center candidate set and the historical hash code drift trajectory model, it generates an adaptive calibration hash code for the current newly added data; Index construction and retrieval module: write the adaptive calibrated hash code of the current newly added data into the vector database to build an accelerated index structure, and retrieve similar data based on the adaptive calibrated hash code to achieve semantic consistency backtracking query of historical data.

[0048] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.

[0049] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or may be transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state hard disk.

[0050] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0051] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0052] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0053] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0054] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0055] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0056] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0057] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A vector database retrieval method accelerated by semantic hashing, characterized in that: The steps include: Acquire multiple historical semantic vector samples, and divide them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Perform embedding model fitting operation on each time partition group, extract corresponding semantic embedding feature set, and generate initial hash code space for each feature set to form multiple staged hash representation sets; For the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic clustering matching operation is performed to obtain the target hash center candidate set of the current newly added data; Based on the target hash center candidate set and the historical hash code drift trajectory model, an adaptive calibration hash code for the current newly added data is generated; The adaptive calibration hash code of the newly added data is written into the vector database to build an accelerated index structure, and similar data is retrieved based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data.

2. A semantic hash accelerated vector database retrieval method according to claim 1, characterized in that: A plurality of historical semantic vector samples are obtained, and based on the original data time period to which each semantic vector sample belongs, they are divided into a plurality of time partition groups, specifically: According to the timestamp of each historical semantic vector sample corresponding to the original data generation, a fixed time interval length is set; According to the length of the fixed time interval, all historical semantic vector samples are divided into a plurality of continuous and non-overlapping time partition groups in sequence; the number of historical semantic vector samples in each time partition group is greater than a preset number threshold.

3. A semantic hash accelerated vector database retrieval method according to claim 2, characterized in that: The embedding model fitting operation is performed on each time partition group, the corresponding semantic embedding feature set is extracted, and the initial hash code space is generated for each feature set to form multiple staged hash representation sets, specifically: Input the pre-trained semantic embedding model into all historical semantic vector samples in each time partition group, and perform forward reasoning calculation to obtain the high-dimensional semantic embedding feature vector corresponding to each historical semantic vector sample; All high-dimensional semantic embedding feature vectors obtained from each time partition group are sequentially subjected to hash function mapping operations to obtain the corresponding binary initial hash code representation results; All binary initial hash code representation results obtained for each time partition group are stored separately to form a phased hash representation set corresponding to each time partition group.

4. The vector database retrieval method accelerated by semantic hashing according to claim 3 is characterized in that: For the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic clustering matching operation is performed to obtain the target hash center candidate set of the current newly added data, specifically: Determine the latest time partition group to which the newly added data belongs based on the timestamp generated by the original data; In the phased hash representation set corresponding to the latest time partition group, by calculating the cosine similarity between the semantic embedding features of the current newly added data and the historical semantic vector samples, the hash codes corresponding to the historical semantic vector samples whose cosine similarity is greater than the similarity threshold are selected as cluster matching candidates; The hash codes in the cluster matching candidates are clustered according to the clustering algorithm to determine the target hash center candidate set corresponding to the current newly added data.

5. A semantic hash accelerated vector database retrieval method according to claim 4, characterized in that: Based on the target hash center candidate set and the historical hash code drift trajectory model, an adaptive calibration hash code for the current newly added data is generated. Specifically: According to the temporal change law of the hash center in each historical time partition group, a historical hash code drift trajectory model is constructed; The candidate hash code with the smallest distance is determined as the adaptive calibration hash code for the current newly added data through the Euclidean distance between each candidate hash code in the target hash center candidate set and the hash code at the current moment predicted by the historical hash code drift trajectory model.

6. A semantic hash accelerated vector database retrieval method according to claim 5, characterized in that: The adaptive calibration hash code of the newly added data is written into the vector database to build an accelerated index structure, and similar data is retrieved based on the adaptive calibration hash code to achieve semantic consistency backtracking query of historical data, specifically: In the vector database, the adaptive calibration hash code of the current newly added data is used as the index primary key, and a mapping relationship between the adaptive calibration hash code and the original semantic vector features of the current newly added data is established; When a retrieval request arrives, an adaptive calibrated hash code corresponding to the retrieval request is generated, and the corresponding hash bucket is located in the vector database using the hash index structure, and the historical semantic vector data matching the adaptive calibrated hash code of the retrieval request is returned to complete the rapid retrieval of similar data.

7. The vector database retrieval method accelerated by semantic hashing according to claim 6 is characterized in that: When a search request arrives, an adaptive calibrated hash code corresponding to the search request is generated, and the corresponding hash bucket is located in the vector database using the hash index structure, and the historical semantic vector data matching the adaptive calibrated hash code of the search request is returned to complete the rapid search of similar data, specifically: Input the search request into the pre-trained semantic embedding model to generate the corresponding high-dimensional semantic embedding feature vector; The high-dimensional semantic embedding feature vector representation corresponding to the retrieval request is input into the historical hash code drift trajectory model, and the drift vector corresponding to the retrieval request time is calculated; Performing vector superposition operation on the drift vector and the high-dimensional semantic embedding feature vector corresponding to the retrieval request to generate an adaptive calibration hash code corresponding to the retrieval request; The adaptive calibration hash code is queried in the hash index structure in the vector database to locate the hash bucket with the same hash code, and the historical semantic vector data matching the adaptive calibration hash code corresponding to the retrieval request is extracted from the hash bucket to complete the rapid retrieval of similar data.

8. A semantic hash accelerated vector database retrieval system, used to implement a semantic hash accelerated vector database retrieval method according to any one of claims 1 to 7, characterized in that: It includes time partitioning module, embedding model fitting module, dynamic clustering matching module, hash code calibration module and index building and retrieval module; Time partition division module: obtains multiple historical semantic vector samples and divides them into multiple time partition groups based on the original data time period to which each semantic vector sample belongs; Embedding model fitting module: performs embedding model fitting operations on each time partition group, extracts the corresponding semantic embedding feature set, and generates an initial hash code space for each feature set to form multiple staged hash representation sets; Dynamic cluster matching module: For the current newly added data, according to the time period to which the current newly added data belongs and the mapping position of the corresponding semantic similarity vector in each stage hash representation set, a dynamic cluster matching operation is performed to obtain the target hash center candidate set of the current newly added data; Hash code calibration module: Based on the target hash center candidate set and the historical hash code drift trajectory model, it generates an adaptive calibration hash code for the current newly added data; Index construction and retrieval module: write the adaptive calibrated hash code of the current newly added data into the vector database to build an accelerated index structure, and retrieve similar data based on the adaptive calibrated hash code to achieve semantic consistency backtracking query of historical data.

Citation Information

Patent Citations

  • Distributed index method based on LSH (Locality Sensitive Hashing)

    CN103744934A

  • Rapid cross-modal retrieval method and system for incremental data carrying new categories

    CN113326289A

  • Dynamic cross-modal hash retrieval method and system based on concept vector learning

    CN119513371A

  • Block chain index storage method and apparatus, computer device and medium

    WO2022143540A1