A Large-Scale Vector Retrieval Method and Device

By using clustering and KD trees to build hash bucket allocation methods in large-scale vector search, the data skew problem is solved, efficient similarity metric retrieval and recall tasks are realized, and retrieval efficiency is improved.

CN111625530BActive Publication Date: 2025-07-18BEIJING QIHOOD TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201910147494.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-02-27
Publication Date
2025-07-18
Estimated Expiration
2039-02-27

AI Technical Summary

Technical Problem

Traditional vector search technology has data skew problem in large-scale data processing, which affects retrieval efficiency and is not conducive to distributed recall tasks.

Method used

The hash bucket allocation method based on clustering and KD tree is adopted to divide the vector set to be retrieved into multiple hash buckets with unique identification, and the hash bucket is constructed through a binary tree structure to conduct efficient similarity metric search.

Benefits of technology

It improves the efficiency of large-scale vector retrieval, can effectively process recall tasks of billions of levels, avoids the impact of data tilt on retrieval, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111625530B_ABST
    Figure CN111625530B_ABST
Patent Text Reader

Abstract

The present invention provides a large-scale vector retrieval method and apparatus. The method includes: receiving a query request from a user for text and / or multimedia data, and mapping the query request to a preset multi-dimensional vector space to generate a query vector corresponding to the query request; obtaining a pre-set set of vectors to be retrieved, and evenly partitioning the set of vectors to be retrieved into multiple hash buckets with unique identifiers; and retrieving at least one similar retrieval vector with a similarity metric greater than a specified threshold to the query vector based on the multiple hash buckets. Based on the solution provided by the present invention, not only can the efficiency of vector retrieval be effectively improved, but also a recall task with hundreds of millions of levels can be efficiently and easily implemented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of retrieval, and particularly to a large-scale vector retrieval method and device. Background Art

[0002] As an efficient vector retrieval method that can support billions or even tens of billions, large-scale vector retrieval technology has always been an important technical requirement in the Internet field, especially playing an important role in business scenarios such as search and e-commerce. In traditional vector retrieval technology, bucket strategies such as LSH (Location Sensitive Hash) or PQ (Product quantization) are mostly used. Although it is suitable for high-precision similarity retrieval and duplicate removal tasks, serious data skew is likely to occur in the buckets, which not only affects the retrieval efficiency but also is not conducive to distributed recall tasks. Summary of the Invention

[0003] The present invention provides a large-scale vector retrieval method and device to overcome or at least partially solve the above problems.

[0004] According to one aspect of the present invention, a large-scale vector retrieval method is provided, including:

[0005] Receiving a query request from a user for text and / or multimedia data, and mapping the query request to a preset multi-dimensional vector space to generate a query vector corresponding to the query request;

[0006] Obtaining a pre-set set of vectors to be retrieved, and evenly dividing the set of vectors to be retrieved into multiple hash buckets with unique identifiers;

[0007] Retrieving at least one similar retrieval vector whose similarity metric with the query vector is greater than a specified threshold based on the multiple hash buckets.

[0008] Optionally, the obtaining a pre-set set of vectors to be retrieved and evenly dividing the set of vectors to be retrieved into multiple hash buckets with unique identifiers includes:

[0009] Obtaining a pre-set set of vectors to be retrieved;

[0010] Clustering each vector to be retrieved in the set of vectors to be retrieved to obtain multiple clustering categories, and selecting the clustering centers of each clustering category;

[0011] Evenly dividing the set of vectors to be retrieved into multiple hash buckets with unique identifiers based on the clustering centers of the multiple clustering categories.

[0012] Optionally, the cluster centers based on the multiple cluster categories evenly divide the vector to be retrieved into multiple hash buckets with unique identifiers, including:

[0013] Obtain the central retrieval vector as the center point in any cluster category, and calculate the similarity measure between each vector to be retrieved in the cluster category and the central retrieval vector;

[0014] Construct a KD tree based on the similarity measure between each vector to be retrieved and the central retrieval vector;

[0015] Take each leaf node of the KD tree as a hash bucket, and set a unique identifier for each hash bucket.

[0016] Optionally, the constructing a KD tree based on the similarity measure between each vector to be retrieved and the central retrieval vector includes:

[0017] Set a standard similarity measure, and compare the similarity measure between each vector to be retrieved in the cluster category and the central retrieval vector with the standard similarity measure;

[0018] Construct a binary tree structure according to the comparison result.

[0019] Optionally, the constructing a binary tree structure according to the comparison result includes:

[0020] For any vector to be retrieved in the cluster category, if the similarity measure between the vector to be retrieved and the central retrieval vector is less than the preset standard similarity measure, then classify the vector to be retrieved into the left subtree;

[0021] If the similarity measure between the vector to be retrieved and the central retrieval vector is greater than the preset standard similarity measure, then classify the vector to be retrieved into the right subtree;

[0022] Among them, the number of elements in the left subtree and the right subtree is the same.

[0023] Optionally, the constructing a binary tree structure according to the comparison result further includes:

[0024] For the similarity measure between any vector to be retrieved and the central retrieval vector, if the absolute value of the difference between the similarity measure and the standard similarity measure is less than a preset value, then retain the vector to be retrieved in both the left subtree and the right subtree of the binary tree.

[0025] Optionally, the taking each leaf node of the KD tree as a hash bucket and setting a unique identifier for each hash bucket includes:

[0026] Construct hash buckets based on the vectors to be retrieved included in each leaf node of the binary tree, and set unique identifiers for the corresponding hash buckets according to the positions of the leaf nodes.

[0027] Optionally, after the balanced partitioning of the hash buckets with unique identifiers based on the clustering centers of the multiple clustering categories, the method further includes:

[0028] For any one hash bucket, it is at least balancedly partitioned once based on the retrieval vectors to be retrieved in the hash bucket, resulting in multiple sub-hash buckets.

[0029] Optionally, the retrieving of at least one similar retrieval vector whose similarity metric with the query vector is greater than a specified threshold based on the multiple hash buckets includes:

[0030] Selecting at least one hash bucket from the multiple hash buckets and performing a retrieval in each of the selected hash buckets;

[0031] If the number of selected hash buckets is one, obtaining at least one similar retrieval vector in the hash bucket whose similarity metric with the query vector is greater than the specified threshold;

[0032] If the number of selected hash buckets is multiple, merging the similar retrieval vectors retrieved from each of the hash buckets and reconfirming at least one similar retrieval vector whose similarity metric with the query vector is greater than the specified threshold.

[0033] Optionally, the selecting at least one hash bucket from the multiple hash buckets and performing a retrieval in each of the selected hash buckets includes:

[0034] Selecting at least one hash bucket from the multiple hash buckets and performing a fine-grained sorting calculation within the selected hash bucket.

[0035] According to another aspect of the present invention, there is also provided a large-scale vector retrieval device, including:

[0036] A generating module, configured to receive a query request from a user for text and / or multimedia data, map the query request to a preset multi-dimensional vector space, and generate a query vector corresponding to the query request;

[0037] A partitioning module, configured to obtain a pre-set set of retrieval vectors to be retrieved, and balance partition the set of retrieval vectors to be retrieved into multiple hash buckets with unique identifiers;

[0038] A retrieval module, configured to retrieve at least one similar retrieval vector whose similarity metric with the query vector is greater than a specified threshold based on the multiple hash buckets.

[0039] Optionally, the partitioning module includes:

[0040] An obtaining unit, configured to obtain a pre-set set of retrieval vectors to be retrieved;

[0041] A clustering unit, configured to cluster each to-be-retrieved vector in the to-be-retrieved vector set to obtain multiple clustering categories, and select the clustering centers of each clustering category;

[0042] A balanced partitioning unit, configured to balance partition the to-be-retrieved vector set into multiple hash buckets with unique identifiers based on the clustering centers of the multiple clustering categories.

[0043] Optionally, the balanced partitioning unit is further configured to:

[0044] Obtain a central retrieval vector as the center point in any clustering category, and calculate the similarity measure between each to-be-retrieved vector in the clustering category and the central retrieval vector;

[0045] Construct a KD tree based on the similarity measure between each to-be-retrieved vector and the central retrieval vector;

[0046] Take each leaf node of the KD tree as a hash bucket, and set a unique identifier for each hash bucket.

[0047] Optionally, the balanced partitioning unit is further configured to:

[0048] Set a standard similarity measure, and compare the similarity measure between each to-be-retrieved vector in the clustering category and the central retrieval vector with the standard similarity measure;

[0049] Construct a binary tree structure according to the comparison result.

[0050] Optionally, the balanced partitioning unit is further configured to:

[0051] For any to-be-retrieved vector in the clustering category, if the similarity measure between the to-be-retrieved vector and the central retrieval vector is less than a preset standard similarity measure, then ascribe the to-be-retrieved vector to the left subtree;

[0052] If the similarity measure between the to-be-retrieved vector and the central retrieval vector is greater than the preset standard similarity measure, then ascribe the to-be-retrieved vector to the right subtree;

[0053] Wherein, the number of elements in the left subtree and the right subtree is the same.

[0054] Optionally, the balanced partitioning unit is further configured to:

[0055] For the similarity measure between any to-be-retrieved vector and the central retrieval vector, if the absolute value of the difference between the similarity measure and the standard similarity measure is less than a preset value, then retain the to-be-retrieved vector in both the left subtree and the right subtree of the binary tree.

[0056] Optionally, the balanced partitioning unit is further configured to:

[0057] Construct hash buckets based on the retrieval vectors included in the leaf nodes of the binary tree, and set unique identifiers for the corresponding hash buckets according to the positions of the leaf nodes.

[0058] Optionally, the partitioning module further includes:

[0059] A secondary partitioning unit, configured to perform at least one balanced partitioning on any one hash bucket based on the retrieval vectors in the hash bucket to obtain a plurality of sub - hash buckets.

[0060] Optionally, the retrieval module is further configured to:

[0061] Select at least one hash bucket from the plurality of hash buckets and perform retrieval in each of the selected hash buckets;

[0062] When the number of selected hash buckets is one, obtain at least one similar retrieval vector in the hash bucket whose similarity metric with the query vector is greater than a specified threshold;

[0063] When the number of selected hash buckets is multiple, then merge the similar retrieval vectors retrieved from each hash bucket, and re - confirm at least one similar retrieval vector whose similarity metric with the query vector is greater than the specified threshold.

[0064] Optionally, the retrieval module is further configured to:

[0065] Select at least one hash bucket from the plurality of hash buckets and perform fine - grained sorting calculations within the selected hash buckets.

[0066] According to another aspect of the present invention, there is also provided a computer storage medium storing computer program code, which, when run on a computing device, causes the computing device to execute the large - scale vector retrieval method described in any one of the above.

[0067] According to another aspect of the present invention, there is also provided a computing device, including:

[0068] A processor;

[0069] A memory storing computer program code;

[0070] When the computer program code is run by the processor, it causes the computing device to execute the large - scale vector retrieval method described in any one of the above.

[0071] The present invention provides a more efficient large-scale vector retrieval method and apparatus. In the solution provided by the present invention, after receiving a user's query request, it is first mapped to a multi-dimensional vector space to generate a corresponding query vector, and then at least one similar retrieval vector with a similarity greater than a specified threshold is retrieved from the set of vectors to be retrieved. When performing retrieval in the multi-dimensional vector space according to the present invention, it will first be evenly divided into multiple hash buckets with unique identifiers, and then retrieval is performed in multiple hash buckets. Based on the method provided by the present invention, not only can the efficiency of vector retrieval be effectively improved, but also an efficient and easy-to-implement recall task of hundreds of millions of levels can be achieved.

[0072] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention.

[0073] According to the following detailed description of the specific embodiments of the present invention in conjunction with the drawings, those skilled in the art will be more clear about the above and other purposes, advantages and features of the present invention. Description of the Drawings

[0074] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0075] Figure 1 is a schematic flowchart of a large-scale vector retrieval method according to an embodiment of the present invention;

[0076] Figure 2 is a schematic structural diagram of a large-scale vector retrieval apparatus according to an embodiment of the present invention;

[0077] Figure 3 is a schematic structural diagram of a large-scale vector retrieval apparatus according to a preferred embodiment of the present invention. Detailed Embodiments

[0078] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0079] Figure 1 is a schematic flowchart of a large-scale vector retrieval method according to an embodiment of the present invention, seeFigure 1 It can be seen that the large-scale vector retrieval method provided by the embodiments of the present invention may include:

[0080] Step S102, receiving a query request from a user for text and / or multimedia data, and mapping the query request to a preset multi-dimensional vector space to generate a query vector corresponding to the query request;

[0081] Step S104, obtaining a pre-set set of vectors to be retrieved, and evenly partitioning the set of vectors to be retrieved into multiple hash buckets with unique identifiers;

[0082] Step S106, retrieving at least one similar retrieval vector whose similarity metric with the query vector is greater than a specified threshold based on multiple hash buckets.

[0083] The embodiments of the present invention provide a more efficient large-scale vector retrieval method. After receiving a query request from a user, it first maps it to a multi-dimensional vector space to generate a corresponding query vector, and then queries at least one similar retrieval vector in the set of vectors to be retrieved whose similarity with it is greater than a specified threshold. When performing retrieval in the multi-dimensional vector space, the embodiments of the present invention will first evenly partition it into multiple hash buckets with unique identifiers, and then perform retrieval in multiple hash buckets. Based on the method provided by the embodiments of the present invention, not only can the efficiency of vector retrieval be effectively improved, but also the impact of data skew on retrieval can be avoided.

[0084] Especially in the fields of e-commerce, search, etc. For example, after receiving a query input by a user at the front end, advertisements will be issued based on the query input by the user. Based on this, it is necessary to perform an extended recall task to obtain the user's advertising demand for the query. At this time, the amount of tasks pulled is relatively large, such as hundreds of millions or billions or even more. Based on the method provided by the embodiments of the present invention, while improving the retrieval efficiency, it can efficiently and easily implement a recall task of hundreds of millions of levels, greatly improving the processing efficiency.

[0085] As mentioned in the above step S104, after obtaining the pre-set set of vectors to be retrieved, it will be evenly partitioned into multiple hash buckets with unique identifiers. Further, it may include: obtaining the pre-set set of vectors to be retrieved; clustering each vector to be retrieved in the set of vectors to be retrieved to obtain multiple clustering categories, and selecting the clustering centers of each clustering category; evenly partitioning multiple hash buckets with unique identifiers based on the clustering centers of multiple clustering categories. The set of vectors to be retrieved in the embodiments of the present invention is a comprehensive vector set composed of different types of information such as pictures, videos, audios, and texts generated based on massive data, or multiple sets of vectors to be retrieved corresponding to different types. In the set of vectors to be retrieved, each vector to be retrieved can be a vector generated based on the same multi-dimensional vector space.

[0086] After obtaining the set of vectors to be retrieved, clustering is first performed to obtain multiple clustering categories and the clustering centers of each clustering category. In the embodiments of the present invention, the K-means clustering algorithm preferably used generates multiple clustering categories based on the vectors to be retrieved. k is a hyperparameter calculated by the algorithm, representing the number of classes; Kmeans can automatically assign samples to different classes, and the number of clustering categories can be set according to different requirements, which is not limited in the present invention. After obtaining multiple clustering categories, the set of vectors to be retrieved can be evenly divided into multiple hash buckets with unique identifiers based on the clustering centers of each clustering category. That is to say, for a given set of vectors to be retrieved M, N central retrieval vectors are obtained through K-means, where N << M, for later space partitioning, and the size of each bucket is approximately log(M / 2 N ).

[0087] In the embodiments of the present invention, when dividing the hash buckets based on the clustering centers of each clustering category, the central retrieval vector as the center point in any clustering category can be obtained, and the similarity measure between each vector to be retrieved in this clustering category and the central retrieval vector can be calculated; a KD tree is constructed based on the similarity measure between each vector to be retrieved and the central retrieval vector; each leaf node of the KD tree is used as a hash bucket, and a unique identifier is set for each hash bucket. For any clustering category, the KD tree can be constructed in the manner described in the above embodiments. In the KD tree, KD represents k-dimension, and each node is a k-dimensional point. Each non-leaf node can be imagined as a splitting hyperplane, which divides the space into two parts with a hyperplane perpendicular to the coordinate axis, and thus recursively divides from the root node continuously until there are no instances. The KD tree in the embodiments of the present invention is preferably a binary tree. When constructing the binary tree structure, a standard similarity measure can be first set, and the similarity measure between each vector to be retrieved in this clustering category and the central retrieval vector is compared with the standard similarity measure; the binary tree structure is constructed according to the comparison results.

[0088] For example, given a central retrieval vector A, the similarity measure F between the central retrieval vector A and other vectors to be retrieved in the clustering category to which the central retrieval vector A belongs can be calculated, such as calculating the cosine distance, Euclidean distance, etc., so as to obtain the similarity measures between each vector to be retrieved in this clustering category and the central retrieval vector, and list all the similarity measures in ascending order, and then divide this list equally according to the distance, and the number of each part is the same, and they are respectively the two nodes of the tree. The same method is used for the central retrieval vectors of other clustering categories.

[0089] A binary tree is a tree structure in which each node has at most two subtrees. Usually, the subtrees are called "left subtree" and "right subtree". Binary trees are often used to implement binary search trees and binary heaps. A binary tree with a depth of k and 2 k-1A binary tree with a certain number of nodes is called a full binary tree. The characteristic of this kind of tree is that the number of nodes on each layer is the maximum number of nodes. In a binary tree, except for the last layer, if the remaining layers are all full, and the last layer is either full or lacks a continuous number of nodes on the right, then this binary tree is a complete binary tree.

[0090] In the embodiment of the present invention, when constructing a binary tree structure based on any clustering category, for any retrieval vector to be retrieved in this clustering category, if the similarity measure between the retrieval vector to be retrieved and the central retrieval vector is less than the preset standard similarity measure, then the retrieval vector to be retrieved is assigned to the left subtree; if the similarity measure between the retrieval vector to be retrieved and the central retrieval vector is greater than the preset standard similarity measure, then the retrieval vector to be retrieved is assigned to the right subtree; wherein, the number of elements in the left subtree and the right subtree is the same. The value of the preset standard similarity measure can be set according to different application scenarios, and the present invention does not make a limitation.

[0091] For example, for any clustering category C, its center point is the central retrieval vector v. For any other retrieval vector x to be retrieved in this clustering category C, the similarity between it and the central retrieval vector v is expressed as F(x, v), and a set standard similarity measure F0 is set. When constructing a KD tree, by balancing the split of F, for example, vectors with F < F0 are all assigned to the left subtree, and those greater than F0 are all assigned to the right subtree, and the number of elements in the left and right subtrees remains the same, and continuous splitting is performed to obtain a balanced binary tree structure.

[0092] Optionally, for the similarity measure between any retrieval vector to be retrieved and the central retrieval vector, if the absolute value of the difference between the similarity measure and the standard similarity measure is less than a preset value, then the retrieval vector to be retrieved is retained in both the left subtree and the right subtree of the binary tree. That is to say, if a certain vector is close to the split boundary, that is, |F(x, v) - F0| < the preset value, and the absolute value of the difference of F(x, v) - F0 is less than the preset value, a copy of this vector is retained in both the left subtree and the right subtree.

[0093] Since during retrieval, it may be necessary to perform a preliminary retrieval in tens of millions of hash buckets, which is a retrieval within a bucket. Of course, only retrieving within a bucket may not be enough. If possible, some similar ones may be in other adjacent hash buckets, and then cross-bucket retrieval is required. And based on the method provided by the embodiment of the present invention, it is judged based on the absolute value of the difference from the standard similarity measure and the preset value. When the absolute value of this difference is less than the preset value, the relevant vectors can be retained in both the left and right subtrees. Then, during subsequent retrieval, similar retrieval vectors can be obtained within a single bucket, and cross-bucket retrieval is ultimately transformed into retrieval within a single bucket, which is conducive to parallel computing. The standard similarity measure and the preset value in the embodiment of the present invention can be set according to different requirements, and the present invention does not make a limitation.

[0094] As mentioned above, each leaf node of the KD tree can also be used as a hash bucket, and a unique identifier is set for each hash bucket. That is, hash buckets are constructed based on the retrieval vectors to be retrieved included in each leaf node of the binary tree, and then a unique identifier is set for the corresponding hash bucket according to the position of the leaf node. Optionally, when setting the bucket label for the hash bucket, the chaining method can be used. For example, for a binary tree, it can be set that going left is 1 and going right is 0. In this way, any path from the root node to the leaf node can be expressed in the form of 101010, and then the combination of 0 and 1 is used as the bucket label. Of course, in practical applications, other methods can also be used to set the bucket label for the hash bucket, such as setting it in combination with the type of the hash bucket, the amount of data in the bucket, and the positional relationship, etc., which will not be elaborated here.

[0095] In addition to the above introduction, in a preferred embodiment of the present invention, for any hash bucket, it is at least evenly divided once based on the retrieval vectors in the hash bucket to obtain a plurality of sub-hash buckets, so as to perform a more accurate and detailed retrieval.

[0096] Optionally, after obtaining a plurality of hash buckets, at least one similar retrieval vector whose similarity metric with the query vector is greater than a specified threshold can be retrieved. Further, it may include: selecting at least one hash bucket from the plurality of hash buckets and performing retrieval in each selected hash bucket; if the number of selected hash buckets is one, obtaining at least one similar retrieval vector in the hash bucket whose similarity metric with the query vector is greater than the specified threshold; if the number of selected hash buckets is multiple, merging the similar retrieval vectors retrieved from each of the hash buckets, and reconfirming at least one similar retrieval vector whose similarity metric with the query vector is greater than the specified threshold. Among them, when performing retrieval in a single hash bucket, fine-grained sorting calculation can be performed within the selected hash bucket. The fine-grained sorting uses PQ (Product quantization) to further divide the data in the hash bucket, and the top-k results are obtained through depth search of bit positions.

[0097] That is to say, first, the hash bucket where the retrieval vector with a higher similarity to the query vector is located can be selected from the plurality of hash buckets. The selected hash bucket may be one or multiple. If the selected hash bucket is one, for in-bucket search, the size of each bucket is approximately log(M / 2 N ). As long as the division is sufficient, the size of each bucket is small enough. We can perform fine-grained sorting calculation within the bucket to obtain the top-k for each query vector within each bucket. From the n numbers in arr[1,n], finding the largest k numbers is the top-k. In the embodiment of the present invention, it is to screen out the k detection vectors with the highest similarity to the query vector in the selected hash bucket as the similar retrieval vectors of the query vector.

[0098] If multiple hash buckets are selected, the top-k with the highest similarity to the query vector in each hash bucket can be obtained first. However, the top-k elements obtained from each bucket are merged, and then sorted again to obtain the true top-k elements.

[0099] For example, there are two hash buckets, 10010 and 00100. The vectors to be retrieved in bucket 10010 are a, b, c, d, and the vectors to be retrieved in bucket 00100 are f, e, g. When the query request based on the user is converted into a query vector, the top-2 in bucket 10010 are a and d; the top-2 in bucket 00100 are f and e. Then, a total of four similar retrieval vectors, a, d, f, and e, are retrieved, and then a, d, f, and e are further sorted, and it is found that a and e are the true global top-2.

[0100] Based on the same inventive concept, an embodiment of the present invention further provides a large-scale vector retrieval device, as Figure 2 shown, the large-scale vector retrieval device provided by the embodiment of the present invention may include:

[0101] A generation module 210, configured to receive a query request from a user for text and / or multimedia data, and map the query request to a preset multi-dimensional vector space to generate a query vector corresponding to the query request;

[0102] A partitioning module 220, configured to obtain a pre-set set of vectors to be retrieved, and evenly partition the set of vectors to be retrieved into multiple hash buckets with unique identifiers;

[0103] A retrieval module 230, configured to retrieve at least one similar retrieval vector whose similarity metric to the above query vector is greater than a specified threshold based on multiple hash buckets.

[0104] In a preferred embodiment of the present invention, the partitioning module 220 may include:

[0105] An obtaining unit 221, configured to obtain a pre-set set of vectors to be retrieved;

[0106] A clustering unit 222, configured to cluster each vector to be retrieved in the set of vectors to be retrieved to obtain multiple clustering categories, and select the clustering center of each clustering category;

[0107] An equal partitioning unit 223, configured to evenly partition the set of vectors to be retrieved into multiple hash buckets with unique identifiers based on the clustering centers of multiple clustering categories.

[0108] In a preferred embodiment of the present invention, the equal partitioning unit 223 may further be configured to:

[0109] Obtain the central retrieval vector that is the center point in any clustering category, and calculate the similarity measure between each retrieval vector to be retrieved in the clustering category and the central retrieval vector;

[0110] Construct a KD-tree based on the similarity measure between each retrieval vector to be retrieved and the central retrieval vector;

[0111] Take each leaf node of the KD-tree as a hash bucket, and set a unique identifier for each hash bucket.

[0112] In a preferred embodiment of the present invention, the balance partitioning unit 223 can also be configured to:

[0113] Set a standard similarity measure, and compare the similarity measure between each retrieval vector to be retrieved in the clustering category and the central retrieval vector with the standard similarity measure;

[0114] Construct a binary tree structure according to the comparison result.

[0115] In a preferred embodiment of the present invention, the balance partitioning unit 223 can also be configured to:

[0116] For any retrieval vector to be retrieved in the clustering category, if the similarity measure between the retrieval vector to be retrieved and the central retrieval vector is less than the preset standard similarity measure, then classify the retrieval vector to be retrieved into the left subtree;

[0117] If the similarity measure between the retrieval vector to be retrieved and the central retrieval vector is greater than the preset standard similarity measure, then classify the retrieval vector to be retrieved into the right subtree;

[0118] Wherein, the number of elements in the left subtree and the right subtree is the same.

[0119] In a preferred embodiment of the present invention, the balance partitioning unit 223 can also be configured to:

[0120] For the similarity measure between any retrieval vector to be retrieved and the central retrieval vector, if the absolute value of the difference between the similarity measure and the standard similarity measure is less than a preset value, then retain the retrieval vector to be retrieved in both the left subtree and the right subtree of the binary tree.

[0121] In a preferred embodiment of the present invention, the balance partitioning unit 223 can also be configured to:

[0122] Construct hash buckets based on the retrieval vectors included in each leaf node of the binary tree, and set a unique identifier for the corresponding hash bucket according to the position of the leaf node.

[0123] In a preferred embodiment of the present invention, the partitioning module 220 can also include:

[0124] A secondary partitioning unit 224, configured to perform at least one balance partitioning on the basis of the retrieval vectors in any one hash bucket to obtain a plurality of sub-hash buckets.

[0125] In a preferred embodiment of the present invention, the retrieval module 230 may further be configured to:

[0126] Select at least one hash bucket from multiple hash buckets and perform retrieval in each selected hash bucket;

[0127] When the number of selected hash buckets is one, obtain at least one similar retrieval vector in the hash bucket whose similarity metric with the query vector is greater than a specified threshold;

[0128] When the number of selected hash buckets is multiple, merge the similar retrieval vectors retrieved from each hash bucket and re-confirm at least one similar retrieval vector whose similarity metric with the query vector is greater than a specified threshold.

[0129] In a preferred embodiment of the present invention, the retrieval module 240 may further be configured to:

[0130] Select at least one hash bucket from multiple hash buckets and perform fine-grained sorting calculation within the selected hash buckets.

[0131] Based on the same inventive concept, an embodiment of the present invention further provides a computer storage medium storing computer program code, which, when running on a computing device, causes the computing device to execute the large-scale vector retrieval method described in any of the above embodiments.

[0132] Based on the same inventive concept, an embodiment of the present invention further provides a computing device, including:

[0133] A processor;

[0134] A memory storing computer program code;

[0135] When the computer program code is run by the processor, it causes the computing device to execute the large-scale vector retrieval method described in any of the above embodiments.

[0136] An embodiment of the present invention provides a more efficient large-scale vector retrieval method and apparatus. After receiving a user's query request, it first maps it to a multi-dimensional vector space to generate a corresponding query vector, and then queries at least one similar retrieval vector in the vector set to be retrieved whose similarity with it is greater than a specified threshold. When performing retrieval in the multi-dimensional vector space, an embodiment of the present invention will first evenly divide it into multiple hash buckets with unique identifiers, and then perform retrieval in multiple hash buckets. Based on the method provided by the embodiment of the present invention, it can not only effectively improve the efficiency of vector retrieval, but also avoid the impact of data skew on retrieval.

[0137] In addition, by retaining the vectors close to the segmentation boundary in both the left and right subtrees simultaneously, when performing subsequent retrieval, the acquisition of similar retrieval vectors can be achieved within a single bucket, transforming the cross-bucket retrieval into retrieval within a single bucket, which is conducive to parallel computing.

[0138] Those skilled in the art can clearly understand the specific working processes of the above-described systems, devices, and units. They can refer to the corresponding processes in the foregoing method embodiments. For the sake of brevity, they will not be described in detail here.

[0139] In addition, in each embodiment of the present invention, the functional units can be physically independent of each other, or two or more functional units can be integrated together, or all the functional units can be integrated in a processing unit. The above-mentioned integrated functional units can be implemented in the form of hardware, or in the form of software or firmware.

[0140] Those of ordinary skill in the art can understand that if the integrated functional unit is implemented in the form of software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, which includes several instructions for causing a computing device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention when running the instructions. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks, etc., which can store program codes.

[0141] Alternatively, all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions (such as a computing device such as a personal computer, a server, or a network device). The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the method described in each embodiment of the present invention.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that within the spirit and principles of the present invention, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of the present invention.

Claims

1. A large-scale vector retrieval method, comprising: Receiving a query request from a user for text and / or multimedia data, and mapping the query request to a preset multi-dimensional vector space to generate a query vector corresponding to the query request; Obtaining a preset set of vectors to be retrieved, and evenly partitioning the set of vectors to be retrieved into a plurality of hash buckets with unique identifiers; Retrieving at least one similar retrieval vector from the plurality of hash buckets, the similarity metric of which with the query vector is greater than a specified threshold; Wherein, the step of obtaining a preset set of vectors to be retrieved and evenly partitioning the set of vectors to be retrieved into a plurality of hash buckets with unique identifiers includes: Obtaining a preset set of vectors to be retrieved; Clustering each vector to be retrieved in the set of vectors to be retrieved to obtain a plurality of clustering categories, and selecting the clustering center of each clustering category; Evenly partitioning the set of vectors to be retrieved into a plurality of hash buckets with unique identifiers based on the clustering centers of the plurality of clustering categories.

2. The method according to claim 1, wherein The step of evenly partitioning the set of vectors to be retrieved into a plurality of hash buckets with unique identifiers based on the clustering centers of the plurality of clustering categories includes: Obtaining a central retrieval vector as the center point in any clustering category, and calculating the similarity metric between each vector to be retrieved in the clustering category and the central retrieval vector; Constructing a KD tree based on the similarity metric between each vector to be retrieved and the central retrieval vector; Regarding each leaf node of the KD tree as a hash bucket, and setting a unique identifier for each hash bucket.

3. The method according to claim 2, wherein, The step of constructing a KD tree based on the similarity metric between each vector to be retrieved and the central retrieval vector includes: Setting a standard similarity metric, and comparing the similarity metric between each vector to be retrieved in the clustering category and the central retrieval vector with the standard similarity metric; Constructing a binary tree structure according to the comparison result.

4. The method according to claim 3, wherein The step of constructing a binary tree structure according to the comparison result includes: For any vector to be retrieved in the clustering category, if the similarity metric between the vector to be retrieved and the central retrieval vector is less than a preset standard similarity metric, then classifying the vector to be retrieved into the left subtree; If the similarity metric between the vector to be retrieved and the central retrieval vector is greater than a preset standard similarity metric, then classifying the vector to be retrieved into the right subtree; Wherein, the number of elements in the left subtree and the right subtree is the same.

5. The method according to claim 3, wherein The step of constructing a binary tree structure according to the comparison result further includes: For the similarity metric between any vector to be retrieved and the central retrieval vector, if the absolute value of the difference between the similarity metric and the standard similarity metric is less than a preset value, then retaining the vector to be retrieved in both the left subtree and the right subtree of the binary tree.

6. The method according to claim 3, wherein, The step of regarding each leaf node of the KD tree as a hash bucket and setting a unique identifier for each hash bucket includes: Constructing hash buckets based on the vectors to be retrieved included in each leaf node of the binary tree, and setting a unique identifier for the corresponding hash bucket according to the position of the leaf node.

7. The method according to any one of claims 2-6, wherein, After evenly partitioning a plurality of hash buckets with unique identifiers based on the clustering centers of the plurality of clustering categories, further included is: For any hash bucket, perform at least one balanced partition based on the vectors to be retrieved in the hash bucket to obtain multiple sub - hash buckets.

8. The method according to claim 1, wherein, The retrieving of at least one similar retrieval vector with a similarity metric greater than a specified threshold to the query vector based on the multiple hash buckets includes: Select at least one hash bucket from the multiple hash buckets and perform retrieval in each of the selected hash buckets; If the number of selected hash buckets is one, obtain at least one similar retrieval vector in the hash bucket with a similarity metric greater than the specified threshold to the query vector; If the number of selected hash buckets is multiple, merge the similar retrieval vectors retrieved from each hash bucket and re - confirm at least one similar retrieval vector with a similarity metric greater than the specified threshold to the query vector.

9. The method according to claim 8, wherein The selecting at least one hash bucket from the multiple hash buckets and performing retrieval in each of the selected hash buckets includes: Select at least one hash bucket from the multiple hash buckets and perform fine - grained sorting calculation within the selected hash bucket.

10. A large - scale vector retrieval device, including: A generating module, configured to receive a query request from a user for text and / or multimedia data, map the query request to a preset multi - dimensional vector space, and generate a query vector corresponding to the query request; A partitioning module, configured to obtain a pre - set set of vectors to be retrieved, and balance - partition the set of vectors to be retrieved into multiple hash buckets with unique identifiers; A retrieving module, configured to retrieve at least one similar retrieval vector with a similarity metric greater than a specified threshold to the query vector based on the multiple hash buckets; Wherein, the partitioning module includes: An obtaining unit, configured to obtain a pre - set set of vectors to be retrieved; A clustering unit, configured to cluster each vector to be retrieved in the set of vectors to be retrieved to obtain multiple clustering categories, and select the clustering centers of each clustering category; A balanced - partitioning unit, configured to balance - partition the set of vectors to be retrieved into multiple hash buckets with unique identifiers based on the clustering centers of the multiple clustering categories.

11. The apparatus according to claim 10, wherein, The balanced - partitioning unit is further configured to: Obtain the central retrieval vector as the center point in any clustering category, and calculate the similarity metric between each vector to be retrieved in the clustering category and the central retrieval vector; Construct a KD - tree based on the similarity metric between each vector to be retrieved and the central retrieval vector; Take each leaf node of the KD - tree as a hash bucket, and set a unique identifier for each hash bucket.

12. The device according to claim 11, wherein, The balanced - partitioning unit is further configured to: Set a standard similarity metric, and compare the similarity metric between each vector to be retrieved in the clustering category and the central retrieval vector with the standard similarity metric; Construct a binary - tree structure according to the comparison result.

13. The device according to claim 12, wherein, The balanced - partitioning unit is further configured to: For any vector to be retrieved in the clustering category, if the similarity metric between the vector to be retrieved and the central retrieval vector is less than a preset standard similarity metric, then assign the vector to be retrieved to the left subtree; If the similarity metric between the vector to be retrieved and the central retrieval vector is greater than a preset standard similarity metric, then assign the vector to be retrieved to the right subtree; Among them, the number of elements in the left subtree and the right subtree is the same.

14. The apparatus according to claim 12, wherein, The balanced partitioning unit is further configured to: For the similarity measure between any one of the to-be-retrieved vectors and the central retrieval vector, if the absolute value of the difference between the similarity measure and the standard similarity measure is less than a preset value, then retain the to-be-retrieved vector in both the left subtree and the right subtree of the binary tree.

15. The apparatus according to claim 12, wherein, The balanced partitioning unit is further configured to: Construct hash buckets based on the to-be-retrieved vectors included in each leaf node of the binary tree, and set a unique identifier for the corresponding hash bucket according to the position of the leaf node.

16. The device according to any one of claims 11-15, wherein, The partitioning module further includes: A secondary partitioning unit, configured to, for any one hash bucket, perform at least one balanced partitioning based on the to-be-retrieved vectors in the hash bucket to obtain a plurality of sub-hash buckets.

17. The apparatus according to claim 10, wherein, The retrieval module is further configured to: Select at least one hash bucket from the plurality of hash buckets and perform retrieval in each of the selected hash buckets; When the number of selected hash buckets is one, obtain at least one similar retrieval vector in the hash bucket whose similarity measure with the query vector is greater than a specified threshold; When the number of selected hash buckets is multiple, then merge the similar retrieval vectors retrieved from each hash bucket, and re-confirm at least one similar retrieval vector whose similarity measure with the query vector is greater than the specified threshold.

18. The apparatus according to claim 17, wherein The retrieval module is further configured to: Select at least one hash bucket from the plurality of hash buckets and perform fine-grained sorting calculations within the selected hash buckets.

19. A computer storage medium storing computer program code that, when run on a computing device, causes the computing device to execute the large-scale vector retrieval method according to any one of claims 1-9.

20. A computing device, comprising: A processor; A memory storing computer program code; When the computer program code is run by the processor, it causes the computing device to execute the large-scale vector retrieval method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Multimedia resource acquisition method, apparatus and system

    CN101382959A

  • LSH (Locality Sensitive Hashing)-based clustering and indexing method and LSH-based clustering and indexing system

    CN103631928A

  • Track query method, electronic equipment and storage medium

    CN108536813A