A data retrieval method, apparatus and device

By grouping the feature vectors into graph indexes and inverted quantization indexes, combining the advantages of both, the problems of large memory usage and low retrieval accuracy in existing data retrieval methods are solved, and more efficient data retrieval is achieved.

CN114356976BActive Publication Date: 2025-07-11BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210028470.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-07-11
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

The existing data retrieval methods cannot meet user needs. The graph indexing algorithm requires large memory, while the inverted quantization indexing algorithm has low retrieval accuracy.

Method used

Compare the feature vector to be queried with each feature vector group, determine the target feature vector grouping to which it belongs, and use the graph index or inverted quantization index to retrieve the matching feature vectors in the corresponding grouping. Combined with the advantages of graph index and inverted quantization index, reduce memory usage and improve retrieval accuracy.

Benefits of technology

Compared with simply using graph indexing algorithm, the memory usage is reduced, and the retrieval accuracy is improved compared with simply using inverted quantized indexing algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114356976B_ABST
    Figure CN114356976B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a data retrieval method, apparatus, and device. The feature vector to be queried is compared with each feature vector group to obtain the target feature vector group to which the feature vector to be queried belongs. Furthermore, when the target feature vector group is a graph index group, the graph index corresponding to the target feature vector group is used to retrieve the feature vector matching the feature vector to be queried in the target feature vector group. When the target feature vector group is an inverted quantization index group, the inverted quantization index corresponding to the target feature vector group is used to retrieve the feature vector matching the feature vector to be queried in the target feature vector group. In this way, compared with simply using the graph index algorithm for retrieval, the method of the present application can reduce the memory occupied by the retrieval algorithm to a certain extent. Compared with simply using the inverted quantization index algorithm for retrieval, the method of the present application can improve the retrieval accuracy to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular to a data retrieval method, apparatus, and device. Background Art

[0002] With the development of machine learning and neural networks, more and more data is stored in a database in the form of vectors. For example, the image features used in face recognition and the voice features used in speech recognition can both be stored in the corresponding vector database in the form of vectors.

[0003] After obtaining a user's data query request, a retrieval algorithm is used to retrieve the required vector data from the massive vector data stored in the vector database. However, the current data retrieval method cannot meet the user's needs. Summary of the Invention

[0004] In view of this, embodiments of this application provide a data retrieval method, apparatus, and device, which reduce the memory occupied by data retrieval to a certain extent and improve the retrieval accuracy.

[0005] To solve the above problems, the technical solutions provided by the embodiments of this application are as follows:

[0006] A data retrieval method, the method including:

[0007] Comparing the feature vector to be queried with each feature vector group to obtain the target feature vector group to which the feature vector to be queried belongs;

[0008] If the target feature vector group is a graph index group, retrieving, through the graph index corresponding to the target feature vector group, the feature vector that matches the feature vector to be queried in the target feature vector group;

[0009] If the target feature vector group is an inverted quantization index group, retrieving, through the inverted quantization index corresponding to the target feature vector group, the feature vector that matches the feature vector to be queried in the target feature vector group.

[0010] A data retrieval apparatus, the apparatus including:

[0011] A first obtaining unit, configured to compare the feature vector to be queried with each feature vector group to obtain the target feature vector group to which the feature vector to be queried belongs;

[0012] A first retrieval unit, configured to, if the target feature vector group is a graph index group, retrieve, through the graph index corresponding to the target feature vector group, the feature vector that matches the feature vector to be queried in the target feature vector group;

[0013] A second retrieval unit, configured to, if the target feature vector group is an inverted quantization index group, retrieve a feature vector matching the to-be-query feature vector in the target feature vector group through the inverted quantization index corresponding to the target feature vector group.

[0014] An electronic device, comprising:

[0015] One or more processors;

[0016] A storage device having stored thereon one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the data retrieval method as described above.

[0018] A computer-readable medium having stored thereon a computer program, wherein when the program is executed by a processor, the data retrieval method as described above is implemented.

[0019] As can be seen, the embodiments of the present application have the following beneficial effects:

[0020] The embodiments of the present application provide a data retrieval method, apparatus and device. The grouping of the massive feature vectors stored in the database includes a graph index group and an inverted quantization index group. When the to-be-query feature vector belongs to the graph index group, the corresponding graph index is used to retrieve a feature vector matching the to-be-query feature vector in the graph index group. When the to-be-query feature vector belongs to the inverted quantization index group, the corresponding inverted quantization index is used to retrieve a feature vector matching the to-be-query feature vector in the inverted quantization index group. Thus, compared with simply using the graph index algorithm for vector retrieval, the method of the embodiments of the present application can, to a certain extent, reduce the memory occupied by the retrieval algorithm. Compared with simply using the inverted quantization index algorithm for vector retrieval, the method of the embodiments of the present application can, to a certain extent, improve the retrieval accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Is a schematic framework diagram of an exemplary application scenario provided by the embodiments of the present application;

[0022] Figure 2 Is a flowchart of a data retrieval method provided by the embodiments of the present application;

[0023] Figure 3 Is a schematic structural diagram of a data retrieval apparatus provided by the embodiments of the present application;

[0024] Figure 4 Is a schematic diagram of the basic structure of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0025] To make the above objects, features, and advantages of the present application more apparent and understandable, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] To facilitate the understanding and explanation of the technical solutions provided by the embodiments of the present application, the background technology of the present application will be described first below.

[0027] With the development of machine learning and neural networks, more and more data is stored in the database in the form of vectors. For example, the image data used in face recognition and the voice data used in voice recognition are both stored in the corresponding vector database in the form of vectors.

[0028] After obtaining the user's data query request, a retrieval algorithm is used to retrieve the required vector data from the massive vector data stored in the vector database. For example, in face recognition, the user's face image is first collected, and the collected face image is processed into a query feature vector to be queried. Then, in the vector database storing the feature vectors corresponding to a large number of face images, the feature vector that matches the query feature vector to be queried is retrieved. It can be understood that the retrieved feature vector that matches the query feature vector is the required vector data. The inventor has found through research that there are currently existing retrieval algorithms, such as the graph index algorithm and the inverted quantization index algorithm, etc. Among them, the graph index algorithm needs to use a graph index for retrieval, and the storage of the graph index requires a large amount of memory. In addition, the inverted quantization index algorithm has the problem of low retrieval accuracy. It can be seen that both the graph index algorithm and the inverted quantization index algorithm cannot meet the needs of users.

[0029] Based on this, the embodiments of the present application provide a data retrieval method, apparatus, and device. First, the feature vector to be queried is compared with each group of feature vectors to obtain the target feature vector group to which the feature vector to be queried belongs. Furthermore, when the target feature vector group is a graph index group, the graph index corresponding to the target feature vector group is used to retrieve the feature vector that matches the feature vector to be queried in the target feature vector group. When the target feature vector group is an inverted quantization index group, the inverted quantization index corresponding to the target feature vector group is used to retrieve the feature vector that matches the feature vector to be queried in the target feature vector group. In the embodiments of the present application, the groups of a large number of feature vectors stored in the database include a graph index group and an inverted quantization index group. When the feature vector to be queried belongs to the graph index group, the corresponding graph index is used to retrieve the feature vector that matches the feature vector to be queried in the graph index group. When the feature vector to be queried belongs to the inverted quantization index group, the corresponding inverted quantization index is used to retrieve the feature vector that matches the feature vector to be queried in the inverted quantization index group. In this way, compared with simply using the graph index algorithm for retrieval, the method of the embodiments of the present application can reduce the memory occupied by the retrieval algorithm to a certain extent. Compared with simply using the inverted quantization index algorithm for retrieval, the method of the embodiments of the present application can improve the retrieval accuracy to a certain extent.

[0030] To facilitate understanding of the data retrieval method provided by the embodiments of the present application, the following Figure 1 is described in conjunction with the Figure 1 scenario example shown. Refer to

[0031] shown. This figure is a schematic framework diagram of an exemplary application scenario provided by the embodiments of the present application.

[0032] In practical applications, the groups of a large number of feature vectors stored in the database include a graph index group and an inverted quantization index group. Based on this, the feature vector to be queried is compared with each group of feature vectors to obtain the target feature vector group to which the feature vector to be queried belongs. In one or more embodiments, the feature vector to be queried is an image feature vector to be queried, a face feature vector to be queried, or an information feature vector to be queried. Among them, the information feature vector to be queried is, for example, a commodity information feature vector to be queried, an advertisement information feature vector to be queried, a multimedia information (such as audio and video) feature vector to be queried, or a user information feature vector to be queried, etc.

[0032] If the target feature vector group is a graph index group, the graph index algorithm is used to retrieve the feature vector that matches the feature vector to be queried. Specifically, the graph index corresponding to the target feature vector group is used to retrieve the feature vector that matches the feature vector to be queried in the target feature vector group.

[0033] If the target feature vector group is an inverted quantization index group, the inverted quantization index algorithm is used to retrieve the feature vectors that match the query feature vector. Specifically, the inverted quantization index corresponding to the target feature vector group is used to retrieve the feature vectors that match the query feature vector in the target feature vector group.

[0034] Those skilled in the art can understand that Figure 1 The illustrated schematic diagram of the framework is only an example in which the embodiments of the present application can be implemented. The scope of application of the embodiments of the present application is not limited by any aspect of this framework.

[0035] In one or more embodiments, first, a large number of feature vectors in the vector database are grouped to determine the graph index group and the inverted quantization index group. Furthermore, the graph index corresponding to each graph index group and the inverted quantization index corresponding to each inverted quantization index group are established respectively. Based on this, it is possible to determine the target feature vector group to which the query feature vector belongs and perform retrieval using the index corresponding to the target feature vector group.

[0036] Thus, in a possible implementation manner, the embodiments of the present application provide a method for determining a feature vector group, including:

[0037] A1: Clustering the first feature vectors to obtain a plurality of feature vector groups, each feature vector group including at least one first feature vector.

[0038] In a possible implementation manner, the first feature vectors are partial feature vectors among the large number of feature vectors in the vector database, and are used to determine the graph index group and the inverted quantization index group. It can be understood that the method of determining the graph index group and the inverted quantization index group by using partial feature vectors can save time and quickly obtain the grouping result. Specifically, the first feature vectors are first clustered to obtain a plurality of feature vector groups, each feature vector group including at least one first feature vector.

[0039] Among them, in one or more embodiments, the first feature vectors are image query feature vectors, face query feature vectors, or information query feature vectors. Among them, the information query feature vectors are, for example, commodity information query feature vectors, advertising information query feature vectors, multimedia information (such as audio and video) query feature vectors, or user information query feature vectors, etc.

[0040] A2: According to the number of first feature vectors included in the feature vector group, the preset number of feature vector groups with the number of first feature vectors ranked in the front are determined as the graph index group, and the other feature vector groups are determined as the inverted quantization index group.

[0041] Furthermore, based on the obtained feature vector groups, the graph index groups and the inverted quantization index groups are determined. Since the retrieval accuracy of the graph index algorithm is relatively high, for the feature vector groups with a large amount of data, the graph index algorithm can be used for indexing to improve the retrieval accuracy. Based on this, in a possible implementation manner, according to the number of the first feature vectors included in the feature vector groups, the preset number of feature vector groups with the first feature vector numbers ranked in the front are determined as the graph index groups, and the other feature vector groups are determined as the inverted quantization index groups. For example, there are a total of 10 feature vector groups. The feature vector groups are sorted according to the number of the first feature vectors included, and the preset number is set to 4. Then, the first 4 sorted feature vector groups are determined as the graph index groups, and the last 6 feature vector groups are determined as the inverted quantization index groups.

[0042] It can be understood that the preset number in the embodiments of the present application is not limited, and the preset number can be set according to the actual situation.

[0043] After the graph index groups and the inverted quantization index groups are determined, the graph indexes corresponding to each graph index group and the inverted quantization indexes corresponding to each inverted quantization index group are established respectively.

[0044] In a possible implementation manner, the embodiments of the present application provide a specific implementation manner for establishing the graph indexes corresponding to each graph index group and the inverted quantization indexes corresponding to each inverted quantization index group, including:

[0045] B1: Obtain the second feature vector, compare the second feature vector with each feature vector group, and obtain the feature vector group to which the second feature vector belongs.

[0046] In order to quickly obtain the graph index groups and the inverted quantization index groups in the vector database, the first feature vectors used in the above embodiments are partial feature vectors in the vector database. On the basis of determining the groups after clustering of the first feature vectors, the remaining feature vectors in the vector database are obtained, and the remaining feature vectors are determined as the second feature vectors.

[0047] Among them, in one or more embodiments, the second feature vector is an image to-be-query feature vector, a face to-be-query feature vector, or an information to-be-query feature vector. Among them, the information to-be-query feature vector is, for example, a commodity information to-be-query feature vector, an advertisement information to-be-query feature vector, a multimedia information (such as audio and video) to-be-query feature vector, or a user information to-be-query feature vector, etc.

[0048] Further, compare the second feature vector with each feature vector group to determine the feature vector group to which the second feature vector belongs. It can be understood that the feature vector group to which the second feature vector belongs may be a graph index group or an inverted quantization index group. For example, if there are a total of 10 feature vector groups, including 3 graph index groups and 7 inverted quantization index groups. Compare the second feature vector with these 10 feature vector groups to obtain the feature vector group to which the second feature vector belongs.

[0049] B2: Add the second feature vector to the feature vector group to which the second feature vector belongs.

[0050] After determining the feature vector group to which the second feature vector belongs, add the second feature vector to the feature vector group to which the second feature vector belongs. Thus, the grouping of all feature vectors in the vector database is completed.

[0051] For example, when there are 3 graph index groups, namely graph index group A, graph index group B, and graph index group C, and the second feature vector belongs to graph index group A among the 3 graph index groups, add the second feature vector to the belonging graph index group A.

[0052] B3: Use the feature vectors in the target graph index group to establish a graph index corresponding to the target graph index group; the feature vectors in the target graph index group include the first feature vector and the second feature vector in the target graph index group; the target graph index groups are each of the graph index groups.

[0053] After all the feature vector groups in the vector database are completed, determine each group in the graph index group as the target graph index group. For example, when the graph index group includes graph index group A, graph index group B, and graph index group C, regard each of graph index group A, graph index group B, and graph index group C as the target graph index group. It can be understood that the feature vectors in the target graph index group include the first feature vector and the second feature vector in the target graph index group.

[0054] Furthermore, use the feature vectors in the target graph index group to establish a graph index corresponding to the target graph index group. In an optional example, use the graph designed in the Hierarchical Navigable Small World Graph (HNSW) algorithm to establish the graph index corresponding to the target graph index group.

[0055] B4: Establish an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vector and the second feature vector in the target inverted quantization index group; each of the target inverted quantization index groups is each of the inverted quantization index groups.

[0056] After all the feature vector groups in the vector database are grouped, determine each group in the inverted quantization index group as the target inverted quantization index group. For example, when the inverted quantization index group includes figure index group D, figure index group E, figure index group F, figure index group G, figure index group H, figure index group I, and figure index group J, each inverted quantization index group in figure index groups D-J is regarded as the target inverted quantization index group. It can be understood that the feature vectors in the target inverted quantization index group include the first feature vector and the second feature vector in the target inverted quantization index group.

[0057] Furthermore, establish an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group.

[0058] In a possible implementation, use the inverted quantization index designed in the inverted product quantization IVFPQ algorithm to establish an inverted quantization index corresponding to the target inverted quantization index group. Specifically, the embodiments of the present application provide a specific implementation for establishing an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group, including:

[0059] B41: Calculate a first residual vector between the feature vectors in the target inverted quantization index group and the cluster center of the target inverted quantization index group.

[0060] Before calculating the first residual vector, it is necessary to first cluster the feature vectors in the target inverted quantization index group to obtain the feature vectors corresponding to the cluster center. Furthermore, then calculate the first residual vector between the feature vectors in the target inverted quantization index group and the cluster center of the target inverted quantization index group. For example, a feature vector in the target inverted quantization index group is [2.1, 2.4], and the cluster center is [2.0, 2.3], then the first residual vector is [2.1 - 2.0, 2.4 - 2.3] = [0.1, 0.1]. Thus, the first residual vector corresponding to each feature vector in the target inverted quantization index group can be obtained.

[0061] B42: Perform product quantization on the first residual vector to establish an inverted quantization index corresponding to the target inverted quantization index group.

[0062] Further, product quantization is performed on each first residual vector. Specifically, when implemented, each first residual vector is divided into M parts, and each first residual vector consists of M sub-vectors. The sub-vectors at the same position in each first residual vector form a vector group, so that M groups of sub-vectors can be formed. Clustering is performed on each group of sub-vectors to obtain multiple cluster centers for each group of sub-vectors. Furthermore, the multiple cluster centers of each group of sub-vectors are represented by IDs. And each sub-vector in each group of sub-vectors is represented by the ID of the cluster center to which it belongs. For example, when M = 4, the first residual vector is divided into 4 parts. If the ID of the cluster center to which the first sub-vector of the first residual vector belongs is 5, then the first sub-vector of the first residual vector is represented by 5; if the ID of the cluster center to which the second sub-vector of the first residual vector belongs is 7, then the second sub-vector of the first residual vector is represented by 7; if the ID of the cluster center to which the third sub-vector of the first residual vector belongs is 18, then the third sub-vector of the first residual vector is represented by 18; if the ID of the cluster center to which the fourth sub-vector of the first residual vector belongs is 200, then the first sub-vector of the first residual vector is represented by 200. Then the first residual vector after product quantization is represented by 5, 7, 18, 200. It can be seen that by using the product quantization method to process the first residual vector, the storage space can be reduced.

[0063] In addition, compared with the feature vectors in the target inverted quantization index group, the value range of the first residual vector is smaller. Then, compared with directly performing product quantization on the feature vectors in the target inverted quantization index group, since the value range of the first residual vector is smaller, performing product quantization on the first residual vector can reduce the quantization error of subsequent product quantization.

[0064] After performing product quantization on the first residual vector, the first residual vector after product quantization is obtained, and the inverted quantization index corresponding to the target inverted quantization index group is established. Among them, the inverted quantization index corresponding to the target inverted quantization index group is the first residual vector after product quantization. It can be understood that each feature vector in the target inverted quantization index group corresponds to a first residual vector after product quantization.

[0065] From the content of B1 - B4, it can be seen that based on the first feature vector and the second feature vector, first determine the feature vector groups of all feature vectors in the vector database, and then establish the graph indexes corresponding to each graph index group and the inverted quantization indexes corresponding to each inverted quantization index group. Thus, each graph index group and the corresponding graph index can be obtained, and each inverted quantization index group and the corresponding inverted quantization index can also be obtained.

[0066] In another possible implementation manner, the embodiments of the present application also provide another specific implementation manner for establishing the graph indexes corresponding to each graph index group and the inverted quantization indexes corresponding to each inverted quantization index group, including:

[0067] C1: Establish a graph index corresponding to the target graph index group by using the feature vectors in the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group; the target graph index groups are each of the graph index groups.

[0068] After clustering the first feature vectors based on A1 - A2 to obtain the graph index groups, each graph index group is determined as a target graph index group. It can be understood that the feature vectors in this target graph index group include the first feature vectors in the target graph index group. In an optional example, a graph designed in the Hierarchical Navigable Small World Graph (HNSW) algorithm is used to establish the graph index corresponding to the target graph index group.

[0069] C2: Establish an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vectors in the target inverted quantization index group; the target inverted quantization index groups are each of the inverted quantization index groups.

[0070] After clustering the first feature vectors based on A1 - A2 to obtain the inverted quantization index groups, each inverted quantization index group is determined as a target inverted quantization index group. It can be understood that the feature vectors in this target inverted quantization index group include the first feature vectors in the target inverted quantization index group. In an optional example, an inverted quantization index designed in the Inverted File Product Quantization (IVFPQ) algorithm is used to establish the inverted quantization index corresponding to the target inverted quantization index group.

[0071] In one or more embodiments, the specific implementation manner of B41 - B42 can be adopted to implement C2.

[0072] C3: Obtain the second feature vectors, compare the second feature vectors with each feature vector group, and obtain the feature vector group to which the second feature vectors belong.

[0073] Based on A1 - A2 and C1 - C2, the graph index groups, the graph indexes corresponding to the graph index groups, the inverted quantization index groups, and the inverted quantization indexes corresponding to the inverted quantization index groups can be obtained from the first feature vectors. Since the first feature vectors are only part of the feature vectors in the vector database, the remaining feature vectors in the vector database also need to be grouped and the corresponding indexes updated.

[0074] Specifically, first determine the remaining feature vectors in the vector database as the second feature vectors, and then compare the second feature vectors with each group of feature vectors to obtain the group of feature vectors to which the second feature vectors belong. Add the second feature vectors to the group of feature vectors to which they belong. Thus, the grouping of all feature vectors in the vector database is completed.

[0075] C4: If the group of feature vectors to which the second feature vectors belong is the graph index group, insert the second feature vectors into the graph index corresponding to the group of feature vectors to which the second feature vectors belong.

[0076] Further, when the group of feature vectors to which the second feature vectors belong is the graph index group, insert the second feature vectors into the graph index corresponding to the group of feature vectors to which the second feature vectors belong to update the graph index corresponding to the group of feature vectors to which the second feature vectors belong. For example, when the group of feature vectors to which the second feature vectors belong is the graph index group A, first obtain the graph index a corresponding to the graph index group A, and then insert the second feature vectors into the graph index a to obtain the updated graph index a.

[0077] C5: If the group of feature vectors to which the second feature vectors belong is the inverted quantization index group, calculate the second residual vector between the second feature vectors and the clustering center of the group of feature vectors to which the second feature vectors belong.

[0078] When the group of feature vectors to which the second feature vectors belong is the inverted quantization index group, first calculate the second residual vector between the second feature vectors and the clustering center of the group of feature vectors to which the second feature vectors belong.

[0079] C6: Update the inverted quantization index corresponding to the group of feature vectors to which the second feature vectors belong using the second residual vector.

[0080] Furthermore, update the inverted quantization index corresponding to the group of feature vectors to which the second feature vectors belong using the second residual vector. For example, when the group of feature vectors to which the second feature vectors belong is the inverted quantization index group D and the inverted quantization index group D corresponds to the inverted quantization index d, first calculate the second residual vector between the second feature vectors and the clustering center of the inverted quantization index group D. Update the inverted quantization index d using the second residual vector.

[0081] In specific implementation, after obtaining the second residual vector, perform product quantization on the second residual vector to obtain the second residual vector after product quantization. The second residual vector after product quantization is an inverted quantization index newly added based on the inverted quantization index corresponding to the eigenvector group to which the existing second eigenvector belongs. It can be understood that the inverted quantization index corresponding to the eigenvector group to which the existing second eigenvector belongs consists of at least one first residual vector after product quantization. Then, the inverted quantization index corresponding to the eigenvector group to which the second eigenvector belongs can be updated with the second residual vector after product quantization, that is, the second residual vector after product quantization is used as the newly added inverted quantization index to supplement the existing inverted quantization index, and the updated inverted quantization index corresponding to the eigenvector group to which the second eigenvector belongs is obtained. Thus, the inverted quantization index corresponding to the eigenvector group can be obtained, that is, the residual vector after product quantization corresponding to each eigenvector in the eigenvector group.

[0082] As can be seen from the content of C1 - C6, first, through A1 - A2, the graph index group and the inverted quantization index group are obtained, and the graph index corresponding to each graph index group and the inverted quantization index corresponding to the inverted quantization index group are initially determined. Then, the second eigenvectors are grouped, and further, the corresponding indexes are updated using the newly added second eigenvectors in each group to obtain the updated graph index and the inverted quantization index.

[0083] Based on the pre - obtained graph index group, the corresponding graph index, the inverted quantization index group, and the corresponding inverted quantization index, a data retrieval method provided by an embodiment of the present application will be described below with reference to the accompanying drawings. Refer to Figure 2 As shown, this figure is a flowchart of a data retrieval method provided by an embodiment of the present application. As Figure 2 shown, the method may include S201 - S203:

[0084] S201: Compare the feature vector to be queried with each eigenvector group to obtain the target eigenvector group to which the feature vector to be queried belongs.

[0085] After determining the feature vector to be queried, compare the feature vector to be queried with each eigenvector group to obtain the target eigenvector group to which the feature vector to be queried belongs. In an optional example, the similarity between the feature vector to be queried and the cluster center of each eigenvector group can be calculated respectively, and the eigenvector group to which the cluster center with the highest similarity belongs is determined as the target eigenvector group to which the feature vector to be queried belongs.

[0086] In one or more embodiments, the feature vector to be queried is an image feature vector to be queried, a face feature vector to be queried, or an information feature vector to be queried. Among them, the information feature vector to be queried is, for example, a commodity information feature vector to be queried, an advertisement information feature vector to be queried, a multimedia information (such as audio and video) feature vector to be queried, or a user information feature vector to be queried, etc. It can be understood that when the feature vector to be queried is an image feature vector to be queried, each feature vector in the feature vector group is an image feature vector.

[0087] S202: If the target feature vector group is a graph index group, retrieve the feature vector matching the feature vector to be queried in the target feature vector group through the graph index corresponding to the target feature vector group.

[0088] When the target feature vector group is a graph index group, retrieve the feature vector matching the feature vector to be queried through the graph index algorithm. Specifically, retrieve the feature vector matching the feature vector to be queried in the graph index group through the graph index corresponding to the target feature vector group. Among them, the graph index corresponding to the target feature vector group is established in advance by the method provided in the above embodiments.

[0089] It can be understood that, compared with directly using the graph index algorithm to retrieve the feature vector matching the feature vector to be queried among all feature vectors, retrieving the feature vector matching the feature vector to be queried in the graph index group is more efficient.

[0090] S203: If the target feature vector group is an inverted quantization index group, retrieve the feature vector matching the feature vector to be queried in the target feature vector group through the inverted quantization index corresponding to the target feature vector group.

[0091] When the target feature vector group is an inverted quantization index group, retrieve the feature vector matching the feature vector to be queried through the inverted quantization index algorithm. Specifically, retrieve the feature vector matching the feature vector to be queried in the inverted quantization index group through the inverted quantization index corresponding to the target feature vector group. Among them, the inverted quantization index corresponding to the target feature vector group is established in advance by the method provided in the above embodiments.

[0092] In a possible implementation manner, the embodiments of the present application provide a specific implementation manner for retrieving the feature vector matching the feature vector to be queried in the target feature vector group through the inverted quantization index corresponding to the target feature vector group if the target feature vector group is an inverted quantization index group, including:

[0093] D1: Calculate the third residual vector between the feature vector to be queried and the clustering center of the target feature vector group.

[0094] After obtaining the target feature vector group to which the feature vector to be queried belongs, determine the clustering center of the target feature vector group, and calculate the third residual vector between the feature vector to be queried and the clustering center of the target feature vector group.

[0095] D2: Search for the fourth residual vector that matches the third residual vector through the inverted quantization index corresponding to the target feature vector group, and determine the feature vector corresponding to the fourth residual vector as the feature vector that matches the feature vector to be queried.

[0096] In specific implementation, perform product quantization on the third residual vector to obtain the product-quantized third residual vector. Furthermore, calculate the similarity between the product-quantized third residual vector and the inverted quantization index corresponding to the target feature group (i.e., the product-quantized residual vectors corresponding to each feature vector in the target feature group), and determine the residual vector that meets the condition of similarity as the fourth residual vector. If the fourth residual vector matches the third residual vector, then determine the feature vector corresponding to the fourth residual vector as the feature vector that matches the feature vector to be queried.

[0097] It can be understood that, compared with directly using the inverted quantization index algorithm to retrieve the feature vector that matches the feature vector to be queried among all feature vectors, it is more efficient to retrieve the feature vector that matches the feature vector to be queried in the inverted quantization index group.

[0098] Based on the content of S201-S203, the embodiments of the present application provide a data retrieval method. First, compare the feature vector to be queried with each feature vector group to obtain the target feature vector group to which the feature vector to be queried belongs. Furthermore, when the target feature vector group is a graph index group, retrieve the feature vector that matches the feature vector to be queried in the target feature vector group through the graph index corresponding to the target feature vector group. When the target feature vector group is an inverted quantization index group, retrieve the feature vector that matches the feature vector to be queried in the target feature vector group through the inverted quantization index corresponding to the target feature vector group. In the embodiments of the present application, the groups of the massive feature vectors stored in the database include a graph index group and an inverted quantization index group. When the feature vector to be queried belongs to the graph index group, use the corresponding graph index to retrieve the feature vector that matches the feature vector to be queried in the graph index group. When the feature vector to be queried belongs to the inverted quantization index group, use the corresponding inverted quantization index to retrieve the feature vector that matches the feature vector to be queried in the inverted quantization index group. In this way, compared with simply using the graph index algorithm for vector retrieval, the method of the embodiments of the present application can reduce the memory occupied by the retrieval algorithm to a certain extent. Compared with simply using the inverted quantization index algorithm for vector retrieval, the method of the embodiments of the present application can improve the retrieval accuracy to a certain extent.

[0099] Based on a data retrieval method provided in the foregoing method embodiments, an embodiment of the present application further provides a data retrieval device, which will be described below with reference to the accompanying drawings.

[0100] See Figure 3 As shown, this figure is a schematic structural diagram of a data retrieval device provided in an embodiment of the present application. As Figure 3 shown, the data retrieval device includes:

[0101] A first acquisition unit 301, configured to compare a to-be-query feature vector with each feature vector group to obtain a target feature vector group to which the to-be-query feature vector belongs;

[0102] A first retrieval unit 302, configured to, if the target feature vector group is a graph index group, retrieve a feature vector matching the to-be-query feature vector in the target feature vector group through a graph index corresponding to the target feature vector group;

[0103] A second retrieval unit 303, configured to, if the target feature vector group is an inverted quantization index group, retrieve a feature vector matching the to-be-query feature vector in the target feature vector group through an inverted quantization index corresponding to the target feature vector group.

[0104] In a possible implementation manner, the device further includes:

[0105] A clustering unit, configured to cluster first feature vectors to obtain multiple feature vector groups, and each feature vector group includes at least one of the first feature vectors;

[0106] A determination unit, configured to determine, according to the number of first feature vectors included in the feature vector group, a preset number of feature vector groups with the number of first feature vectors ranked in the front as graph index groups, and determine other feature vector groups as inverted quantization index groups.

[0107] In a possible implementation manner, the device further includes:

[0108] A second acquisition unit, configured to acquire a second feature vector, compare the second feature vector with each of the feature vector groups to obtain a feature vector group to which the second feature vector belongs;

[0109] An addition unit, configured to add the second feature vector to the feature vector group to which the second feature vector belongs.

[0110] In a possible implementation manner, the device further includes:

[0111] A first establishment unit, configured to establish a graph index corresponding to the target graph index group by using the feature vectors in the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group and the second feature vectors in the target graph index group; the target graph index groups are respectively each of the graph index groups.

[0112] A second establishment unit, configured to establish an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vectors in the target inverted quantization index group and the second feature vectors in the target inverted quantization index group; the target inverted quantization index groups are respectively each of the inverted quantization index groups.

[0113] In a possible implementation manner, the apparatus further includes:

[0114] A third establishment unit, configured to establish a graph index corresponding to the target graph index group by using the feature vectors in the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group; the target graph index groups are respectively each of the graph index groups.

[0115] A fourth establishment unit, configured to establish an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vectors in the target inverted quantization index group; the target inverted quantization index groups are respectively each of the inverted quantization index groups.

[0116] In a possible implementation manner, the second establishment unit or the fourth establishment unit includes:

[0117] A first calculation subunit, configured to calculate a first residual vector between the feature vectors in the target inverted quantization index group and the cluster center of the target inverted quantization index group.

[0118] An establishment subunit, configured to perform product quantization on the first residual vector to establish an inverted quantization index corresponding to the target inverted quantization index group.

[0119] In a possible implementation manner, the apparatus further includes:

[0120] A third acquisition unit, configured to acquire second feature vectors, compare the second feature vectors with each of the feature vector groups, and obtain the feature vector group to which the second feature vectors belong.

[0121] An insertion unit, configured to insert the second feature vector into a graph index corresponding to the feature vector group to which the second feature vector belongs if the feature vector group to which the second feature vector belongs is a graph index group;

[0122] A calculation unit, configured to calculate a second residual vector between the second feature vector and a cluster center of the feature vector group to which the second feature vector belongs if the feature vector group to which the second feature vector belongs is an inverted quantization index group;

[0123] An update unit, configured to update an inverted quantization index corresponding to the feature vector group to which the second feature vector belongs by using the second residual vector.

[0124] In a possible implementation manner, the second retrieval unit 303 includes:

[0125] A second calculation subunit, configured to calculate a third residual vector between the to-be-query feature vector and a cluster center of the target feature vector group;

[0126] A search subunit, configured to search for a fourth residual vector matching the third residual vector through an inverted quantization index corresponding to the target feature vector group, and determine a feature vector corresponding to the fourth residual vector as a feature vector matching the to-be-query feature vector.

[0127] Based on a data retrieval method provided in the foregoing method embodiment, the present application further provides an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the data retrieval method described in any one of the foregoing embodiments.

[0128] Next, referring to Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present application. The terminal device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable android devices, tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs (televisions) and desktop computers. Figure 4 The shown electronic device is merely an example and should not impose any limitation on the functions and usage scopes of the embodiments of the present application.

[0129] AsFigure 4 As shown in Figure 4 , the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 406 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0130] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 406 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0131] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 409, or installed from the storage device 406, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the method of the embodiment of the present application are executed.

[0132] The electronic device provided by the embodiment of the present application and the data retrieval method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0133] Based on the data retrieval method provided by the above method embodiment, an embodiment of the present application provides a computer-readable medium, on which a computer program is stored, wherein when the program is executed by a processor, the data retrieval method described in any of the above embodiments is implemented.

[0134] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0135] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0136] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0137] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the above data retrieval method.

[0138] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0140] The units involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the unit / module does not, in some cases, constitute a limitation on the unit itself. For example, the voice data acquisition module can also be described as the "data acquisition module".

[0141] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0142] In the context of the present application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0143] According to one or more embodiments of the present application, [Example 1] provides a data retrieval method, the method comprising:

[0144] Comparing the to-be-query feature vector with each feature vector group to obtain the target feature vector group to which the to-be-query feature vector belongs;

[0145] If the target feature vector group is a graph index group, retrieving, through the graph index corresponding to the target feature vector group, the feature vector that matches the to-be-query feature vector in the target feature vector group;

[0146] If the target feature vector group is an inverted quantization index group, retrieving, through the inverted quantization index corresponding to the target feature vector group, the feature vector that matches the to-be-query feature vector in the target feature vector group.

[0147] According to one or more embodiments of the present application, [Example 2] provides a data retrieval method, the method further comprising:

[0148] Clustering the first feature vectors to obtain a plurality of feature vector groups, each of the feature vector groups including at least one of the first feature vectors;

[0149] Determining, according to the number of the first feature vectors included in the feature vector group, the preset number of feature vector groups with the number of the first feature vectors ranked in the front as the graph index groups, and determining the other feature vector groups as the inverted quantization index groups.

[0150] According to one or more embodiments of the present application, [Example 3] provides a data retrieval method, the method further comprising:

[0151] Obtain a second feature vector, compare the second feature vector with each of the feature vector groups, and obtain the feature vector group to which the second feature vector belongs;

[0152] Add the second feature vector to the feature vector group to which the second feature vector belongs.

[0153] According to one or more embodiments of the present application, [Example Four] provides a data retrieval method, and the method further includes:

[0154] Use the feature vectors in the target graph index group to establish a graph index corresponding to the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group and the second feature vectors in the target graph index group; the target graph index groups are each of the graph index groups;

[0155] Use the feature vectors in the target inverted quantization index group to establish an inverted quantization index corresponding to the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vectors in the target inverted quantization index group and the second feature vectors in the target inverted quantization index group; the target inverted quantization index groups are each of the inverted quantization index groups.

[0156] According to one or more embodiments of the present application, [Example Five] provides a data retrieval method, and the method further includes:

[0157] Use the feature vectors in the target graph index group to establish a graph index corresponding to the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group; the target graph index groups are each of the graph index groups;

[0158] Use the feature vectors in the target inverted quantization index group to establish an inverted quantization index corresponding to the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vectors in the target inverted quantization index group; the target inverted quantization index groups are each of the inverted quantization index groups.

[0159] According to one or more embodiments of the present application, [Example Six] provides a data retrieval method, and using the feature vectors in the target inverted quantization index group to establish an inverted quantization index corresponding to the target inverted quantization index group includes:

[0160] Calculate a first residual vector between the feature vectors in the target inverted quantization index group and the cluster center of the target inverted quantization index group;

[0161] Perform product quantization on the first residual vector to establish an inverted quantization index corresponding to the target inverted quantization index group.

[0162] According to one or more embodiments of the present application, [Example Seven] provides a data retrieval method, and the method further includes:

[0163] Obtain a second feature vector, compare the second feature vector with each of the feature vector groups, and obtain the feature vector group to which the second feature vector belongs;

[0164] If the feature vector group to which the second feature vector belongs is a graph index group, insert the second feature vector into the graph index corresponding to the feature vector group to which the second feature vector belongs;

[0165] If the feature vector group to which the second feature vector belongs is an inverted quantization index group, calculate a second residual vector between the second feature vector and the cluster center of the feature vector group to which the second feature vector belongs;

[0166] Use the second residual vector to update the inverted quantization index corresponding to the feature vector group to which the second feature vector belongs.

[0167] According to one or more embodiments of the present application, [Example Eight] provides a data retrieval method. If the target feature vector group is an inverted quantization index group, retrieve a feature vector matching the query feature vector in the target feature vector group through the inverted quantization index corresponding to the target feature vector group, including:

[0168] Calculate a third residual vector between the query feature vector and the cluster center of the target feature vector group;

[0169] Find a fourth residual vector matching the third residual vector through the inverted quantization index corresponding to the target feature vector group, and determine the feature vector corresponding to the fourth residual vector as the feature vector matching the query feature vector.

[0170] According to one or more embodiments of the present application, [Example Nine] provides a data retrieval device, and the device includes:

[0171] A first acquisition unit, configured to compare a query feature vector with each feature vector group to obtain a target feature vector group to which the query feature vector belongs;

[0172] A first retrieval unit, configured to, if the target feature vector group is a graph index group, retrieve, in the target feature vector group, a feature vector that matches the to-be-query feature vector through the graph index corresponding to the target feature vector group;

[0173] A second retrieval unit, configured to, if the target feature vector group is an inverted quantization index group, retrieve, in the target feature vector group, a feature vector that matches the to-be-query feature vector through the inverted quantization index corresponding to the target feature vector group.

[0174] According to one or more embodiments of the present application, [Example Ten] provides a data retrieval device, and the device further includes:

[0175] A clustering unit, configured to cluster first feature vectors to obtain a plurality of feature vector groups, and each of the feature vector groups includes at least one of the first feature vectors;

[0176] A determining unit, configured to determine, according to the number of first feature vectors included in the feature vector group, a preset number of feature vector groups with the number of first feature vectors ranked in the front as graph index groups, and determine other feature vector groups as inverted quantization index groups.

[0177] According to one or more embodiments of the present application, [Example Eleven] provides a data retrieval device, and the device further includes:

[0178] A second obtaining unit, configured to obtain a second feature vector, compare the second feature vector with each of the feature vector groups, and obtain the feature vector group to which the second feature vector belongs;

[0179] An adding unit, configured to add the second feature vector to the feature vector group to which the second feature vector belongs.

[0180] According to one or more embodiments of the present application, [Example Twelve] provides a data retrieval device, and the device further includes:

[0181] A first establishing unit, configured to establish a graph index corresponding to the target graph index group by using the feature vectors in the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group and the second feature vectors in the target graph index group; the target graph index group is each of the graph index groups;

[0182] A second building unit, configured to build an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vector in the target inverted quantization index group and the second feature vector in the target inverted quantization index group; the target inverted quantization index groups are respectively each of the inverted quantization index groups.

[0183] According to one or more embodiments of the present application, [Example XIII] provides a data retrieval device, and the device further includes:

[0184] A third building unit, configured to build a graph index corresponding to the target graph index group by using the feature vectors in the target graph index group; the feature vectors in the target graph index group include the first feature vector in the target graph index group; the target graph index groups are respectively each of the graph index groups;

[0185] A fourth building unit, configured to build an inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vector in the target inverted quantization index group; the target inverted quantization index groups are respectively each of the inverted quantization index groups.

[0186] According to one or more embodiments of the present application, [Example XIV] provides a data retrieval device, and the second building unit or the fourth building unit includes:

[0187] A first calculation subunit, configured to calculate a first residual vector between the feature vectors in the target inverted quantization index group and the cluster center of the target inverted quantization index group;

[0188] A building subunit, configured to perform product quantization on the first residual vector to build an inverted quantization index corresponding to the target inverted quantization index group.

[0189] According to one or more embodiments of the present application, [Example XV] provides a data retrieval device, and the device further includes:

[0190] A third acquisition unit, configured to acquire a second feature vector, compare the second feature vector with each of the feature vector groups, and obtain the feature vector group to which the second feature vector belongs;

[0191] An insertion unit, configured to insert the second feature vector into the graph index corresponding to the feature vector group to which the second feature vector belongs if the feature vector group to which the second feature vector belongs is a graph index group.

[0192] A calculation unit, configured to calculate a second residual vector between the second feature vector and a clustering center of the feature vector group to which the second feature vector belongs if the feature vector group to which the second feature vector belongs is an inverted quantization index group;

[0193] An update unit, configured to update an inverted quantization index corresponding to the feature vector group to which the second feature vector belongs by using the second residual vector.

[0194] According to one or more embodiments of the present application, [Example XVI] provides a data retrieval device, and the second retrieval unit 303 includes:

[0195] A second calculation subunit, configured to calculate a third residual vector between the to-be-query feature vector and a clustering center of the target feature vector group;

[0196] A search subunit, configured to search for a fourth residual vector matching the third residual vector through the inverted quantization index corresponding to the target feature vector group, and determine the feature vector corresponding to the fourth residual vector as the feature vector matching the to-be-query feature vector.

[0197] According to one or more embodiments of the present application, [Example XVII] provides an electronic device, including:

[0198] One or more processors;

[0199] A storage device, on which one or more programs are stored,

[0200] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the data retrieval method as described in any of the above.

[0201] According to one or more embodiments of the present application, [Example XVIII] provides a computer-readable medium, on which a computer program is stored, wherein when the program is executed by a processor, the data retrieval method as described in any of the above is implemented.

[0202] It should be noted that the various embodiments in this specification are described in a progressive manner, with the key points of each embodiment being the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method part.

[0203] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0204] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0205] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0206] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data retrieval method, characterized in that, The method includes: Clustering the first feature vectors to obtain a plurality of feature vector groups, each of the feature vector groups including at least one of the first feature vectors; Determining a preset number of feature vector groups with the first feature vectors sorted in the front in terms of the number of the first feature vectors included in the feature vector groups as graph index groups, and determining the other feature vector groups as inverted quantization index groups; Comparing the to-be-query feature vector with each of the feature vector groups to obtain the target feature vector group to which the to-be-query feature vector belongs; If the target feature vector group is a graph index group, retrieving, in the target feature vector group, the feature vectors matching the to-be-query feature vector through the graph index corresponding to the target feature vector group; If the target feature vector group is an inverted quantization index group, retrieving, in the target feature vector group, the feature vectors matching the to-be-query feature vector through the inverted quantization index corresponding to the target feature vector group.

2. The method according to claim 1, wherein The method further includes: Obtaining a second feature vector, comparing the second feature vector with each of the feature vector groups to obtain the feature vector group to which the second feature vector belongs; Adding the second feature vector to the feature vector group to which the second feature vector belongs.

3. The method according to claim 2, characterized in that, The method further includes: Establishing the graph index corresponding to the target graph index group by using the feature vectors in the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group and the second feature vectors in the target graph index group; the target graph index groups are each of the graph index groups; Establishing the inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vectors in the target inverted quantization index group and the second feature vectors in the target inverted quantization index group; the target inverted quantization index groups are each of the inverted quantization index groups.

4. The method according to claim 1, characterized in that, The method further includes: Establishing the graph index corresponding to the target graph index group by using the feature vectors in the target graph index group; the feature vectors in the target graph index group include the first feature vectors in the target graph index group; the target graph index groups are each of the graph index groups; Establishing the inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group; the feature vectors in the target inverted quantization index group include the first feature vectors in the target inverted quantization index group; the target inverted quantization index groups are each of the inverted quantization index groups.

5. The method according to claim 3 or 4, characterized in that, The establishing the inverted quantization index corresponding to the target inverted quantization index group by using the feature vectors in the target inverted quantization index group includes: Calculating a first residual vector between the feature vectors in the target inverted quantization index group and the clustering center of the target inverted quantization index group; Perform product quantization on the first residual vector to establish an inverted quantization index corresponding to the target inverted quantization index group.

6. The method according to claim 4, wherein The method further includes: Obtain a second feature vector, compare the second feature vector with each of the feature vector groups, and obtain the feature vector group to which the second feature vector belongs; If the feature vector group to which the second feature vector belongs is a graph index group, insert the second feature vector into the graph index corresponding to the feature vector group to which the second feature vector belongs; If the feature vector group to which the second feature vector belongs is an inverted quantization index group, calculate a second residual vector between the second feature vector and the cluster center of the feature vector group to which the second feature vector belongs; Update the inverted quantization index corresponding to the feature vector group to which the second feature vector belongs by using the second residual vector.

7. The method according to claim 5, wherein The step of, if the target feature vector group is an inverted quantization index group, retrieving, in the target feature vector group, a feature vector matching the to-be-query feature vector through the inverted quantization index corresponding to the target feature vector group includes: Calculate a third residual vector between the to-be-query feature vector and the cluster center of the target feature vector group; Find a fourth residual vector matching the third residual vector through the inverted quantization index corresponding to the target feature vector group, and determine the feature vector corresponding to the fourth residual vector as the feature vector matching the to-be-query feature vector.

8. The method according to claim 6, characterized in that, The step of, if the target feature vector group is an inverted quantization index group, retrieving, in the target feature vector group, a feature vector matching the to-be-query feature vector through the inverted quantization index corresponding to the target feature vector group includes: Calculate a third residual vector between the to-be-query feature vector and the cluster center of the target feature vector group; Find a fourth residual vector matching the third residual vector through the inverted quantization index corresponding to the target feature vector group, and determine the feature vector corresponding to the fourth residual vector as the feature vector matching the to-be-query feature vector.

9. A data retrieval device, characterized in that, The apparatus includes: A clustering unit, configured to cluster first feature vectors to obtain a plurality of feature vector groups, where each feature vector group includes at least one of the first feature vectors; A determining unit, configured to determine, according to the number of first feature vectors included in the feature vector group, a preset number of feature vector groups with the number of first feature vectors ranked ahead as graph index groups, and determine other feature vector groups as inverted quantization index groups; A first obtaining unit, configured to compare a to-be-query feature vector with each of the feature vector groups to obtain a target feature vector group to which the to-be-query feature vector belongs; A first retrieving unit, configured to, if the target feature vector group is a graph index group, retrieve, in the target feature vector group, a feature vector matching the to-be-query feature vector through the graph index corresponding to the target feature vector group; A second retrieval unit, configured to retrieve, from the target feature vector group, a feature vector that matches the to-be-query feature vector by using the inverted quantization index corresponding to the target feature vector group if the target feature vector group is an inverted quantization index group.

10. An electronic device, characterized in that, Comprising: One or more processors; A storage device storing one or more programs thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the data retrieval method according to any one of claims 1-8.

11. A computer-readable medium, characterized in that, A computer program is stored thereon, wherein when the program is executed by a processor, the data retrieval method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Spark-oriented remote sensing data indexing method and system and electronic equipment

    CN110083598A

  • Data processing method and device, electronic equipment and storage medium

    CN113609313A