A data retrieval method and apparatus
By constructing topology graph index and sorting similarity, the problems of large amount of calculation, high memory usage and low accuracy in vectorized data retrieval are solved, and more efficient and accurate data retrieval is achieved, and applicable scenarios are expanded.
Patent Information
- Application Number
- CN202110291851.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-03-18
AI Technical Summary
The existing vectorized data retrieval methods have large calculation volume, large memory occupancy, low retrieval accuracy, low retrieval efficiency and few applicable scenarios.
By constructing a topology graph index, the similarity between the included angle threshold, quantity threshold and the similarity between the eigenvector nodes is determined, and the similarity sort is performed to determine the search results.
It reduces the amount of computing, reduces the memory footprint, improves the retrieval accuracy and efficiency, and expands the retrieval applicable scenarios.
Smart Images

Figure CN113076447B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a data retrieval method and apparatus. Background Art
[0002] Vectorized data retrieval can be divided into exact retrieval and approximate retrieval. The essence of exact retrieval is linear search, which traverses all vectors in the entire vector space to find the vector information that best matches the target vector; approximate retrieval converts the search that originally needs to be performed in the entire high-dimensional vector space into a search in a small range space or a relatively low-dimensional space through methods such as clustering, dimensionality reduction, or encoding.
[0003] There are at least the following problems in the prior art:
[0004] In the existing vectorized data retrieval methods, there are technical problems such as large computational complexity, large memory occupancy, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a data retrieval method and apparatus, which can reduce the computational complexity, reduce the memory occupancy space, improve the retrieval accuracy, improve the retrieval efficiency, and expand the applicable scenarios of the retrieval.
[0006] To achieve the above object, according to the first aspect of the embodiments of the present invention, a data retrieval method is provided, including:
[0007] Receiving retrieval data, performing feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data;
[0008] Determining a navigation node from the topological graph index, querying the topological graph index starting from the navigation node, and determining a set of target nodes according to the similarity between the query nodes queried from the navigation node and the retrieval feature vector; wherein, the topological graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes;
[0009] Sorting the similarities between each target node in the set of target nodes and the retrieval feature vector, and determining a retrieval result according to the sorting result.
[0010] Further, the step of constructing a topological graph index according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes includes:
[0011] Obtaining a plurality of stored data, performing feature extraction on the plurality of stored data to obtain a plurality of feature vector nodes corresponding to the plurality of stored data, and determining the similarity between the plurality of feature vector nodes;
[0012] Determine candidate nodes corresponding to each eigenvector node according to the included angle threshold, the second quantity threshold, and the similarity between multiple eigenvector nodes; and determine the connection graph between each eigenvector node and its corresponding candidate node; wherein, the included angle threshold indicates the threshold corresponding to the included angle formed by the connection lines between each eigenvector node and any two eigenvector nodes respectively;
[0013] Construct a topological graph index according to the connection graph corresponding to each eigenvector node.
[0014] Further, determining candidate nodes corresponding to each eigenvector node according to the included angle threshold, the second quantity threshold, and the similarity between multiple eigenvector nodes further includes:
[0015] Determine multi-level adjacent nodes corresponding to each eigenvector node according to the similarity between multiple eigenvector nodes;
[0016] Determine candidate nodes corresponding to each eigenvector node from the multi-level adjacent nodes according to the included angle threshold and the second quantity threshold.
[0017] Further, determining the similarity between multiple eigenvector nodes further includes:
[0018] Perform clustering processing on multiple eigenvector nodes, and determine the similarity between multiple eigenvector nodes according to the clustering processing result.
[0019] Further, receive node operation information, and determine the node operation type corresponding to the node operation information;
[0020] If the node operation type is adding a node, determine the connection graph between the added node and its corresponding candidate node, and update the topological graph index;
[0021] If the node operation type is deleting a node, update the deleted node in the topological graph index to a pseudo-node; wherein, the pseudo-node is not placed in the target node set;
[0022] If the node operation type is modifying a node, set the node before modification as a deleted node, set the node after modification as a newly added node, and update the topological graph index.
[0023] Further, query the topological graph index starting from the navigation node, and determine the target node set according to the first quantity threshold and the similarity between the query node queried by the navigation node and the retrieval eigenvector, further including:
[0024] Construct a target node set;
[0025] Starting from the navigation node, query the topological graph index according to the greedy search algorithm, and place the queried nodes in the target node set; update the target nodes in the target node set according to the first quantity threshold and the similarity between the queried nodes and the retrieval feature vector until the sum of the similarities between each target node in the target node set and the retrieval feature vector no longer increases.
[0026] Furthermore, the number of navigation nodes is at least one, and at least one navigation node connects all the feature vector nodes in the topological graph index.
[0027] According to the second aspect of the embodiments of the present invention, there is provided a data retrieval device, including:
[0028] A receiving module, configured to receive retrieval data, perform feature processing on the retrieval data, and obtain a retrieval feature vector corresponding to the retrieval data;
[0029] A target node set determination module, configured to determine a navigation node from the topological graph index, query the topological graph index starting from the navigation node, and determine a target node set according to the first quantity threshold and the similarity between the queried nodes obtained by the navigation node query and the retrieval feature vector; wherein, the topological graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes.
[0030] A sorting module, configured to sort the similarities between each target node in the target node set and the retrieval feature vector, and determine a retrieval result according to the sorting result.
[0031] According to the third aspect of the embodiments of the present invention, there is provided an electronic device, including:
[0032] One or more processors;
[0033] A storage device, configured to store one or more programs,
[0034] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above data retrieval methods.
[0035] According to the fourth aspect of the embodiments of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, it implements any of the above data retrieval methods.
[0036] One embodiment of the above invention has the following advantages or beneficial effects: By receiving retrieval data, performing feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data; determining a navigation node from a topological graph index, querying the topological graph index starting from the navigation node, and determining a target node set according to the similarity between the query nodes queried based on the first quantity threshold and the navigation node and the retrieval feature vector; where the topological graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes; sorting the similarity between each target node in the target node set and the retrieval feature vector, and determining the retrieval result according to the sorting result, the technical problems existing in the existing vectorized data retrieval methods, such as large computational complexity, large memory occupation, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios, are overcome, and further the technical effects of reducing the computational complexity, reducing the memory occupation space, improving the retrieval accuracy, improving the retrieval efficiency, and expanding the applicable scenarios of the retrieval are achieved.
[0037] The further effects of the above non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0039] Figure 1 is a schematic diagram of the main process of the data retrieval method provided by the first embodiment of the present invention;
[0040] Figure 2a is a schematic diagram of the main process of the data retrieval method provided by the second embodiment of the present invention;
[0041] Figure 2b is Figure 2a a schematic diagram of the main process of updating the topological graph index in the method shown;
[0042] Figure 2c is Figure 2a a schematic diagram of the process of performing data retrieval in the method shown;
[0043] Figure 3 is a schematic diagram of the main modules of the data retrieval device provided by the embodiment of the present invention;
[0044] Figure 4 is an exemplary system architecture diagram to which the embodiment of the present invention can be applied;
[0045] Figure 5 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0047] Figure 1 is a schematic diagram of the main process of the data retrieval method provided according to the first embodiment of the present invention; as Figure 1 shown, the data retrieval method provided by the embodiments of the present invention mainly includes:
[0048] Step S101: Receive retrieval data, perform feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data.
[0049] Specifically, the above-mentioned retrieval data is user input data, and the search device determines retrieval result data that is the same as or similar to the input data according to the input data.
[0050] In the process of performing feature processing on the retrieval data to obtain the corresponding feature vector, any one of the existing feature processing methods can be used. Through the above settings, it helps to determine the retrieval result according to the similarity between the retrieval feature vector corresponding to the retrieval data and the feature vector corresponding to the stored data, improving the retrieval efficiency and accuracy.
[0051] Step S102: Determine a navigation node from the topological graph index, query the topological graph index starting from the navigation node, and determine a target node set according to the first quantity threshold and the similarity between the query nodes queried from the navigation node and the retrieval feature vector; wherein, the topological graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes.
[0052] Specifically, according to the embodiments of the present invention, the similarity between any two feature vectors can be measured by existing methods such as cosine angle, inner product, Hamming distance, and Euclidean distance.
[0053] Through the above settings, according to the retrieval feature vector corresponding to the retrieval data, query the topological graph index, and then determine the target node according to the number of target nodes (the first quantity threshold) and the similarity, overcoming the technical problems of large computational complexity, large memory occupation, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios in the existing vectorized data retrieval methods, and thus achieving the reduction of computational complexity, the reduction of memory occupation space, the improvement of retrieval accuracy, the improvement of retrieval efficiency, and the expansion of retrieval applicable scenarios.
[0054] Specifically, according to an embodiment of the present invention, the step of constructing a topological graph index based on the angle threshold, the second quantity threshold, and the similarity between multiple feature vector nodes includes:
[0055] Obtain multiple stored data, perform feature extraction on the multiple stored data to obtain multiple feature vector nodes corresponding to the multiple stored data, and determine the similarity between the multiple feature vector nodes;
[0056] According to the angle threshold, the second quantity threshold, and the similarity between the multiple feature vector nodes, determine the candidate nodes corresponding to each feature vector node; and determine the connection graph between each feature vector node and the corresponding candidate node; wherein, the angle threshold indicates the threshold corresponding to the angle formed by the connection lines between each feature vector node and any two feature vector nodes respectively;
[0057] Construct a topological graph index according to the connection graph corresponding to each feature vector node.
[0058] With the above settings, a topological graph index is constructed based on the stored data in the database; first, according to the angle threshold, the second quantity threshold, and the similarity between the multiple feature vector nodes; the candidate nodes indicate the nodes with higher similarity to the feature vector node. In the process of determining the candidate nodes of each feature vector node, the second quantity threshold and the angle threshold are used to screen the nodes with higher similarity corresponding to each feature vector node, thereby reducing the number of nodes in the subsequent topological graph index and helping to improve the subsequent retrieval efficiency.
[0059] Furthermore, according to an embodiment of the present invention, the step of determining the candidate nodes corresponding to each feature vector node according to the angle threshold, the second quantity threshold, and the similarity between the multiple feature vector nodes further includes:
[0060] Determine the multi-level adjacent nodes corresponding to each feature vector node according to the similarity between the multiple feature vector nodes;
[0061] Determine the candidate nodes corresponding to each feature vector node from the multi-level adjacent nodes according to the angle threshold and the second quantity threshold.
[0062] According to an embodiment of the present invention, the candidate nodes are determined based on the multi-level adjacent nodes. The idea of selecting the multi-level adjacent nodes is: first determine multiple first adjacent nodes (i.e., the nodes with higher similarity to the feature vector node) of a certain feature vector node; then respectively determine the second adjacent nodes corresponding to the multiple first adjacent nodes; until the multi-level adjacent nodes are determined. With the above settings, the coverage range of the candidate nodes is ensured, which helps to improve the retrieval efficiency.
[0063] Exemplarily, according to an embodiment of the present invention, the above determining the similarity between multiple feature vector nodes further includes:
[0064] Performing clustering processing on multiple feature vector nodes, and determining the similarity between multiple feature vector nodes according to the clustering result.
[0065] According to a specific implementation manner of an embodiment of the present invention, the k-nearest neighbor method can be used to perform clustering processing on multiple feature vector nodes to determine a KNN graph, and the KNN graph indicates that the distances of each feature vector node can be used to represent the similarity between multiple feature vector nodes.
[0066] Preferably, according to an embodiment of the present invention, the above method further includes:
[0067] Receiving node operation information and determining the corresponding node operation type of the node operation information;
[0068] If the node operation type is adding a node, determining the connection graph between the added node and the corresponding candidate node, and updating the topological graph index;
[0069] If the node operation type is deleting a node, updating the deleted node in the topological graph index to a pseudo-node; wherein, the pseudo-node is not placed in the target node set;
[0070] If the node operation type is modifying a node, setting the node before modification as a deleted node, setting the node after modification as a newly added node, and updating the topological graph index.
[0071] Through the above settings, for the scenarios of adding, modifying, and deleting stored data; only the original topological index graph needs to be updated, and there is no need to reconstruct the index graph, which reduces the large overhead caused by reconstructing the index due to changes in stored data, and the efficiency of obtaining the updated topological graph index is higher.
[0072] Further, according to an embodiment of the present invention, the above querying the topological graph index starting from the navigation node, and determining the target node set according to the first quantity threshold and the similarity between the query node queried by the navigation node and the retrieval feature vector further includes:
[0073] Constructing a target node set;
[0074] Taking the navigation node as the starting point, querying the topological graph index according to the greedy search algorithm, placing the queried query nodes in the target node set; updating the target nodes in the target node set according to the first quantity threshold and the similarity between the queried query nodes and the retrieval feature vector until the sum of the similarities between each target node in the target node set and the retrieval feature vector no longer increases.
[0075] Specifically, the number of target nodes that can be placed in the target node set is limited by a first quantity threshold. When replacing existing target nodes in the target node set (that is, when the number of target nodes placed in the target node set reaches the first quantity threshold, and subsequent newly added target nodes are added, a corresponding number of original target nodes need to be deleted), it is necessary to judge according to the sum of the similarities between each target node in the target node set and the retrieval feature vector. Through the above settings, the calculation amount is reduced, the retrieval accuracy is improved, and the retrieval efficiency is improved.
[0076] Specifically, according to an embodiment of the present invention, the number of navigation nodes is at least one, and at least one navigation node connects all the feature vector nodes in the topological graph index.
[0077] According to a specific implementation manner of an embodiment of the present invention, the navigation nodes can be randomly selected from the topological graph index or fixedly selected. The number can be one or more, as long as at least one selected navigation node can connect all the feature vector nodes in the topological graph index, so as to ensure that the navigation nodes can query all the feature vector nodes of the topological graph index and improve the retrieval efficiency.
[0078] Step S103, sort the similarities between each target node in the target node set and the retrieval feature vector, and determine the retrieval result according to the sorting result.
[0079] Specifically, after sorting the similarities between each target node in the target node set and the retrieval feature vector, determine the stored data corresponding to one or more target nodes with higher similarities as the retrieval result according to the retrieval quantity threshold.
[0080] According to the technical solution of the embodiment of the present invention, by receiving retrieval data, performing feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data; determining navigation nodes from the topological graph index, querying the topological graph index starting from the navigation nodes, and determining a target node set according to the first quantity threshold and the similarity between the query nodes queried by the navigation nodes and the retrieval feature vector; wherein, the topological graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarities between multiple feature vector nodes; sorting the similarities between each target node in the target node set and the retrieval feature vector, and determining the retrieval result according to the sorting result, the technical problems existing in the existing vectorized data retrieval method, such as large calculation amount, large memory occupation, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios, are overcome, and thus the technical effects of reducing the calculation amount, reducing the memory occupation space, improving the retrieval accuracy, improving the retrieval efficiency, and expanding the applicable scenarios of the retrieval are achieved.
[0081] Figure 2a is a schematic diagram of the main process of the data retrieval method provided by the second embodiment of the present invention; asFigure 2a As shown in the figure, the data retrieval method provided by the embodiments of the present invention mainly includes:
[0082] Step S201: Obtain multiple stored data, perform feature extraction on the multiple stored data, and obtain multiple feature vector nodes corresponding to the multiple stored data.
[0083] Specifically, any one of the existing methods can be used to perform feature processing on the mass stored data to determine multiple feature vector nodes, so as to facilitate subsequent vectorized retrieval.
[0084] Step S202: Perform clustering processing on the multiple feature vector nodes, and determine the similarity between the multiple feature vector nodes according to the clustering processing result.
[0085] Specifically, according to a specific implementation manner of the embodiments of the present invention, the k-nearest neighbor method can be used to perform clustering processing on the multiple feature vector nodes to obtain a KNN (k-Nearest Neighbor, a classification algorithm, also known as the proximity algorithm) graph corresponding to the multiple stored data. The KNN graph indicates that the distances of each feature vector node can be used to represent the similarity between the multiple feature vector nodes.
[0086] Step S203: Determine the multi-level adjacent nodes corresponding to each feature vector node according to the similarity between the multiple feature vector nodes; determine the candidate nodes corresponding to each feature vector node from the multi-level adjacent nodes according to the included angle threshold and the second quantity threshold.
[0087] Through the above settings, according to the spatial distribution of the satellites indicated by the SSG (Satellite System Graph), taking the feature vector node A as an example, the feature vector node A is analogized to a ground measurement station, and the nodes adjacent to the feature vector node A are analogized to satellites. Among them, the second quantity threshold is analogized to the constraint condition of the number of satellites, and the included angle threshold is analogized to the constraint condition of the included angle formed by any two satellites and the ground measurement station respectively; the determined candidate nodes are analogized to the observation satellites.
[0088] According to the embodiments of the present invention, the candidate nodes are determined based on the multi-level adjacent nodes. The selection idea of the multi-level adjacent nodes is: first determine multiple first adjacent nodes of a certain feature vector node (that is, the nodes with a higher similarity to the feature vector node); then respectively determine the second adjacent nodes corresponding to the multiple first adjacent nodes; until the multi-level adjacent nodes are determined. Through the above settings, the coverage range of the candidate nodes is ensured, which helps to improve the retrieval efficiency.
[0089] Through the above settings, a topological graph index is constructed according to the stored data in the database. First, according to the included angle threshold, the second quantity threshold, and the similarity between multiple feature vector nodes; the candidate nodes indicate the nodes with a relatively high similarity to the feature vector node. In the process of determining the candidate nodes for each feature vector node, the screening of the nodes with a relatively high similarity corresponding to each feature vector node is realized according to the second quantity threshold and the included angle threshold, thereby reducing the number of nodes in the subsequent topological graph index and helping to improve the subsequent retrieval efficiency.
[0090] According to a specific implementation manner of an embodiment of the present invention, the feature vectors adjacent to the feature vector node A are 500, the second quantity threshold is 200, and the included angle threshold is 30°. Then, among the 500 feature vector nodes adjacent to the feature vector node A, the feature vector nodes that meet the condition that the included angle between any two feature vector nodes and A is greater than or equal to 30° and the total quantity does not exceed 200 are the candidate nodes corresponding to the feature vector node A. Among them, the above values are only examples, and the specific settings can be adjusted according to the actual situation.
[0091] Step S204, determine the connection graph between each feature vector node and the corresponding candidate node; construct a topological graph index according to the connection graph corresponding to each feature vector node.
[0092] Specifically, connect the above candidate nodes to the feature vector node A respectively to obtain a connection graph, and determine a topological graph index according to the connection graph corresponding to each feature vector node.
[0093] Step S205, receive node operation information, determine the node operation type corresponding to the node operation information, and update the topological graph index according to the node operation type.
[0094] According to Figure 2b As shown, the above steps of receiving node operation information, determining the node operation type corresponding to the node operation information, and updating the topological graph index according to the node operation type include the following steps:
[0095] Step S2051, receive node operation information, and determine the node operation type corresponding to the node operation information. If the node operation type is adding a node, execute step S2052; if the node operation type is deleting a node, execute step S2053; if the node operation type is modifying a node, execute step S2054.
[0096] Step S2052, determine the connection graph between the added node and the corresponding candidate node, and update the topological graph index.
[0097] Specifically, determine the candidate nodes corresponding to the newly added node, and then determine the connection graph corresponding to the newly added node and its corresponding candidate nodes. Next, mount the connection graph corresponding to the newly added node in the original topology graph index to obtain the updated topology graph index.
[0098] According to a specific implementation manner of an embodiment of the present invention, the newly added node can be retrieved from the original topology graph index, and N (the second quantity threshold) nodes that are closest (with the highest similarity) to the newly added node are retrieved as candidate nodes. At the same time, if the newly added node can be used as a candidate node corresponding to other nodes, it is also necessary to update the connection graph corresponding to other nodes according to the newly added node, and then update the topology graph index with the updated connection graph.
[0099] Step S2053: Update the deleted node in the topology graph index to a pseudo-node; where the pseudo-node is not placed in the target node set.
[0100] For the node operation type of the deleted node, only need to set the deleted node as a pseudo-node; during the retrieval process, the navigation node can still query the pseudo-node, but regardless of whether the pseudo-node meets the requirements, and it is not placed in the target node set.
[0101] Step S2054: Set the node before modification as the deleted node, set the node after modification as the newly added node, and update the topology graph index.
[0102] Through the above settings, for scenarios of adding, modifying, and deleting stored data, only need to update the original topology index graph, without having to reconstruct the index graph, reducing the large overhead caused by reconstructing the index due to changes in stored data, and the efficiency of obtaining the updated topology graph index is higher.
[0103] Step S206: Receive the retrieval data, perform feature processing on the retrieval data to obtain the retrieval feature vector corresponding to the retrieval data.
[0104] Step S207: Determine the navigation node from the topology graph index and construct the target node set; starting from the navigation node, query the topology graph index according to the greedy search algorithm, and place the queried nodes in the target node set.
[0105] Specifically, according to an embodiment of the present invention, the similarity between any two feature vectors can be measured by existing methods such as cosine angle, inner product, Hamming distance, and Euclidean distance.
[0106] The number of target nodes that can be placed in the target node set is limited by a first quantity threshold. The replacement of existing target nodes in the target node set (that is, when the number of target nodes placed in the target node set reaches the first quantity threshold, and subsequent new target nodes are added, a corresponding number of original target nodes need to be deleted) needs to be determined according to the sum of the similarities between each target node in the target node set and the retrieval feature vector. Through the above settings, the computational amount is reduced, the retrieval accuracy is improved, and the retrieval efficiency is improved.
[0107] As Figure 2c shown, the black dots are navigation nodes set fixedly, the dark gray dots are randomly determined navigation nodes, the light gray dots are candidate nodes placed in the target node set during the query process, the white dots are the finally determined target nodes, and the pentagrams are the nodes corresponding to the retrieved data.
[0108] Starting from the navigation nodes (black dots and dark gray dots), query the topological graph index according to the greedy search algorithm, and place the queried query nodes (light gray dots) in the target node set; update the target nodes in the target node set according to the first quantity threshold and the similarity between the queried query nodes and the retrieval feature vector until the sum of the similarities between each target node in the target node set and the retrieval feature vector no longer increases.
[0109] Step S208: Update the target nodes in the target node set according to the first quantity threshold and the similarity between the queried query nodes and the retrieval feature vector until the sum of the similarities between each target node in the target node set and the retrieval feature vector no longer increases.
[0110] Through the above settings, query the topological graph index according to the retrieval feature vector corresponding to the retrieved data, and then determine the target nodes according to the target node quantity (the first quantity threshold) and the similarity, overcoming the technical problems of large computational amount, large memory occupation, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios in the existing vectorized data retrieval methods. Furthermore, the computational amount is reduced, the memory occupation space is reduced, the retrieval accuracy is improved, the retrieval efficiency is improved, and the applicable scenarios of the retrieval are expanded.
[0111] Step S209: Sort the similarities between each target node in the target node set and the retrieval feature vector, and determine the retrieval result according to the sorting result.
[0112] Specifically, after sorting the similarities between each target node in the target node set and the retrieval feature vector, determine the stored data corresponding to one or more target nodes with higher similarities as the retrieval result according to the retrieval quantity threshold.
[0113] According to the technical solution of the embodiment of the present invention, by receiving retrieval data, performing feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data; determining a navigation node from the topological graph index, querying the topological graph index starting from the navigation node, and determining a target node set according to the first quantity threshold and the similarity between the query nodes queried by the navigation node and the retrieval feature vector; wherein, the topological graph index is constructed according to the included angle threshold, the second quantity threshold, and the similarity between multiple feature vector nodes; sorting the similarities between each target node in the target node set and the retrieval feature vector, and determining the retrieval result according to the sorting result, the technical problems existing in the existing vectorized data retrieval method, such as large computational amount, large memory occupation, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios, are overcome, and thus the technical effects of reducing the computational amount, reducing the memory occupation space, improving the retrieval accuracy, improving the retrieval efficiency, and expanding the applicable scenarios of the retrieval are achieved.
[0114] Figure 3 is a schematic diagram of the main modules of the data retrieval device provided by the embodiment of the present invention; as Figure 3 shown, the data retrieval device 300 provided by the embodiment of the present invention mainly includes:
[0115] A receiving module 301, configured to receive retrieval data, perform feature processing on the retrieval data, and obtain a retrieval feature vector corresponding to the retrieval data.
[0116] Specifically, the above-mentioned retrieval data is user input data, and the search device determines retrieval result data that is the same as or similar to the input data according to the input data.
[0117] In the process of performing the above-mentioned feature processing on the retrieval data to obtain the corresponding feature vector, any one of the existing feature processing methods can be used. Through the above settings, it is helpful to determine the retrieval result according to the similarity between the retrieval feature vector corresponding to the retrieval data and the feature vector corresponding to the stored data, and improve the retrieval efficiency and retrieval accuracy.
[0118] A target node set determination module 302, configured to determine a navigation node from the topological graph index, query the topological graph index starting from the navigation node, and determine a target node set according to the first quantity threshold and the similarity between the query nodes queried by the navigation node and the retrieval feature vector; wherein, the topological graph index is constructed according to the included angle threshold, the second quantity threshold, and the similarity between multiple feature vector nodes.
[0119] Specifically, according to the embodiment of the present invention, the similarity between any two feature vectors can be measured by existing methods such as cosine angle, inner product, Hamming distance, and Euclidean distance.
[0120] With the above settings, according to the retrieval feature vectors corresponding to the retrieval data, the topological graph index is queried, and then the target nodes are determined according to the number of target nodes (the first quantity threshold) and the similarity, overcoming the technical problems of large computational complexity, large memory occupation, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios in the existing vectorized data retrieval methods, thereby achieving the reduction of computational complexity, the reduction of memory occupation space, the improvement of retrieval accuracy, the improvement of retrieval efficiency, and the expansion of retrieval applicable scenarios.
[0121] Specifically, according to an embodiment of the present invention, the above data retrieval device 300 further includes a topological graph index construction module for:
[0122] Obtain a plurality of stored data, perform feature extraction on the plurality of stored data to obtain a plurality of feature vector nodes corresponding to the plurality of stored data, and determine the similarity between the plurality of feature vector nodes;
[0123] According to the included angle threshold, the second quantity threshold, and the similarity between the plurality of feature vector nodes, determine the candidate nodes corresponding to each feature vector node; and determine the connection graph between each feature vector node and the corresponding candidate node; wherein, the included angle threshold indicates the threshold corresponding to the included angle formed by the connection lines between each feature vector node and any two feature vector nodes respectively;
[0124] Construct a topological graph index according to the connection graph corresponding to each feature vector node.
[0125] With the above settings, a topological graph index is constructed according to the stored data in the database; first, according to the included angle threshold, the second quantity threshold, and the similarity between the plurality of feature vector nodes; the candidate nodes indicate the nodes with higher similarity to the feature vector node. In the process of determining the candidate nodes of each feature vector node, the screening of the nodes with higher similarity corresponding to each feature vector node is realized according to the second quantity threshold and the included angle threshold, thereby reducing the number of nodes in the subsequent topological graph index and helping to improve the subsequent retrieval efficiency.
[0126] Further, according to an embodiment of the present invention, the above topological graph index construction module is further used for:
[0127] Determine the multi-level adjacent nodes corresponding to each feature vector node according to the similarity between the plurality of feature vector nodes;
[0128] Determine the candidate nodes corresponding to each feature vector node from the multi-level adjacent nodes according to the included angle threshold and the second quantity threshold.
[0129] According to an embodiment of the present invention, candidate nodes are determined based on multi-level adjacent nodes. The idea of selecting multi-level adjacent nodes is as follows: first, determine multiple first adjacent nodes of a certain feature vector node (i.e., nodes with a relatively high similarity to the feature vector node); then, respectively determine the second adjacent nodes corresponding to the multiple first adjacent nodes; until multi-level adjacent nodes are determined. Through the above settings, the coverage range of candidate nodes is ensured, which helps to improve the retrieval efficiency.
[0130] Exemplarily, according to an embodiment of the present invention, the above topological graph index construction module is further configured to:
[0131] Perform clustering processing on multiple feature vector nodes, and determine the similarity between the multiple feature vector nodes according to the clustering processing result.
[0132] According to a specific implementation manner of an embodiment of the present invention, the k-nearest neighbor method can be used to perform clustering processing on multiple feature vector nodes to determine a KNN graph, and the KNN graph indicates that the distances of the feature vector nodes can be used to represent the similarity between the multiple feature vector nodes.
[0133] Preferably, according to an embodiment of the present invention, the above data retrieval device 300 further includes an update module, which is configured to:
[0134] Receive node operation information, and determine the node operation type corresponding to the node operation information;
[0135] If the node operation type is adding a node, determine the connection graph between the added node and the corresponding candidate nodes, and update the topological graph index;
[0136] If the node operation type is deleting a node, update the deleted node in the topological graph index to a pseudo-node; wherein, the pseudo-node is not placed in the target node set;
[0137] If the node operation type is modifying a node, set the node before modification as a deleted node, set the node after modification as a newly added node, and update the topological graph index.
[0138] Through the above settings, for scenarios of adding, modifying, and deleting stored data; only the original topological index graph needs to be updated, and there is no need to reconstruct the index graph, which reduces the large overhead caused by reconstructing the index due to changes in stored data, and the efficiency of obtaining the updated topological graph index is higher.
[0139] Further, according to an embodiment of the present invention, the above target node set determination module 302 is further configured to:
[0140] Construct a target node set;
[0141] Starting from the navigation node, query the topological graph index according to the greedy search algorithm, and place the queried query nodes in the target node set; update the target nodes in the target node set according to the first quantity threshold and the similarity between the queried query nodes and the retrieval feature vector until the sum of the similarities between each target node in the target node set and the retrieval feature vector no longer increases.
[0142] Specifically, the number of target nodes that can be placed in the target node set is limited by the first quantity threshold. The replacement of existing target nodes in the target node set (that is, when the number of target nodes placed in the target node set reaches the first quantity threshold, when adding subsequent new target nodes, a corresponding number of original target nodes need to be deleted) needs to be judged according to the sum of the similarities between each target node in the target node set and the retrieval feature vector. Through the above settings, the calculation amount is reduced, the retrieval accuracy is improved, and the retrieval efficiency is improved.
[0143] Specifically, according to the embodiment of the present invention, the number of navigation nodes is at least one, and at least one navigation node connects all the feature vector nodes in the topological graph index.
[0144] According to a specific implementation manner of the embodiment of the present invention, the navigation node can be randomly selected from the topological graph index or fixedly selected, and the number can be one or more, as long as at least one selected navigation node can connect all the feature vector nodes in the topological graph index, so as to ensure that the navigation node can query all the feature vector nodes of the topological graph index and improve the retrieval efficiency.
[0145] The sorting module 303 is used to sort the similarities between each target node in the target node set and the retrieval feature vector, and determine the retrieval result according to the sorting result.
[0146] Specifically, after sorting the similarities between each target node in the target node set and the retrieval feature vector, determine the stored data corresponding to one or more target nodes with higher similarities as the retrieval result according to the retrieval quantity threshold.
[0147] According to the technical solution of the embodiment of the present invention, by receiving retrieval data, performing feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data; determining a navigation node from the topological graph index, querying the topological graph index starting from the navigation node, and determining a target node set according to a first quantity threshold and the similarity between the query nodes queried from the navigation node and the retrieval feature vector; wherein, the topological graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes; sorting the similarities between each target node in the target node set and the retrieval feature vector, and determining the retrieval result according to the sorting result, the technical problems existing in the existing vectorized data retrieval method, such as large computational complexity, large memory occupation, low retrieval accuracy, low retrieval efficiency, and few applicable scenarios, are overcome, and thus the technical effects of reducing the computational complexity, reducing the memory occupation space, improving the retrieval accuracy, improving the retrieval efficiency, and expanding the applicable scenarios of the retrieval are achieved.
[0148] Figure 4 FIG. 400 shows an exemplary system architecture to which the data retrieval method or data retrieval device according to the embodiment of the present invention can be applied.
[0149] As Figure 4 shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405 (this architecture is only an example, and the components included in the specific architecture can be adjusted according to the specific situation of the application). The network 404 is used to provide a medium for communication links between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0150] Users can use the terminal devices 401, 402, 403 to interact with the server 405 through the network 404 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 401, 402, 403, such as shopping applications, web browser applications, search applications, data processing tools, data retrieval clients, social platform software, etc. (only examples).
[0151] The terminal devices 401, 402, 403 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0152] The server 405 may be a server that provides various services, such as a server (for example only) that performs (data retrieval / data processing) on the user using the terminal devices 401, 402, and 403. The server may analyze the received data retrieval, etc., and feed back the processing results (such as the target node set, the retrieval result - only an example) to the terminal device.
[0153] It should be noted that the data retrieval method provided in the embodiment of the present invention is generally executed by the server 405 , and accordingly, the data retrieval device is generally disposed in the server 405 .
[0154] It should be understood that Figure 4 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0155] Reference below Figure 5 , which shows a schematic diagram of the structure of a computer system 500 of a terminal device or a server suitable for implementing an embodiment of the present invention. Figure 5 The terminal device or server shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0156] like Figure 5 As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0157] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that a computer program read therefrom is installed into the storage section 508 as needed.
[0158] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 509 and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-described functions defined in the system of the present invention are performed.
[0159] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0161] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a receiving module, a target node set determination module, and a sorting module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the receiving module can also be described as "a module for receiving retrieval data, performing feature processing on the retrieval data, and obtaining a retrieval feature vector corresponding to the retrieval data".
[0162] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: receiving retrieval data, performing feature processing on the retrieval data, and obtaining a retrieval feature vector corresponding to the retrieval data; determining a navigation node from a topology graph index, querying the topology graph index starting from the navigation node, and determining a target node set according to a first quantity threshold and the similarity between the query nodes queried from the navigation node and the retrieval feature vector; wherein the topology graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes; sorting the similarity between each target node in the target node set and the retrieval feature vector, and determining a retrieval result according to the sorting result.
[0163] According to the technical solution of the embodiment of the present invention, by receiving retrieval data, performing feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data; determining a navigation node from the topological graph index, querying the topological graph index starting from the navigation node, and determining a target node set according to the similarity between the query nodes queried according to the first quantity threshold and the navigation node and the retrieval feature vector; wherein, the topological graph index is constructed according to an included angle threshold, a second quantity threshold and the similarity between multiple feature vector nodes; sorting the similarity between each target node in the target node set and the retrieval feature vector, and determining the retrieval result according to the sorting result, the technical problems existing in the existing vectorized data retrieval method, such as large computational complexity, large memory occupation, low retrieval accuracy, low retrieval efficiency and few applicable scenarios, are overcome, and further the technical effects of reducing the computational complexity, reducing the memory occupation space, improving the retrieval accuracy, improving the retrieval efficiency and expanding the applicable scenarios of the retrieval are achieved.
[0164] The above specific embodiments do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data retrieval method, characterized in that, Applied to a server, including: Receiving retrieval data, performing feature processing on the retrieval data to obtain a retrieval feature vector corresponding to the retrieval data; Determining a navigation node from a topological graph index, querying the topological graph index starting from the navigation node, and determining a target node set according to a first quantity threshold and the similarity between the query nodes queried from the navigation node and the retrieval feature vector; wherein, the topological graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes; Among them, determining the similarity between multiple feature vector nodes includes any of the following methods: calculating the cosine angle, inner product, Hamming distance, Euclidean distance, and clustering the multiple feature vector nodes to obtain a KNN graph corresponding to multiple stored data, and using the distances of each feature vector node indicated by the KNN graph to represent the similarity between multiple feature vector nodes; Taking a certain feature vector node as an analogy to a ground measurement station, and taking the nodes adjacent to the feature vector node as analogies to satellites. Among them, the second quantity threshold is analogous to the constraint condition of the number of satellites, and the included angle threshold is analogous to the constraint condition of the included angle formed by any two satellites and the ground measurement station respectively; Sorting the similarity between each target node in the target node set and the retrieval feature vector, and determining the retrieval result according to the sorting result; Feeding back the target node set and the retrieval result to the terminal device.
2. The data retrieval method according to claim 1, wherein The step of constructing the topological graph index according to the included angle threshold, the second quantity threshold, and the similarity between multiple feature vector nodes includes: Obtaining multiple stored data, performing feature extraction on the multiple stored data to obtain multiple feature vector nodes corresponding to the multiple stored data, and determining the similarity between the multiple feature vector nodes; According to the included angle threshold, the second quantity threshold, and the similarity between the multiple feature vector nodes, determining candidate nodes corresponding to each feature vector node; and determining a connection graph between each feature vector node and the corresponding candidate node; wherein, the included angle threshold indicates the threshold corresponding to the included angle formed by the connection lines between each feature vector node and any two feature vector nodes respectively; Constructing the topological graph index according to the connection graph corresponding to each feature vector node.
3. The data retrieval method according to claim 2, wherein The determining of the candidate nodes corresponding to each feature vector node according to the included angle threshold, the second quantity threshold, and the similarity between the multiple feature vector nodes further includes: Determining multi-level adjacent nodes corresponding to each feature vector node according to the similarity between the multiple feature vector nodes; Determining candidate nodes corresponding to each feature vector node from the multi-level adjacent nodes according to the included angle threshold and the second quantity threshold.
4. The data retrieval method according to claim 2, wherein Receiving node operation information, and determining the node operation type corresponding to the node operation information; If the node operation type is adding a node, determining the connection graph between the added node and the corresponding candidate node, and updating the topological graph index; If the node operation type is to delete a node, update the deleted node in the topology graph index to a pseudo-node; wherein, the pseudo-node is not placed in the target node set; If the node operation type is to modify a node, set the node before modification as a deleted node, set the node after modification as a newly added node, and update the topology graph index.
5. The data retrieval method according to claim 1, wherein The querying the topology graph index starting from the navigation node and determining the target node set according to the first quantity threshold and the similarity between the query nodes queried by the navigation node and the retrieval feature vector further includes: Construct a target node set; Starting from the navigation node, query the topology graph index according to the greedy search algorithm, and place the queried query nodes in the target node set; update the target nodes in the target node set according to the first quantity threshold and the similarity between the queried query nodes and the retrieval feature vector until the sum of the similarities between the target nodes in the target node set and the retrieval feature vector no longer increases.
6. The data retrieval method according to claim 1, wherein The number of the navigation nodes is at least one, and the at least one navigation node connects all the feature vector nodes in the topology graph index.
7. A data retrieval device, characterized in that, Applied to a server, it includes: A receiving module, configured to receive retrieval data, perform feature processing on the retrieval data, and obtain a retrieval feature vector corresponding to the retrieval data; A target node set determination module, configured to determine a navigation node from the topology graph index, query the topology graph index starting from the navigation node, and determine a target node set according to the first quantity threshold and the similarity between the query nodes queried by the navigation node and the retrieval feature vector; wherein, the topology graph index is constructed according to an included angle threshold, a second quantity threshold, and the similarity between multiple feature vector nodes; Wherein, determining the similarity between multiple feature vector nodes includes any of the following methods: calculating a cosine angle, an inner product, a Hamming distance, a Euclidean distance, and performing clustering processing on the multiple feature vector nodes to obtain a KNN graph corresponding to multiple stored data, and using the distances of the feature vector nodes indicated by the KNN graph to represent the similarity between the multiple feature vector nodes; Analogize a certain feature vector node to a ground survey point, and analogize the nodes adjacent to the feature vector node to satellites. Among them, the second quantity threshold is analogized to the constraint condition of the number of satellites, and the included angle threshold is analogized to the constraint condition of the included angle formed by any two satellites and the ground survey point respectively; A sorting module, configured to sort the similarities between each target node in the target node set and the retrieval feature vector, and determine a retrieval result according to the sorting result; The device is further configured to feed back the target node set and the retrieval result to a terminal device.
8. An electronic device, characterized in that, It includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method of searching a data base, navigation device and method of generating an index structure
CN102693266A