Query method, processor, processing system, storage medium and program product
By compressing and simplifying the query data, and using the index shape diagram to query the target index vector, the problems of low query efficiency and large storage space of high-dimensional vector data are solved, and the effect of saving storage space and improving query efficiency is achieved.
Patent Information
- Application Number
- CN202510555654.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-29
AI Technical Summary
When facing vector data with high-dimensional and large data volume, the existing index query method has low query efficiency, large storage space occupies a lot, and high index update and maintenance overhead, making it difficult to meet real-time processing requirements.
The simplified vector is obtained by compressing the query data corresponding to the query instruction. The target index vector matching the simplified vector is queried from the index shape diagram, and the K index vectors are represented by the index shape diagram. The initial source data corresponding to the target index vector after being compressed is determined, and the target source data is obtained based on the similarity.
Reduced the query vector dimension, reduce the amount of calculation, save storage space, improve query efficiency, narrow the query scope, and improve index construction speed.
Smart Images

Figure CN120067177B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data query technology, and more specifically to a query method, a processor, a processing system, a storage medium, and a program product. Background Art
[0002] Index query is a method of retrieving data through an index structure. The index structure is usually pre-built. When a user initiates a query request, the index structure is used to locate the data that meets the query request from the index library, and then the data that meets the query request is returned to the user.
[0003] In the process of realizing the concept of this application, the inventors found that there are at least the following problems in the related technology: the related index query method usually traverses and matches the data to be queried with the data items in the index library, for example, comparing the data to be queried with each data item stored in the index library one by one until a qualified result is found. However, when the amount of data in the index library is large, this traversal query method will result in low query efficiency and the data items in the index library will occupy a large amount of storage space. Summary of the Invention
[0004] In view of the above problems, the present application provides a query method, a processor, a processing system, a storage medium and a program product.
[0005] According to the first aspect of the present application, a query method is provided, comprising receiving a query instruction; compressing query data corresponding to the query instruction to obtain a simplified vector; obtaining a target index vector matching the simplified vector from an index shape graph; wherein the index shape graph represents K index vectors, and the K index vectors are obtained by compressing S source data in an index library, S and K are positive integers, and S≥K; determining at least one preliminary source data corresponding to the target index vector after compression from the S source data; and obtaining target source data matching the query data based on the similarity between the query data and the at least one preliminary source data.
[0006] A second aspect of the present application provides a processor, comprising: an on-chip computing component for executing a query method; and an on-chip storage unit for storing an index shape graph.
[0007] A third aspect of the present application provides a processing system, including: a first processor, configured to query a target index vector matching the simplified vector from an index shape graph; wherein, the index shape graph represents K index vectors, and the K index vectors are obtained by compressing S source data in an index library, S and K are positive integers, and S≥K; a second processor, configured to compress query data corresponding to a query instruction to obtain a simplified vector, determine at least one primary selected source data corresponding to the target index vector after compression from the S source data, and obtain target source data matching the query data according to the similarity between the query data and the at least one primary selected source data.
[0008] A fourth aspect of the present application further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0009] A fifth aspect of the present application further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0010] According to an embodiment of the present application, by compressing query data corresponding to a query instruction to obtain a simplified vector, the dimension of the query vector can be reduced, the amount of calculation can be reduced, and the query efficiency can be improved. By compressing source data into index vectors, the storage overhead can be reduced and the index construction speed can be improved. By representing index vectors with an index shape graph, querying a target index vector matching the simplified vector from the index shape graph, and then determining at least one primary selected source data corresponding to the target index vector after compression from the S source data, the storage space can be saved, the query range can be narrowed, and the query efficiency can be improved. Thus, at least partially, the technical problems of low query efficiency and large storage overhead in related query methods are solved, and the technical effects of saving storage space and improving query efficiency are achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other objects, features and advantages of the present application will become clearer. In the drawings:
[0012] Figure 1 The application scenario diagrams of the query method, processor, processing system, storage medium and program product according to the embodiments of the present application are shown.
[0013] Figure 2 The flowchart of the query method according to the embodiments of the present application is shown.
[0014] Figure 3 The schematic diagram of constructing an index binary tree based on K index vectors according to the embodiments of the present application is shown.
[0015] Figure 4 A schematic diagram showing various subtree shapes in an index binary tree according to an embodiment of the present application.
[0016] Figure 5 A schematic diagram showing the correspondence between subtree shapes and shape encodings according to an embodiment of the present application.
[0017] Figure 6 A schematic diagram showing the generation of an index shape diagram according to an embodiment of the present application.
[0018] Figure 7 A schematic diagram showing the association relationship between leaf nodes and a source data set according to an embodiment of the present application.
[0019] Figure 8 A schematic diagram showing the generation of an index shape diagram carrying offset information according to an embodiment of the present application.
[0020] Figure 9 A schematic diagram showing the structures of an on-chip computing component and an on-chip storage unit according to an embodiment of the present application.
[0021] Figure 10 A schematic diagram showing the system interaction diagram between a first processor and a second processor according to an embodiment of the present application.
[0022] Figure 11 A flowchart showing a query method according to another embodiment of the present application.
[0023] Figure 12 A schematic diagram showing the system interaction diagram between a first processor and a second processor according to another embodiment of the present application.
[0024] Figure 13 A block diagram showing an electronic device suitable for implementing the query method according to an embodiment of the present application. Detailed implementation manners
[0025] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present application. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present application.
[0026] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted to have a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0028] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0029] Index query is a method of retrieving data through an index structure. Index query methods include, for example, index queries on vector data. For example: Unstructured data such as text, images, voice, video, etc. can be converted into vector data through feature vector extraction processing. Vector data is a mathematical expression form that characterizes an object or data point through a set of ordered numerical values. Vector data can include one-dimensional arrays, and the elements in the array generally appear in numerical form, for example, in the form of floating-point numbers. These numerical values can accurately reflect the position, features, and attributes of the object or data point in a multi-dimensional space. After converting the unstructured data into feature vectors, the feature vectors can be stored in an index library. When a user needs to query certain data, several feature vectors corresponding to the query content can be found from the index library.
[0030] Methods related to constructing an index structure for vector data include tree-based methods, hash-based methods, quantization-based methods, and graph-based methods, etc.
[0031] Specifically, tree-based index methods include, for example, the KD-tree method (K - Dimensional Tree, abbreviated as KD-Tree), which recursively divides data points into different subspaces, continuously selects a dimension, and divides the data points into left and right subtrees according to the numerical values on that dimension, so that the data points in each subtree have a certain order in a certain dimension.
[0032] Hash-based index methods include, for example, Locality Sensitive Hashing (abbreviated as LSH), etc. It converts continuous real values into discrete values through a hash function, and uses the similarity between vectors as the key metric index to perform segmentation processing on vectors, so that each vector is classified into different hash buckets according to the degree of similarity.
[0033] Quantization-based indexing methods include, for example, Product Quantization (PQ for short), Inverted Index Product Quantization (IVFPQ for short), etc. These methods decompose the original vector into smaller sub-vectors and then simplify the representation of each block to narrow the potential range of values.
[0034] Graph-based indexing methods include, for example, Navigation Net (NN-Net for short), Hierarchical Navigable Small World Graphs (HNSW for short), etc. These methods establish navigation relationships between data points and search for data based on these relationships.
[0035] However, for tree-based indexing methods, as the dimension increases, the number of tree nodes will rapidly expand, leading to an increase in the memory overhead of the index. Therefore, it is not suitable for data retrieval of high-dimensional vectors. Hash-based indexing methods rely on the design of hash functions, and hash collisions may occur, causing multiple groups of vectors with low similarity to be mapped to the same hash bucket, resulting in a decrease in retrieval efficiency. Moreover, the selection and calculation of hash functions increase the computational overhead of index construction. The recall rate of quantization-based indexing methods is relatively low. Specifically, the retrieval accuracy of quantization-based indexing methods is low, and some truly similar vectors may not be detected. The construction process of graph-based indexing methods is relatively slow, not suitable for large-scale data sets, and it is difficult to tune parameters and has a high usage difficulty.
[0036] In addition, in the process of implementing the concept of this application, the inventors also found that there are at least the following problems in the related art: Vector data usually has high-dimensional characteristics. For example, the feature vectors extracted by a deep learning model may have hundreds or even thousands of dimensions. And vector data usually has a large amount of data. For example, for an image library containing millions of images, each image needs to be extracted as a feature vector. In addition, vector data usually needs to be processed in real time, that is, queries and calculations need to be completed in a short time. However, the related index query methods usually traverse and match the query data with the vector data in the index library, and each data stored in the index library needs to be compared one by one with the data to be queried. When facing high-dimensional and large-volume vector data, the query efficiency of this traversal query method is relatively low, which is difficult to meet the requirements of real-time processing, and the data items in the index library occupy a large amount of storage space. In addition, the update frequency of vector data is relatively high, and the index structure adapted to vector data is usually quite complex. Whenever the vector data is updated, it is often necessary to reconstruct the index structure, thereby increasing the overhead of index update and maintenance. At the same time, the distribution of vector data often shows an uneven situation. For example, there are high-density and sparse regions. The uneven situation will cause the query performance to fluctuate greatly, and it is difficult to maintain a stable and efficient query state.
[0037] In view of this, an embodiment of this application provides a query method, including receiving a query instruction; compressing the query data corresponding to the query instruction to obtain a simplified vector; querying a target index vector that matches the simplified vector from an index shape graph; wherein, the index shape graph represents K index vectors, and the K index vectors are obtained by compressing S source data in an index library, S and K are positive integers, and S≥K; determining at least one primary selected source data corresponding to the target index vector after compression from the S source data; obtaining a target source data that matches the query data according to the similarity between the query data and the multiple at least one primary selected source data.
[0038] Figure 1 The application scenario diagram of the query method, processor, processing system, storage medium and program product according to the embodiment of this application is shown.
[0039] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first processor 101 and a second processor 102. Data interaction can be carried out between the first processor 101 and the second processor 102 to execute the query method.
[0040] For example, after receiving a query instruction, the second processor 102 may compress the query data corresponding to the query instruction to obtain a simplified vector, so that the first processor 101 can query the target index vector matching the simplified vector from the index shape graph. Further, the second processor 102 may determine at least one primary selected source data corresponding to the target index vector after compression from the S source data, and the second processor 102 may also obtain the target source data matching the query data according to the similarity between the query data and the at least one primary selected source data.
[0041] The following will be based on Figure 1 the described scenario, and will describe in detail the query method of the embodiments of the present application through Figures 2 to 11 the following.
[0042] Figure 2 FIG. shows a flowchart of the query method according to an embodiment of the present application.
[0043] As Figure 2 shown, the query method of this embodiment includes operation S210 to operation S250.
[0044] In operation S210, a query instruction is received.
[0045] In operation S220, the query data corresponding to the query instruction is compressed to obtain a simplified vector.
[0046] In operation S230, a target index vector matching the simplified vector is queried from the index shape graph. Among them, the index shape graph represents K index vectors, and the K index vectors are obtained by compressing the S source data in the index library. S and K are positive integers, and S≥K.
[0047] In operation S240, at least one primary selected source data corresponding to the target index vector after compression is determined from the S source data.
[0048] In operation S250, the target source data matching the query data is obtained according to the similarity between the query data and the at least one primary selected source data.
[0049] According to an embodiment of the present application, in order to improve the efficiency of querying data from a database, an index structure may be constructed for the data in the database to obtain an index library. When a user needs to query target data from the database, the index library can be used to locate the position of the target data.
[0050] For example, in the field of target tracking, it is necessary to extract the feature information of the target object from an image. For a database containing 1 million images, an index structure can be built for the 1 million images. Specifically, each image in the database can be converted into an image feature vector, and the image feature vectors can be stored in the index library in a predetermined form. When a user needs to query for images that are the same as or similar to Image A, the images that are the same as or similar to Image A can be located through the index library. Specifically, an image usually includes multi-dimensional information, such as color information, texture information, etc. After converting the images into vectors, the multi-dimensional information in the images may be directly mapped to certain dimensions of the vectors. Therefore, the image feature vectors in the index library can represent the multi-dimensional information in the images. When querying for images that are the same as or similar to Image A, the images that are the same as or similar to Image A can be determined through the multi-dimensional information in the images represented by the image feature vectors in the index library.
[0051] For another example, in the field of intelligent customer service, text retrieval or voice retrieval is usually performed to find the answer that best matches the user's question from a text database or a voice database. Multiple text data in the text database can be converted into multiple text feature vectors, and the text feature vectors can be stored in a pre-constructed index library. When it is necessary to retrieve text from the text database that is relatively relevant to the user's question, multiple text feature vectors can be read from the index library, and the text feature vectors that are relatively relevant to the user's question can be located. Then, based on the located text feature vectors, the text data that is relatively relevant to the user's question can be determined.
[0052] Specifically, in operation S210, the user can input a query instruction, and the query instruction is used to represent the content that the user wants to query. For example, the user uploads an Image A and needs to query for images that are the same as or similar to Image A. In this case, the query instruction can represent that the user needs to query for images that are the same as or similar to Image A.
[0053] In operation S220, the query instruction can be converted into the form of vector data to obtain the query data corresponding to the query instruction. For example, the query instruction can be converted into vector data through a pre-trained language model, but it is not limited to this. The method of converting the query instruction into vector data can be set according to actual needs and is not limited here.
[0054] Specifically, the query instruction can reflect information in N dimensions, such as information in multiple dimensions of an image, speech, or text. The query data can be in the form of an N-dimensional vector. The N-dimensional vector is the eigenvalue of the query instruction reflected in N dimensions, and the eigenvalue includes at least one of the following: image feature, speech feature, text feature. For example, when the query instruction is used to reflect the multi-dimensional information of an image, the N-dimensional vector includes at least one of the following: image texture, image color, image brightness, the shape of the target object in the image, image depth, image grayscale. Or, when the query instruction is used to reflect the multi-dimensional information of speech, the N-dimensional vector includes at least one of the following: speech volume, speech rhythm, speech speed, speech intonation, speech timbre. Or, when the query instruction is used to reflect the multi-dimensional information of text, the N-dimensional vector includes at least one of the following: text semantics, text length, text paragraph structure, the word frequency of the target word in the text.
[0055] Furthermore, since the vector dimension of the query data is usually high and the computational complexity is high, the query data can be compressed to obtain a simplified vector. Compressing the query data, for example, includes reducing the dimension of the query vector, performing binary processing on the data in each dimension of the query vector, etc. By compressing the query data to obtain a simplified vector, the amount of computation can be reduced and the query efficiency can be improved.
[0056] In operation S230, multiple data in the database can be converted into vector data form to obtain S source data. The method of converting multiple data in the database into vector data can be set according to actual needs and is not limited here. Since the database usually contains a large amount of data, a large amount of vector data will be obtained after conversion, and the dimension of the vector data is usually high. Directly constructing a vector index for a large amount of high-dimensional vector data has a large computational cost, and the index data of the high-dimensional vector data also occupies a large storage space. Therefore, to reduce the storage overhead and improve the index construction speed, the S source data can be compressed to obtain K index vectors.
[0057] Specifically, compressing the S source data, for example, can include: first performing dimensionality reduction processing on the S source data to obtain S vector data after dimensionality reduction processing; then performing binarization processing on the S vector data after dimensionality reduction processing to obtain S index vectors. In this case, S is equal to K.
[0058] Among them, the dimensionality reduction process maps high-dimensional data to a low-dimensional space through a predetermined mathematical transformation method to improve the sample density and computational efficiency. The predetermined mathematical transformation methods adopted in the dimensionality reduction process include, for example, Principal Component Analysis (abbreviated as PCA), Multi-Dimensional Scaling (abbreviated as MDS), Isometric Mapping (abbreviated as ISOMAP), Linear Discriminant Analysis (abbreviated as LDA), etc. The predetermined mathematical transformation method can be set according to actual needs and is not limited herein.
[0059] Quantitative coding is performed on the vector data after dimensionality reduction through binarization processing. The binarization processing includes converting the data in multiple dimensions in the vector data after dimensionality reduction into a first value or a second value respectively according to a preset value. Specifically, it may include converting the data with a value greater than the preset value into the first value, and converting the data with a value less than or equal to the preset value into the second value. The first value can be 1, for example, and the second value can be 0, for example. The vector obtained after binarization processing can be regarded as a numerical sequence composed of 0 and 1.
[0060] The preset value can be calculated based on the multiple-dimensional numerical values in the vector data after dimensionality reduction. For example, the preset value can include the median of the multiple-dimensional numerical values, or the mean of the multiple-dimensional numerical values, or can also be calculated from the multiple-dimensional numerical values in other ways, which is not limited herein.
[0061] Specifically, compressing S source data can also include: first performing dimensionality reduction processing on the S source data to obtain S vector data after dimensionality reduction; then performing binarization processing on the S vector data after dimensionality reduction to obtain S binarized vectors; and then performing the same vector merging processing on the S binarized vectors to obtain K index vectors, where S is greater than K in this case. Among them, the methods of dimensionality reduction processing and binarization processing refer to the methods of dimensionality reduction processing and binarization processing in the embodiment of operation S230 and will not be elaborated herein.
[0062] Among them, performing the same vector merging processing on the S binarized vectors includes: extracting the same binarized vectors in the S binarized vectors and using the same binarized vectors as an index vector, and the value of the index vector is the same as the value of the same binarized vector. For example, among the S binarized vectors, if the values of N binarized vectors are all (0, 0, 1), then these N binarized vectors can be used as an index vector, and the value of the index vector is (0, 0, 1).
[0063] Among the S source data, there may be identical data. After separately performing dimensionality reduction and binarization processing on these identical data, duplicate binarized vectors will be obtained. If these duplicate binarized vectors are directly processed, it will cause unnecessary waste of computing and storage resources. By performing identical vector merging processing on the S binarized vectors, the identical binarized vectors can be processed as one index vector, thereby reducing the computational amount and saving storage resources.
[0064] Furthermore, after obtaining the K index vectors, the K index vectors can be represented in the form of an index shape graph. The index shape graph includes multiple shape nodes, and the multiple shape nodes are connected by directed edges. The directed edges can represent numerical values of different dimensions, and the numerical values of multiple dimensions in the index vector can be represented by multiple directed edges.
[0065] For example: If an index vector is (1, 0, 0), it can be represented by the directed edge between shape node 1 and shape node 2, the directed edge between shape node 2 and shape node 3, and the directed edge between shape node 3 and shape node 4 in the index shape graph.
[0066] Specifically, obtaining the target index vector that matches the simplified vector from the index shape graph may include: traversing and querying the directed edges between multiple shape nodes, matching the numerical values of different dimensions represented by the directed edges with the corresponding dimension numerical values in the simplified vector, and taking the index vector that completely matches the simplified vector as the target index vector. Among them, completely matching the simplified vector may include that the numerical values of each dimension of the target index vector are exactly the same as the corresponding dimension numerical values of the simplified vector.
[0067] For example: If the simplified vector is (1, 0, 0), it can traverse and query the directed edges to determine that the index vector represented by the directed edge between shape node 1 and shape node 2, the directed edge between shape node 2 and shape node 3, and the directed edge between shape node 3 and shape node 4 is also (1, 0, 0), which is exactly the same as the values of each dimension of the simplified vector. Therefore, this index vector can be taken as the target index vector that matches the simplified vector.
[0068] Specifically, obtaining the target index vector that matches the simplified vector from the index shape graph may also include: taking the index vector that partially matches the simplified vector as the target index vector. Specifically, there may be no vector in the K index vectors that is exactly the same as the simplified vector. In this case, the index vector that partially matches the simplified vector can be taken as the target index vector. Among them, partially matching the simplified vector may include that the similarity between the target index vector and the simplified vector satisfies a preset numerical range, and the preset numerical range can be set according to actual needs.
[0069] For example: the simplified vector is (1, 0, 0), and an index vector (1, 0, 1) is queried from the index shape diagram. Since the index vector is highly similar to the simplified vector, this index vector can be used as the target index vector that matches the simplified vector.
[0070] Furthermore, in operation S240, there is a correspondence between the source data and the index vector obtained by compressing the source data, and one index vector corresponds to at least one source data. For example: in the case of obtaining K index vectors by successively performing dimensionality reduction processing and binarization processing on S source data, one index vector can correspond to one source data; while in the case of obtaining K index vectors by successively performing dimensionality reduction processing, binarization processing, and identical vector merging processing on S source data, one index vector may correspond to multiple source data. After determining the target index vector that matches the simplified vector, the source data corresponding to the target index vector is used as the primary selected source data.
[0071] Since the target index vector is obtained by compressing at least one initial source data, the simplified vector is obtained by compressing the query data, and the target index vector matches the simplified vector, there is also a matching relationship between at least one initial source data and the query data.
[0072] For example: the target index vector is obtained by compressing source data v1 and v2, that is, the initial source data includes source data v1 and v2, and the simplified vector is obtained by compressing the query data. Since the target index vector matches the simplified vector, for example, the values of the target index vector and the simplified vector are exactly the same, the query data also matches source data v1 and v2, for example, the values of the query data and source data v1 and v2 are exactly the same.
[0073] On the one hand, in the case of a large number of index vectors, directly storing the index vectors will occupy a large storage space. By compressing the source data to obtain index vectors and then representing the index vectors through an index shape diagram, the index shape diagram occupies less storage space than the index vectors, thus saving storage space.
[0074] On the other hand, if the query data is directly traversed and matched with the source data in the index library, the traversal process will increase the computational overhead when the number of index vectors is large, resulting in low query efficiency. By compressing the query data into a simplified vector, compressing the source data into index vectors, and representing the index vectors through an index shape graph, when querying, first find the target index vector matching the simplified vector from the index shape graph, and then determine the initial selected source data corresponding to the target index vector, so that a small number of initial selected source data corresponding to the query data can be preliminarily screened out from a large amount of source data. Subsequently, only the source data matching the query data needs to be determined from the small number of initial selected source data. Thus, compared with directly traversing and matching the query data with the data in the index library, the matching range can be reduced and the query efficiency can be improved.
[0075] According to an embodiment of the present application, in operation S250, in order to further screen out the source data matching the query data from at least one initial selected source data, the vector distance between the query data and each of the at least one initial selected source data can be calculated to obtain the similarity between the query data and the at least one initial selected source data. The closer the vector distance is, the higher the similarity between the query data and the initial selected source data. The method for calculating the vector distance between the query data and the initial selected source data can, for example, include calculating the Euclidean distance, Manhattan distance, Chebyshev distance, Minkowski distance, and can also include calculating the cosine similarity and distance, haversine distance, Hamming distance, Jaccard index and distance, etc. The specific method for calculating the vector distance between the query data and the initial selected source data can be set according to actual needs and is not limited herein. The initial selected source data with a similarity greater than a preset threshold to the query data can be used as the target source data.
[0076] By compressing the query data corresponding to the query instruction to obtain a simplified vector, the dimension of the query vector can be reduced, the amount of calculation can be reduced, and the query efficiency can be improved. By compressing the source data into index vectors, the storage overhead can be reduced and the index construction speed can be improved. By representing the index vectors through an index shape graph, querying the target index vector matching the simplified vector from the index shape graph, and then determining at least one initial selected source data corresponding to the target index vector after compression from S source data, the storage space can be saved, the query range can be reduced, and the query efficiency can be improved.
[0077] According to an embodiment of the present application, compressing the query data corresponding to the query instruction to obtain a simplified vector includes: performing dimensionality reduction processing on the query data to obtain dimensionality-reduced data; performing binary encoding processing on the dimensionality-reduced data to obtain a simplified vector.
[0078] Specifically, the query data is usually a high-dimensional vector. Directly processing high-dimensional vectors has a relatively high computational complexity and a time-consuming calculation process. The query data can be dimensionally reduced. By dimensional reduction, the high-dimensional query data is mapped into a low-dimensional space, reducing the number of dimensions to be considered during calculation, thereby reducing the computational complexity and improving the computational efficiency. The query data can be dimensionally reduced by methods such as principal component analysis, multidimensional scaling, isometric feature mapping, linear discriminant analysis, etc., or other methods can also be used according to actual needs to dimensionally reduce the query data. The method for dimensionally reducing the query data is not limited herein. Binarizing the dimensionally reduced data can include converting the values of multiple dimensions in the dimensionally reduced data into binary forms, for example, it can be converted into 0 or 1.
[0079] By dimensionally reducing the query data to obtain dimensionally reduced data, and then binarizing the dimensionally reduced data to obtain a simplified vector, the dimension of the query data can be reduced and the data complexity can be reduced. Thus, when querying the target index vector matching the simplified vector from the index shape graph, the amount of calculation can be reduced and the query efficiency can be improved.
[0080] According to an embodiment of the present application, the simplified vector includes a first value and a second value; binarizing the dimensionally reduced data includes: performing numerical calculation on the multi-dimensional sub-data included in the dimensionally reduced data to obtain a reference value; generating a first value or a second value corresponding to each of the multi-dimensional sub-data based on the numerical deviation between each multi-dimensional sub-data and the reference value.
[0081] Specifically, performing numerical calculation on the multi-dimensional sub-data included in the dimensionally reduced data can specifically include calculating the median of the multi-dimensional sub-data or calculating the average value of the multi-dimensional sub-data, and the mean or average value of the multi-dimensional sub-data can be used as the reference value. Further, calculate the difference between each multi-dimensional sub-data and the reference value respectively to obtain the data deviation. When the data deviation is positive, a first value corresponding to the sub-data can be generated, and when the data deviation is zero or negative, a second value corresponding to the sub-data can be generated.
[0082] For example: the first value is 1 and the second value is 0. For the n-dimensional dimensionally reduced data, the mean value of the values on n dimensions can be calculated, and the calculated mean value is used as the reference value. For the values on n dimensions, when the numerical difference between the value on a certain dimension and the reference value is positive, that is, when the value on this dimension is greater than the mean value, the value on this dimension is encoded as 1, and when the numerical difference between the value on this dimension and the reference value is zero or negative, that is, when the value on this dimension is less than or equal to the mean value, the value on this dimension is encoded as 0. Thus, the dimensionally reduced data can be converted into a numerical sequence composed of 0 and 1.
[0083] Through binarization processing, the dimensionality-reduced data can be converted into a simpler structure, thereby reducing the complexity of the dimensionality-reduced data, and further improving the query efficiency when querying for target index vectors that match the simplified vectors from the index shape graph.
[0084] According to an embodiment of the present application, the query method further includes constructing an index binary tree based on K index vectors; and performing node compression on the index binary tree to obtain an index shape graph.
[0085] Figure 3 The schematic diagram shows constructing an index binary tree based on K index vectors according to an embodiment of the present application.
[0086] As Figure 3 shown, for 3 n-dimensional source data v1, v2, v3, where n is greater than or equal to 4, first, dimensionality reduction processing can be performed to obtain 4-dimensional vector data after dimensionality reduction processing, including dimensionality-reduced data v1′, v2′, v3′. Then, binarization processing is performed on v1′, v2′, v3′ to obtain index vectors v1′′, v2′′, v3′′. An index binary tree is constructed based on the index vectors v1′′, v2′′, v3′′. The index binary tree can, for example, include multiple nodes, and the edges between the nodes can represent the values of a certain dimension in the index vector. For example, Figure 3 the multiple edges on the left side of the root node of the index binary tree in
[0087] can successively represent: the value of the first dimension is 0, the value of the second dimension is 0, the value of the third dimension is 1, and the value of the fourth dimension is 1. Figure 3 The number of nodes in the index binary tree constructed based on multiple multi-dimensional vectors is significantly more than the number of multi-dimensional vectors. For example,
[0088] in
[0089] the number of index vectors is 3, and the number of nodes in the index binary tree constructed based on the index vectors is 11. Moreover, the number of nodes in the index binary tree will increase significantly with the increase in the number of index vectors and the dimension of a single index vector. Therefore, the memory overhead required to construct and store the index binary tree for a large amount of high-dimensional vector data is quite large, making it difficult for a processor with limited storage resources to accommodate a huge number of index binary tree nodes. Therefore, node compression can be performed on the index binary tree to obtain an index shape graph, and the number of nodes in the index shape graph is significantly reduced, reducing the memory occupancy.
[0088] By constructing an index binary tree based on K index vectors and then performing node compression on the index binary tree to obtain an index shape graph, an index shape graph with fewer nodes can be obtained, thereby reducing the memory occupancy.
[0089] According to an embodiment of the present application, each of the K index vectors includes N - dimensional first numerical values, where N is a positive integer; the index binary tree is divided into multiple tree nodes at N levels and associated edges between the tree nodes, and the associated edges represent the N - dimensional first numerical values; the index shape graph is divided into multiple shape nodes at N levels and directed edges between the shape nodes, and the directed edges represent the N - dimensional first numerical values; one shape node in the index shape graph represents multiple tree nodes in the index binary tree that have the same subtree shape.
[0090] For example: The index binary tree includes multiple tree nodes and associated edges between the tree nodes. Among them, the multiple tree nodes are divided into N levels. For example, the root tree node is the first - level tree node, the tree nodes directly connected to the root tree node are the second - level tree nodes, the tree nodes directly connected to the second - level tree nodes are the third - level tree nodes, and so on. Among them, the associated edge between the first - level tree node and the second - level tree node is used to represent the first - dimensional numerical value in the index vector, the associated edge between the second - level tree node and the third - level tree node is used to represent the second - dimensional numerical value in the index vector, and so on.
[0091] For example Figure 3 As shown in, there are two second - level tree nodes connected to the first - level tree node. The two associated edges between the two second - level tree nodes and the first - level tree node can both represent the first - dimensional numerical value in the index vector. For example, the associated edge between the first - level tree node and the left second - level tree node represents the first - dimensional numerical value in the index vector as 0, and the associated edge between the first - level tree node and the right second - level tree node represents the first - dimensional numerical value in the index vector as 1. Further, the associated edge between the left second - level tree node and the third - level tree node represents the second - dimensional numerical value in the index vector as 0, and the associated edge between the right second - level tree node and the third - level tree node represents the second - dimensional numerical value in the index vector as 1.
[0092] Further, there are multiple subtree shapes in the index binary tree.
[0093] Figure 4 The schematic diagram shows multiple subtree shapes in the index binary tree according to an embodiment of the present application.
[0094] As Figure 4 shown, the index binary tree includes subtree shapes 1 to 5.
[0095] Further, multiple tree nodes in the index binary tree may correspond to the same subtree shape. For example Figure 4 in the index binary tree, the left - most third - level tree node and the right - most third - level tree node both correspond to subtree shape 1. Multiple tree nodes with the same subtree shape can be represented by one shape node in the index shape graph. For example, represented by one shape node in the index shape graph Figure 4The third-level tree nodes on the leftmost side and the third-level tree nodes on the rightmost side of the index binary tree.
[0096] Specifically, the index shape graph includes multiple shape nodes, which are connected by directed edges. Among them, the multiple shape nodes are divided into N levels. For example, the root shape node is the first-level shape node, the shape nodes directly connected to the root shape node are the second-level shape nodes, and the shape nodes directly connected to the second-level shape nodes are the third-level shape nodes, and so on. Among them, the directed edge between the first-level shape node and the second-level shape node is used to represent the first-dimensional value in the index vector, and the directed edge between the second-level shape node and the third-level shape node is used to represent the second-dimensional value in the index vector, and so on.
[0097] By using a shape node in the index shape graph to represent multiple tree nodes with the same subtree shape in the index binary tree, node compression can be performed on the index binary tree, thereby reducing the storage space occupied by the index shape graph.
[0098] According to an embodiment of the present application, performing node compression on the index binary tree to obtain an index shape graph includes: determining the shape encoding of the tree node, where the shape encoding is used to represent the subtree shape category of the tree node; merging the tree nodes with the same shape encoding to generate an index shape graph.
[0099] Figure 5 A schematic diagram showing the correspondence between the subtree shape and the shape encoding according to an embodiment of the present application is shown.
[0100] As Figure 5 shown, the index binary tree may include different subtree shapes, and different subtree shapes correspond to different shape encodings. In addition, the nodes at certain positions in the index binary tree may not exist, that is, there are empty nodes. Further, a number can be uniformly set for the empty nodes in the index binary tree to obtain the shape encoding corresponding to the empty nodes, or the empty tree nodes are not numbered.
[0101] Further, different subtree shapes may correspond to the same or different subtree heights. The subtree height is used to represent, for a subtree of a certain shape, the number of nodes on the longest path from the root node of the subtree to the farthest node of the subtree. For example, the subtree heights of the subtree shapes numbered 2 and 3 are 2, the subtree heights of the subtree shapes numbered 4 and 5 are 3, and the subtree height of the subtree shape numbered 6 is 4.
[0102] Further, the tree nodes with the same shape encoding can be merged to generate the same shape node in the index shape graph, and the number of the shape node is the same as the shape encoding. For example: combining Figure 4 and Figure 5 it can be known that Figure 4The third-level tree nodes on the leftmost and rightmost sides of the index binary tree both have sub-tree shapes with a shape code of 2. Therefore, the third-level tree nodes on the leftmost and rightmost sides can be merged into the same shape node in the index shape graph, and the number of this shape node is also 2. Further, the last-level tree nodes in the index binary tree can be merged into the same shape node in the index shape graph, and the number of this shape node is the same as the shape code of the tree node. For example, all the fourth-level tree nodes in Figure 4 the index binary tree can be merged into the same shape node in the index shape graph. From Figure 5 it can be seen that the shape code of the tree node is 1, so the number of this shape node is also 1.
[0103] Figure 6 FIG. shows a schematic diagram of generating an index shape graph according to an embodiment of the present application.
[0104] As Figure 6 shown, the binarized vectors include identical vectors, such as r1′′ and r6′′, r2′′ and r7′′, r4′ and r8′′. The identical binarized vectors can be used as one index vector to obtain index vectors v1′′′ to v5′′′. An index binary tree is constructed for the index vectors v1′′′ to v5′′′, and according to Figure 5 the different sub-tree shapes in the shown index binary tree and their corresponding shape codes, the tree nodes with the same shape code in the index binary tree are merged to generate an index shape graph.
[0105] From Figure 6 it can be seen that the index binary tree contains a total of 11 tree nodes. The number of shape nodes in the index shape graph obtained after node compression of the index binary tree is 6 (such as the shape nodes with shape codes 6, 4, 5, 3, 2, 1 in Figure 6 ), that is, the number of nodes is compressed to about half of the number of tree nodes in the index binary tree. It can be seen that by compressing the index binary tree into an index shape graph according to the sub-tree shape, the generated index shape graph only occupies a small storage space.
[0106] In addition, the higher the height of the index binary tree and the closer it is to a full binary tree, the higher the compression ratio of compressing the index binary tree into an index shape graph. Among them, a full binary tree can include a binary tree in which the number of nodes at each level reaches the maximum value. For a full binary tree constructed based on n index vectors with a dimension of D, the height of the full binary tree is D + 1, and the total number of nodes is 2 D+1 -1. When compressing the index binary tree into an index shape graph, only 2 D+1 -1 nodes of the full binary tree need to be traversed. Therefore, the time complexity of compressing the index binary tree into an index shape graph is O(2 D+1-1), indicating that only 2 D+1 traversals of the nodes need to be performed when compressing the index binary tree into an index shape graph. Therefore, the time complexity of compressing the index binary tree into an index shape graph is only related to the number of dimensions D of the index vector, and has nothing to do with the number n of index vectors.
[0107] According to an embodiment of the present application, multiple source data corresponding to the same index vector after compression correspond to a source data set; the index binary tree includes K leaf nodes corresponding to K index vectors, and the query method further includes: associating the K source data sets corresponding to the K index vectors with the K leaf nodes respectively.
[0108] Specifically, since the index vector may be obtained by compressing multiple source data, the index vector may correspond to multiple source data, and the multiple source data corresponding to one index vector can be used as a source data set. For example: a certain index vector is obtained by compressing source data v1 and source data v2, so the source data set corresponding to this index vector includes source data v1 and source data v2.
[0109] Specifically, the leaf nodes may include the tree nodes at the last level in the index binary tree.
[0110] Figure 7 shows a schematic diagram of the association relationship between the leaf nodes and the source data sets according to an embodiment of the present application. As Figure 7 the index binary tree in includes a total of 5 leaf nodes.
[0111] Since the index binary tree is divided into multiple tree nodes at N levels and the association edges between the tree nodes, and the association edges can represent the first value of a certain dimension in the index vector, therefore, according to the multiple association edges between the multiple tree nodes connected in sequence through the association edges, the first values of multiple dimensions in the index vector can be obtained.
[0112] For example, for Figure 7 the index binary tree in, the index binary tree includes multiple association edges. Among them, the association edge between the first-level tree node and the second-level tree node on the left represents that the first dimension value in the index vector is 0, the association edge between the second-level tree node on the left and the third-level tree node on the leftmost represents that the second dimension value in the index vector is 0, and the association edge between the third-level tree node on the leftmost and the fourth-level tree node on the leftmost represents that the third dimension value in the index vector is 0, so as to represent the index vector (0, 0, 0). From Figure 7 it can be seen that each leaf node can correspond to an index vector. For example, the leftmost leaf node corresponds to the index vector (0, 0, 0).
[0113] Furthermore, since one index vector can correspond to one source data set and one leaf node can correspond to one index vector, the leaf node can be associated with the source data set. For example Figure 7 As shown, the leftmost leaf node corresponds to the index vector (0, 0, 0), and the index vector (0, 0, 0) corresponds to the source data set 1. Therefore, the leftmost leaf node can be associated with the source data set 1.
[0114] Specifically, by associating the K source data sets corresponding to the K index vectors with the K leaf nodes respectively, the relevant source data can be located more quickly, thereby improving the query efficiency.
[0115] According to an embodiment of the present application, in the index shape graph, the target shape node in the target path corresponding to the target index vector carries offset information, and the offset information is used to represent: in the index binary tree, the position offset of the target leaf node corresponding to the target index vector relative to the first leaf node in the horizontal direction; determining at least one primary source data corresponding to the target index vector after compression from S source data includes: obtaining the offset information carried by the target shape node during the process of querying the target index vector node by node along the target path; determining the target leaf node from the K leaf nodes based on the offset information; and determining the data in the target source data set associated with the target leaf node as multiple primary source data.
[0116] Specifically, although compressing the index binary tree into an index shape graph can reduce the number of nodes and save storage space, the index shape graph will also lose some information in the index binary tree. For example, the position information of the last-level leaf nodes will be lost (the last-level leaf nodes in the index binary tree only correspond to one shape node in the index shape graph). The restoration of this part of the lost information can be achieved by adding offset information, and then the source data corresponding to the query path can be located.
[0117] Figure 8 FIG. shows a schematic diagram of generating an index shape graph carrying offset information according to an embodiment of the present application.
[0118] As Figure 8As shown, the index binary tree can represent the association relationship between leaf nodes and the source data set. For example, the leftmost leaf node is associated with the source data set 1. However, since all leaf nodes in the index binary tree are merged into the same shape node, that is, the shape node with a shape code of 1, all query paths pass through the shape node with a shape code of 1. Therefore, based on the index shape diagram alone, it is impossible to determine which source data set the target index vector or its query path is specifically associated with. Therefore, offset information can be added to multiple shape nodes in the index shape diagram. The offset information represents the horizontal position offset of the leaf node corresponding to the index vector relative to the first leaf node in the index binary tree corresponding to the index shape diagram. Among them, the offset information of each shape node can be expressed as the number of leaf nodes included in the left subtree of the tree node corresponding to the index binary tree corresponding to the shape node.
[0119] The offset information is, for example Figure 8 as shown in the content in the brackets of the shape nodes of the index shape diagram. As Figure 8 shown, the offset information carried by the shape node with a shape code of 6 can be expressed as (3). Correspondingly, in the index binary tree, the number of leaf nodes at the last level of the left subtree of the tree node corresponding to the shape node with a shape code of 6 is 3; for another example, the offset information carried by the shape node with a shape code of 4 can be expressed as (2). Correspondingly, in the index binary tree, the number of leaf nodes at the last level of the left subtree of the tree node corresponding to the shape node with a shape code of 4 is 2.
[0120] According to an embodiment of the present application, during the process of querying the target index vector node by node along the target path, the offset information carried by the target shape node can be obtained; and based on the offset information, the target leaf node can be determined from the K leaf nodes, and then the source data set associated with the target leaf node can be located.
[0121] Further, the target index vector queried from the index shape diagram that matches the simplified vector can be the target index vector that completely matches the simplified vector; or it can be the target index vector that partially matches the simplified vector.
[0122] Based on this, there are also two cases for determining the target leaf node from the K leaf nodes.
[0123] The first case: In the case where the target index vector that completely matches the simplified vector is queried, there is only one target index vector, and there is only one target path corresponding to the target index vector. Based on the offset information carried by the target shape node in this path, one target leaf node can be determined from the K leaf nodes.
[0124] Specifically, it can be to sum the offset information carried by the respective upper - level shape nodes of multiple edges with the value of the first value (for example, "1") in the target path to obtain the total offset. This total offset characterizes the position offset of the leaf node corresponding to the index vector in the horizontal direction relative to the first leaf node in the corresponding index binary tree. Therefore, the target leaf node can be determined from the K leaf nodes according to this total offset, that is, starting from the first leaf node, the target leaf node is located at the position offset by the corresponding total offset to the right.
[0125] For example: the simplified vector is (0, 1, 1). For the Figure 8 index shape graph as shown, it can be traversed and queried sequentially starting from the root node, that is, the shape node with shape code 6. First, it is determined that the first - dimension value represented by the directed edge between the shape node with shape code 6 and the shape node with shape code 4 is 0, which matches the first - dimension value 0 of the simplified vector. Further, it can be determined that the second - dimension value represented by the directed edge between the shape node with shape code 4 and the shape node with shape code 3 is 1, which matches the second - dimension value 1 of the simplified vector. Then, it can be determined that the third - dimension value represented by the directed edge between the shape node with shape code 3 and the shape node with shape code 1 is 1, which matches the third - dimension value 1 of the simplified vector. Thus, the target index vector (0, 1, 1) that matches the simplified vector (0, 1, 1) can be obtained, and the target path includes the target shape nodes with shape codes 6, 4, 3, and 1 respectively.
[0126] Sum the multiple offset information corresponding to multiple upper - level shape nodes of multiple edges with the value of 1 in the target path. As Figure 8 shown: the offset information of the shape node with shape code 4 is 2, and the offset information of the shape node with shape code 3 is 0. Calculate the sum of the offset information of the shape nodes with offset information 4 and 3 respectively, and the total offset information is 2.
[0127] As Figure 8 shown, when the target index vector is (0, 1, 1) and the total offset information is 2, the first leaf node is the left - most leaf node in the index binary tree, that is, the leaf node associated with the source data set 1. Since the total offset information is 2, the leaf node with a position offset of 2 relative to the first leaf node can be determined, that is, the leaf node at the third position from left to right, and the target leaf node is obtained. Specifically, the target leaf node is the leaf node associated with the source data set 3. The source data set associated with the target leaf node can be used as the target source data set, and the target source data set includes at least one initial source data.
[0128] The second case: When target index vectors that partially match the simplified vector are found, and there are multiple target index vectors and multiple corresponding target paths, multiple target leaf nodes can be determined from the K leaf nodes based on the offset information carried by the target shape nodes of these multiple paths.
[0129] Specifically, it can be to sum up the offset information carried by the respective upper-level shape nodes corresponding to multiple edges with the value of the first value (for example, "1") in each target path, to obtain multiple total offsets corresponding one by one to the multiple target paths. Multiple target leaf nodes are determined one by one according to these total offsets, and these multiple target leaf nodes all belong to the retrieval range.
[0130] For example, if the simplified vector is (1, 1, 1), then querying from Figure 8 two target index vectors that partially match this vector can be found: (1, 0, 0) and (1, 0, 1). Then, two sets of total offsets corresponding to the two target paths where these two target index vectors are located are calculated, which are 3 and 4 respectively. Then, according to the offset positions of 3 and 4, 2 target leaf nodes are determined from the multiple leaf nodes of the index binary tree, which are the nodes that are offset 3 positions and 4 positions to the right starting from the first leaf node.
[0131] In this case, another method for determining multiple target leaf nodes from the K leaf nodes can be: When a target index vector that partially matches the simplified vector is found, that is, the last-level leaf node is not matched and a null node is turned to on the intermediate path. Assume that when reaching the longest matching node (i.e., the shape node where the next step will turn to a null node), the offset position of the corresponding leaf node in the index binary tree is i, and the total number of leaf nodes corresponding to the longest matching node is r. Then, the leaf nodes from the i-th to the (i + (r - 1))-th are all within the matching range.
[0132] According to the embodiments of the present application, when no target index vector that completely matches the simplified vector is found, by returning the data sets corresponding to multiple target leaf nodes, a matching query to the greatest extent is realized.
[0133] According to the embodiments of the present application, by pre-determining the association relationship between the source data set and the leaf nodes in the index binary tree, and adding offset information to the shape nodes in the index shape graph, and using the offset information to represent the position information of the corresponding index binary tree node in the index binary tree in the index shape graph, the association relationship between the leaf nodes and the source data set can be represented in the index shape graph, thus avoiding the information loss problem caused by compressing the index binary tree into the index shape graph.
[0134] Moreover, by adding the position offset information of the index binary tree nodes to the shape nodes and pre - constructing the association relationship between the source data and the index binary tree leaf nodes as the association information of multiple primary source data, the association information of multiple primary source data can be synchronously obtained during the data query process, so that the relevant source data can be located relatively quickly without the need for subsequent data matching operations, thereby improving the query efficiency.
[0135] According to an embodiment of the present application, the query method further includes: after obtaining the index shape graph by compressing the nodes of the index binary tree, deleting the index binary tree.
[0136] Specifically, since the storage space occupied by the index binary tree is relatively large, if both the index binary tree and the index shape graph are stored, it will occupy a large amount of storage space. The index shape graph can already represent the information that the index binary tree is to represent. For example, the offset information in the index shape graph can already represent the association relationship between the leaf nodes and the source data set. During the process of querying the target index vector node by node along the target path, the offset information carried by the target shape node can be obtained, so as to determine the association relationship between the leaf nodes and the source data set. Therefore, based on the index shape graph, the primary source data corresponding to the query data can be determined. In this case, the index binary tree can be deleted and only the index shape graph is stored, thereby releasing the storage space.
[0137] According to an embodiment of the present application, the query method further includes: recording multiple source data corresponding to the same index vector after compression in a data list, and storing the mapping relationship between the K index vectors and the K data lists; determining at least one primary source data corresponding to the target index vector after compression from the S source data includes: based on the mapping relationship, determining the target data list corresponding to the target index vector from the K data lists.
[0138] For example: for an index vector, multiple source data corresponding to the index vector can be determined from the S source data, and these multiple source data are stored in a data list. Thus, K data lists corresponding to the K index vectors can be obtained respectively.
[0139] Specifically, in order to facilitate the association between the index vector and its corresponding data list, the mapping relationship between the K index vectors and the K data lists can be determined and stored. Specifically, vector identifiers can be set for the K index vectors respectively, and list identifiers can be set for the K data lists respectively. The mapping relationship between the index vector and the data list can be represented by the mapping relationship between the vector identifier and the list identifier. For example: there is a mapping relationship between the vector identifier v1′′′ and the list identifier T1, indicating that the index vector with the vector identifier v1′′′ corresponds to the data list with the list identifier T1.
[0140] Further, based on the mapping relationship, determining the target data list corresponding to the target index vector from the K data lists may include: determining the target vector identifier of the target index vector; determining the target list identifier having a mapping relationship with the target vector identifier; determining the data list corresponding to the target list identifier to obtain the target data list.
[0141] By recording multiple source data corresponding to the same index vector after compression in one data list, storing the mapping relationship between the K index vectors and the K data lists, and then determining the target data list corresponding to the target index vector from the K data lists based on the mapping relationship, the primary source data corresponding to the target index vector can be determined simply and accurately, reducing the query complexity.
[0142] This application also provides a processor, including: an on-chip computing component for executing the query method; an on-chip storage unit for storing the index shape diagram.
[0143] Specifically, the processor includes, for example, a Field-Programmable Gate Array (FPGA for short), and may also include a Graphics Processing Unit (GPU for short), etc.
[0144] Specifically, the query method can be executed by a processor and the index shape diagram can be stored, thereby completing the query task. Specifically, on the one hand, by storing the index data in the on-chip storage, the index shape diagram can be directly read from the on-chip storage unit when querying data. Compared with obtaining the index shape diagram from an external storage during query, the latency of obtaining index data from the external storage during query can be reduced, thereby improving the query efficiency.
[0145] On the other hand, the storage space of the on-chip storage unit is limited, and the storage space occupied by the index shape diagram is small, so it can be stored in the on-chip storage unit, thereby reducing the memory overhead.
[0146] According to an embodiment of the present application, the on-chip computing component includes M groups of pipelined processing units; the on-chip storage unit includes M sub-storage areas communicatively connected to the M groups of pipelined processing units, where the N-dimensional first values included in each of the K index vectors corresponding to the index shape diagram are stored in N target sub-storage areas among the M sub-storage areas, M and N are positive integers, and M is greater than or equal to N; N groups of target pipelined processing units among the M groups of pipelined processing units are used to perform the following operations: reading the N-dimensional first values from the N target sub-storage areas in parallel; matching the N-dimensional second values included in the simplified vector with the N-dimensional first values in parallel to obtain the target index vector that matches the simplified vector.
[0147] Figure 9The structural schematic diagram of the on-chip computing component and the on-chip storage unit according to an embodiment of the present application is shown.
[0148] As Figure 9 shown, the on-chip computing component includes, for example, pipelined processing unit 1, pipelined processing unit 2... pipelined processing unit N, and pipelined processing unit M, and the on-chip storage unit includes, for example, sub-storage area 1, sub-storage area 2... sub-storage area N, and sub-storage area M.
[0149] As Figure 9 shown, each sub-storage area in the on-chip storage unit can store the first value of one dimension in the index vector. For example, sub-storage area 1 can store the first value of the first dimension in K index vectors, sub-storage area 2 can store the first value of the second dimension in K index vectors, and the Nth dimension first value is stored in sub-storage area N. The sub-storage area storing the first value can be used as the target sub-storage area. For example, in the case where the index vector includes N-dimensional first values, N target sub-storage areas are included in M sub-storage areas. The on-chip computing component includes M groups of pipelined processing units, and each group of pipelined processing units can be communicatively connected to the sub-storage area, so as to read the first value stored in the sub-storage area to execute the query method. For example, N-dimensional first values can be read from N target sub-storage areas through N groups of target pipelined processing units.
[0150] Specifically, N groups of target pipelined processing units can read N-dimensional first values from N target sub-storage areas in parallel. For example, N groups of target pipelined processing units simultaneously read N-dimensional first values from N target sub-storage areas. Compared with sequentially reading N-dimensional first values from N target sub-storage areas, the parallel reading method can improve the reading speed.
[0151] Further, after reading the N-dimensional first values, N groups of target pipelined processing units can match the N-dimensional second values included in the simplified vector with the N-dimensional first values in parallel, so as to obtain the target index vector that matches the simplified vector. For example, N groups of target pipelined processing units simultaneously match the N-dimensional second values included in the simplified vector with the N-dimensional first values. By the method of parallel matching by N groups of target pipelined processing units, the matching speed can be improved.
[0152] The present application also provides a processing system, including a first processor and a second processor. The query task can also be completed jointly by the first processor and the second processor.
[0153] Among them, the first processor is used to query a target index vector that matches the simplified vector from the index shape graph. Among them, the index shape graph represents K index vectors, and the K index vectors are obtained by compressing S source data in the index library. S and K are positive integers, and S≥K. The second processor is used to compress the query data corresponding to the query instruction to obtain a simplified vector, determine at least one primary selected source data corresponding to the target index vector after compression from the S source data, and obtain a target source data that matches the query data according to the similarity between the query data and the at least one primary selected source data.
[0154] Specifically, the first processor includes, for example, an FPGA or a GPU. Among them, the index shape graph can be stored in the first processor. For example, it can be stored in the on-chip storage space of the FPGA. The on-chip storage space of the FPGA includes, for example, the static random access memory (SRAM) of the FPGA. In addition, the mapping relationship between the K index vectors and the K data lists can be stored in the off-chip storage space of the FPGA. The off-chip storage space of the FPGA includes, for example, the dynamic random access memory (DRAM). The second processor includes, for example, a central processing unit (CPU) and can also include a GPU.
[0155] Figure 10 The system interaction diagram of the first processor and the second processor according to an embodiment of the present application is shown.
[0156] As Figure 10 shown, the query data can be input into the second processor, and the second processor compresses the query data to obtain a simplified vector. Compressing the query data includes, for example, performing dimensionality reduction processing on the query data to obtain dimensionality-reduced data, and then performing binary encoding processing on the dimensionality-reduced data.
[0157] The simplified vector is input into the first processor. The index shape graph is stored in the on-chip storage space in the first processor. In addition, the mapping relationship between the index vector and the data list is stored in the off-chip storage space. The first processor also includes multiple groups of pipelined processing units. The multiple groups of pipelined processing units can perform parallel processing. The parallel processing includes parallelly reading N-dimensional first numerical values from the on-chip storage space, and parallelly matching the N-dimensional second numerical values included in the simplified vector with the N-dimensional first numerical values to obtain a target index vector that matches the simplified vector.
[0158] The target index vector can be returned to the second processor, so that the second processor can determine at least one primary selected source data corresponding to the target index vector after compression from S source data according to the target index vector, and obtain the target source data matching the query data according to the similarity between the query data and the multiple at least one primary selected source data.
[0159] Further, the second processor can also construct an index binary tree based on K index vectors, perform node compression on the index binary tree to obtain an index shape graph, and determine the mapping relationship between the K index vectors and K data lists. The index shape graph and the mapping relationship generated by the second processor can be written into the first processor.
[0160] Figure 11 The flowchart of a query method according to another embodiment of the present application is shown.
[0161] As Figure 11 shown, the query method of this embodiment includes operation S1110 to operation S1160.
[0162] In operation S1110, S source data are obtained.
[0163] In operation S1120, dimensionality reduction processing is performed on the source data to obtain vector data after dimensionality reduction processing.
[0164] In operation S1130, binarization processing is performed on the vector data after dimensionality reduction processing to obtain an index vector.
[0165] In operation S1140, an index binary tree is constructed based on the index vector, and the mapping relationship is determined.
[0166] In operation S1150, node compression is performed on the index binary tree to obtain an index shape graph.
[0167] In operation S1160, the index shape graph and the mapping relationship are written into the first processor.
[0168] Figure 12 The system interaction diagram of the first processor and the second processor according to another embodiment of the present application is shown.
[0169] According to Figure 11 and Figure 12It can be seen that the second processor can compress S source data in the index library into K index vectors. The compression process may include first performing dimensionality reduction processing on the source data to obtain vector data after dimensionality reduction processing, and then performing binarization processing on the vector data after dimensionality reduction processing to obtain index vectors. The second processor can also construct an index binary tree based on the index vectors, perform node compression on the index binary tree to obtain an index shape graph, and can also generate a mapping relationship between the index vectors and the data list. The index shape graph can be written into the first processor. For example, the index shape graph can be stored in the on-chip storage space.
[0170] In addition, the mapping relationship can be written into the first processor. For example, the mapping relationship can be stored in the off-chip storage space of the first processor.
[0171] Specifically, since the first processor includes multiple groups of pipelined processing units, it has good parallel processing capabilities and relatively small storage resources. Therefore, the first processor can be used to store the index shape graph and the mapping relationship, and the multiple groups of pipelined processing units in the first processor can perform parallel processing to obtain the target index vectors, thereby giving full play to the advantages of the first processor in processing large-scale data parallel query tasks. Since the second processor can efficiently process various types of tasks that require complex logic control, the second processor can be responsible for executing tasks such as compressing query data into simplified vectors, constructing the index graph, and determining the mapping relationship, which require complex logic control, thereby giving full play to the advantages of the second processor in processing various types of computing tasks. Therefore, by combining the first processor and the second processor to complete the query task, the respective advantages of the first processor and the second processor can be fully utilized, and the query efficiency is improved.
[0172] Figure 13 FIG. shows a block diagram of an electronic device suitable for implementing the query method according to an embodiment of the present application.
[0173] As Figure 13 shown, the electronic device 1300 according to an embodiment of the present application includes a processor 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage section 1308 into a random access memory (RAM) 1303. The processor 1301 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1301 may also include on-board memory for caching purposes. The processor 1301 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0174] In the RAM 1303, various programs and data required for the operation of the electronic device 1300 are stored. The processor 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. The processor 1301 performs various operations of the method flow according to the embodiments of the present application by executing programs in the ROM 1302 and / or the RAM 1303. It should be noted that the programs can also be stored in one or more memories other than the ROM 1302 and the RAM 1303. The processor 1301 can also perform various operations of the method flow according to the embodiments of the present application by executing programs stored in one or more memories.
[0175] According to an embodiment of the present application, the electronic device 1300 may further include an input / output (I / O) interface 1305, and the input / output (I / O) interface 1305 is also connected to the bus 1304. The electronic device 1300 may further include one or more of the following components connected to the input / output (I / O) interface 1305: an input portion 1306 including a keyboard, a mouse, etc.; an output portion 1307 including, for example, a cathode ray tube (CRT), a liquid crystal display (KCD), etc., and a speaker, etc.; a storage portion 1308 including a hard disk, etc.; and a communication portion 1309 including a network interface card such as a KAN card, a modem, etc. The communication portion 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the input / output (I / O) interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed so that a computer program read from it can be installed into the storage portion 1308 as needed.
[0176] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0177] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer-readable storage medium may include the above-described ROM 1302 and / or RAM 1303 and / or one or more memories other than ROM 1302 and RAM 1303.
[0178] An embodiment of the present application further includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present application.
[0179] When the computer program is executed by the processor 1301, it executes the above functions defined in the system / apparatus of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0180] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 1309, and / or installed from the removable medium 1311. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0181] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1309, and / or installed from the removable medium 1311. When the computer program is executed by the processor 1301, it executes the above functions defined in the system of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0182] In accordance with embodiments of the present application, program code for executing the computer programs provided by the embodiments of the present application can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0183] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0184] Those skilled in the art can understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0185] The above describes the embodiments of the present application. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.
Claims
1. A query method, characterized in that, The method includes: Receiving a query instruction; Compressing the query data corresponding to the query instruction to obtain a simplified vector; Querying a target index vector matching the simplified vector from an index shape graph; wherein, the index shape graph represents K index vectors, and the K index vectors are obtained by compressing S source data in an index library, S and K are positive integers, and S≥K; Determining at least one primary source data corresponding to the target index vector after compression from the S source data; Obtaining target source data matching the query data according to the similarity between the query data and the at least one primary source data; The method further includes: constructing an index binary tree based on the K index vectors; performing node compression on the index binary tree to obtain the index shape graph; Wherein, each of the K index vectors includes N-dimensional first numerical values, N is a positive integer; the index binary tree is divided into multiple tree nodes at N levels and associated edges between the tree nodes, and the associated edges represent the N-dimensional first numerical values; the index shape graph is divided into multiple shape nodes at N levels and directed edges between the shape nodes, and the directed edges represent the N-dimensional first numerical values; one shape node in the index shape graph represents multiple tree nodes in the index binary tree having the same subtree shape.
2. The method according to claim 1, wherein Compressing the query data corresponding to the query instruction to obtain a simplified vector includes: Performing dimensionality reduction processing on the query data to obtain dimensionality-reduced data; Performing binary encoding processing on the dimensionality-reduced data to obtain the simplified vector.
3. The method according to claim 2, characterized in that, The simplified vector includes a first value and a second value; Performing binary encoding processing on the dimensionality-reduced data includes: Performing numerical calculation on multi-dimensional sub-data included in the dimensionality-reduced data to obtain a reference numerical value; Generating the first value or the second value corresponding to each of the multi-dimensional sub-data based on the numerical deviation between each of the multi-dimensional sub-data and the reference numerical value.
4. The method according to claim 1, wherein Performing node compression on the index binary tree to obtain the index shape graph includes: Determining a shape encoding of the tree node, and the shape encoding is used to characterize the subtree shape category of the tree node; Merging tree nodes having the same shape encoding to generate the index shape graph.
5. The method according to claim 1, characterized in that Multiple source data corresponding to the same index vector after compression correspond to a source data set; the index binary tree includes K leaf nodes corresponding to the K index vectors; The method further includes: Associating the K source data sets corresponding to the K index vectors with the K leaf nodes respectively.
6. The method according to claim 5, wherein In the index shape graph, the target shape node in the target path corresponding to the target index vector carries offset information, and the offset information is used to characterize: in the index binary tree, the position offset of the target leaf node corresponding to the target index vector in the horizontal direction relative to the first leaf node; Determining at least one primary source data corresponding to the target index vector after compression from the S source data includes: During the process of querying the target index vector node by node along the target path, obtaining the offset information carried by the target shape node; Determine a target leaf node from the K leaf nodes based on the offset information; Determine the data in the target source data set associated with the target leaf node as the at least one primary selected source data.
7. The method according to claim 1, characterized in that, The method further includes: After obtaining the index shape graph by performing node compression on the index binary tree, delete the index binary tree.
8. The method according to claim 1, wherein: The method further includes: recording multiple source data records corresponding to the same index vector after compression in a data list, and storing the mapping relationship between the K index vectors and the K data lists; Determining at least one primary selected source data corresponding to the target index vector after compression from the S source data includes: based on the mapping relationship, determining a target data list corresponding to the target index vector from the K data lists.
9. A processor, characterized in that, Includes: An on-chip computing component for executing the query method according to any one of claims 1 to 8; An on-chip storage unit for storing the index shape graph.
10. The processor according to claim 9, wherein: The on-chip computing component includes M groups of pipelined processing units; the on-chip storage unit includes M sub-storage areas communicatively connected to the M groups of pipelined processing units, wherein the N-dimensional first values included in each of the K index vectors corresponding to the index shape graph are stored in N target sub-storage areas among the M sub-storage areas, and M and N are positive integers, and M is greater than or equal to N; N groups of target pipelined processing units among the M groups of pipelined processing units are used to perform the following operations: Read the N-dimensional first values from the N target sub-storage areas in parallel; Match the N-dimensional second values included in the simplified vector with the N-dimensional first values in parallel to obtain a target index vector that matches the simplified vector.
11. A processing system, characterized in that, Includes: A first processor for querying from the index shape graph to obtain a target index vector that matches the simplified vector; wherein the index shape graph represents K index vectors, and the K index vectors are obtained by compressing S source data in an index library, and S and K are positive integers, and S≥K; A second processor for compressing the query data corresponding to the query instruction to obtain the simplified vector, determining at least one primary selected source data corresponding to the target index vector after compression from the S source data, and obtaining target source data that matches the query data according to the similarity between the query data and the at least one primary selected source data; The second processor is further used to construct an index binary tree based on the K index vectors; perform node compression on the index binary tree to obtain the index shape graph; Among them, each of the K index vectors includes N-dimensional first numerical values, where N is a positive integer; the index binary tree is divided into multiple tree nodes at N levels and associated edges between the tree nodes, and the associated edges represent the N-dimensional first numerical values; the index shape graph is divided into multiple shape nodes at N levels and directed edges between the shape nodes, and the directed edges represent the N-dimensional first numerical values; one shape node in the index shape graph represents multiple tree nodes in the index binary tree that have the same subtree shape.
12. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Methods and devices for binary tree construction, compression and lookup
CN102405622A
Vector range retrieval method and device, equipment, medium and program product
CN115129949A