Vector query method, electronic equipment, storage medium and program product
By dimensional reduction and encoding processing of high-dimensional vector data, a lightweight adaptive index structure is built, which solves the problem of low efficiency of high-dimensional vector query and achieves fast and resource-efficient query results.
Patent Information
- Application Number
- CN202510849498.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The prior art is inefficient and consumes a lot of resources in high-dimensional vector data query, making it difficult to meet the efficiency and quality requirements of large-scale data processing.
By dimensionality reduction processing on the original high-dimensional vector data, the target vector data set is generated, and the encoded vector set is generated. An efficient vector index structure is constructed based on the association relationship of the encoded vectors to achieve fast matching and positioning.
It significantly improves the speed and resource utilization of high-dimensional vector queries, reduces calculation and storage overhead, and ensures the accuracy of the query.
Smart Images

Figure CN120371838A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data query, and particularly to a vector query method, an electronic device, a storage medium, and a program product. Background Art
[0002] With the development of artificial intelligence, unstructured data such as texts and images are usually represented as high-dimensional feature vectors. In this context, queries based on vector similarity become the core operations.
[0003] However, in the face of a large amount of high-dimensional vector data, if an exhaustive similarity calculation is performed in the original space, it will not only consume a large amount of computing resources, but also greatly reduce the query efficiency. In addition, storing these original high-dimensional vectors will also bring significant storage overhead. Summary of the Invention
[0004] This application provides a vector query method, an electronic device, a storage medium, and a program product to at least solve the problem of low efficiency and high resource consumption in high-dimensional vector similarity queries in related technologies.
[0005] This application provides a vector query method, including: obtaining an original vector data set; performing dimensionality reduction processing on the original vector data set to generate a target vector data set; performing encoding processing on the target vector data set to generate an encoded vector set; generating a vector index for each encoded vector based on the association relationship of the encoded vectors in the encoded vector set; obtaining a vector to be queried, and determining a target vector list that matches the vector to be queried from the vector array corresponding to the vector index according to the vector index.
[0006] This application also provides a vector query device, including: an obtaining module, configured to obtain an original vector data set; a dimensionality reduction module, configured to perform dimensionality reduction processing on the original vector data set to generate a target vector data set; an encoding module, configured to perform encoding processing on the target vector data set to generate an encoded vector set; a generating module, configured to generate a vector index for each encoded vector based on the association relationship of the encoded vectors in the encoded vector set; a query module, configured to obtain a vector to be queried, and determine a target vector list that matches the vector to be queried from the vector array corresponding to the vector index according to the vector index.
[0007] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above vector query methods when executing the computer program.
[0008] This application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program implements the steps of any of the above vector query methods when executed by a processor.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above vector query methods when executed by a processor.
[0010] Through the present application, since the original high-dimensional vector data is dimension-reduced to reduce the dimensional complexity, then the dimension-reduced data is encoded and compressed to significantly reduce the storage and computing overhead, and further an efficient index structure is constructed based on the correlation relationship between the encoded vectors, thereby avoiding the exhaustive similarity calculation in the original high-dimensional space. Therefore, the technical problem of low efficiency in large-scale high-dimensional vector similarity query and excessive consumption of computing and storage resources can be solved, and the technical effects of significantly improving the query speed and significantly reducing the system resource occupancy while ensuring acceptable query accuracy can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic flowchart of a vector query method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of another vector query method provided by an embodiment of the present application; Figure 3 It is a schematic diagram of classification according to numerical distribution density provided by an embodiment of the present application; Figure 4 It is a schematic flowchart of yet another vector query method provided by an embodiment of the present application; Figure 5 It is a schematic diagram of quantization coding provided by an embodiment of the present application; Figure 6 It is a schematic diagram of binary tree representation construction provided by an embodiment of the present application; Figure 7 It is a schematic diagram of shape coding and shape map of binary tree representation provided by an embodiment of the present application; Figure 8 It is a schematic diagram of 4-equal division processing for the first 2 dimensions provided by an embodiment of the present application; Figure 9 It is a schematic flowchart of still another vector query method provided by an embodiment of the present application; Figure 10 It is a schematic flowchart of yet another vector query method provided by an embodiment of the present application; Figure 11Schematic diagram of the vector index construction process provided by an embodiment of the present application; Figure 12 Schematic diagram of the index query process provided by an embodiment of the present application; Figure 13 Block diagram of the structure of the vector query device provided by an embodiment of the present application; Figure 14 Schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0015] In artificial intelligence applications, unstructured data such as text and images are often converted into high-dimensional and massive vector data. The high real-time processing requirements of such data pose severe challenges to traditional storage and computing models. Vector similarity search (such as based on Euclidean distance or cosine similarity) is a core operation, but building an efficient index faces problems such as uneven data distribution and large reconstruction overhead caused by frequent updates, and there is an urgent need for a better index scheme.
[0016] In related dimensionality reduction techniques, Principal Components Analysis (PCA) has relatively high computational efficiency (time complexity is approximately O(nd²)), but it is difficult to balance the information retention degree after dimensionality reduction (to avoid distortion) and the low-dimensional space size (to adapt to hardware limitations); Linear Discriminant Analysis (LDA) has good efficiency (about O(nd)), but it depends on class labels and has limited applications; while Multidimensional Scaling (MDS) and Isometric Mapping (ISOMAP) can preserve distance or manifold structure, but their computational amount surges with the increase in data volume (time complexity reaches O(n³)), making it difficult to scale to large vector data sets. Therefore, related techniques are difficult to meet the comprehensive requirements of large-scale vector data processing for efficiency and quality.
[0017] In view of this, the technical solution of the present invention collaboratively solves the balance problem of efficiency and quality through phased processing and an associated indexing mechanism. Dimensionality reduction is performed on the original high-dimensional vectors, optimizing the information retention degree and hardware adaptability while avoiding the existing cubic computational complexity; subsequently, the dimensionality-reduced vectors are encoded and compressed to improve storage efficiency; the core lies in constructing a lightweight adaptive index based on the topological relationship between the encoded vectors, which does not require supervised labels and significantly reduces the update and maintenance costs; finally, fast matching and positioning of the vector to be queried are achieved through the index, thereby achieving efficient and stable similarity search in large-scale vector data scenarios and comprehensively overcoming the limitations of related techniques.
[0018] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the following further elaborates on the present application in conjunction with the accompanying drawings and specific embodiments.
[0019] According to an embodiment of the present invention, an embodiment of a vector query method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0020] In this embodiment, a vector query method is provided, which can be used in an electronic device, such as a server, Figure 1 is a flowchart of the vector query method according to an embodiment of the present invention, as Figure 1 shown, and this process includes the following steps.
[0021] Step S101, obtain the original vector data set.
[0022] The original vector dataset refers to a collection of original high-dimensional vectors. Specifically, the original vector dataset is obtained by performing feature extraction on unstructured data (such as text, images, speech, videos, etc.). After these data are transformed by the feature extraction algorithm, a set of high-dimensional vectors is obtained, where each vector represents the features and position of an object in a multi-dimensional space, forming a high-dimensional mathematical vector collection.
[0023] Step S102: Perform dimensionality reduction on the original vector dataset to generate a target vector dataset.
[0024] The target vector dataset refers to a collection of low-dimensional vectors after dimensionality reduction of the original dataset. Specifically, dimensionality reduction maps the high-dimensional vector dataset to a lower-dimensional space by using a preset dimensionality reduction algorithm or method, thereby reducing the complexity of the data while trying to retain the important features in the original data. The data after dimensionality reduction forms the target vector dataset, which is a collection of low-dimensional vectors and can reduce the storage and calculation overhead.
[0025] Step S103: Perform encoding on the target vector dataset to generate an encoded vector set.
[0026] The encoded vector set refers to a collection of binary vectors generated by quantized encoding of the target vector set. Specifically, the encoding process uses a preset quantization technique to convert each vector in the target vector dataset into a binary code. These encoded vectors are represented by binary sequences, thereby achieving data compression, reducing storage space, and facilitating fast retrieval.
[0027] Step S104: Generate a vector index for each encoded vector based on the association relationship between the encoded vectors in the encoded vector set.
[0028] An encoded vector refers to a single vector in the encoded vector set, which is a binary sequence (such as [0, 1, 0, 1]). The association relationship refers to the structural relationship between the encoded vectors. The vector index refers to the compressed index data structure. Specifically, the generation of the vector index is based on the structural relationship between the encoded vectors, and the encoded vectors are stored in an efficient data structure. By establishing the similarity relationship between the vectors, it is possible to more quickly locate similar encoded vectors during the query operation.
[0029] Step S105: Obtain the vector to be queried, and determine a list of target vectors that match the vector to be queried from the vector array corresponding to the vector index according to the vector index.
[0030] The vector to be queried refers to the vector input by the user that needs to retrieve similar items. Specifically, the vector to be queried is homologous to the original vector data, that is, a high-dimensional vector generated from the query object (such as a picture) by the same feature extraction method.
[0031] The target vector list refers to the target vectors that are found to match the query vector from the vector array through the vector indexing mechanism. Specifically, during the query process, according to the index information of the query vector, a set of matching vectors, that is, the target vector list, is read from the vector array corresponding to the vector index.
[0032] The vector query method provided by the embodiments of the present invention significantly reduces the computational complexity through dimensionality reduction, compresses the storage space and accelerates the distance calculation by using encoding, and realizes efficient nearest neighbor search with the help of the index structure based on the association relationship, thereby greatly improving the speed and resource efficiency of retrieving a large number of high-dimensional vectors while ensuring the query accuracy.
[0033] In this embodiment, a vector query method is provided, which can be used in an electronic device, such as a server. Figure 2 It is a flowchart of the vector query method according to the embodiments of the present invention, as Figure 2 shown, and the process includes the following steps.
[0034] Step S201, obtain the original vector data set. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.
[0035] Step S202, perform dimensionality reduction processing on the original vector data set to generate a target vector data set.
[0036] Specifically, the above step S202 includes: Step S2021, obtain the block storage resource capacity and pipeline stage parameters of the programmable logic device.
[0037] The block storage resource capacity refers to the available space size of the block random access memory (Block RAM, BRAM) on the field programmable gate array (FPGA) chip. The pipeline stage parameter refers to the preset FPGA pipeline processing depth threshold. Specifically, the block storage resource capacity of the current FPGA is obtained through the resource analysis report of the FPGA development tool, and at the same time, the preset safe pipeline stage number (SAFE_PIPELINE) and the maximum pipeline stage number (MAX_PIPELINE) are used as the pipeline stage parameters.
[0038] Step S2022, perform dimensionality reduction processing on the original vector data set to obtain the dimensionality reduction parameters corresponding to the original vector data set.
[0039] The dimensionality reduction parameter refers to the intermediate data generated by a preset dimensionality reduction algorithm. Specifically, the principal component analysis (PCA) method or the fast priority dimensionality reduction method can be used to perform dimensionality reduction on the original vector data set. For the principal component analysis (PCA) method, the covariance matrix of the original vector data set is calculated and eigenvalue decomposition is performed to obtain the eigenvectors (principal components) sorted by eigenvalue magnitude. The dimensionality reduction parameter corresponding to the original vector data set is determined based on the eigenvectors and eigenvalues. For the fast priority dimensionality reduction method, a priority index is calculated for each dimension, and the dimension sequence is sorted from high to low according to the priority. This sorting result is the dimensionality reduction parameter without eigenvalue decomposition.
[0040] Step S2023, based on the block storage resource capacity and the pipeline-level parameter, process the dimensionality reduction parameter to obtain the target dimension number.
[0041] The target dimension number refers to the number of dimensions of the final determined low-dimensional space after dimensionality reduction, which is adaptively determined by resource constraints. Specifically, combining the block storage resource capacity of the FPGA, the preset safe pipeline stage number, and the maximum pipeline stage number, process the dimensionality reduction parameter, and the maximum dimension value that meets the resource constraints is used as the target dimension number.
[0042] Step S2024, use the target dimension number to extract data from the original vector data set to obtain the target vector data set.
[0043] If PCA dimensionality reduction is adopted, take the eigenvectors corresponding to the first k (k = target dimension number) principal components to form a projection matrix, and project the original data into the low-dimensional space. If fast dimensionality reduction is adopted, retain the first k dimensions according to the priority and discard the low-priority dimensions. Directly intercept the values of the original vectors on the selected dimensions to form a k-dimensional target vector data set.
[0044] In some alternative embodiments, the above step S2022 includes: Step a1, determine the eigenvalues of the original vector data set.
[0045] The eigenvalue refers to the value representing the data variance size after the covariance matrix decomposition in principal component analysis. Specifically, the principal component analysis (PCA) algorithm is used to calculate the eigenvalues. For the original vector data set containing n d-dimensional vectors perform centering, and the specific formula is as follows:
[0046] Calculate a covariance matrix containing × numerical values .
[0047] Perform eigenvalue decomposition on the covariance matrix to obtain Eigenvalues sorted from largest to smallest and the corresponding eigenvectors of the eigenvalues.
[0048] Step a2: Determine the initial dimension number that meets the preset conditions corresponding to the original vector dataset by using the cumulative distribution of the eigenvalues.
[0049] Among them, the dimensionality reduction parameters include the initial dimension number and the eigenvectors corresponding to the eigenvalues.
[0050] The cumulative distribution refers to the cumulative proportion of the eigenvalues in order. The preset condition refers to the threshold that the cumulative distribution needs to meet. The initial dimension number refers to the dimensionality reduction dimension initially estimated through the cumulative distribution of the eigenvalues. Specifically, the sum of all eigenvalues is Starting from the first eigenvalue and accumulating until the th one, making
[0051] Just meet, where the threshold is a percentage value (such as 90%), and thus the first eigenvalues can be obtained, where is the initial dimension number.
[0052] In some alternative embodiments, the above step S2023 includes: adjusting the initial dimension number based on the block storage resource capacity and the pipeline stage parameters to obtain the target dimension number.
[0053] For an original vector dataset containing n d-dimensional vectors, after quantization coding, vector tree representation, and shape-based compression, the shape node number can be estimated through the vector tree node number and the shape-based compression ratio (i.e., shape node number / vector tree node number), so as to determine whether the entire shape graph can be accommodated in the on-chip storage space (SRAM) of the FPGA. After obtaining the maximum number of pipeline stages MAX_PIPELINE and the safe number of pipeline stages SAFE_PIPELINE designed on the FPGA, the adaptive dimensionality reduction processing program can determine the dimensionality of the vectors after dimensionality reduction, that is, the target dimension number, according to the estimated shape node number, the first eigenvalues, the on-chip storage available space distribution of the FPGA, MAX_PIPELINE, and SAFE_PIPELINE.
[0054] In the above embodiments, the information integrity of the dimension-reduced data is ensured through the cumulative distribution of eigenvalues, and the target dimension is dynamically optimized based on the hardware resources, realizing the collaborative optimization of algorithm accuracy and hardware efficiency. The initial dimension is scientifically determined using the cumulative distribution of eigenvalues to retain the core features of the data, and then the initial dimension is dynamically calibrated in combination with the block storage capacity and pipeline depth of the programmable logic device, enabling the dimension reduction result to strictly match the hardware resource limitations, maximizing the utilization of the hardware parallel computing efficiency while ensuring the information representation ability, and significantly improving the real-time performance and resource utilization rate of subsequent encoding, indexing, and query processes.
[0055] In some alternative embodiments, SAFE_PIPELINE and MAX_PIPELINE can be determined according to the BRAM capacity of the FPGA.
[0056] Assuming that the size of each shape node data structure is S (such as 16 bytes), and the number of all possible shapes of a binary tree with height h is m, it can be obtained that: when h = 2, m = 3; when h = 3, m = 15; when h = 4, m = 255; when h = 5, m = 65535. That is, assuming that the number of all possible shapes of a binary tree with height h is m, then for a binary tree with height h + 1, the number of all possible shapes it has is (m + 1) 2 - 1.
[0057] On the other hand, assuming a binary tree with height 16, that is, max(h) = 16, its root node has only 1, that is, h = 16, and the actual number of shape nodes = 1; when h = 15, there are at most 2 nodes, that is, the corresponding number of shape nodes <= 2; and so on, when h = 14, the actual number of shape nodes <= 4; when h = 13, the actual number of shape nodes <= 8; when h = 12, the actual number of shape nodes <= 16; when h = 11, the actual number of shape nodes <= 32; when h = 10, the actual number of shape nodes <= 64. That is, for a binary tree with height max(h), the number of shape nodes in the layer with height i is at most 2 max(h)-i .
[0058] Through the above, it can be estimated how many shape nodes at most there are in the i-th level of a shape graph with height max(h). By multiplying the number of shape nodes by the size of the node data structure, it can be obtained how much BRAM should be reserved at the i-th level to be safe. When ensuring that a shape graph with height max(h) can be safely written into the BRAM, any value less than or equal to the maximum value of max(h) can be used as the safe pipeline stage number SAFE_PIPELINE. When the height is greater than the maximum value of max(h), there may be a situation where the space of a certain level is insufficient, that is, there is a risk of BRAM space overflow. Generally, it is the middle several layers of the tree that are more difficult to allocate appropriate BRAM.
[0059] The maximum number of pipeline stages MAX_PIPELINE is an empirical value, which has the risk of BRAM overflow and is also related to the BRAM partition layout of each pipeline stage. Specifically, assuming that the minimum block size of BRAM is 1KB and the safe number of pipeline stages is SAFE_PIPELINE, when 1KB of BRAM is allocated to the stage with the corresponding highest height (h = SAFE_PIPELINE + 1), then MAX_PIPELINE <= SAFE_PIPELINE + 6 (that is, when the height is increased by 6 levels, the number of shapes at this stage may fill up 1KB); when 2KB of BRAM is allocated to the stage with the corresponding highest height, then MAX_PIPELINE <= SAFE_PIPELINE + 7; and so on.
[0060] Through the differential resource allocation strategy, while ensuring system security, the utilization rate of hardware resources is maximized: it combines the distribution law of binary tree-shaped nodes (the middle layer is the densest, and the top layer / bottom layer is sparse), breaks through the redundant scheme of allocating BRAM according to the maximum theoretical value in the traditional way, and instead dynamically allocates storage space according to the hierarchical characteristics; at the same time, the maximum number of pipeline stages threshold is set through an empirical formula (such as SAFE_PIPELINE + 6), which not only avoids the risk of BRAM overflow in the middle layer, but also fully exploits the parallel potential of FPGA to achieve the optimal balance between the storage resource efficiency and system stability.
[0061] In the above implementation manner, through the differential resource allocation strategy, while ensuring system security, the utilization rate of hardware resources is maximized. Combining the distribution law of binary tree-shaped nodes (the middle layer is the densest, and the top layer / bottom layer is sparse), breaks through the redundant scheme of allocating BRAM according to the maximum theoretical value in the traditional way, and instead dynamically allocates storage space according to the hierarchical characteristics. At the same time, the maximum number of pipeline stages threshold is set through an empirical formula, which not only avoids the risk of BRAM overflow in the middle layer, but also fully exploits the parallel potential of FPGA to achieve the optimal balance between the storage resource efficiency and system stability.
[0062] In some optional implementation manners, the above step S2022 includes: Step b1, obtaining the priority index of each first dimension in the original vector dataset.
[0063] Step b2, sorting each first dimension according to the priority index to obtain the dimension sorting result corresponding to the original vector dataset.
[0064] Among them, the dimensionality reduction parameter includes the dimension sorting result.
[0065] The first dimension refers to the directions of each coordinate axis in the original vector dataset. The priority index refers to the numerical value quantifying the importance of the dimension, which is jointly determined by the numerical value diversity and the numerical distribution range of the dimension. The dimension sorting result refers to the sequence of the original dimension order arranged from high to low according to the priority index. Specifically, in addition to the above-mentioned principal component analysis (PCA) algorithm, a preset fast dimensionality reduction method can also be used to reduce the dimensionality of the original vector dataset. Sort the input d-dimensional original vector dataset by dimension, and the dimension order is sorted from high to low according to the priority, and the first dimensions ( <= d) of data are intercepted to generate and output a low-dimensional vector set. The considerations for the dimension priority index include the number of different numerical values on each dimension and the relative value range of the numerical values. Among them, the number of different numerical values on each dimension is the main factor, and the relative value range of the numerical values is the secondary factor. Other strategies can also be used for the dimension priority index. For example, on each dimension, classification is performed according to the numerical distribution density (that is, a single numerical value dense area is divided into one class), and the number of classifications is used as the dimension priority, as shown in Figure 3 . For clustering processing on one-dimensional linear data, an adjacent point distance algorithm based on sorting can be used. Its main process includes: first, sort the one-dimensional data by numerical value, second, calculate the distance between adjacent data points after sorting, and finally set a distance threshold (a ratio threshold of the front and back distance changes can also be set). If the distance between adjacent points is greater than , they are divided into different classes; otherwise, they are classified into the same class. The time complexity of the algorithm for processing each dimension is O(nlogn), and the total complexity for processing d-dimensional data is O(dnlogn).
[0066] In addition, the dimension priority index includes: the number of different numerical values on each dimension, denoted as w (w >= 1); and the relative value range of the numerical values, denoted as u (0 <= u <= 1). The dimension priority index q = w + u. When q is small, such as q <= 2, the discrimination degree of the numerical values of this dimension is very low and can basically be ignored. Therefore, a threshold Q can be set, and the dimensions with q < Q can be discarded during the dimensionality reduction process. Q can adopt a dynamically adjusted value. First, calculate the priority index q of all dimensions, sort these q from large to small, obtain the cumulative distribution of q, and satisfy the condition:
[0067] where S( q ) is the sum of all q, is a fixed threshold (such as taking 90%), and Q = .
[0068] In some alternative embodiments, step S2023 includes: based on the block storage resource capacity and pipeline-level parameters, performing dimension truncation on the dimension sorting result to obtain the target number of dimensions.
[0069] During dimensionality reduction processing, according to the limit parameters MAX_PIPELINE and SAFE_PIPELINE, the target number of dimensions needs to satisfy the condition of SAFE_PIPELINE <= <= MAX_PIPELINE. At the same time, when determining the value, the available space on the SRAM corresponding to each stage of the pipeline should be considered to avoid the situation where the SRAM cannot accommodate the newly generated shape nodes. After determining the target number of dimensions , for the original vector data set, according to the sorted dimension sorting result, the first dimensional data is intercepted to generate a new vector list.
[0070] In the above embodiments, the importance of vector dimensions is sorted through priority metrics, and key dimensions are dynamically intercepted based on hardware resource constraints, achieving a dual optimization of high discriminative feature retention and hardware computing efficiency. Sort the original vector features according to the priority metrics of each dimension to ensure that high-value features are preferentially retained, enhancing the representational ability of the data after dimensionality reduction. Based on the block storage capacity and pipeline depth of the programmable logic device, the high-priority dimensions at the head are intercepted from the sorting result as the target dimensions, making the dimensionality reduction result strictly match the hardware storage and parallel computing limitations. Combining the feature importance sorting and the hardware interception mechanism, the discriminative features are maximally retained under limited resources, avoiding low-value dimensions from occupying computing power, and significantly improving the real-time performance of subsequent encoding, indexing, and query processes.
[0071] Step S203, performing encoding processing on the target vector data set to generate an encoded vector set. For details, please refer to Figure 1 step S103 of the illustrated embodiment, which will not be elaborated here.
[0072] Step S204, generating a vector index for each encoded vector based on the association relationship of the encoded vectors in the encoded vector set. For details, please refer to Figure 1 step S104 of the illustrated embodiment, which will not be elaborated here.
[0073] Step S205, obtaining the vector to be queried, and determining the target vector list that matches the vector to be queried from the vector array corresponding to the vector index according to the vector index. For details, please refer to Figure 1 step S105 of the illustrated embodiment, which will not be elaborated here.
[0074] The vector query method provided by the embodiment of the present invention realizes hardware-aware adaptive dimensionality reduction by dynamically adapting to the hardware resource constraints of a programmable logic device. While ensuring that the vector data after dimensionality reduction is completely loaded into the high-speed on-chip memory, it precisely matches the parallel processing capabilities of the pipelined computing architecture, thereby maximizing the utilization of hardware resources, eliminating storage and computing bottlenecks, and significantly improving the real-time performance and end-to-end query efficiency of vector processing.
[0075] In this embodiment, a vector query method is provided, which can be used in an electronic device, such as a server. Figure 4 It is a flowchart of the vector query method according to the embodiment of the present invention, as Figure 4 shown, and the process includes the following steps: Step S301, obtain the original vector data set. For details, please refer to Figure 2 step S201 of the embodiment shown, which will not be elaborated here.
[0076] Step S302, perform dimensionality reduction processing on the original vector data set to generate a target vector data set. For details, please refer to Figure 2 step S202 of the embodiment shown, which will not be elaborated here.
[0077] Step S303, perform encoding processing on the target vector data set to generate an encoded vector set.
[0078] Specifically, the above step S303 includes: Step S3031, determine the target values corresponding to multiple vector data on each second dimension in the target vector data set, and the target value is any one of the median, the average value, and the value corresponding to the bisection of the number of clusters.
[0079] The second dimension refers to the dimension of the vector after dimensionality reduction processing. Vector data refers to a single vector value in the target vector dataset. The median refers to the statistical median of all vector values on a certain second dimension. The average value refers to the arithmetic mean of all vector values on a certain second dimension. The value corresponding to the bisection of the number of clusters refers to dividing the vector values on a certain second dimension into multiple classes through a one-dimensional clustering algorithm, selecting a value to make the number of classes on both sides of the value as balanced as possible (equal or differing by 1), and finally obtaining the optimal division point value. Specifically, for the target vector dataset after dimensionality reduction, calculate the median independently for each dimension. Traverse all vector values on each dimension (the second dimension), sort them, and take the middle value (if the data volume is even, take the average of the two middle numbers). For example, if a dimension contains the values [1.2, 0.5, 3.7], after sorting it is [0.5, 1.2, 3.7], and the median is 1.2. Or, for the target vector dataset after dimensionality reduction, calculate the average value independently for each dimension. Traverse all vector values on each dimension (the second dimension), and calculate the arithmetic mean of all vector values for each dimension. Or, sort the vector values on each dimension (the second dimension), calculate the distance between adjacent data points; then set a distance threshold or a distance change ratio threshold. If the distance between adjacent points exceeds the threshold, divide them into different clusters, otherwise classify them into the same cluster. After division, select the value point that can make the number of clusters as evenly divided as possible (the number of clusters on both sides is the same or differs by 1) as the division point, and on the premise of satisfying the balance of the number of clusters, preferentially select the division position with a more balanced numerical distribution to obtain the value corresponding to the bisection of the number of clusters.
[0080] Step S3032, for any vector data on any second dimension, if the vector data is less than the target value, encode the vector data as the first value; if the vector data is greater than or equal to the target value, encode the vector data as the second value.
[0081] The first value refers to the value encoded when the vector data is less than the target value, for example, it can be 0. The second value refers to the value encoded when the vector data is greater than or equal to the target value, for example, it can be 1. Specifically, for each vector value of each second dimension, if it is less than the target value corresponding to that dimension, encode it as the first value, and if it is greater than or equal to it, encode it as the second value. For example, if the median of a dimension is 1.2, if the vector value is 0.5 (<1.2), it is encoded as 0, and if the vector value is 3.7 (≥1.2), it is encoded as 1.
[0082] Step S3033, combine the first values or second values encoded for the target vector dataset on all second dimensions into an encoded vector set.
[0083] Concatenate the encoding results of each vector in all dimensions into a binary sequence to form an encoding vector set. For example, if the encodings of a 3D vector in three dimensions are [0, 1, 1], then its encoding vector is "011", and the set of such binary sequences of all vectors is the encoding vector set.
[0084] As Figure 5 shown, the target vector dataset V contains 3 four-dimensional vectors. First, calculate the median vector M in all dimensions. Compare each value in the target vector dataset V with the value (i.e., the median) in the corresponding dimension of M. If it is less than the median, the corresponding position is encoded as 0; if it is greater than or equal to the median, the corresponding position is encoded as 1, thus generating a quantized and encoded vector set V' composed of 0s and 1s.
[0085] The vector query method provided by the embodiments of the present invention uses dimension-level thresholds for binary encoding, calculates the target value independently for each dimension as an adaptive encoding threshold, effectively avoiding the problem of mismatch between the preset threshold and the data distribution. At the same time, it converts the floating-point vector into binary values, compresses the storage space, and significantly reduces the storage overhead. The generated binary encoding directly supports fast comparison operations, significantly simplifying the complexity of subsequent similarity calculations. Therefore, through the cooperation of dynamic thresholds and binary conversion, this method systematically optimizes the storage, calculation, and retrieval efficiency while ensuring data discrimination.
[0086] Step S304: Generate vector indexes for each encoding vector based on the association relationship of the encoding vectors in the encoding vector set.
[0087] Specifically, the above step S304 includes: Step S3041: Construct a vector tree structure for each encoding vector based on the association relationship of the encoding vectors in the encoding vector set.
[0088] The vector tree structure refers to the logical representation of a binary tree. Specifically, regard the encoding vector as a path from the root node to the leaf node (e.g., "0" = left branch, "1" = right branch), and recursively insert it into the binary tree. Vectors with the same path prefix share intermediate nodes, and different paths create new branches. The leaf node stores the list of corresponding original vector ID identifiers. For example, the vector with the encoding "011" will create nodes along the path of root node → left child → right child → right child.
[0089] In some alternative embodiments, the above step S3041 includes: Step c1: For any encoding vector in the encoding vector set, use each bit of the encoding vector as a hierarchical node to construct left and right branch paths according to the first value or the second value.
[0090] Step c2: Generate a vector tree structure using the left and right branch paths.
[0091] A hierarchical node refers to the tree node corresponding to each bit (i.e., each dimension) of the encoding vector in a binary tree structure. The left and right branch paths refer to the paths divided in the binary tree according to the encoding value (the first numerical value or the second numerical value), where the first numerical value represents a branch to the left subtree and the second numerical value represents a branch to the right subtree. Specifically, each bit of the encoding vector is used as a hierarchical node of the tree (the root node corresponds to the 1st bit, and the child nodes correspond to the subsequent bits). If this bit is the first numerical value (such as 0), a left branch path is created; if it is the second numerical value (such as 1), a right branch path is created. By recursively traversing each bit of all encoding vectors, the path branches are gradually expanded, and finally a binary tree (vector tree) is formed. Each leaf node represents a unique encoding sequence, and the path from the root to the leaf corresponds to the complete encoding value.
[0092] Step c3: Obtain the list of identifiers of the original vector data associated with each leaf node in the vector tree structure.
[0093] A leaf node refers to the node at the bottommost layer of the binary tree, which has no child nodes. The list of identifiers of the original vector data refers to the set of original vector IDs associated with the leaf node (such as vector unique identifiers). Specifically, when constructing the vector tree, each leaf node is associated with a list that stores all the original vector IDs with the same encoding (i.e., the vectors sharing this path). For example, if three original vectors are encoded as "001", they will be included in the identifier list of the leaf node at the end of this path.
[0094] Step c4: Aggregate the identifier lists associated with each leaf node to generate an array of leaf nodes.
[0095] The array of leaf nodes refers to the array formed by aggregating all the leaf node identifier lists in the preorder traversal order. Specifically, perform a preorder traversal on the vector tree, collect all the leaf node identifier lists in the order of access in sequence, and store them in a consecutive array (array of leaf nodes) in order. The index position of this array corresponds one-to-one with the traversal order of the leaf nodes, which is convenient for quickly locating through the offset later.
[0096] As Figure 6 shown, each vector in the set V' of quantization-encoded encoding vectors can be regarded as a numerical sequence composed of "0" and "1". According to the "0" / "1" numerical sequences corresponding to these vectors, a binary tree is constructed, that is, these vectors are represented by a binary tree. Each edge in the binary tree represents "0" or "1", and the leaf node points to a list, and each element in this list is a vector ID, and this vector is the original vector before dimensionality reduction and quantization encoding. As Figure 6 shown, a data containing 3 d-dimensional (d >= 4) vectors , a 4D data V is obtained through dimensionality reduction processing. V generates data V´ through quantization encoding, and a binary tree representation T is constructed based on 3 vectors composed of "0" / "1" in V´.
[0097] In the above embodiment, a hierarchical tree structure is directly generated based on the bit value characteristics of binary encoding, realizing efficient positioning of candidate vectors and pre-organization of the result set. Each bit of the encoded vector is used as a hierarchical node, and the left and right branch paths are automatically generated according to the first value / second value, and the tree-shaped index can be built without complex calculations. The tree structure can be quickly traversed through bit-by-bit matching to reach the target leaf node, avoiding global scanning. The leaf nodes are directly bound to the original vector identification list, and the candidate data is aggregated in advance. The scattered identification lists are continuously stored as a leaf node array to improve the data loading efficiency.
[0098] Step S3042, perform shape compression processing on the vector tree structure to generate an index shape graph.
[0099] Among them, the vector index includes the index shape graph.
[0100] The index shape graph refers to the compressed binary tree index structure. Specifically, the subtrees with the same topological structure are merged. A unique shape number is assigned to each type of subtree structure, and a shape node (including the shape numbers of the left and right subtrees and the number of leaf nodes information) is established. Finally, the mapping relationship between shape nodes (i.e., the index shape graph) is used to replace the original vector tree nodes, and the vector ID list is stored through the leaf node array. For example, if two subtrees have the same structure, they are mapped to the same shape node, avoiding duplicate storage.
[0101] In some alternative embodiments, the above step S3042 includes: Step d1, identify the vector tree structure to determine each subtree structure in the vector tree structure.
[0102] The subtree structure refers to the local tree structure (including the node and all its descendant nodes) rooted at a certain node in the binary tree. Specifically, by recursively traversing the vector tree, the subtrees with the same topological structure are classified as the same "shape". For example, if two subtrees both have a left branch (0) and a right branch (1), and the branch depth and node connection methods are exactly the same, they are regarded as the same subtree structure, regardless of the specific vector IDs they store.
[0103] Step d2, assign shape numbers to each subtree structure, and create a shape node corresponding to the subtree structure by using the number of leaf nodes and the shape number in the subtree structure.
[0104] The shape number refers to a unique identifier (such as an integer value) assigned to a sub - tree with a unique structure. The number of leaf nodes refers to the total number of leaf nodes contained in the sub - tree. Shape nodes store metadata of the sub - tree structure. Specifically, a separate shape number is assigned to each unique sub - tree structure. Each shape node includes metadata such as the shape number of the left sub - tree, the shape number of the right sub - tree, the total number of leaf nodes contained in the current sub - tree, and the number of leaf nodes contained in the left sub - tree of the current sub - tree.
[0105] Step d3, construct an index shape graph using the reference relationships between each shape node.
[0106] When organizing shape nodes into a graph structure by level, the root shape node represents the root node of the entire tree. Non - leaf shape nodes refer to the shape nodes of the next level through the stored sub - shape numbers, while leaf shape nodes point to the position of the identifier list in the leaf node array. Finally, this organization forms a directed shape graph with references, achieving the compression and sharing of the tree structure. The index shape graph will be loaded into the on - chip memory (SRAM) of the FPGA.
[0107] In addition, a mapping dictionary will be generated to record which shape node each node in the vector tree structure corresponds to, and it can be stored outside the chip.
[0108] As Figure 7 shown, construct a binary tree representation T for the encoded vector set V´, and then compress the binary tree T into a shape graph G. Each shape node in G contains two values. The number on the left represents the shape number, and the value in the parentheses on the right represents the number of leaf nodes contained in the left sub - tree of the corresponding binary tree. Starting from the entry node of the shape graph, initialize a subscript offset of 0 and jump according to the "0" / "1" sequence. When passing through the edge representing "1", the offset is added with the value in the parentheses of the current shape node. When passing through the edge representing "0", the offset remains unchanged; and so on. When reaching the leaf node shape, an offset can be calculated, and through this offset, the corresponding leaf node information (such as the vector list) can be obtained from the leaf node array. The leaf node array is an array constructed by pre - order traversing the binary tree and arranging the accessed leaf nodes in sequence. Each leaf node contains all the vector IDs that match the "0" / "1" sequence corresponding to the leaf node.
[0109] In the above - mentioned embodiment, the extreme compression and efficient retrieval of the tree - shaped index are achieved by identifying and reusing repeated sub - tree structures. Identify sub - tree structures with the same topology and assign unique shape numbers to avoid repeated storage of the same tree shape. Replace the complete sub - tree structure with lightweight shape nodes, significantly reducing memory occupancy. Reference the same sub - structure through the shape number, making the scale of the index graph only depend on the type of topology rather than the total number of nodes. The compressed index shape graph retains the hierarchical reference relationship and maintains the traversal efficiency of the tree structure.
[0110] In some alternative embodiments, step S303 further includes: Step e1, obtaining the target eigenvalue of the first preset number of dimensions of the target vector dataset.
[0111] After dimensionality reduction by principal component analysis (PCA), perform eigenvalue decomposition on the generated covariance matrix, and select the first m (preset value) largest eigenvalues arranged in descending order as the target eigenvalues. If a fast dimensionality reduction method (such as priority sorting) is adopted, then generate a priority index by calculating the numerical distribution characteristics (such as the number of unique values, value range) of each dimension, and select the feature index (such as the number of clusters) corresponding to the first m dimensions with the highest priority as the target eigenvalue.
[0112] Step e2, determining the target quantization strategy corresponding to the target vector dataset based on the distribution state of each target eigenvalue.
[0113] Analyze the relative differences of the first m-dimensional target eigenvalues. Specifically, if the first eigenvalue is significantly greater than the second eigenvalue (such as the ratio exceeds the threshold ), then adopt a fine partitioning method for the first dimension (such as 4 equal parts); if the first two eigenvalues are close and both are large, then adopt a fine partitioning method for the first two dimensions simultaneously. The strategy selection depends on the degree of difference of the eigenvalues, and the goal is to improve the classification granularity by finely partitioning the high-information dimensions.
[0114] Step e3, using the target quantization strategy, perform partitioning processing on the first preset number of dimensions of the target vector dataset to generate multiple subsets corresponding to the target vector dataset.
[0115] Obtain the target eigenvalues (such as the eigenvalues of principal component analysis or dimension priority indicators) of the first m dimensions (such as m = 2). Select the partitioning method (such as 4 equal parts or 8 equal parts) according to the eigenvalue distribution state. Taking 4 equal parts as an example, for each target dimension, calculate the median of the full-dimensional data, divide the data into two parts, calculate the median for each part again, form 3 segmentation points, and divide the data into 4 subintervals (such as [min, Q1), [Q1, median), [median, Q3), [Q3, max]). Each combined interval of the first m dimensions corresponds to a subset (such as 4 equal parts in two dimensions generates = 16 subsets), and each subset contains the data in the original vector that falls into this multi-dimensional interval.
[0116] Step e4, determining the remaining dimensions in each subset except for the first preset number of dimensions.
[0117] Suppose the original vector dimension is d, the first m processed dimensions are used to generate subset partitioning, and the remaining dimension number is d - m. For example, when the first 2 dimensions (m = 2) of the original 128-dimensional vector are segmented, the remaining 126 dimensions are used as the subsequent processing objects.
[0118] Step e5, perform binary quantization encoding on the remaining dimensions of each subset to generate a set of encoded vectors.
[0119] Operate independently on each of the remaining d - m dimensions. Specifically, calculate the median of the dimension within the subset. Traverse all vectors within the subset. If the value of a vector in this dimension is less than the median, encode it as the first value; if it is greater than or equal to the median, encode it as the second value. Each vector generates a binary encoding sequence of length d - m, and all sequences form the set of encoded vectors for the subset.
[0120] Step e6, for the encoded vectors of the remaining dimensions of each subset, use the encoded sequences after binary quantization encoding to construct a target vector tree structure corresponding to the remaining dimensions.
[0121] For the remaining dimensions (d - m dimensions) of each subset, independently calculate the median of each dimension. Convert the remaining dimension values of each vector within the subset to binary. Use the binary encoding sequence as the path to generate tree nodes bit by bit. Each bit serves as a tree level (e.g., the first dimension is the root node, and the second dimension is the second - level node). The first value points to the left subtree, and the second value points to the right subtree. The leaf nodes store the list of original vector IDs corresponding to this path.
[0122] Step e7, perform shape compression processing on the target vector tree structures of each subset to generate a shared index shape graph.
[0123] Traverse the binary trees of all subsets, and extract subtrees with the same topological structure. Assign a globally unique ID to each unique subtree (e.g., Shape_ID = 5), and create shape nodes. Link all shape nodes into a graph according to the reference relationship. The same subtrees in different subsets reuse the same shape node (e.g., the 0 - 1 subtrees of subset 1 and subset 2 are both mapped to Shape_ID = 5), significantly reducing the storage overhead. The shape graph is written into the on - chip storage of the FPGA. The entry shape numbers and leaf node arrays of each subtree are stored in off - chip storage.
[0124] As Figure 8 shown, assume that d eigenvalues from large to small are generated by PCA dimensionality reduction . When (e.g., , that is, the threshold represents that the information amount or difference of the first dimension is much higher than that of the second dimension), then the values of the eigenvectors on the first dimension are divided into four equal parts (for example, first calculate the median of the values of the first - dimension vectors, and then take the median of the two parts of values divided by this median. Through these three medians, the value range of this dimension data can be divided into four equal parts). When performing the above operations, the values of the feature vectors in the second dimension are also equally divided into four parts. Assume that at most the above four-way equal division process is allowed for the first m (e.g., m = 2) dimensions. As Figure 8 shown, when the first two dimensions adopt the four-way equal division process, the vector set can be divided into at most 16 subsets according to the values of the first two dimensions. Then, a binary tree representation and shape-based compression are respectively constructed for each subset, and 1 summary shape graph, 16 shape graph entries, and 16 corresponding leaf node arrays can be obtained. Among them, when the current m-dimensional vectors are processed using a non-binary method (such as four-way or eight-way equal division), the first m dimensions do not participate in the subsequent binary tree and shape graph compression processing. During the process of constructing the vector index, multiple subsets are generated from the original vector set according to the division of the first m dimensions, and starting from the m + 1 dimension, these subsets are quantized and encoded using the "0" / "1" binary method, a binary tree is constructed, and compressed into a shape graph. These subsets share a shape graph, and only the entry shape numbers (i.e., shape graph entries) of each subset need to be saved.
[0125] In the above embodiment, through the feature dimension-guided adaptive data segmentation and hierarchical compression, the collaborative optimization of storage and computing efficiency is achieved. The quantization strategy is dynamically selected based on the distribution of the feature values of the first preset number of dimensions to ensure that the segmentation method adapts to the data characteristics. The remaining dimensions are independently processed for each segmented subset, and structured encoding is generated through binary quantization to maintain the high efficiency of processing each subset. After constructing a vector tree for each subset, the repeated topological structures across subsets are identified through shape compression, and a shared index shape graph is generated. The shared shape graph eliminates redundant storage of subtrees, making the index size depend on the topological types rather than the total amount of data. This method realizes exponential storage compression and cross-subset computing reuse while maintaining the retrieval accuracy through the three-level collaboration of "feature segmentation - subset quantization - topological reuse".
[0126] Step S305: Obtain the vector to be queried, and determine a list of target vectors that match the vector to be queried from the vector array corresponding to the vector index. For details, please refer to Figure 2 step S205 of the embodiment shown, which will not be elaborated here.
[0127] In this embodiment, a vector query method is provided, which can be used in an electronic device, such as a server. Figure 9 is a flowchart of the vector query method according to an embodiment of the present invention. As Figure 9 shown, the process includes the following steps: Step S401: Obtain the original vector data set. For details, please refer to Figure 4 step S301 of the embodiment shown, which will not be elaborated here.
[0128] Step S402: Perform dimensionality reduction processing on the original vector data set to generate a target vector data set. For details, please refer to Figure 4Step S302 of the illustrated embodiment will not be elaborated herein.
[0129] Step S403: Perform encoding processing on the target vector dataset to generate an encoded vector set. For details, please refer to Figure 4 Step S303 of the illustrated embodiment will not be elaborated herein.
[0130] Step S404: Generate vector indexes for each encoded vector based on the association relationship of the encoded vectors in the encoded vector set. For details, please refer to Figure 4 Step S304 of the illustrated embodiment will not be elaborated herein.
[0131] Step S405: Obtain the vector to be queried, and determine a list of target vectors matching the vector to be queried from the vector array corresponding to the vector index according to the vector index.
[0132] Specifically, the above Step S405 includes: Step S4051: Obtain the vector to be queried, perform dimensionality reduction processing and quantization encoding processing on the vector to be queried, and generate a query encoded vector.
[0133] The query encoded vector refers to the encoded vector generated after dimensionality reduction and quantization encoding processing of the vector to be queried. Specifically, the vector to be queried is input by the user (such as the feature vector of the search request). The dimensionality reduction processing reuses the parameters generated in the index construction phase (such as the covariance matrix of PCA or the dimension priority order of fast dimensionality reduction) to project the high-dimensional vector into a low-dimensional space. The quantization encoding reuses the median threshold of each dimension. For each dimension value of the vector after dimensionality reduction, if it is less than the median of this dimension, it is encoded as the first value, otherwise as the second value, to generate a binary sequence. This process is completed by the CPU to ensure consistency with the logic during index construction.
[0134] Step S4052: Use a preset quantization encoding strategy to determine the index entry corresponding to the query encoded vector.
[0135] The preset quantization encoding strategy refers to a preset customized encoding rule. The first m high-information dimensions adopt a non-binary method (such as 4 / 8 equal division), and the subsequent dimensions still use the binary method (0 / 1 encoding). The index entry refers to the starting node number of the subtree in the shape graph. Specifically, if a non-binary method (such as 4 equal division) is used for the first m dimensions during index construction, the subset entry needs to be located according to the values of the first m dimensions of the query encoded vector. For example, assuming that the first 2 dimensions adopt 4 equal division, the encoding combinations of the first 2 dimensions (such as 00, 01, 10, 11) correspond to the entry number of one of the 16 subsets in the shape graph. Through a pre-stored mapping table (such as {dimension combination: entry number}) for quick matching, it is ensured that the query jump starts from the correct subset root node.
[0136] Step S4053: Starting from the index entry as the starting node, perform jump matching along the target encoding sequence of the query encoding vector in the vector index to obtain a matching result.
[0137] The target encoding sequence refers to the binary sequence corresponding to the query encoding vector. The matching result refers to the intermediate data generated when jumping to the leaf node. Specifically, perform step-by-step jumps on the shared shape graph according to the encoding sequence. Starting from the index entry (shape node), the offset = 0. Traverse according to the binary sequence of the query encoding vector (starting from the m+1 dimension): If the bit value is the first value, jump to the left subtree shape node of the current node, and the offset remains unchanged. If the bit value is the second value, jump to the right subtree shape node, and the offset is accumulated by the number of leaf nodes in the left subtree (the value pre-stored in the node). When reaching the leaf shape node, output the final offset as the matching result. Among them, the jump process is hardware-accelerated by the FPGA multi-stage pipeline, and each stage of the pipeline accesses the shape node data stored on the chip.
[0138] Step S4054: Use the matching result to determine the target vector list that matches the vector to be queried.
[0139] Extract the target vector list from the pre-stored leaf node array through the offset. Specifically, the offset obtained by matching points to the position in the leaf node array (this array stores all leaf nodes in the preorder traversal order). The leaf node at this position contains a list of vector IDs, corresponding to all the original vectors that match the query encoding vector path. Output this list as the target vector list for subsequent similarity calculations (such as calculating the Top-k similar vectors).
[0140] In some alternative embodiments, the above step S4054 includes: Step f1: When the current encoding bit of the target encoding sequence is the first value, jump along the index left subtree corresponding to the current index node.
[0141] When the current encoding bit of the query vector is the first value, perform a jump according to the left subtree shape number of the current shape node. Specifically, obtain the left subtree shape number corresponding to the current node, and use this number as the node address for the next-stage pipeline processing. This process only updates the node position without modifying the offset, and essentially descends along the left path of the binary tree.
[0142] Step f2: When the current encoding bit of the target encoding sequence is the second value, jump along the index right subtree corresponding to the current index node, and accumulate the offset of the number of leaf nodes in the index left subtree to obtain the target offset.
[0143] The target offset refers to the accumulated offset value during the jump process, which is used to locate the leaf node array. Specifically, if the current encoding bit is the second value, obtain the right subtree shape number of the current node and jump to the corresponding node. Accumulate the number of leaf nodes in the left subtree of the current node to the total offset. For example, when the root node (shape 1) encounters "1", accumulate the number of its left subtree leaves (value 2), and then jump to the node of shape 3. This design is because the vector array is stored in pre-order traversal, and the positions of the right subtree leaves in the array need to span all the leaves of the left subtree.
[0144] Step f3, when reaching the leaf index node, use the target offset to extract the target vector list from the vector array corresponding to the vector index.
[0145] When jumping to the leaf node (with a special mark for the shape number such as 1), the process terminates. The total offset accumulated at this time directly corresponds to the starting index in the leaf node array. According to this index, extract the list of vector IDs continuously stored in the vector array. For example, the offset of the path "1→0" is 2, that is, read the vector list starting from the 2nd bit of the array.
[0146] The vector query method provided by the embodiment of the present invention realizes the rapid positioning of the query path and the accurate extraction of the result through the collaborative jump mechanism of encoding alignment and tree-shaped indexing. The vector to be queried adopts the same dimensionality reduction and quantization encoding process as that for building the library to ensure the alignment of the encoding rules. Directly determine the index entry node according to the quantization strategy to avoid global search. Dynamically select the left / right subtree path according to each bit value (the first value / the second value) of the query encoding sequence to achieve hierarchical jumps with bit-by-bit matching. During the jump process, accumulate the offset of the number of leaf nodes in the left subtree to convert the tree structure into the physical address of a continuous array. After reaching the leaf node, use the calculated target offset to extract the target vector list at one time from the continuous storage area of the vector array. This method converts the logical traversal of the tree-shaped index into efficient random access of the array, completely avoiding backtracking operations while ensuring the accuracy of the result, and significantly improving the query response speed.
[0147] Below this embodiment, a specific application embodiment will be used to exemplarily illustrate the above vector query method.
[0148] As Figure 10 shown, the vector index construction includes four processing stages, which are: vector dimensionality reduction, quantization encoding, binary tree representation, and shape-based binary tree compression. The vector similarity query includes three stages, which are pre-processing before index query (including query vector dimensionality reduction and quantization), index query based on FPGA (i.e., FPGA multi-stage pipelined matching), and post-processing after index query (including calculating the similarity between the query vector and the vectors in the matched vector list and returning the top (k) vectors).
[0149] For the vector index construction process, asFigure 11 As shown, the vector index construction process includes four strictly sequential CPU processing modules: First, perform adaptive vector dimensionality reduction, compressing high-dimensional data into a low-dimensional space through PCA or a fast priority dimensionality reduction algorithm (sorted according to the dimensional numerical diversity / range of values). The number of dimensions is dynamically determined by the on-chip storage resources (BRAM capacity) of the FPGA and needs to meet the constraints of the preset SAFE_PIPELINE (number of safe pipeline stages, e.g., 16) and MAX_PIPELINE (maximum number of pipeline stages, e.g., 64) to avoid the subsequent shape graph exceeding the storage capacity. Second, execute quantization encoding, converting the vector into a binary sequence (0 / 1) based on the median of each dimension. Then, construct a binary tree representation, mapping the binary sequence to a tree structure (edges represent 0 / 1, and leaf nodes store the original vector IDs). Finally, perform shape-based binary tree compression, abstracting the tree topology into a shape graph (including shape nodes, leaf node arrays, etc.), significantly reducing the storage overhead by sharing subtree shapes. In addition, a customized quantization strategy is supported. If the eigenvalue differences in the first m dimensions are significant (e.g., in PCA ), the dimension can be divided into 4 / 8 equal parts for processing, generating multiple subsets to independently construct subtrees and share the shape graph, and the entry information is stored in off-chip storage. The entire process needs to ensure that the number of dimensions (k) after dimensionality reduction balances the index compression rate and query efficiency - if k is too small, the vector density of leaf nodes increases, increasing the similarity calculation amount, and if k is too large, the shape nodes surge, possibly overflowing the on-chip SRAM of the FPGA (avoided through pre-allocation checking and dimensionality reduction fallback mechanisms).
[0150] For the vector similarity query process, as Figure 12 shown, the vector similarity query is a process that combines CPU and FPGA collaborative processing. The CPU first performs dimensionality reduction and quantization encoding on the input query vector (using specific parameters) and determines the shape graph entry according to the customized quantization encoding strategy (e.g., 4-equal-part mapping based on the first m-dimensional values). Subsequently, the CPU uses the shape graph entry and the binary sequence vector generated after dimensionality reduction and quantization as parameters to call the multi-stage pipelined processing unit of the FPGA vector index query engine to execute the core vector index query operation. After the FPGA engine completes the query, it returns a list of matching vectors to the CPU. The post-processing program on the CPU side is responsible for calculating the actual similarity between each vector in the list and the original query vector and finally screening out the top K vectors with the highest similarity as the query result output.
[0151] The vector query method provided by the embodiments of the present invention designs an adaptive dimensionality reduction method based on PCA, which can dynamically determine the optimal dimensionality of the low-dimensional space according to the cumulative distribution of eigenvalues, preset pipeline-level thresholds (MAX_PIPELINE, SAFE_PIPELINE), and the available SRAM space. At the same time, a fast dimensionality reduction processing technology is proposed, which significantly reduces the computational complexity and latency, and also realizes the adaptive decision of the dimensionality reduction dimension based on the dimension priority and the aforementioned thresholds and SRAM constraints. Secondly, a customized quantization coding strategy is formulated. According to the eigenvalue distribution or priority change of the input vector set in the first few dimensions, it is divided into four equal parts or eight equal parts, so as to achieve a more refined division in the dimensions with large amount of information or differences. Finally, a vector index structure sharing a shape graph is constructed. Based on the above quantization strategy, the vector set is divided into subsets, and a shared shape mapping dictionary is used to index these subsets. This not only greatly improves the compression ratio from binary tree nodes to shape nodes by compressing the tree height and reusing shape nodes, effectively reducing the SRAM storage overhead of the shape graph, but also enables the query vector after dimensionality reduction to be mapped to multiple different entries of the shared shape graph according to the quantization strategy during vector query. This design of further reducing the tree height and the number of pipeline stages on the basis of dimensionality reduction significantly reduces the risk that the shape graph cannot be accommodated by the SRAM.
[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0153] An embodiment of the present application also provides a vector query device, as Figure 13 shown, including: An acquisition module 501, configured to acquire an original vector data set; A dimensionality reduction module 502, configured to perform dimensionality reduction processing on the original vector data set to generate a target vector data set; An encoding module 503, configured to perform encoding processing on the target vector data set to generate an encoded vector set; A generation module 504, configured to generate a vector index of each encoded vector based on the association relationship of each encoded vector in the encoded vector set; A query module 505, configured to acquire a vector to be queried, and determine a target vector list matching the vector to be queried from the vector array corresponding to the vector index according to the vector index.
[0154] In some alternative embodiments, the dimensionality reduction module 502 includes: A first acquisition sub-module, configured to acquire the block storage resource capacity and pipeline-level parameters of a programmable logic device; A dimensionality reduction sub-module, configured to perform dimensionality reduction processing on the original vector dataset to obtain dimensionality reduction parameters corresponding to the original vector dataset; A processing sub-module, configured to process the dimensionality reduction parameters based on the block storage resource capacity and the pipeline-level parameters to obtain the target dimension number; An extraction sub-module, configured to extract data from the original vector dataset by using the target dimension number to obtain a target vector dataset.
[0155] In some alternative embodiments, the dimensionality reduction sub-module includes: A first determination unit, configured to determine the eigenvalues of the original vector dataset; A second determination unit, configured to determine an initial dimension number corresponding to the original vector dataset that satisfies a preset condition by using the cumulative distribution of the eigenvalues; wherein, the dimensionality reduction parameters include the initial dimension number.
[0156] In some alternative embodiments, the processing sub-module includes: An adjustment unit, configured to adjust the initial dimension number based on the block storage resource capacity and the pipeline-level parameters to obtain the target dimension number.
[0157] In some alternative embodiments, the dimensionality reduction sub-module includes: A first acquisition unit, configured to acquire the priority metrics of each first dimension in the original vector dataset; A sorting unit, configured to sort each first dimension according to the priority metrics to obtain a dimension sorting result corresponding to the original vector dataset; wherein, the dimensionality reduction parameters include the dimension sorting result.
[0158] In some alternative embodiments, the processing sub-module includes: An intercepting unit, configured to perform dimension interception on the dimension sorting result based on the block storage resource capacity and the pipeline-level parameters to obtain the target dimension number.
[0159] In some alternative embodiments, the encoding module 503 includes: A first determination sub-module, configured to determine the target values corresponding to multiple vector data on each second dimension in the target vector dataset, and the target value is any one of the median, the average value, and the value corresponding to the bisection of the number of clusters; An encoding sub-module, configured to, for any vector data on any second dimension, if the vector data is less than the target value, encode the vector data as a first value; if the vector data is greater than or equal to the target value, encode the vector data as a second value; A combination sub-module, configured to combine the first values or the second values encoded on all the second dimensions of the target vector dataset into an encoded vector set.
[0160] In some alternative embodiments, the generating module 504 includes: A first constructing sub-module, configured to construct a vector tree structure of each encoding vector based on the association relationship of each encoding vector in the encoding vector set; A first compressing sub-module, configured to perform shape compression processing on the vector tree structure to generate an index shape graph; wherein, the vector index includes the index shape graph.
[0161] In some alternative embodiments, the constructing sub-module includes: A first constructing unit, configured to, for any one encoding vector in the encoding vector set, use each bit of the encoding vector as a hierarchical node to construct left and right branch paths according to a first value or a second value; A generating unit, configured to generate a vector tree structure by using the left and right branch paths.
[0162] In some alternative embodiments, the compressing sub-module includes: An identifying unit, configured to identify the vector tree structure to determine each subtree structure in the vector tree structure; A creating unit, configured to assign a shape number to each subtree structure, and create a shape node corresponding to the subtree structure by using the number of leaf nodes and the shape number in the subtree structure; A second constructing unit, configured to construct an index shape graph by using the reference relationship between each shape node.
[0163] In some alternative embodiments, the encoding module 503 further includes: A second obtaining sub-module, configured to obtain the target feature values of the first preset number of dimensions of the target vector data set; A second determining sub-module, configured to determine a target quantization strategy corresponding to the target vector data set based on the distribution state of each target feature value; A splitting sub-module, configured to perform splitting processing on the first preset number of dimensions of the target vector data set by using the target quantization strategy to generate a plurality of subsets corresponding to the target vector data set; A third determining sub-module, configured to determine the remaining dimensions other than the first preset number of dimensions in each subset; A first generating sub-module, configured to perform binary quantization encoding processing on the remaining dimensions of each subset to generate an encoding vector set.
[0164] In some alternative embodiments, the encoding module 503 further includes: A second constructing sub-module, configured to, for the encoding vectors of the remaining dimensions of each subset, construct a target vector tree structure corresponding to the remaining dimensions by using the encoding sequence after binary quantization encoding; A second compression sub-module, configured to perform shape compression processing on the target vector tree structures of each subset to generate a shared index shape graph.
[0165] In some alternative embodiments, the query module 505 includes: A second generation sub-module, configured to perform dimensionality reduction processing and quantization coding processing on the query vector to be queried to generate a query coding vector; A fourth determination sub-module, configured to determine an index entry corresponding to the query coding vector by using a preset quantization coding strategy; A matching sub-module, configured to start from the index entry as the starting node, perform jump matching along the target coding sequence of the query coding vector in the vector index to obtain a matching result; A fifth determination sub-module, configured to determine a target vector list matching the query vector to be queried by using the matching result.
[0166] In some alternative embodiments, the fifth determination sub-module includes: A first jump unit, configured to jump along the index left subtree corresponding to the current index node when the current coding bit of the target coding sequence is a first value; A second jump unit, configured to jump along the index right subtree corresponding to the current index node when the current coding bit of the target coding sequence is a second value, and accumulate the leaf node number offset of the index left subtree to obtain a target offset; An extraction unit, configured to extract the target vector list from the vector array corresponding to the vector index by using the target offset when reaching the leaf index node.
[0167] For the description of the features in the embodiments corresponding to the vector query device, reference may be made to the relevant description of the embodiments corresponding to the vector query method, which will not be elaborated here one by one.
[0168] An embodiment of the present application further provides an electronic device, as Figure 14 shown, including a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any one of the embodiments of the above vector query method.
[0169] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the embodiments of the above vector query method when running.
[0170] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0171] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described vector query method embodiments.
[0172] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described vector query method embodiments.
[0173] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0174] The above has introduced in detail a vector query method, an electronic device, a storage medium, and a program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A vector query method, characterized in that, Including: Obtain an original vector dataset; Perform dimensionality reduction processing on the original vector dataset to generate a target vector dataset; Perform encoding processing on the target vector dataset to generate an encoded vector set; Generate vector indexes for each of the encoded vectors based on the association relationships of the encoded vectors in the encoded vector set; Obtain a vector to be queried, and determine a target vector list that matches the vector to be queried from the vector array corresponding to the vector index according to the vector index.
2. The vector query method according to claim 1, wherein The performing dimensionality reduction processing on the original vector dataset to generate a target vector dataset includes: Obtain the block storage resource capacity and pipeline stage parameters of a programmable logic device; Perform dimensionality reduction processing on the original vector dataset to obtain dimensionality reduction parameters corresponding to the original vector dataset; Process the dimensionality reduction parameters based on the block storage resource capacity and the pipeline stage parameters to obtain a target dimension number; Extract data from the original vector dataset using the target dimension number to obtain the target vector dataset.
3. The vector query method according to claim 2, wherein The performing dimensionality reduction processing on the original vector dataset to obtain dimensionality reduction parameters corresponding to the original vector dataset includes: Determine the eigenvalues of the original vector dataset; Determine an initial dimension number that meets a preset condition and corresponds to the original vector dataset using the cumulative distribution of the eigenvalues; wherein, the dimensionality reduction parameters include the initial dimension number and the eigenvectors corresponding to the eigenvalues; The processing the dimensionality reduction parameters based on the block storage resource capacity and the pipeline stage parameters to obtain a target dimension number includes: Adjust the initial dimension number based on the block storage resource capacity and the pipeline stage parameters to obtain the target dimension number.
4. The vector query method according to claim 2, wherein The performing dimensionality reduction processing on the original vector dataset to obtain dimensionality reduction parameters corresponding to the original vector dataset includes: Obtain the priority indicators of each first dimension in the original vector dataset; Sort each of the first dimensions according to the priority indicators to obtain a dimension sorting result corresponding to the original vector dataset; wherein, the dimensionality reduction parameters include the dimension sorting result; The processing the dimensionality reduction parameters based on the block storage resource capacity and the pipeline stage parameters to obtain a target dimension number includes: Perform dimension truncation on the dimension sorting result based on the block storage resource capacity and the pipeline stage parameters to obtain the target dimension number.
5. The vector query method according to claim 1, wherein The performing encoding processing on the target vector dataset to generate an encoded vector set includes: Determine target values corresponding to multiple vector data on each second dimension in the target vector dataset, and the target value is any one of the median, average value, and the value corresponding to the bisection of the number of clusters; For any vector data on any one of the second dimensions, if the vector data is less than the target value, encode the vector data as a first value; if the vector data is greater than or equal to the target value, encode the vector data as a second value; Combining the first numerical value or the second numerical value after encoding the target vector dataset in all second dimensions into the encoded vector set.
6. The vector query method according to claim 1 or 5, characterized in that, Generating vector indexes for each of the encoded vectors based on the association relationships of the encoded vectors in the encoded vector set, including: Constructing a vector tree structure for each of the encoded vectors based on the association relationships of the encoded vectors in the encoded vector set; Performing shape compression processing on the vector tree structure to generate an index shape graph; Wherein, the vector index includes the index shape graph.
7. The vector query method according to claim 6, wherein Constructing a vector tree structure for each of the encoded vectors based on the association relationships of the encoded vectors in the encoded vector set, including: For any one encoded vector in the encoded vector set, using each bit of the encoded vector as a hierarchical node to construct left and right branch paths according to the first numerical value or the second numerical value; Generating the vector tree structure using the left and right branch paths.
8. The vector query method according to claim 6, wherein Performing shape compression processing on the vector tree structure to generate an index shape graph, including: Identifying the vector tree structure to determine each subtree structure in the vector tree structure; Assigning a shape number to each of the subtree structures, and creating a shape node corresponding to the subtree structure using the number of leaf nodes and the shape number in the subtree structure; Constructing the index shape graph using the reference relationships between the shape nodes.
9. The vector query method according to claim 1 or 5, characterized in that Further including: Obtaining the target feature values of the first preset number of dimensions of the target vector dataset; Determining a target quantization strategy corresponding to the target vector dataset based on the distribution states of the target feature values; Using the target quantization strategy to perform segmentation processing on the first preset number of dimensions of the target vector dataset to generate a plurality of subsets corresponding to the target vector dataset; Determining the remaining dimensions of each subset except for the first preset number of dimensions; Performing binary quantization encoding processing on the remaining dimensions of each subset to generate the encoded vector set.
10. The vector query method according to claim 9, wherein Further including: For the encoded vectors of the remaining dimensions of each subset, constructing a target vector tree structure corresponding to the remaining dimensions using the encoded sequence after binary quantization encoding; Performing shape compression processing on the target vector tree structures of each subset to generate a shared index shape graph.
11. The vector query method according to claim 1, wherein Determining a target vector list matching the vector to be queried from the vector array corresponding to the vector index according to the vector index, including: Performing dimensionality reduction processing and quantization encoding processing on the vector to be queried to generate a query encoded vector; Determining the index entry corresponding to the query encoded vector using a preset quantization encoding strategy; Starting from the index entry as the starting node, performing jump matching along the target encoded sequence of the query encoded vector in the vector index to obtain a matching result; Determining a target vector list matching the vector to be queried using the matching result.
12. The vector query method according to claim 11, wherein Determining a target vector list matching the vector to be queried using the matching result, including: When the current encoded bit of the target encoded sequence is the first numerical value, jumping along the index left subtree corresponding to the current index node; When the current encoding bit of the target encoding sequence is the second numerical value, jump along the index right subtree corresponding to the current index node, and accumulate the leaf node number offset of the index left subtree to obtain the target offset; When reaching the leaf index node, extract the target vector list from the vector array corresponding to the vector index by using the target offset.
13. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the vector query method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the vector query method according to any one of claims 1 to 12 when executed by a processor.
15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the vector query method according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
High-dimensional vector query method and device, computer equipment and readable storage medium
CN118093633A
Large-scale high-dimensional vector nearest neighbor data retrieval method and device
CN119089005A
Aviation data bus signal encryption method and system based on FPGA
CN119420580A
Query method, processor, processing system, storage medium and program product
CN120067177A
Remapping locality-sensitive hash vectors to compact bit vectors
US20130204905A1