Vector query method, electronic device, storage medium and program product

By performing dimensionality reduction and encoding on high-dimensional vector data and building a lightweight adaptive index, the problem of low query efficiency of high-dimensional vector data is solved, and fast matching and efficient storage are achieved.

CN120371838BActive Publication Date: 2025-09-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510849498.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-05
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In high-dimensional vector data queries, existing technologies have problems such as low query efficiency and high resource consumption, especially in massive data scenarios, where it is difficult to achieve efficient similarity search.

Method used

By performing dimensionality reduction processing on the original high-dimensional vector data, a target vector data set is generated, and then it is encoded to generate a coded vector set. A vector index is constructed based on the association relationship of the coded vectors to achieve fast matching and positioning.

Benefits of technology

It significantly improves the speed and resource efficiency of high-dimensional vector data queries, reduces computing and storage overhead, and ensures query accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371838B_ABST
    Figure CN120371838B_ABST
Patent Text Reader

Abstract

The present application discloses a vector query method, electronic device, storage medium and program product, which relate to the field of data query technology. The method effectively reduces the data dimension by performing dimensionality reduction processing on the original high-dimensional vector data set; performs encoding processing on the target vector data set after dimensionality reduction to generate a compact encoding vector set; and constructs an efficient vector index structure based on the correlation between the encoding vectors. When the vector to be queried is received, the constructed index is used to quickly locate and retrieve the most relevant target vector list. The present application reduces the computational complexity and storage overhead of high-dimensional vector queries by combining dimensionality reduction with encoding, and uses an index structure based on correlation to achieve a fast response of approximate nearest neighbor search, solving the technical problems of low efficiency and excessive resource consumption in large-scale high-dimensional vector data queries, and achieves the technical effect of greatly improving query speed and reducing system resource usage while ensuring query accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data query technology, and in particular to a vector query method, electronic device, storage medium, and program product. Background Art

[0002] With the development of artificial intelligence, unstructured data such as text and images are often represented as high-dimensional feature vectors. In this context, vector similarity-based queries have become a core operation.

[0003] However, when faced with massive amounts of high-dimensional vector data, exhaustive similarity calculations in the original space not only consume a large amount of computing resources but also significantly reduce query efficiency. Furthermore, storing these raw high-dimensional vectors also incurs significant storage overhead. Summary of the Invention

[0004] The present application provides a vector query method, electronic device, storage medium, and program product to at least solve the problems of low efficiency and high resource consumption in high-dimensional vector similarity query in related technologies.

[0005] The present application provides a vector query method, comprising: obtaining an original vector dataset; performing dimensionality reduction processing on the original vector dataset to generate a target vector dataset; performing encoding processing on the target vector dataset to generate an encoding vector set; generating a vector index for each encoding vector based on the association relationship between the encoding vectors in the encoding vector set; obtaining a vector to be queried, and determining, according to the vector index, a list of target vectors that match the vector to be queried from a vector array corresponding to the vector index.

[0006] The present application also provides a vector query device, including: an acquisition module for acquiring an original vector data set; a dimensionality reduction module for performing dimensionality reduction processing on the original vector data set to generate a target vector data set; an encoding module for performing encoding processing on the target vector data set to generate an encoding vector set; a generation module for generating a vector index of each encoding vector based on the association relationship between each encoding vector in the encoding vector set; and a query module for acquiring a vector to be queried, and determining a list of target vectors that match the vector to be queried from a vector array corresponding to the vector index according to the vector index.

[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned vector query methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned vector query methods are implemented.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned vector query methods when executed by a processor.

[0010] Through this application, the original high-dimensional vector data is subjected to dimensionality reduction processing to reduce dimensional complexity, and the reduced-dimensional data is then encoded and compressed to significantly reduce storage and computing overhead, and an efficient index structure is further constructed based on the correlation between the encoded vectors, thereby avoiding exhaustive similarity calculation in the original high-dimensional space. Therefore, the technical problems of low efficiency and excessive consumption of computing and storage resources in large-scale high-dimensional vector similarity queries can be solved, and the technical effect of greatly improving query speed and significantly reducing system resource usage while ensuring acceptable query accuracy can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 A schematic diagram of a flow chart of a vector query method provided in an embodiment of the present application;

[0013] Figure 2 A schematic diagram of a flow chart of another vector query method provided in an embodiment of the present application;

[0014] Figure 3 A schematic diagram of classification according to numerical distribution density provided in an embodiment of the present application;

[0015] Figure 4 A schematic diagram of a flow chart of another vector query method provided in an embodiment of the present application;

[0016] Figure 5 A schematic diagram of quantization coding provided in an embodiment of the present application;

[0017] Figure 6 A schematic diagram of a binary tree representation constructed in accordance with an embodiment of the present application;

[0018] Figure 7 Schematic diagram of shape coding and shape graph represented by a binary tree provided in an embodiment of the present application;

[0019] Figure 8 A schematic diagram of the first two dimensions provided in the embodiment of the present application being divided into four equal parts;

[0020] Figure 9A schematic diagram of a flow chart of another vector query method provided in an embodiment of the present application;

[0021] Figure 10 A schematic diagram of a flow chart of another vector query method provided in an embodiment of the present application;

[0022] Figure 11 A schematic diagram of the vector index construction process provided in an embodiment of the present application;

[0023] Figure 12 A schematic diagram of the index query process provided in an embodiment of the present application;

[0024] Figure 13 A structural block diagram of a vector query device provided in an embodiment of the present application;

[0025] Figure 14 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0028] In AI applications, unstructured data such as text and images is often converted into high-dimensional, massive vector data. The demanding real-time processing of this data poses a significant challenge to traditional storage and computing models. Vector similarity search (such as that based on Euclidean distance or cosine similarity) is a core operation, but building efficient indexes faces challenges such as uneven data distribution and high rebuild overhead due to frequent updates. A better indexing solution is urgently needed.

[0029] Among related dimensionality reduction techniques, principal component analysis (PCA) offers relatively high computational efficiency (with a time complexity of approximately O(nd²)), but struggles to balance information retention after dimensionality reduction (to avoid distortion) with the size of the low-dimensional space (to accommodate hardware limitations). Linear discriminant analysis (LDA) offers good efficiency (approximately O(nd)), but its reliance on class labels limits its application. Multidimensional scaling (MDS) and isometric mapping (ISOMAP), while capable of preserving distance or manifold structure, experience computational complexity that surges with data volume (with a time complexity of up to O(n³)), making them difficult to scale to large vector datasets. Consequently, these techniques struggle to meet the combined efficiency and quality requirements of large-scale vector data processing.

[0030] In view of this, the technical solution of the present invention solves the problem of balancing efficiency and quality through phased processing and an associated indexing mechanism. Dimensionality reduction is performed on the original high-dimensional vectors to optimize information retention and hardware adaptability while avoiding the existing cubic-level computational complexity; the reduced-dimensional vectors are then encoded and compressed to improve storage efficiency; the core lies in constructing a lightweight adaptive index based on the topological relationship between the encoded vectors. This structure does not require supervised labels and significantly reduces update and maintenance costs; ultimately, the index is used to quickly match and locate the vector to be queried, thereby achieving efficient and stable similarity search in large-scale vector data scenarios, fully overcoming the limitations of related technologies.

[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] According to an embodiment of the present invention, an embodiment of a vector query method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0033] In this embodiment, a vector query method is provided, which can be used in electronic devices, such as servers. Figure 1 is a flow chart of a vector query method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps.

[0034] Step S101: Obtain an original vector data set.

[0035] A raw vector dataset refers to a collection of raw high-dimensional vectors. Specifically, a raw vector dataset is obtained by performing feature extraction on unstructured data (such as text, images, speech, and video). After being transformed using a feature extraction algorithm, this data is transformed into a set of high-dimensional vectors. Each vector represents the characteristics and position of an object in multidimensional space, forming a high-dimensional mathematical vector set.

[0036] Step S102: performing dimensionality reduction processing on the original vector dataset to generate a target vector dataset.

[0037] The target vector dataset is the low-dimensional vector set obtained by dimensionality reduction from the original dataset. Specifically, dimensionality reduction uses a pre-defined dimensionality reduction algorithm or method to map a high-dimensional vector dataset to a lower-dimensional space, thereby reducing data complexity while preserving the important features of the original data. The reduced data forms the target vector dataset, a low-dimensional vector set that reduces storage and computational overhead.

[0038] Step S103: Encode the target vector data set to generate an encoded vector set.

[0039] A coded vector set is a set of binary vectors generated by quantizing and encoding a target vector set. Specifically, the encoding process uses a preset quantization technique to convert each vector in the target vector set into a binary code. These coded vectors are represented as binary sequences, which compresses data, reduces storage space, and facilitates fast retrieval.

[0040] Step S104: Generate a vector index of each coding vector based on the association relationship between each coding vector in the coding vector set.

[0041] An encoding vector is a single vector in a set of encoding vectors, a binary sequence (e.g., [0, 1, 0, 1]). An association relationship refers to the structural relationship between encoding vectors. A vector index is a compressed index data structure. Specifically, vector indexes are generated based on the structural relationship between encoding vectors, storing them in an efficient data structure. By establishing similarity relationships between vectors, queries can more quickly locate similar encoding vectors.

[0042] Step S105 , obtaining a vector to be queried, and determining a target vector list matching the vector to be queried from a vector array corresponding to the vector index according to the vector index.

[0043] The query vector is the vector entered by the user for which similar items need to be retrieved. Specifically, the query vector and the original vector data are derived from the same source: a high-dimensional vector generated from the query object (e.g., an image) using the same feature extraction method.

[0044] The target vector list refers to the target vectors that match the query vector, found from the vector array using the vector indexing mechanism. Specifically, during the query process, the target vector list is retrieved from the vector array corresponding to the query vector's index.

[0045] The vector query method provided by the embodiments of the present invention significantly reduces computational complexity through dimensionality reduction, uses encoding to compress storage space and accelerate distance calculations, and implements efficient neighbor search with the help of an index structure based on association relationships. This significantly improves the speed and resource efficiency of searching massive high-dimensional vectors while ensuring query accuracy.

[0046] In this embodiment, a vector query method is provided, which can be used in electronic devices, such as servers. Figure 2 is a flow chart of a vector query method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps.

[0047] Step S201: Get the original vector dataset. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0048] Step S202: Perform dimensionality reduction processing on the original vector dataset to generate a target vector dataset.

[0049] Specifically, the above step S202 includes:

[0050] Step S2021: Obtain the block storage resource capacity and pipeline level parameters of the programmable logic device.

[0051] Block storage resource capacity refers to the available space in the block random access memory (BRAM) on the Field Programmable Gate Array (FPGA) chip. Pipeline level parameters refer to the preset FPGA pipeline processing depth threshold. Specifically, the current FPGA block storage resource capacity is obtained from the resource analysis report of the FPGA development tool. The preset safe pipeline stage number (SAFE_PIPELINE) and maximum pipeline stage number (MAX_PIPELINE) are used as pipeline level parameters.

[0052] Step S2022: Perform dimensionality reduction processing on the original vector data set to obtain dimensionality reduction parameters corresponding to the original vector data set.

[0053] Dimensionality reduction parameters refer to the intermediate data generated by a preset dimensionality reduction algorithm. Specifically, principal component analysis (PCA) or fast priority dimensionality reduction can be used to reduce the dimensionality of the original vector dataset. For PCA, the covariance matrix of the original vector dataset is calculated and eigenvalue decomposition is performed to obtain eigenvectors (principal components) sorted by eigenvalue size. Based on the eigenvectors and eigenvalues, the dimensionality reduction parameters corresponding to the original vector dataset are determined. For fast priority dimensionality reduction, a priority index is calculated for each dimension, and the dimension sequence is sorted from high to low priority. This sorting result is the dimensionality reduction parameter, without the need for eigendecomposition.

[0054] Step S2023: Process the dimensionality reduction parameters based on the block storage resource capacity and the pipeline level parameters to obtain the target number of dimensions.

[0055] The target number of dimensions refers to the final number of low-dimensional space dimensions after dimensionality reduction, which is adaptively determined by resource constraints. Specifically, the dimensionality reduction parameters are processed based on the FPGA's block storage resource capacity, the preset safe number of pipeline stages, and the maximum number of pipeline stages. The target number of dimensions is the maximum value that meets the resource constraints.

[0056] Step S2024: extract data from the original vector dataset using the target number of dimensions to obtain a target vector dataset.

[0057] If PCA is used for dimensionality reduction, the eigenvectors corresponding to the first k principal components (k = the number of target dimensions) are taken to form a projection matrix, which projects the original data into a lower-dimensional space. If fast dimensionality reduction is used, the first k dimensions are retained based on priority, while lower-priority dimensions are discarded. The values ​​of the original vectors along the selected dimensions are directly truncated to form a k-dimensional target vector dataset.

[0058] In some optional implementations, the above step S2022 includes:

[0059] Step a1: determine the eigenvalues ​​of the original vector data set.

[0060] Eigenvalue refers to the value that represents the variance of the data after the covariance matrix is ​​decomposed in principal component analysis. Specifically, the principal component analysis (PCA) algorithm is used to calculate the eigenvalue. Centralization is performed, and the specific formula is as follows:

[0061]

[0062] Calculate a × The covariance matrix of the values .

[0063] Pair covariance matrix Perform eigenvalue decomposition and obtain eigenvalues ​​sorted from largest to smallest , and the eigenvectors corresponding to the eigenvalues.

[0064] Step a2: using the cumulative distribution of eigenvalues, determine the number of initial dimensions that meet preset conditions and correspond to the original vector data set.

[0065] The dimensionality reduction parameters include the number of initial dimensions and the eigenvectors corresponding to the eigenvalues.

[0066] Cumulative distribution refers to the cumulative proportion of eigenvalues ​​in order. The preset condition refers to the threshold that the cumulative distribution must meet. The initial number of dimensions refers to the number of dimensions reduced by the initial estimate of the cumulative distribution of eigenvalues. Specifically, the sum of all eigenvalues ​​is , starting from the first eigenvalue and accumulating until the individual, make

[0067]

[0068] Just meet the threshold is a percentage value (for example, 90%), from which we can get the previous eigenvalues ,in is the initial number of dimensions.

[0069] In some optional implementations, the above step S2023 includes: adjusting the initial number of dimensions based on the block storage resource capacity and the pipeline level parameters to obtain the target number of dimensions.

[0070] For an original vector data set containing n d-dimensional vectors, after quantization encoding, vector tree representation and shape-based compression, the number of shape nodes can be estimated by the number of vector tree nodes and the shape-based compression ratio (i.e., the number of shape nodes / the number of vector tree nodes), thereby determining whether the entire shape graph can be accommodated in the FPGA on-chip storage space (SRAM). After obtaining the maximum number of pipeline stages MAX_PIPELINE and the safe number of pipeline stages SAFE_PIPELINE designed on the FPGA, the adaptive dimensionality reduction processing program can be used to calculate the number of shape nodes and the number of previous stages. The eigenvalues, the distribution of available on-chip storage space of the FPGA, MAX_PIPELINE, and SAFE_PIPELINE determine the dimension of the vector after dimensionality reduction, that is, the target number of dimensions.

[0071] In the above implementation, the information integrity of the dimensionality reduction data is guaranteed through the cumulative distribution of eigenvalues, and the target dimension is dynamically optimized based on hardware resources, achieving a coordinated optimization of algorithm accuracy and hardware efficiency. The initial dimension is scientifically determined using the cumulative distribution of eigenvalues ​​to retain the core characteristics of the data. The initial dimension is then dynamically calibrated in combination with the block storage capacity and pipeline depth of the programmable logic device, ensuring that the dimensionality reduction results strictly match the hardware resource constraints. This maximizes the use of hardware parallel computing efficiency while ensuring information representation capabilities, significantly improving the real-time performance and resource utilization of subsequent encoding, indexing, and query processes.

[0072] In some optional implementations, SAFE_PIPELINE and MAX_PIPELINE may be determined based on the BRAM capacity of the FPGA.

[0073] Assuming that the size of each shape node data structure is S (e.g. 16 bytes), and the number of all possible shapes of a binary tree of height h is m, we can conclude that: when h=2, m=3; when h=3, m=15; when h=4, m=255; when h=5, m=65535. In other words, assuming that the number of all possible shapes of a binary tree of height h is m, then the number of all possible shapes of a binary tree of height h+1 is (m+1). 2 -1.

[0074] On the other hand, suppose a binary tree with a height of 16, that is, max(h)=16, has only one root node, that is, h=16, the actual number of shape nodes = 1; when h=15, there are at most 2 nodes, that is, the corresponding number of shape nodes <= 2; and so on, when h=14, the actual number of shape nodes <= 4; when h=13, the actual number of shape nodes <= 8; when h=12, the actual number of shape nodes <= 16; when h=11, the actual number of shape nodes <= 32; when h=10, the actual number of shape nodes <= 64. That is, for a binary tree with a height of max(h), the number of shape nodes at the layer with height i is at most 2 max(h)-i .

[0075] The above method estimates the maximum number of shape nodes at level i for a shape graph of height max(h). By multiplying the number of shape nodes by the size of the node data structure, we can determine the safe BRAM reserve for level i. To ensure that a shape graph of height max(h) can be safely written to BRAM, values ​​less than or equal to max(h) can be used as the safe pipeline level SAFE_PIPELINE. However, if the height exceeds max(h), insufficient space may exist at a certain level, creating the risk of BRAM overflow. This is typically the case in the middle layers of the tree, where appropriate BRAM allocation is more difficult.

[0076] The maximum number of pipeline stages, MAX_PIPELINE, is an empirical value. There's a risk of BRAM overflow, and it's also related to the BRAM allocation at each pipeline stage. Specifically, assuming the minimum BRAM block size is 1KB and the safe number of pipeline stages is SAFE_PIPELINE, if the highest level (h = SAFE_PIPELINE + 1) is allocated 1KB of BRAM, then MAX_PIPELINE <= SAFE_PIPELINE + 6 (that is, if the height is increased by 6, the number of shapes at that level may reach 1KB). If the highest level is allocated 2KB of BRAM, then MAX_PIPELINE <= SAFE_PIPELINE + 7, and so on.

[0077] Through differentiated resource allocation strategies, hardware resource utilization is maximized while ensuring system security. This strategy incorporates the distribution pattern of binary tree-shaped nodes (dense in the middle layer and sparse in the top / bottom layers), breaking away from the traditional redundant BRAM allocation scheme based on maximum theoretical values ​​and instead dynamically allocating storage space based on hierarchical characteristics. Furthermore, by setting the maximum pipeline stage threshold through empirical formulas (such as SAFE_PIPELINE+6), this strategy avoids the risk of BRAM overflow in the middle layer while fully tapping the parallel potential of the FPGA, achieving an optimal balance between storage resource efficiency and system stability.

[0078] In this implementation, a differentiated resource allocation strategy is employed to maximize hardware resource utilization while ensuring system security. By combining the node distribution pattern of a binary tree (dense in the middle layer and sparse in the top and bottom layers), this approach breaks away from the traditional BRAM redundancy scheme of allocating BRAM based on maximum theoretical values ​​and instead dynamically allocates storage space based on hierarchical characteristics. Furthermore, an empirical formula is used to set the maximum pipeline stage threshold, mitigating the risk of BRAM overflow in the middle layer while fully exploiting the parallel potential of the FPGA, achieving an optimal balance between storage resource efficiency and system stability.

[0079] In some optional implementations, the above step S2022 includes:

[0080] Step b1: Obtain the priority index of each first dimension in the original vector dataset.

[0081] Step b2: sort each first dimension according to the priority index to obtain the dimension sorting result corresponding to the original vector data set.

[0082] The dimensionality reduction parameters include dimension sorting results.

[0083] The first dimension refers to the directions of each coordinate axis in the original vector dataset. The priority index refers to the value quantifying the importance of the dimension, which is jointly determined by the numerical diversity of the dimension and the range of numerical distribution. The dimension sorting result refers to the sequence of the original dimension order arranged from high to low according to the priority index. Specifically, in addition to the above-mentioned principal component analysis (PCA) algorithm, a preset fast dimensionality reduction method can also be used to reduce the dimensionality of the original vector dataset. Sort the input d-dimensional original vector dataset by dimension, and the dimension order is sorted from high to low according to the priority, and the first dimensions ( <= d) of the data are intercepted to generate and output a low-dimensional vector set. The considerations for the dimension priority index include the number of different values on each dimension and the relative value range of the values. Among them, the number of different values on each dimension is the main factor, and the relative value range of the values is the secondary factor. Other strategies can also be used for the dimension priority index. For example, classify according to the numerical distribution density on each dimension (that is, divide a single numerical dense area into 1 class), and use the number of classifications as the dimension priority, as shown in Figure 3 . For clustering processing on one-dimensional linear data, an adjacent point distance algorithm based on sorting can be used. Its main process includes: first, sort the one-dimensional data by numerical size, then calculate the distance between adjacent data points after sorting, and finally set a distance threshold (a threshold for the change ratio of the front and back distances can also be set). If the distance between adjacent points is greater than , they are divided into different classes; otherwise, they are classified into the same class. The time complexity of the algorithm for processing each dimension is O(nlogn), and the total complexity for processing d-dimensional data is O(dnlogn).

[0084] In addition, the dimension priority index includes: the number of different values on each dimension, denoted as w (w >= 1); and the relative value range of the values, denoted as u (0 <= u <= 1). The dimension priority index q = w + u. When q is small, such as q <= 2, the discrimination degree of the values on this dimension is very low and can basically be ignored. Therefore, a threshold Q can be set, and the dimensions with q < Q can be discarded during the dimensionality reduction process. Q can adopt a dynamically adjusted value. First, calculate the priority index q of all dimensions, sort these q from large to small, obtain the cumulative distribution of q, and satisfy the condition:

[0085]

[0086] Among them, S( q ) is the sum of all q, is a fixed threshold (such as taking 90%), and Q = .

[0087] In some optional implementations, the above step S2023 includes: performing dimension truncation on the dimension sorting result based on the block storage resource capacity and the pipeline level parameter to obtain the target number of dimensions.

[0088] During dimensionality reduction, the number of target dimensions is determined by the limit parameters MAX_PIPELINE and SAFE_PIPELINE. SAFE_PIPELINE<= <= MAX_PIPELINE, and also in determining When selecting the value, the available space on the SRAM corresponding to each level of pipeline should be considered to avoid the situation where the SRAM cannot accommodate the newly generated shape nodes. After that, sort the original vector dataset according to the sorted dimensions and intercept the dimensional data, generating a new vector list.

[0089] In the above implementation, the importance of vector dimensions is sorted by priority indicators, and key dimensions are dynamically intercepted based on hardware resource constraints to achieve dual optimization of high-discriminative feature retention and hardware computing efficiency. The original vector features are sorted according to the priority indicators of each dimension to ensure that high-value features are retained first and enhance the characterization ability of the data after dimensionality reduction. Based on the block storage capacity and pipeline depth of the programmable logic device, the head high-priority dimension is intercepted from the sorting result as the target dimension, so that the dimensionality reduction result strictly matches the hardware storage and parallel computing limitations. Combined with the feature importance sorting and hardware interception mechanism, the discriminative features are retained to the maximum extent under limited resources, and the low-value dimensions are avoided from occupying computing power, which significantly improves the real-time performance of subsequent encoding, indexing and query processes.

[0090] Step S203: Encode the target vector data set to generate an encoded vector set. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0091] Step S204: Generate a vector index for each code vector based on the correlation between the code vectors in the code vector set. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0092] Step S205: Get the query vector and determine the target vector list that matches the query vector from the vector array corresponding to the vector index according to the vector index. Figure 1 Step S105 of the illustrated embodiment will not be described in detail here.

[0093] The vector query method provided by the embodiment of the present invention realizes hardware-aware adaptive dimensionality reduction by dynamically adapting to the hardware resource constraints of programmable logic devices. While ensuring that the vector data after dimensionality reduction is completely loaded into high-speed on-chip storage, it accurately matches the parallel processing capabilities of the pipeline computing architecture, thereby maximizing the utilization of hardware resources, eliminating storage and computing bottlenecks, and significantly improving the real-time performance of vector processing and end-to-end query efficiency.

[0094] In this embodiment, a vector query method is provided, which can be used in electronic devices, such as servers. Figure 4 is a flow chart of a vector query method according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:

[0095] Step S301: Obtain the original vector dataset. Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0096] Step S302: Perform dimensionality reduction on the original vector dataset to generate a target vector dataset. Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.

[0097] Step S303: Encode the target vector data set to generate an encoded vector set.

[0098] Specifically, the above step S303 includes:

[0099] Step S3031, determining target values ​​corresponding to multiple vector data on each second dimension in the target vector data set, where the target value is any one of the median, the average, and the value corresponding to the bisection of the number of clusters.

[0100] The second dimension refers to the vector dimension after dimensionality reduction. Vector data refers to individual vector values ​​in the target vector dataset. The median is the statistical median of all vector values ​​along a particular second dimension. The mean is the arithmetic mean of all vector values ​​along a particular second dimension. The value corresponding to the bisection of the number of clusters is the value obtained by using a one-dimensional clustering algorithm to divide the vector values ​​along a particular second dimension into multiple clusters, selecting a value that ensures the number of clusters on both sides of the value is as balanced as possible (equal or differing by one). Specifically, for the target vector dataset after dimensionality reduction, the median is calculated independently for each dimension. All vector values ​​along each dimension (the second dimension) are traversed, sorted, and the median value is taken (if the number of vector values ​​is even, the median of the two middle values ​​is taken). For example, if a dimension contains the values ​​[1.2, 0.5, 3.7], after sorting, they become [0.5, 1.2, 3.7], with a median of 1.2. Alternatively, for the target vector dataset after dimensionality reduction, the mean is calculated independently for each dimension. Traverse all vector values ​​along each dimension (the second dimension) and calculate the arithmetic mean of all vector values ​​along each dimension. Alternatively, sort the vector values ​​along each dimension (the second dimension) and calculate the distance between adjacent data points. Then, set a distance threshold or distance change ratio threshold. If the distance between adjacent points exceeds the threshold, they are divided into different clusters; otherwise, they are grouped together. After the division, select the value point that achieves the most balanced number of clusters (the number of clusters on both sides is the same or differs by one) as the split point. Preferentially selecting the split position with a more balanced distribution of values, provided that the number of clusters is balanced, to obtain the value corresponding to the bisection of the number of clusters.

[0101] Step S3032: for any vector data on any second dimension, if the vector data is less than the target value, the vector data is encoded as the first value; if the vector data is greater than or equal to the target value, the vector data is encoded as the second value.

[0102] The first value is the value encoded when the vector data is less than the target value, for example, it can be 0. The second value is the value encoded when the vector data is greater than or equal to the target value, for example, it can be 1. Specifically, for each vector value in each second dimension, if it is less than the target value corresponding to that dimension, it is encoded as the first value, and if it is greater than or equal to the target value, it is encoded as the second value. For example, if the median of a dimension is 1.2, a vector value of 0.5 (<1.2) is encoded as 0, and a vector value of 3.7 (≥1.2) is encoded as 1.

[0103] Step S3033: Combining the first values ​​or the second values ​​of the target vector data set after encoding on all second dimensions into a coding vector set.

[0104] The encoding results of each vector in all dimensions are concatenated into a binary sequence to form a coded vector set. For example, if a 3D vector is encoded as [0, 1, 1] in each of the three dimensions, its coded vector is "011". The set of such binary sequences for all vectors is the coded vector set.

[0105] like Figure 5 As shown in the figure, the target vector dataset V contains three 4-dimensional vectors. First, the median vector M of all dimensions is calculated. Each value in the target vector dataset V is compared with the value of the corresponding dimension in M ​​(i.e., the median). If the value is less than the median, the corresponding position is encoded as 0, and if it is greater than or equal to the median, the corresponding position is encoded as 1. This generates a quantized encoded encoding vector set V´ consisting of 0 and 1.

[0106] The vector query method provided by an embodiment of the present invention uses dimension-level thresholds for binary encoding. By independently calculating a target value for each dimension as an adaptive encoding threshold, this method effectively avoids the mismatch between the preset threshold and the data distribution. Floating-point vectors are simultaneously converted into binary values, compressing storage space and significantly reducing storage overhead. The generated binary encoding directly supports fast comparison operations, significantly simplifying the complexity of subsequent similarity calculations. Therefore, by synergizing dynamic thresholds with binary conversion, this method systematically optimizes storage, computation, and retrieval efficiency while ensuring data differentiation.

[0107] Step S304: Generate a vector index of each coding vector based on the association relationship between each coding vector in the coding vector set.

[0108] Specifically, the above step S304 includes:

[0109] Step S3041: construct a vector tree structure of each coding vector based on the association relationship between each coding vector in the coding vector set.

[0110] The vector tree structure is a logical representation of a binary tree. Specifically, the encoded vector is treated as a path from the root node to the leaf node (e.g., "0" = left branch, "1" = right branch), and is recursively inserted into the binary tree. Vectors with the same path prefix share intermediate nodes, while different paths create new branches. Leaf nodes store a list of the corresponding original vector IDs. For example, a vector encoding "011" would create nodes along the path from root node to left child, then right child, and finally right child.

[0111] In some optional implementations, step S3041 includes:

[0112] Step c1: for any coding vector in the coding vector set, taking each bit of the coding vector as a hierarchical node, constructing a left or right branch path according to the first value or the second value.

[0113] Step c2: Generate a vector tree structure using the left and right branch paths.

[0114] A hierarchical node refers to a tree node corresponding to each bit (i.e., each dimension) of the encoding vector in a binary tree structure. Left and right branch paths refer to paths in the binary tree divided according to the encoding value (first value or second value). The first value represents a branch to the left subtree, and the second value represents a branch to the right subtree. Specifically, each bit of the encoding vector is used as a hierarchical node of the tree (the root node corresponds to the first bit, and the child nodes correspond to subsequent bits). If the bit is the first value (such as 0), a left branch path is created; if it is the second value (such as 1), a right branch path is created. By recursively traversing each bit of all encoding vectors, the path branches are gradually expanded, and finally a binary tree (vector tree) is formed. Each leaf node represents a unique encoding sequence, and the path from root to leaf corresponds to the complete encoding value.

[0115] Step c3: Obtain a list of identifiers of original vector data associated with each leaf node in the vector tree structure.

[0116] A leaf node is the bottom-most node in a binary tree and has no child nodes. The raw vector data identification list refers to the set of raw vector IDs (e.g., vector unique identifiers) associated with the leaf node. Specifically, when constructing the vector tree, each leaf node is associated with a list storing all raw vector IDs with the same encoding (i.e., vectors sharing the same path). For example, if three raw vectors are all encoded as "001," they will be included in the identification list of the leaf node at the end of the path.

[0117] Step c4: Aggregate the identifier lists associated with each leaf node to generate a leaf node array.

[0118] The leaf node array is an array formed by aggregating the identifier lists of all leaf nodes in pre-order traversal order. Specifically, a pre-order traversal of the vector tree collects the identifier lists of all leaf nodes in the order they were accessed and stores them sequentially in a contiguous array (the leaf node array). The index positions in this array correspond one-to-one with the order in which the leaf nodes were traversed, facilitating quick location via offsets.

[0119] like Figure 6 As shown, each vector in the quantized coded vector set V' can be regarded as a numerical sequence composed of "0" and "1". According to the "0" / "1" numerical sequences corresponding to these vectors, a binary tree is constructed, that is, these vectors are represented by a binary tree. Each edge in the binary tree represents "0" or "1", and the leaf node points to a list. Each element in the list is a vector ID, which is the original vector before dimensionality reduction and quantization coding. Figure 6 As shown, a data set containing 3 d-dimensional (d>=4) vectors After dimensionality reduction, a 4-dimensional data V is obtained. V is quantized and encoded to generate data V´. The binary tree representation T is constructed based on the three vectors composed of "0" / "1" in V´.

[0120] In the above implementation, a hierarchical tree structure is directly generated based on the bit-value characteristics of binary encoding, enabling efficient candidate vector location and result set pre-organization. Each bit of the encoded vector is used as a hierarchical node, and left and right branch paths are automatically generated based on the first and second values, completing the tree index without complex calculations. Bit-by-bit matching allows for rapid traversal of the tree structure, directly reaching the target leaf node, avoiding global scans. Leaf nodes are directly bound to the original vector identifier list, allowing for pre-aggregation of candidate data. Distributed identifier lists are stored contiguously as leaf node arrays, improving data loading efficiency.

[0121] Step S3042: perform shape compression processing on the vector tree structure to generate an index shape graph.

[0122] The vector index includes an index shape map.

[0123] The indexed shape graph is a compressed binary tree index structure. Specifically, subtrees with the same topology are merged. Each subtree is assigned a unique shape number, and a shape node (containing the shape numbers of the left and right subtrees and the number of leaf nodes) is created. Finally, the mapping between shape nodes (i.e., the indexed shape graph) replaces the original vector tree nodes, and the leaf node array stores the vector ID list. For example, if two subtrees have the same structure, they are mapped to the same shape node to avoid duplicate storage.

[0124] In some optional implementations, step S3042 includes:

[0125] Step d1: Identify the vector tree structure and determine each subtree structure in the vector tree structure.

[0126] A subtree structure refers to the local tree structure rooted at a node in a binary tree (including that node and all its descendant nodes). Specifically, by recursively traversing the vector tree, subtrees with the same topological structure are classified as having the same "shape." For example, if two subtrees both have a left branch (0) and a right branch (1), and their branch depths and node connections are identical, they are considered to have the same subtree structure, regardless of whether the specific vector IDs they store are the same.

[0127] Step d2: assigning a shape number to each subtree structure, and creating a shape node corresponding to the subtree structure using the number of leaf nodes in the subtree structure and the shape number.

[0128] A shape number is a unique identifier (e.g., an integer value) assigned to a subtree with a unique structure. The leaf node count refers to the total number of leaf nodes in the subtree. A shape node stores metadata about the subtree structure. Specifically, a separate shape number is assigned to each unique subtree structure. Each shape node includes metadata such as the shape number of the left subtree, the shape number of the right subtree, the total number of leaf nodes in the current subtree, and the number of leaf nodes in the left subtree within the current subtree.

[0129] Step d3: construct an index shape graph using the reference relationship between each shape node.

[0130] When shape nodes are organized hierarchically into a graph structure, the root shape node represents the root of the entire tree. Non-leaf shape nodes reference the shape nodes in the next layer by storing their child shape numbers, while leaf shape nodes point to the position of the identifier list in the leaf node array. Ultimately, this organization forms a directed shape graph with references, achieving compression and sharing of the tree structure. The indexed shape graph is then loaded into the FPGA's on-chip memory (SRAM).

[0131] In addition, a mapping dictionary is generated to record which shape node each node in the vector tree structure corresponds to, which can be stored off-chip.

[0132] like Figure 7 As shown, a binary tree representation T is constructed for the encoded vector set V´, and the binary tree T is then compressed into a shape graph G. Each shape node in G contains two values: the number on the left represents the shape number, and the number in parentheses on the right represents the number of leaf nodes in the left subtree of the corresponding binary tree. Starting from the entry node of the shape graph, an index offset is initialized to 0. Jumps are made according to the "0" / "1" sequence. When passing through an edge representing "1", the offset is increased by the value in the parentheses of the current shape node. When passing through an edge representing "0", the offset remains unchanged. This process continues. When a leaf node shape is reached, an offset is calculated, which can be used to retrieve the corresponding leaf node information (such as a vector list) from the leaf node array. The leaf node array is constructed by traversing the binary tree in pre-order, sorting the leaf nodes visited in order. The leaf nodes contain all the vector IDs that match the "0" / "1" sequence corresponding to the leaf node.

[0133] In the above implementation, extreme compression and efficient retrieval of tree-based indexes are achieved by identifying and reusing repeated subtree structures. Subtree structures with the same topology are identified and assigned unique shape numbers, avoiding repeated storage of identical tree shapes. Lightweight shape nodes replace complete subtree structures, significantly reducing memory usage. By referencing identical substructures via shape numbers, the size of the index graph depends solely on the type of topology, not the total number of nodes. The compressed index shape graph retains hierarchical reference relationships, maintaining efficient tree traversal.

[0134] In some optional implementations, the above step S303 further includes:

[0135] Step e1: Obtain target feature values ​​of a preset number of dimensions of a target vector dataset.

[0136] After dimensionality reduction using principal component analysis (PCA), the resulting covariance matrix is ​​subjected to eigenvalue decomposition. The maximum eigenvalues ​​of the first m (preset values) arranged in descending order are selected as the target eigenvalues. If a fast dimensionality reduction method (such as priority sorting) is used, priority indices are generated by calculating the numerical distribution characteristics of each dimension (such as the number of unique values ​​and the range of values). The characteristic indices corresponding to the first m highest priority dimensions (such as the number of clusters) are selected as the target eigenvalues.

[0137] Step e2: determining the target division and quantization strategy corresponding to the target vector data set based on the distribution state of each target eigenvalue.

[0138] Analyze the relative difference of the target eigenvalues ​​in the first m dimensions. Specifically, if the eigenvalue of the first dimension is significantly greater than that of the second dimension (e.g., the ratio exceeds the threshold ), a finer partitioning method (e.g., quartering) is used for the first dimension. If the eigenvalues ​​of the first two dimensions are close and both large, a finer partitioning method is used for both dimensions. The choice of strategy depends on the degree of difference in the eigenvalues. The goal is to improve classification granularity by finely partitioning high-information dimensions.

[0139] In step e3, a target segmentation quantization strategy is used to segment the target vector data set into a predetermined number of dimensions to generate a plurality of subsets corresponding to the target vector data set.

[0140] Get the target eigenvalues ​​(such as the eigenvalues ​​of principal component analysis or dimension priority indicators) of the first m dimensions (such as m=2). Select the division method (such as 4 equal divisions or 8 equal divisions) according to the eigenvalue distribution. Taking 4 equal divisions as an example, for each target dimension, calculate the median of the full-dimensional data, divide the data into two parts, and calculate the median of each part to form 3 split points, and divide the data into 4 sub-intervals (such as [min, Q1), [Q1, median), [median, Q3), [Q3, max]). Each combined interval of the first m dimensions corresponds to a subset (such as 4 equal divisions of two dimensions to generate = 16 subsets), each subset contains the data in the original vector that falls into the multidimensional interval.

[0141] Step e4: determine the remaining dimensions in each subset except the pre-set number of dimensions.

[0142] Let the original vector dimension be d, the first m dimensions processed are used to generate the subset partitioning, and the remaining dimensions are dm. For example, after the original 128-dimensional vector is partitioned by the first 2 dimensions (m=2), the remaining 126 dimensions are used for subsequent processing.

[0143] Step e5: perform binary quantization coding on the remaining dimensions of each subset to generate a coding vector set.

[0144] Operate independently on each of the remaining dm dimensions. Specifically, calculate the median of that dimension within the subset. Iterate over all vectors in the subset. If a vector's value in that dimension is less than the median, encode it as the first value; if it is greater than or equal to the median, encode it as the second value. Each vector generates a binary code sequence of length dm, and all these sequences constitute the set of encoded vectors for the subset.

[0145] Step e6: For the coding vectors of the remaining dimensions of each subset, a target vector tree structure corresponding to the remaining dimensions is constructed using the coding sequence after binary quantization coding.

[0146] For each subset's remaining dimensions (dm dimensions), independently calculate the median of each dimension. Convert the remaining dimension values ​​of each vector in the subset to binary. Generate a tree node bit by bit, using the binary code sequence as a path. Each bit represents a tree level (e.g., dimension 1 is the root node, dimension 2 is the second-level node). The first value points to the left subtree, and the second value points to the right subtree. Leaf nodes store a list of the original vector IDs corresponding to the path.

[0147] Step e7: performing shape compression processing on the target vector tree structure of each subset to generate a shared index shape graph.

[0148] Traverse the binary trees of all subsets and extract subtrees with the same topology. Assign a globally unique ID (e.g., Shape_ID = 5) to each unique subtree and create a shape node. Link all shape nodes together into a graph based on reference relationships. Identical subtrees in different subsets reuse the same shape node (e.g., the 0-1 subtrees of both subsets 1 and 2 are mapped to Shape_ID = 5), significantly reducing storage overhead. The shape graph is written to on-chip FPGA storage. The entry shape number and leaf node array for each subtree are stored off-chip.

[0149] like Figure 8 As shown, assuming that PCA dimensionality reduction is used to generate d eigenvalues ​​from large to small .when (like, , that is, the threshold If the information volume or difference of the first dimension is much higher than that of the second dimension), the value of the eigenvector on the first dimension is divided into four equal parts (for example, the median of each value of the first-dimensional vector is calculated first, and then the median of the two parts of the value divided by the median is taken respectively. Through the above three medians, the value range of the data in this dimension can be divided into four equal parts). When , the values ​​of the eigenvectors in the second dimension are also divided into four equal parts. Assume that at most the first m (for example, m=2) dimensions can be divided into four equal parts. Figure 8 As shown, if the current two dimensions are divided into four equal parts, the vector set can be divided into up to 16 subsets based on the values ​​of the first two dimensions. A binary tree representation and shape-based compression are then constructed for each subset, resulting in a summary shape graph, 16 shape graph entries, and 16 corresponding leaf node arrays. If the current m-dimensional vector is processed using a non-binary partitioning method (such as a 4- or 8-partitioning method), the first m dimensions do not participate in the subsequent binary tree and shape graph compression. During the vector index construction process, multiple subsets are generated from the original vector set based on the partitioning of the first m dimensions. Starting from the m+1 dimension, these subsets are quantized using a "0" / "1" binary partitioning method, a binary tree is constructed, and the resulting shape graph is compressed. These subsets share a shape graph, requiring only the shape entry numbers (i.e., shape graph entries) to be stored.

[0150] In the above-mentioned implementation, the coordinated optimization of storage and computing efficiency is achieved through adaptive data segmentation and hierarchical compression guided by feature dimensions. The quantization strategy is dynamically selected based on the distribution of eigenvalues ​​of a preset number of dimensions to ensure that the segmentation method adapts to the data characteristics. The remaining dimensions of the segmented subsets are processed independently, and structured codes are generated through binary quantization to maintain the efficiency of processing each subset. After constructing a vector tree for each subset, the repeated topological structure across subsets is identified through shape compression to generate a shared index shape graph. The shared shape graph eliminates redundant storage of subtrees, so that the index scale depends on the type of topology rather than the total amount of data. This method achieves exponential storage compression and cross-subset computing reuse while maintaining retrieval accuracy through the three-level collaboration of "feature segmentation-subset quantization-topology reuse".

[0151] Step S305: Get the query vector and determine the target vector list that matches the query vector from the vector array corresponding to the vector index according to the vector index. Figure 2 Step S205 of the illustrated embodiment will not be described in detail here.

[0152] In this embodiment, a vector query method is provided, which can be used in electronic devices, such as servers. Figure 9 is a flow chart of a vector query method according to an embodiment of the present invention. Figure 9 As shown, the process includes the following steps:

[0153] Step S401: Obtain the original vector dataset. Figure 4 Step S301 of the illustrated embodiment will not be described in detail here.

[0154] Step S402: Perform dimensionality reduction on the original vector dataset to generate a target vector dataset. Figure 4 Step S302 of the illustrated embodiment will not be described in detail here.

[0155] Step S403: Encode the target vector data set to generate an encoded vector set. Figure 4 Step S303 of the illustrated embodiment will not be described in detail here.

[0156] Step S404: Generate a vector index for each code vector based on the correlation between the code vectors in the code vector set. Figure 4 Step S304 of the illustrated embodiment will not be described in detail here.

[0157] Step S405 : obtaining a vector to be queried, and determining a list of target vectors that match the vector to be queried from a vector array corresponding to the vector index according to the vector index.

[0158] Specifically, the above step S405 includes:

[0159] Step S4051: Obtain a vector to be queried, perform dimensionality reduction processing and quantization encoding processing on the vector to be queried, and generate a query encoding vector.

[0160] The query encoding vector refers to the encoding vector generated after the query vector is processed by dimensionality reduction and quantization encoding. Specifically, the query vector is input by the user (such as the feature vector of the search request). The dimensionality reduction process reuses the parameters generated in the index construction phase (such as the covariance matrix of PCA or the dimension priority order of fast dimensionality reduction) to project the high-dimensional vector into a low-dimensional space. Quantization encoding reuses the median threshold of each dimension. For each dimension value of the vector after dimensionality reduction, if it is less than the median of the dimension, it is encoded as the first value, otherwise it is encoded as the second value, generating a binary sequence. This process is completed by the CPU to ensure consistency with the logic during index construction.

[0161] Step S4052: Determine the index entry corresponding to the query code vector using a preset quantization coding strategy.

[0162] The preset quantization encoding strategy refers to a preset customized encoding rule that uses a non-binary method (e.g., a 4 / 8 split) for the first m high-information dimensions, while subsequent dimensions still use a binary method (0 / 1 encoding). The index entry refers to the starting node number of the subtree in the shape graph. Specifically, if a non-binary method (e.g., a 4-way split) is used for the first m dimensions during index construction, the subset entry must be located based on the values ​​of the first m dimensions of the query encoding vector. For example, assuming a 4-way split for the first two dimensions, the encoding combination of the first two dimensions (e.g., 00, 01, 10, 11) corresponds to the shape graph entry number of one of the 16 subsets. A pre-stored mapping table (e.g., {dimension combination: entry number}) allows for quick matching, ensuring that the query jump starts at the correct subset root node.

[0163] Step S4053 , taking the index entry as the starting node, performing jump matching along the target code sequence of the query code vector in the vector index to obtain a matching result.

[0164] The target coding sequence refers to the binary sequence corresponding to the query coding vector. The matching result refers to the intermediate data generated when jumping to the leaf node. Specifically, jump step by step according to the coding sequence on the shared shape graph. Starting from the index entry (shape node), the offset offset = 0. Traverse the binary sequence of the query coding vector (starting from the m+1 dimension): if the bit value is the first numerical value, jump to the left subtree shape node of the current node, and the offset remains unchanged. If the bit value is the second numerical value, jump to the right subtree shape node, and the offset is accumulated by the number of leaves in the left subtree (the value pre-stored in the node). When reaching the leaf shape node, the final offset is output as the matching result. Among them, the jump process is accelerated by the FPGA multi-stage pipeline hardware, and each stage of the pipeline accesses the shape node data stored on the chip.

[0165] Step S4054: Determine a list of target vectors that match the query vector using the matching results.

[0166] Extract the target vector list from the pre-stored leaf node array using the offset. Specifically, the matching offset points to a position in the leaf node array (which stores all leaf nodes in pre-order traversal order). The leaf node at that position contains a list of vector IDs corresponding to all original vectors that match the query encoded vector path. This list is output as the target vector list for subsequent similarity calculations (such as calculating the top-k similar vectors).

[0167] In some optional implementations, the above step S4054 includes:

[0168] Step f1: when the current coding bit of the target coding sequence is the first value, jump along the index left subtree corresponding to the current index node.

[0169] When the current coded bit of the query vector is the first value, a jump is performed based on the shape number of the left subtree of the current shape node. Specifically, the shape number of the left subtree corresponding to the current node is obtained and used as the node address for the next level of pipeline processing. This process only updates the node position without modifying the offset, essentially descending along the left path of the binary tree.

[0170] Step f2: When the current coding bit of the target coding sequence is the second value, jump along the index right subtree corresponding to the current index node, accumulate the offset of the number of leaf nodes of the index left subtree, and obtain the target offset.

[0171] The target offset refers to the offset value accumulated during the jump process, which is used to locate the leaf node array. Specifically, if the current encoding bit is the second value, obtain the shape number of the right subtree of the current node and jump to the corresponding node. Accumulate the number of child nodes of the left subtree of the current node to the total offset. For example, when the root node (shape 1) encounters "1", accumulate the number of leaves of its left subtree (value is 2), and then jump to the shape 3 node. This design is because the vector array is stored in pre-order traversal, and the position of the right subtree leaves in the array must span all the leaves of the left subtree.

[0172] Step f3: When a leaf index node is reached, the target vector list is extracted from the vector array corresponding to the vector index using the target offset.

[0173] The process terminates when a jump reaches a leaf node (specially marked with a shape number such as 1). The accumulated total offset at this point directly corresponds to the starting index in the leaf node array. Based on this index, a list of consecutively stored vector IDs is retrieved from the vector array. For example, the offset for the path "1→0" is 2, meaning the vector list starting at position 2 in the array is read.

[0174] The vector query method provided by the embodiment of the present invention realizes rapid positioning of the query path and accurate extraction of results through the collaborative jump mechanism of code alignment and tree index. The query vector is treated with the same dimensionality reduction and quantization coding process as the database construction to ensure the alignment of the coding rules. The index entry node is directly determined according to the quantization strategy to avoid global search. The left / right subtree path is dynamically selected according to each bit value (first value / second value) of the query coding sequence to achieve bit-by-bit matching hierarchical jump. During the jump, the offset of the number of child nodes of the left subtree is accumulated to convert the tree structure into the physical address of a continuous array. After arriving at the leaf node, the target vector list is extracted from the continuous storage area of ​​the vector array at one time by using the calculated target offset. This method converts the logical traversal of the tree index into efficient random access to the array, completely avoiding backtracking operations while ensuring the accuracy of the results, and significantly improving the query response speed.

[0175] In the following embodiment, the above-mentioned vector query method will be exemplified by taking a specific application embodiment as an example.

[0176] like Figure 10 As shown in the figure, vector index construction includes four processing stages: vector dimensionality reduction, quantization encoding, binary tree representation, and shape-based binary tree compression. Vector similarity querying includes three stages: index query pre-processing (including query vector dimensionality reduction and quantization), FPGA-based index querying (i.e., FPGA multi-stage pipeline matching), and index query post-processing (including calculating the similarity between the query vector and the list of matched vectors and returning the top(k) vectors).

[0177] For the vector index construction process, such as Figure 11 As shown, the vector index construction process includes four strictly serial CPU processing modules: first, adaptive vector dimensionality reduction is performed, and high-dimensional data is compressed into a low-dimensional space through PCA or fast priority dimensionality reduction algorithm (sorted according to the diversity / value range of dimension values). The number of dimensions is dynamically determined by the FPGA on-chip storage resources (BRAM capacity) and must meet the preset SAFE_PIPELINE (safe pipeline level, such as 16) and MAX_PIPELINE (maximum pipeline level, such as 64) constraints to prevent the subsequent shape graph from exceeding the storage capacity. Secondly, quantization encoding is performed to convert the vector into a binary sequence (0 / 1) based on the median of each dimension. Then, a binary tree representation is constructed to map the binary sequence into a tree structure (edges represent 0 / 1, and leaf nodes store the original vector ID). Finally, shape-based binary tree compression is performed to abstract the tree topology into a shape graph (including shape nodes, leaf node arrays, etc.), significantly reducing storage overhead by sharing subtree shapes. In addition, customized quantization strategies are supported. If the eigenvalues ​​of the first m dimensions are significantly different (such as in PCA), the vector will be converted into a binary sequence (0 / 1). ), this dimension can be divided equally into 4 / 8, generating multiple subsets to independently construct subtrees and share the shape graph. The entry information is stored off-chip. The entire process must ensure that the number of dimensions (k) after dimensionality reduction balances index compression and query efficiency. If k is too small, leaf node vectors will be densely populated, increasing the complexity of similarity calculations. If k is too large, the proliferation of shape nodes may overflow the FPGA's on-chip SRAM (this can be avoided through pre-allocation checks and dimensionality reduction fallback mechanisms).

[0178] For the vector similarity query process, such as Figure 12As shown, vector similarity query is a process that combines CPU and FPGA collaborative processing. The CPU first reduces the dimension of the input query vector and performs quantization encoding (using specific parameters), and determines the shape graph entry based on a customized quantization encoding strategy (for example, a four-equal mapping based on the values ​​of the first m dimensions). Subsequently, the CPU uses the shape graph entry and the binary sequence vector generated after dimensionality reduction and quantization as parameters, and calls the multi-stage pipeline processing unit of the FPGA vector index query engine to perform the core vector index query operation. After the FPGA engine completes the query, it returns a list of matching vectors to the CPU. The post-processing program on the CPU side is responsible for calculating the actual similarity between each vector in the list and the original query vector, and finally selects the top K vectors with the highest similarity as the query result output.

[0179] The vector query method provided in the embodiment of the present invention designs an adaptive dimensionality reduction method based on PCA, which can dynamically determine the optimal low-dimensional space dimension based on the cumulative distribution of eigenvalues, preset pipeline level thresholds (MAX_PIPELINE, SAFE_PIPELINE) and available SRAM space. At the same time, a fast dimensionality reduction processing technology is proposed, which significantly reduces the computational complexity and latency, and also realizes adaptive decision-making of dimensionality reduction based on dimension priority and the aforementioned thresholds and SRAM constraints. Secondly, a customized quantization coding strategy is formulated, which divides the input vector set into 4 or 8 equal parts according to the eigenvalue distribution or priority changes in the first few dimensions, thereby achieving finer divisions in dimensions with large amounts of information or differences. Finally, a vector index structure of a shared shape graph is constructed, and the vector set is divided into subsets based on the above-mentioned quantization strategy, and a shared shape mapping dictionary is used to index these subsets. This not only significantly improves the compression ratio from binary tree nodes to shape nodes by compressing the tree height and reusing shape nodes, effectively reducing the SRAM storage overhead of the shape graph, but also enables vector queries to map the reduced-dimensionality query vectors to multiple different entries in the shared shape graph based on the quantization strategy. This design, which further reduces the tree height and the number of pipeline stages based on dimensionality reduction, significantly reduces the risk of SRAM being unable to accommodate the shape graph.

[0180] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0181] The embodiment of the present application also provides a vector query device, such as Figure 13 Shown, including:

[0182] An acquisition module 501 is used to acquire an original vector data set;

[0183] A dimensionality reduction module 502 is used to perform dimensionality reduction processing on the original vector data set to generate a target vector data set;

[0184] The encoding module 503 is used to perform encoding processing on the target vector data set to generate an encoded vector set;

[0185] A generating module 504, configured to generate a vector index of each code vector based on an association relationship between each code vector in the code vector set;

[0186] The query module 505 is configured to obtain a vector to be queried, and determine a target vector list matching the vector to be queried from a vector array corresponding to the vector index according to the vector index.

[0187] In some optional implementations, the dimensionality reduction module 502 includes:

[0188] A first acquisition submodule is used to obtain the block storage resource capacity and pipeline level parameters of the programmable logic device;

[0189] The dimensionality reduction submodule is used to perform dimensionality reduction processing on the original vector data set to obtain the dimensionality reduction parameters corresponding to the original vector data set;

[0190] The processing submodule is used to process the dimensionality reduction parameters based on the block storage resource capacity and pipeline level parameters to obtain the target number of dimensions;

[0191] The extraction submodule is used to extract data from the original vector dataset using the target dimension number to obtain the target vector dataset.

[0192] In some optional embodiments, the dimensionality reduction submodule includes:

[0193] A first determining unit, configured to determine a characteristic value of an original vector data set;

[0194] The second determining unit is used to determine the number of initial dimensions that meet preset conditions and correspond to the original vector data set by using the cumulative distribution of the eigenvalues; wherein the dimensionality reduction parameter includes the number of initial dimensions.

[0195] In some optional implementations, the processing submodule includes:

[0196] The adjustment unit is used to adjust the initial number of dimensions based on the block storage resource capacity and the pipeline level parameters to obtain the target number of dimensions.

[0197] In some optional embodiments, the dimensionality reduction submodule includes:

[0198] A first acquisition unit is used to obtain the priority index of each first dimension in the original vector data set;

[0199] The sorting unit is used to sort each first dimension according to the priority index to obtain the dimension sorting result corresponding to the original vector data set; wherein the dimensionality reduction parameter includes the dimension sorting result.

[0200] In some optional implementations, the processing submodule includes:

[0201] The truncation unit is used to perform dimension truncation on the dimension sorting result based on the block storage resource capacity and pipeline level parameters to obtain the target number of dimensions.

[0202] In some optional implementations, the encoding module 503 includes:

[0203] A first determination submodule is used to determine a target value corresponding to a plurality of vector data on each second dimension in the target vector data set, where the target value is any one of the median, the average, and the value corresponding to the bisection of the number of clusters;

[0204] an encoding submodule, configured to encode any vector data on any second dimension into a first value if the vector data is less than a target value; and to encode the vector data into a second value if the vector data is greater than or equal to the target value;

[0205] The combining submodule is used to combine the first values ​​or the second values ​​of the target vector data set after encoding on all second dimensions into a coding vector set.

[0206] In some optional implementations, the generating module 504 includes:

[0207] A first construction submodule is configured to construct a vector tree structure of each coding vector based on an association relationship between each coding vector in the coding vector set;

[0208] The first compression submodule is used to perform shape compression processing on the vector tree structure to generate an index shape graph; wherein the vector index includes the index shape graph.

[0209] In some optional embodiments, the building block includes:

[0210] A first construction unit is configured to construct, for any one of the coding vectors in the coding vector set, a left or right branch path according to the first value or the second value, using each bit of the coding vector as a hierarchical node;

[0211] The generation unit is used to generate a vector tree structure using left and right branch paths.

[0212] In some optional embodiments, the compression submodule includes:

[0213] An identification unit, configured to identify the vector tree structure and determine each subtree structure in the vector tree structure;

[0214] A creation unit, configured to assign a shape number to each subtree structure, and create a shape node corresponding to the subtree structure using the number of leaf nodes in the subtree structure and the shape number;

[0215] The second construction unit is used to construct an index shape graph by using the reference relationship between each shape node.

[0216] In some optional implementations, the encoding module 503 further includes:

[0217] The second acquisition submodule is used to obtain target feature values ​​of a preset number of dimensions of the target vector data set;

[0218] The second determination submodule is used to determine the target division and quantization strategy corresponding to the target vector data set based on the distribution state of each target eigenvalue;

[0219] A segmentation submodule is used to segment the target vector data set into a predetermined number of dimensions using a target segmentation quantization strategy to generate multiple subsets corresponding to the target vector data set;

[0220] The third determining submodule is used to determine the remaining dimensions in each subset except for the pre-set number of dimensions;

[0221] The first generating submodule is used to perform binary quantization coding processing on the remaining dimensions of each subset to generate a coding vector set.

[0222] In some optional implementations, the encoding module 503 further includes:

[0223] The second construction submodule is used to construct a target vector tree structure corresponding to the remaining dimensions of the coding vectors of the remaining dimensions of each subset using the coding sequence after binary quantization coding;

[0224] The second compression submodule is used to perform shape compression processing on the target vector tree structure of each subset to generate a shared index shape graph.

[0225] In some optional implementations, the query module 505 includes:

[0226] The second generation submodule is used to perform dimensionality reduction processing and quantization encoding processing on the query vector to generate a query encoding vector;

[0227] A fourth determination submodule, configured to determine an index entry corresponding to the query encoding vector using a preset quantization encoding strategy;

[0228] The matching submodule is used to perform jump matching along the target code sequence of the query code vector in the vector index with the index entry as the starting node to obtain a matching result;

[0229] The fifth determining submodule is configured to determine a list of target vectors that match the query vector using the matching results.

[0230] In some optional implementations, the fifth determining submodule includes:

[0231] A first jump unit, configured to jump along the index left subtree corresponding to the current index node when the current coding bit of the target coding sequence is a first value;

[0232] A second jump unit is configured to jump along the index right subtree corresponding to the current index node when the current code bit of the target code sequence is a second value, and accumulate an offset of the number of leaf nodes of the index left subtree to obtain a target offset;

[0233] The extraction unit is used to extract the target vector list from the vector array corresponding to the vector index using the target offset when reaching the leaf index node.

[0234] For the description of the features in the embodiment corresponding to the vector query device, please refer to the relevant description of the embodiment corresponding to the vector query method, and will not be repeated here.

[0235] The embodiment of the present application also provides an electronic device, such as Figure 14 As shown, it includes a memory 10 and a processor 20. The memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above-mentioned vector query method embodiments.

[0236] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned vector query method embodiments when executed.

[0237] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0238] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned vector query method embodiments are implemented.

[0239] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned vector query method embodiments are implemented.

[0240] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0241] The above describes in detail the vector query method, electronic device, storage medium, and program product provided by this application. This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is intended only to facilitate understanding of the method and core concepts of this application. It should be noted that those skilled in the art may make various improvements and modifications to this application without departing from the principles of this application, and such improvements and modifications fall within the scope of protection of the claims of this application.

Claims

1. A vector query method, characterized in that: include: Obtaining an original vector dataset, where the original vector dataset refers to an original high-dimensional vector set, and the original vector dataset is obtained by performing feature extraction processing on unstructured data; Performing dimensionality reduction processing on the original vector data set to generate a target vector data set, including: obtaining the block storage resource capacity and pipeline level parameters of the programmable logic device; performing dimensionality reduction processing on the original vector data set to obtain dimensionality reduction parameters corresponding to the original vector data set, including: determining the eigenvalue of the original vector data set, the eigenvalue refers to the value representing the size of the data variance after the covariance matrix decomposition in the principal component analysis; using the cumulative distribution of the eigenvalue, determining the initial number of dimensions corresponding to the original vector data set that meet the preset conditions, the cumulative distribution refers to the cumulative proportion of the eigenvalues ​​in order, the preset condition refers to the threshold that the cumulative distribution needs to meet, and the initial number of dimensions refers to the number of reduced dimensions preliminarily estimated by the cumulative distribution of the eigenvalues; wherein the dimensionality reduction parameters include the initial number of dimensions and the eigenvectors corresponding to the eigenvalues; based on the block storage resource capacity and the pipeline level parameters, processing the dimensionality reduction parameters to obtain the target number of dimensions; using the target number of dimensions to extract data from the original vector data set to obtain the target vector data set; Performing encoding processing on the target vector data set to generate an encoded vector set; generating a vector index of each of the code vectors based on an association relationship between the code vectors in the code vector set; A vector to be queried is obtained, and according to the vector index, a target vector list matching the vector to be queried is determined from a vector array corresponding to the vector index.

2. The vector query method according to claim 1, characterized in that: The processing of the dimensionality reduction parameter based on the block storage resource capacity and the pipeline level parameter to obtain a target number of dimensions includes: Based on the block storage resource capacity and the pipeline level parameter, the initial number of dimensions is adjusted to obtain the target number of dimensions.

3. The vector query method according to claim 1, characterized in that: The performing dimensionality reduction processing on the original vector data set to obtain dimensionality reduction parameters corresponding to the original vector data set includes: Obtaining a priority index of each first dimension in the original vector dataset; Sort each of the first dimensions according to the priority index to obtain a dimension sorting result corresponding to the original vector data set; wherein the dimensionality reduction parameter includes the dimension sorting result; The processing of the dimensionality reduction parameter based on the block storage resource capacity and the pipeline level parameter to obtain a target number of dimensions includes: Based on the block storage resource capacity and the pipeline level parameter, dimension truncation is performed on the dimension sorting result to obtain the target number of dimensions.

4. The vector query method according to claim 1, characterized in that: The encoding process of the target vector data set to generate an encoded vector set includes: Determine a target value corresponding to a plurality of vector data on each second dimension in the target vector data set, where the target value is any one of a median, an average, and a value corresponding to a bisection of the number of clusters; For any vector data on any second dimension, if the vector data is less than the target value, the vector data is encoded as a first value; if the vector data is greater than or equal to the target value, the vector data is encoded as a second value; The first values ​​or the second values ​​of the target vector data set encoded on all second dimensions are combined into the encoded vector set.

5. The vector query method according to claim 1 or 4, characterized in that: Generating a vector index of each encoding vector based on an association relationship between the encoding vectors in the encoding vector set includes: Constructing a vector tree structure of each encoding vector based on an association relationship between the encoding vectors in the encoding vector set; Performing shape compression processing on the vector tree structure to generate an index shape graph; The vector index includes the index shape graph.

6. The vector query method according to claim 5, characterized in that: The constructing of a vector tree structure of each coding vector based on the association relationship between each coding vector in the coding vector set includes: For any one code vector in the code vector set, taking each bit of the code vector as a hierarchical node, constructing a left or right branch path according to the first value or the second value; The vector tree structure is generated using the left and right branch paths.

7. The vector query method according to claim 5, characterized in that: The step of performing shape compression on the vector tree structure to generate an index shape graph includes: Identifying the vector tree structure and determining each subtree structure in the vector tree structure; Allocating a shape number to each of the subtree structures, and creating a shape node corresponding to the subtree structure using the number of leaf nodes in the subtree structure and the shape number; The index shape graph is constructed by utilizing the reference relationship between the shape nodes.

8. The vector query method according to claim 1 or 4, characterized in that: Also includes: Obtain target feature values ​​of a first preset number of dimensions of the target vector data set; Determining a target division quantization strategy corresponding to the target vector data set based on a distribution state of each of the target eigenvalues; Using the target division quantization strategy, segmenting the target vector data set into a predetermined number of dimensions to generate a plurality of subsets corresponding to the target vector data set; Determine the remaining dimensions in each of the subsets except for the pre-set number of dimensions; Perform binary quantization coding processing on the remaining dimensions of each of the subsets to generate the coding vector set.

9. The vector query method according to claim 8, characterized in that: Also includes: For the coding vectors of the remaining dimensions of each subset, constructing a target vector tree structure corresponding to the remaining dimensions using a coding sequence obtained by binary quantization coding; Shape compression processing is performed on the target vector tree structure of each subset to generate a shared index shape graph.

10. The vector query method according to claim 1, characterized in that: Determining, according to the vector index, a target vector list matching the query vector from a vector array corresponding to the vector index includes: Performing dimensionality reduction and quantization encoding processing on the query vector to generate a query encoding vector; Determine the index entry corresponding to the query encoding vector using a preset quantization encoding strategy; Taking the index entry as a starting node, performing jump matching along the target code sequence of the query code vector in the vector index to obtain a matching result; The matching results are used to determine a list of target vectors that match the query vector.

11. The vector query method according to claim 10, characterized in that: The determining a list of target vectors matching the query vector using the matching results includes: When the current code bit of the target code sequence is the first value, jump along the index left subtree corresponding to the current index node; When the current code bit of the target code sequence is the second value, jump along the index right subtree corresponding to the current index node, accumulate the offset of the number of leaf nodes of the index left subtree, and obtain the target offset; When a leaf index node is reached, the target vector list is extracted from the vector array corresponding to the vector index using the target offset.

12. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the vector query method according to any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the vector query method according to any one of claims 1 to 11 are implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the vector query method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • High-dimensional vector query method and device, computer equipment and readable storage medium

    CN118093633A

  • Aviation data bus signal encryption method and system based on FPGA

    CN119420580A