Vector database index recommendation method and device, equipment and medium
By performing cluster analysis and adaptive index recommendation on vector databases, the problem of index type selection in the existing technology depends on manual experience, and efficient retrieval and accuracy improvement of vector databases are achieved, adapting to the diversity of data internal structure and query patterns.
Patent Information
- Application Number
- CN202510827705.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The selection of existing vector database index types depends on manual experience and static configuration, making it difficult to adapt to the diversity of data internal structure and query patterns, resulting in limited retrieval efficiency and accuracy.
By performing cluster analysis on vector data sets, data distribution characteristics are extracted, and an adaptive index recommendation strategy is adopted to recommend appropriate index types based on the characteristics of the center distance between clusters, head and tail structure, such as IVF, HNSW, PQ, etc., combining clustering effect and data scale optimization index construction process.
The search efficiency and accuracy of vector database are improved, query load and system resources are balanced, and data distribution changes and dynamic adjustment of query mode is adapted.
Smart Images

Figure CN120336331A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vector database retrieval, and particularly relates to a method, device, equipment and medium for recommending vector database indexes. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] With the wide application of vector databases in various fields, the engineering importance of their storage capabilities has become increasingly prominent. The performance, efficiency, and stability of storage have become key considerations in software system design. At the same time, with the widespread use of large language models (LLMs) and retrieval-augmented generation (RAG) technologies, high-dimensional vectors encoded based on models such as BERT (e.g., 768-dimensional or 1024-dimensional) have become the most common storage objects. To improve the retrieval efficiency and accuracy of similar vectors, vector storage indexes have become an important part of vector database design.
[0004] Most of the existing selection of vector database index types relies on manual experience and static configuration. However, the retrieval efficiency and accuracy are affected not only by obvious features such as data scale and data dimension, but also by characteristics such as the internal structure of the data and the query pattern. It is difficult to cover all possible scenarios based on manual experience and static configuration. Summary of the Invention
[0005] In view of this, the present invention provides a method, device, equipment and medium for recommending vector database indexes to adaptively recommend appropriate index types according to the internal data distribution characteristics of the vector data set.
[0006] The first aspect of the present invention provides a method for recommending vector database indexes, including the following steps: Perform clustering analysis on the vector data set, extract data distribution characteristics based on the clustering results, and obtain a vector data portrait; the data distribution characteristics include the inter-cluster center distance and the presence or absence of a head-tail structure; Determine the recommended index according to the preset index recommendation strategy based on the vector data portrait; the index recommendation strategy includes the strategy for selecting index types under different data distribution characteristics: If the inter-cluster center distance exceeds the upper limit of the set threshold interval, the IVF index is adopted as a whole; If the inter-cluster center distance is within the set threshold interval, if there is no head-tail structure, the IVF index is adopted as a whole; if there is a head-tail structure, the IVF index is adopted for the main data part, and the HNSW index is adopted for the sparse data part; If the inter-cluster center distance is less than the lower limit of the set threshold interval, the recommended index is HNSW.
[0007] In some embodiments, after performing clustering analysis on the vector data set, the number of clusters in the clustering result is optimized based on the data distribution within each cluster: Taking the number of clusters as a variable, setting its change interval, performing clustering operations sequentially for multiple numbers of clusters within the change interval, calculating the data distribution parameters within the clusters, and determining the corrected value of the number of clusters through the elbow method; According to the corrected value, re-perform the clustering operation, conduct a rationality test on the cluster distribution of the clustering result, obtain the number of data samples within each cluster, and further adjust the number of clusters according to the distribution of the sample numbers to obtain the optimal number of clusters.
[0008] In some embodiments, the rationality test of the cluster distribution is used to test the balance of the number of data samples within each cluster, including the following test contents: whether there are too many low-density clusters, whether there are too many fragmented clusters, whether the overall distribution of the data samples is unbalanced, and the inter-class separability.
[0009] In some embodiments, when the whole or part of the vector data set is recommended to adopt the IVF index, if the amount of data recommended to adopt the IVF index exceeds the set threshold, the in-cluster index is recommended as the FLAT index; otherwise, the in-cluster index is recommended as the PQ index.
[0010] In some embodiments, the data distribution feature further includes the proportion of covariance principal components. If the proportion of covariance principal components exceeds the set threshold, the vector data set is dimensionally reduced and then the index recommendation is performed.
[0011] In some embodiments, the data distribution feature further includes the noise point ratio. If the noise point ratio exceeds the upper limit of the set threshold interval, it is prompted that the vector data set needs to be cleaned; if the noise point ratio is within the set threshold interval, the main data part is recommended to adopt the IVF index, and the sparse data part adopts the HNSW index; if the noise point ratio is lower than the lower limit of the set threshold, it is prompted to remove the noise vectors and then perform the index recommendation.
[0012] In some embodiments, the method further includes: performing index construction on the vector data set based on the recommended index, and evaluating the current index type according to the index construction cost index and the query performance index.
[0013] The second aspect of the present invention provides a vector database index recommendation device, including: A feature extraction module, configured to: perform clustering analysis on the vector data set, extract data distribution features based on the clustering result, and obtain a vector data portrait; the data distribution features include the inter-cluster center distance and whether there is a head-tail structure; An index recommendation module, configured to: determine the recommended index according to the preset index recommendation strategy based on the vector data portrait; the index recommendation strategy includes: If the center distance between clusters exceeds the upper limit of the set threshold range, the IVF index is adopted as a whole; If the center distance between clusters is within the set threshold range, if there is no head-tail structure, the IVF index is adopted as a whole; if there is a head-tail structure, the IVF index is adopted for the main data part, and the HNSW index is adopted for the sparse data part; If the center distance between clusters is less than the lower limit of the set threshold range, the recommended index is HNSW.
[0014] The third aspect of the present invention provides an electronic device, including a processor and a memory, and a computer instruction is stored on the memory. When the computer instruction is executed by the processor, the electronic device executes the method.
[0015] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the program realizes the method when executed by a processor.
[0016] In the above one or more technical solutions, by constructing an index recommendation strategy, it is possible to adaptively recommend a suitable index type according to the internal data distribution characteristics of the vector data set.
[0017] In addition, for the recommended index type, it is also possible to evaluate according to the creation cost and query performance of the index, thereby guiding the optimization of the index type and its parameters in the construction process, and balancing the query load and system resource constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0019] Figure 1 Shows the flowchart of the vector database index recommendation method provided by the embodiment of the present invention; Figure 2 Shows the flowchart of the vector data portrait construction provided by the embodiment of the present invention; Figure 3 Shows the program module diagram of the vector database index recommendation device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Hereinafter, embodiments of the present application will be described in more detail with reference to the drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0021] In the description of the embodiments of the present application, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on".
[0022] The data distribution characteristics involved in one or more of the following embodiments include the number of clusters, the average distance within clusters, the distance between cluster centers, the proportion of noise points, and the proportion of covariance principal components. The definitions of each data distribution characteristic and their uses in the specific implementation are shown in Table 1.
[0023] Table 1 Description of data distribution characteristic items
[0024] (1) Average distance within clusters After clustering the vector data set, the evaluation of the data distribution mainly focuses on dimensions such as the compactness within clusters and the Silhouette Score. Among them, the compactness within clusters is used to evaluate the distribution of data samples within clusters, that is, to measure the sparsity and density of data within each cluster, whether there are too many or too few central data samples in a certain cluster, etc. The average distance within clusters is a linear index describing the compactness within clusters and is strongly correlated with the Inertia in terms of trend. Therefore, the Inertia is used to measure the average distance within clusters.
[0025] The Inertia is a measure of the clustering compactness and is defined as the sum of the squared distances from all sample points to the center of their respective clusters:
[0026] where k represents the number of clustering clusters (i.e., the candidate nlist), represents the i-th cluster, represents the center of the i-th cluster, represents the sample vector belonging to this cluster.
[0027] (2) Distance between cluster centers The Silhouette Score is used to measure the intra-class aggregation and inter-class separation of each sample. The nearest neighbor cluster distance term in the Silhouette Score is closely related to the distance between cluster centers. Therefore, the Silhouette Score is used to measure the distance between cluster centers. The calculation formula is:
[0028] where, represents the average distance between sample i and other samples within its cluster, represents the average distance between sample i and samples in the nearest neighbor cluster, represents the Silhouette Score of the i-th sample, and the distance between cluster centers is represented by the average Silhouette Score of all samples.
[0029] (3)Proportion of noise points The sources of noise points are diverse. For example, in table segmentation, there are blank lines, pure number columns, and punctuation lines. In business logs, there are UUIDs, timestamps, etc. They have no direct semantic information.
[0030] Use DBSCAN to detect noise points. Based on the theory of density reachability, by setting the neighborhood radius and the minimum number of points to identify high-density regions and noise samples in the semantic space. According to the proportion of samples marked as noise points in the clustering results, calculate the noise point proportion noise_rate, and based on this, decide whether to perform recommended strategies such as noise elimination and fallback. Effectively avoid the interference of formatted abnormal samples on the index structure, and improve the quality of the subsequent clustering structure and the stability of index compression.
[0031] Exemplarily, the neighborhood radius is set to 0.3, and the minimum number of points is set to 5.
[0032] (4)Proportion of covariance principal components Use the GMM model to fit the distribution of the vector set. Extract the explanatory ability of the principal components of each dimension by eigenvalue decomposition of the covariance matrix, and calculate the cumulative variance proportion to judge whether there is a compressible principal direction.
[0033] The vector set is , and GMM models it as a mixture of
[0034] where represents the weight of the th Gaussian component, satisfying ; represents the mean vector of the th component; represents the covariance matrix of the th component; represents the multi-dimensional Gaussian probability density function.
[0035] Perform eigenvalue decomposition on each covariance matrix. The system constructs a global cumulative explained variance ratio based on the principal component eigenvalues of all components:
[0036] Finally, represents the retained dimension of the principal component; represents the original vector dimension; represents the The i-th principal component dimension in a Gaussian component. The numerator represents the total explained variance of the first dimensional total explained variance, and the denominator represents the total variance.
[0037] Exemplarily, is set to 20, and the number of retained principal components is 256.
[0038] The index types involved in one or more of the following embodiments include: IVF, HNSW, PQ, Flat, IVF_FLAT, IVF_PQ, IVF_FLOAT, etc. Table 2 shows the meanings of these index types.
[0039] Table 2 Index Types and Their Meanings
[0040] As mentioned in the background art, the selection of existing vector database index types mostly relies on manual experience and static configuration. Although manual experience has certain reference value, it is often difficult to cover all possible scenarios. If the selected index type is not suitable for the current data characteristics or query pattern, it may lead to a significant decrease in retrieval efficiency.
[0041] Figure 1 The flowchart of the vector database index recommendation method provided by the embodiment of the present invention is shown. The method includes the following steps: S110: Perform clustering analysis on the vector data set, extract data distribution characteristics based on the clustering results, and obtain a vector data portrait. The purpose of clustering analysis is to divide the data into multiple clusters to minimize the distance within the clusters, facilitating the extraction of data distribution structure characteristics. Specifically, common clustering analysis methods include K-Means clustering, Hierarchical Clustering, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), Gaussian Mixture Model (GMM), etc., and one or more of them can be selected as needed according to the characteristics of the vector data set.
[0042] S120: Determine the recommended index according to the vector data profile based on the preset index recommendation strategy. The preset index recommendation strategy has two implementation methods: Rule-Based and ML-Based. Among them, the rule engine maps the input features based on a set of manually set empirical rules. The algorithm model uses a lightweight classifier (such as XGBoost) to learn the relationship between features and index strategies. As an example of the rule engine, the index recommendation strategy includes strategies for selecting index types under different data distribution characteristics: if the distance between cluster centers exceeds the upper limit of the set threshold interval, the IVF index is adopted overall; if the distance between cluster centers is within the set threshold interval, and there is no head-tail structure, the IVF index is adopted overall; if there is a head-tail structure, the IVF index is adopted for the backbone data part, and the HNSW index is adopted for the sparse data part; if the distance between cluster centers is less than the lower limit of the set threshold interval, the recommended index is HNSW.
[0043] Based on the above steps, the clustering effect of the data can be measured by the distance between cluster centers, and then used as a basis for measuring the distribution balance of the overall data in the vector database, adaptively recommend indexes, and ensure retrieval efficiency.
[0044] S130: Perform index construction on the vector data set based on the recommended index, and evaluate the current index type according to the index construction cost index and query performance index. The index cost index includes index construction time, index size, etc., and the query performance index includes recall rate, query latency, etc. According to the evaluation results, it can be judged whether it is necessary to further optimize the index type.
[0045] In the optional embodiment based on Figure 1 as shown in Figure 2 the above step S110 can be implemented as steps S111 to S113.
[0046] S111: Perform clustering analysis on the vector data set based on unsupervised clustering analysis. Exemplarily, the unsupervised clustering analysis uses the Kmean algorithm. The purpose of clustering analysis is to divide the index (IVF) partitions. The key lies in dividing appropriate clusters according to the data distribution characteristics so as to quickly locate the content during vector retrieval. Generally speaking, the richer the data sample size, the more partitions should be. The relationship between the total data sample size and the initial number of clusters can be constructed, for example, defined as , where N represents the total data sample size. Of course, the default initial clustering number can also be specified for different magnitudes of sample sizes. Table 3 shows the default configurations of the number of clusters when clustering with the Kmean algorithm for different magnitudes of sample sizes.
[0047] Table 3 Default configuration table of the number of clusters
[0048] S112: Optimize the number of clusters in the clustering result based on the data distribution within each type of cluster. The above steps determine the initial number of clusters based on the principle that the richer the data sample size, the more partitions should be. However, in some scenarios, although the sample size is large, the distribution characteristics are not obvious. Therefore, this initial number of clusters may not be the optimal number of clusters, and it is necessary to optimize the number of clusters based on the data distribution within the clusters and the distribution of the data sample quantities within each cluster.
[0049] As a specific implementation method, first, take the number of clusters as a variable, set its change interval, for multiple numbers of clusters within the change interval, perform clustering operations in sequence, and calculate the data distribution parameters within the clusters. Determine the correction value of the number of clusters through the elbow method; then, according to the correction value, re-perform the clustering operation, conduct a rationality test on the cluster distribution of the clustering result, obtain the data sample quantity within each cluster, and further adjust the number of clusters according to the distribution of the sample quantities to obtain the optimal number of clusters.
[0050] Specifically, the data distribution within the clusters is measured by the clustering cost function (Inertia). The clustering cost function represents the sum of the squares of the distances from all data points to the centers of their respective clusters. The smaller the Inertia, the closer the data points are clustered around the cluster centers, and the better the clustering effect. The principle of the elbow method is: when k increases, Inertia will gradually decrease, but when k exceeds a certain value, the rate of decrease of Inertia will significantly decrease, and this point is the correction value of the number of clusters. Exemplarily, set the change interval of the number of clusters, traverse this change interval at an increment of 12%, record the Inertia after performing the clustering operation based on each number of clusters, and analyze its decreasing trend through the elbow method; if the rate of decrease of Inertia under the current number of clusters is less than 12% compared to the initial value or the limit value, it is considered that the convergence state has been reached, and the current number of clusters k can be used as the correction value of the number of clusters. It can be understood that the number of clusters should be an integer value, and rounding operations are also performed after calculating the number of clusters by setting the increment.
[0051] The rationality test of the cluster distribution mainly examines the balance of the data sample quantities within each cluster. The imbalance of the data sample quantities within each cluster is reflected in: there are too many low-density clusters, there are too many fragmented clusters, the overall distribution of the data samples is uneven, and the separation between classes is poor. When at least one of the above problems exists, it is considered that the current number of clusters is too large and should be adjusted downward; adjust the number of clusters downward according to the set increment (such as 12%), and re-perform the clustering operation until the rationality of the cluster distribution meets the requirements.
[0052] Exemplarily, if the number of data samples in more than 5% of the clusters in the total sample is less than 10, it is considered that there are too many low-density clusters; if the number of samples in the smallest cluster is less than 10% of the average number of samples per cluster, it is considered that there are fragmented clusters. Specifically, if the number of samples in a certain cluster is 5 and the overall average number of samples per cluster is 90, then this cluster is a fragmented cluster.
[0053] The overall imbalance in the distribution of data samples is measured by the ratio of the number of data samples in the largest cluster to the smallest cluster. If the ratio of the largest cluster to the smallest cluster exceeds 10 times, it is determined that the sample distribution is unbalanced.
[0054] The evaluation of structural separability is performed using the distance between cluster centers. Specifically, the distance between cluster centers is measured based on the silhouette coefficient. The value range of the silhouette coefficient is [-1, 1]. The closer it is to 1, the higher the degree of separation between classes and the better the clustering effect. The closer it is to -1, the lower the degree of separation between classes and the worse the clustering effect. Therefore, when the silhouette coefficient is lower than the set threshold, it is considered that the structural separability does not meet the requirements.
[0055] Based on this, the optimal number of clustering clusters is obtained, that is, the size of nlist in the IVF index, which meets the conditions for constructing the IVF series index.
[0056] S113: Perform clustering analysis based on the optimized number of clusters, extract the data distribution characteristics according to the clustering results, and obtain the vector data portrait. Exemplarily, the data distribution characteristics include the distance between cluster centers (inter_dist), the proportion of noise points (noise_rate), the proportion of covariance principal components (var_ratio), and the existence of head-tail structure. According to the above multiple data distribution characteristics, VDataProfile is obtained.
[0057] The identification of the head-tail structure is used to judge whether there is structural skewness in the data samples. After counting the number of data samples in each cluster and sorting them, the "head-tail structure" in the vector set is identified: if the top 10% of the clusters cover 50% of the total samples, it indicates strong head aggregation; if the number of samples in more than 20% of the clusters < 20, or the proportion of single-sample clusters exceeds 10%, it indicates the existence of a significant long-tail structure. Both of the above situations are considered to have a head-tail structure.
[0058] In step S120, by summarizing the influence of different data characteristics on the index, it can be found that: when the clustering effect is good, the IVF index has better performance than the HNSW; when the vector distribution is relatively dense, that is, the average intra-cluster distance is small, since this situation itself will affect the retrieval accuracy. For example, if the PQ index is used, it will lead to information loss and further aggravate the impact on the retrieval accuracy; if the data is evenly distributed within each cluster and the difference between clusters is small, the HNSW index can more effectively utilize its hierarchical graph structure for navigation; when the proportion of the principal component of the intra-cluster covariance is relatively high, it indicates that the data has obvious correlation or trend in certain directions. In this case, using the L2 distance (Euclidean distance) as the metric standard may be more appropriate than using the Cosine similarity. Because the L2 distance can better capture the differences in these directions of the data; when there is the curse of dimensionality, reducing the dimension through principal component analysis (PCA) first and then building the index can improve the retrieval efficiency.
[0059] Based on this, the following recommended strategies can be obtained: when the data shows an obvious clustering structure in the vector space, it is suitable to select an index based on centroid partitioning (such as IVF_FLAT or IVF_PQ); for data with a relatively uniform distribution, the small-world graph structure (HNSW) can better play its advantage of jump ability; if the data has a long-tail structure but good local aggregation, IVF_FLAT / PQ + HNSW can achieve a better balance. When the dimension is very high (such as d > 2048), reducing the dimension through PCA first and then performing the IVF_Flat index can avoid the graph construction degradation caused by high-dimensional sparsity while ensuring the accuracy.
[0060] The above step S120 can be implemented as steps S121 to S124.
[0061] S121: For the vector data set, calculate the inter-cluster center distance, the proportion of the principal component of the covariance, and the proportion of noise points respectively, and determine whether there is a head-tail structure; S122: Judge whether the proportion of the principal component of the covariance exceeds the set threshold. If so, reduce the dimension of the vector data set. If not, directly execute step S123; S123: Recommend the index type according to the inter-cluster center distance: If the inter-cluster center distance exceeds the upper limit of the set threshold interval, it is considered that the clustering effect is good, and it is recommended to use the IVF index as a whole. Further recommend whether the intra-cluster index is the FLAT index or the PQ index according to the data scale: if the number of vectors in the vector data set exceeds the set threshold, the intra-cluster index is recommended to be the FLAT index, otherwise, the intra-cluster index is recommended to be the PQ index; If the inter-cluster center distance is within the set threshold range, the clustering effect is considered average. If there is a head-tail structure, it is recommended to use the IVF index for the entire main data part and the HNSW index for the sparse data part. And among them, according to the data scale of the main data part, it is recommended whether the intra-cluster index is the FLAT index or the PQ index, and the recommended idea is the same as above. If there is no head-tail structure, it is recommended to use the IVF index as a whole. Further, according to the data scale, it is recommended whether the intra-cluster index is the FLAT index or the PQ index, and the recommended idea is the same as above If the inter-cluster center distance is less than the lower limit of the set threshold range, the clustering effect is considered poor, and the recommended index is HNSW
[0062] S124: If the noise point ratio exceeds the upper limit of the set threshold range, it is prompted that the vector data set needs to be cleaned. If the noise point ratio is within the set threshold range, it is recommended to use the IVF index for the main data part and the HNSW index for the sparse data part. If the noise point ratio is lower than the lower limit of the set threshold, it is prompted to remove the noise vectors and then perform indexing
[0063] Exemplarily, the above step S120 specifically includes: (1) Determine whether the average silhouette coefficient S satisfies S ≥0.45. If so, execute step (2). If not, execute step (3); (2) Use the IVF index as a whole. Determine whether the proportion of covariance principal components satisfies var_ratio≥0.9. If so, use PCA to reduce the dimension to 256 or 128. If var_ratio<0.90, do not reduce the dimension. Further, select the IVF type according to the data scale and compression requirements: if the number of vectors < 1 million, recommend IVF_FLAT; if the number of vectors ≥ 1 million, recommend IVF_PQ
[0064] (3) If 0.30≤ S <0.45, it means that the clustering structure is medium. In this case, if there is a head-tail structure, use IVF_FLAT or IVF_PQ for the main data part and the HNSW index for the sparse data part. If there is no head-tail structure, recommend IVF_FLAT or IVF_PQ according to the data scale (4) If S <0.30, it indicates that the clustering is inseparable or there is overlap, and the recommended index is HNSW, using the default configuration (5) If noise_rate≥0.20, it is considered that this data sample needs to be cleaned or intervened manually before it can be used. If noise_rate∈[0.10,0.20), use the IVF index for the main data part and the HNSW index for the sparse data part. If noise_rate<0.10, remove the noise vectors and then perform indexing
[0065] 100,000, 1,000,000, and 10,000,000 are respectively selected as the sample data volumes to participate in the experiment. The sample data involves various modalities such as documents, pictures, and videos. The feature dimension is usually 768 or 1024, which is different from common recommendation systems (the dimension is between 32 and 256). Exemplarily, for 1,000,000 text vector sample data, the dimension of the text vector data is 768 (this dimension is usually used as the encoding dimension of the Chinese text model Bge model), and the initial number of clusters k = 1024. Based on the processing flow of step S110, Kmeans clustering analysis is performed on the vector data set based on the initial number of clusters, and the data distribution characteristics are extracted to optimize the number of clusters.
[0066] After optimizing the number of clusters, the data distribution characteristics are calculated again. According to the preset index recommendation strategy, the recommended index is determined. The recommendation process is as follows: (1) Calculate the data distribution characteristics: S = 0.47, indicating that the clustering structure is clear; the overall is balanced, but there is a significant tail; (2) var_ratio = 91.3, first use PCA compression to reduce the dimension to 256; (3) The proportion of noise points is low: noise_rate = 3%, directly remove the noise vectors. The final recommended index type is to use the IVF_PQ index for the main data part and the HNSW index for the sparse data part.
[0067] The above recommendation process takes the distance between cluster centers as the judgment benchmark, and gives different judgment logics for different clustering effects. For example, when the clustering effect is good, IVF is generally used, and whether to reduce the dimension is selected according to the main component distribution, and different construction strategies FLAT / PQ are recommended for different data volume scales; when the clustering effect is average, index recommendation is performed according to whether there is a head-tail structure, and the default index is used when the clustering effect is poor, realizing personalized recommendation based on different dimensions such as data scale, data dimension, and data internal structure distribution characteristics in the vector database. In the judgment process, dimension reduction, analysis of the proportion of noise points, etc. are also involved. Based on this, the storage of the vector database is optimized and the sample quality is improved.
[0068] In step S130, the index cost metrics use the index construction time and index size, and the query performance metrics use the recall rate and average query latency. Among them, the index construction time is used to measure the construction efficiency of the vector database under different configurations, the index size determines the index deployment cost, the recall rate is used to measure the accuracy of the query. If dimension reduction is performed on the vector database in step S120, the recall rate can also be used to measure whether there is excessive compression, and the average query latency is used to measure the query efficiency.
[0069] When the index construction time exceeds the set threshold, one or more of the following optimization strategies can be selected according to the data scale: (1) reduce the number of clusters; (2) if the current combination of IVF and other index types is used, adjust the in-cluster index type; (3) when the proportion of covariance principal components meets the set standard, perform dimensionality reduction on the vector data set.
[0070] When the index size exceeds the set threshold, one or more of the following optimization strategies can be selected according to the data scale: (1) enable the PQ index type; (2) when the proportion of covariance principal components meets the set standard, perform dimensionality reduction on the vector data set.
[0071] When the recall rate is lower than the set threshold, one or more of the following optimization strategies can be selected according to the data scale: (1) increase the number of clusters; (2) if the current combination of IVF and other index types is used, adjust the in-cluster index type and preferentially use the Float index type; (3) cancel dimensionality reduction; (4) disable the PQ index type.
[0072] When the average query latency exceeds the set threshold, one or more of the following optimization strategies can be selected according to the data scale: (1) if the current combination of IVF and other index types is used, replace the in-cluster index type with FLAT; (2) when the proportion of covariance principal components meets the set standard, perform dimensionality reduction on the vector data set; (2) reduce the number of clusters and limit the number of concurrent threads during query; (3) use HNSW.
[0073] For different data scales, when different index cost metrics and query performance metrics exceed the upper or lower limits, the specific index callback optimization strategies are shown in Table 4.
[0074] Table 4 Index Callback Optimization Strategies for Different Data Scales
[0075] One or more of the above embodiments provide a method for index recommendation based on the data distribution characteristics in a vector database and reverse verification through performance metrics. However, the above method still constructs an index for the current data set stored in the vector database and cannot be dynamically adjusted according to data changes or query pattern changes. Based on this, the vector database is partitioned. After constructing an index for the vector database based on the recommended index type, incremental data is received and stored on a separate shard. When the incremental data scale exceeds the set standard, the above index recommendation method is re-executed for the vector database, and a new index is established for the vector database. Among them, the set standard for the incremental data scale can be the proportion of the total amount (for example, the increment exceeds 15% of the total amount), or the number or size of incremental data samples, etc., which are not specifically limited here.
[0076] Based on the above method, one or more embodiments of the present invention further provide a vector database index recommendation device, as Figure 3 shown, including: a feature extraction module 401, configured to: perform clustering analysis on the vector data set, extract data distribution features based on the clustering result, and obtain a vector data portrait; the data distribution features include the inter-cluster center distance and whether there is a head-tail structure; an index recommendation module 402, configured to: according to a preset index recommendation strategy, determine a recommended index according to the vector data portrait; the index recommendation strategy includes: if the inter-cluster center distance exceeds the upper limit of the set threshold interval, the IVF index is adopted as a whole; if the inter-cluster center distance is within the set threshold interval, if there is no head-tail structure, the IVF index is adopted as a whole; if there is a head-tail structure, the IVF index is adopted for the backbone data part, and the HNSW index is adopted for the sparse data part; if the inter-cluster center distance is less than the lower limit of the set threshold interval, the recommended index is HNSW.
[0077] To verify and optimize the recommended index type, the device further includes an index evaluation module 403, configured to: perform index construction on the vector data set based on the recommended index, and evaluate the current index type according to the index construction cost index and the query performance index.
[0078] One or more embodiments of the present invention further provide an electronic device, which can be used to implement the method in the above embodiments. The electronic device includes one or more processors, one or more memories coupled to the processor, and a communication module coupled to the processor.
[0079] One or more embodiments of the present invention further provide an electronic device, which can be used to implement the method in the above embodiments. The electronic device includes one or more processors, one or more memories coupled to the processor, and a communication module coupled to the processor.
[0080] The memory in the embodiments of the present invention is used to store various types of data to support the execution of the method as Figure 1 shown.
[0081] It can be understood that the memory can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. The memory in the embodiments of the present invention can store computer programs corresponding to each step in the method as Figure 1 shown. Among them, the operating system includes various system programs, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program can include various application programs.
[0082] As an example, the processor can be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0083] In particular, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing Figure 1 the method shown. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by the central processing unit, various functions defined in the device of the present application are executed.
[0084] Among them, Figure 1 the computer program instructions corresponding to the method shown can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the processes Figure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.
[0085] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A vector database indexing and recommendation method, characterized in that, It includes the following steps: Perform clustering analysis on the vector data set, extract data distribution characteristics based on the clustering results, and obtain a vector data portrait; The data distribution characteristics include the inter-cluster center distance and the presence or absence of a head-tail structure; According to a preset index recommendation strategy, determine a recommended index based on the vector data portrait; The index recommendation strategy includes: If the inter-cluster center distance exceeds the upper limit of the set threshold interval, the IVF index is adopted as a whole; If the inter-cluster center distance is within the set threshold interval, if there is no head-tail structure, the IVF index is adopted as a whole; if there is a head-tail structure, the IVF index is adopted for the backbone data part, and the HNSW index is adopted for the sparse data part; If the inter-cluster center distance is less than the lower limit of the set threshold interval, the recommended index is HNSW.
2. The vector database index recommendation method according to claim 1, wherein After performing clustering analysis on the vector data set, the number of clusters in the clustering result is also optimized based on the data distribution within each cluster: Taking the number of clusters as a variable, set its change interval, perform clustering operations in sequence for multiple numbers of clusters within the change interval, calculate the intra-cluster data distribution parameters, and determine the corrected value of the number of clusters by the elbow method; According to the corrected value, re-perform the clustering operation, conduct a rationality test on the cluster distribution of the clustering result, obtain the number of data samples within each cluster, and further adjust the number of clusters according to the distribution of the sample numbers to obtain the optimal number of clusters.
3. The vector database indexing and recommendation method according to claim 2, wherein The rationality test of the cluster distribution is used to test the balance of the number of data samples within each cluster, including the following test contents: the existence of too many low-density clusters, the existence of too many fragmented clusters, whether the overall distribution of the data samples is uneven, and the inter-class separability.
4. The vector database indexing and recommendation method according to claim 1, wherein, When the IVF index is recommended for the whole or part of the vector data set, if the amount of data for which the IVF index is recommended exceeds the set threshold, the intra-cluster index is recommended as the FLAT index; otherwise, the intra-cluster index is recommended as the PQ index.
5. The vector database index recommendation method according to claim 1, wherein The data distribution characteristics also include the proportion of covariance principal components. If the proportion of covariance principal components exceeds the set threshold, the vector data set is dimensionally reduced and then the index recommendation is performed.
6. The vector database index recommendation method according to claim 1, wherein The data distribution characteristics also include the noise point ratio. If the noise point ratio exceeds the upper limit of the set threshold interval, it is prompted that the vector data set needs to be cleaned; if the noise point ratio is within the set threshold interval, the IVF index is recommended for the backbone data part, and the HNSW index is recommended for the sparse data part; if the noise point ratio is lower than the lower limit of the set threshold, it is prompted to remove the noise vectors and then perform the index recommendation.
7. The vector database index recommendation method according to claim 1, wherein The method further includes: performing index construction on the vector data set based on the recommended index, and evaluating the current index type according to the index construction cost index and the query performance index.
8. A vector database index recommendation device, characterized in that, It includes: A feature extraction module, configured to: perform clustering analysis on the vector data set, extract data distribution characteristics based on the clustering results, and obtain a vector data portrait; The data distribution characteristics include the inter-cluster center distance and the presence or absence of a head-tail structure; An index recommendation module, configured to: determine a recommended index according to a preset index recommendation strategy based on the vector data portrait; the index recommendation strategy includes: If the inter-cluster center distance exceeds the upper limit of the set threshold interval, the IVF index is adopted as a whole; If the inter-cluster center distance is within the set threshold range, if there is no head-tail structure, the overall uses the IVF index; if there is a head-tail structure, the main data part uses the IVF index and the sparse data part uses the HNSW index; If the inter-cluster center distance is less than the lower limit of the set threshold range, the recommended index is HNSW.
9. An electronic device, characterized in that, It includes a processor and a memory, and computer instructions are stored on the memory. When the computer instructions are executed by the processor, the electronic device is enabled to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for searching high-dimensional vector combining clustering and city block distances
CN103514264A
Clustering separation distributive indexing method
CN105868414A
Data indexing method and system
CN115640426A
A vector database retrieval method and system
CN119782315A
Intelligent document retrieval and generation system based on metadata driving
CN120104624A
Cited By
ViT model and Faiiss database-based image searching method and system
CN120950723A
Vector database-oriented online incremental learning index system and construction method thereof
CN121681531A