Method and device for constructing vector index based on sub-neighbor network, medium and equipment

Through the combination of clustering and sub-neighbor networks, the problems of high computational burden and low query efficiency under high precision requirements are solved, and more efficient and accurate vector retrieval is achieved.

CN120086218APending Publication Date: 2025-06-03BEIJING HAIPU WANGJU TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510160083.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When dealing with high-precision requirements, traditional cluster indexing methods have problems such as high computational burden and low query efficiency, which is difficult to meet the user needs of both immediate feedback and accurate results.

Method used

By clustering the dataset, dividing it into multiple clusters, and creating a sub-neighbor network within each cluster, recording the information of each vector and its neighbor vectors to achieve a finer index structure.

Benefits of technology

Reduces invalid calculations, improves the accuracy and efficiency of vector retrieval, and can significantly improve search speed and user experience, especially when processing large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086218A_ABST
    Figure CN120086218A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for constructing a vector index based on a sub-neighbor network, a storage medium and equipment, and the method comprises the steps: carrying out the clustering of a whole data set, dividing the data set into a plurality of clusters, enabling each cluster to comprise a plurality of vectors which are close to each other, creating a sub-neighbor network in each cluster, and carrying out the clustering of the sub-neighbor network, and the sub-neighbor network performs retrieval according to the to-be-queried vectors of the cluster and the sub-neighbor network by calculating the relative distance between each vector and other vectors in the cluster and recording the information of each vector and the neighbor vector thereof in the sub-neighbor network, and returns a retrieval result. According to the method, through clustering, the search space is effectively reduced, target-free search in the whole data set is avoided, and meanwhile, a sub-neighbor network (SubNN) is constructed in each cluster, so that when query is performed in the clusters, potential nearest vectors can be found more quickly, unnecessary distance calculation is reduced, and the search efficiency is improved during query. By preferentially accessing adjacent vectors, unnecessary amount of calculation can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of vector databases and information retrieval, and particularly to a method, apparatus, medium, and device for constructing a vector index based on a sub-neighborhood network. Background Art

[0002] With the rapid progress of deep learning technology and the increasing popularity of big data analysis, vector data has become the cornerstone of modern information retrieval systems and is widely used in many cutting-edge fields such as search engines, e-commerce platforms, personalized recommendation systems, and natural language processing. These application scenarios all rely on the efficient and accurate representation of diverse data types such as videos, texts, images, and audios, and the core of all this lies in the fast nearest neighbor search ability for high-dimensional vector data.

[0003] Currently, in order to address the dual challenges of timeliness and accuracy in high-dimensional vector data retrieval, mainstream vector database systems generally adopt a clustering-based indexing strategy. Such methods effectively reduce the search space by intelligently partitioning the vector space into several clusters, significantly improving the retrieval efficiency. However, although these technologies have alleviated the problem of large-scale data retrieval to a certain extent, traditional clustering indexing methods still struggle to handle high-precision requirements. Specifically, due to the overly broad clustering granularity, a large number of unnecessary distance calculations need to be processed during the query process, which not only increases the computational burden but also directly weakens the query efficiency, making it difficult to meet the growing user demands for both instant feedback and accurate results. Summary of the Invention

[0004] In view of this, the present invention provides a method, apparatus, medium, and device for constructing a vector index based on a sub-neighborhood network, which can reduce invalid calculations and thus achieve a better balance between the accuracy and efficiency of vector retrieval.

[0005] In a first aspect, an embodiment of the present invention provides a method for constructing a vector index based on a sub-neighborhood network, the method comprising:

[0006] Clustering the entire data set, dividing the data set into multiple clusters, and each cluster contains multiple vectors with similar distances;

[0007] Creating a sub-neighborhood network within each cluster, the sub-neighborhood network calculates the relative distances between each vector and other vectors within the cluster, and records the information of each vector and its neighbor vectors in the sub-neighborhood network;

[0008] Retrieving according to the cluster and the sub-neighbor network for the vector to be queried.

[0009] Further, clustering the entire data set, dividing the data set into multiple clusters, and each cluster contains multiple vectors with similar distances includes:

[0010] Preprocess the data in the entire dataset and convert the data into vector form;

[0011] Determine several cluster centers, calculate and save the distance from each vector to the cluster centers, and allocate the vectors to the clusters corresponding to the nearest cluster centers according to the distances, thereby dividing the dataset into multiple clusters.

[0012] Further, determining several cluster centers includes:

[0013] Randomly select K data points as the initial cluster centers;

[0014] Allocate the vectors to the nearest clusters according to the distances between each vector and the cluster centers;

[0015] Take the average value of all points within the cluster as the new cluster center;

[0016] Repeat the steps of allocating the vectors to the nearest clusters and calculating the new centers until the change in the cluster center values is within a preset range or the preset number of iterations is reached.

[0017] Further, determining several cluster centers includes:

[0018] Randomly select a data point as the initial cluster center;

[0019] Select the data point farthest from the initial cluster center as the second cluster center;

[0020] Successively select the data points farthest from the previous cluster centers as the new cluster centers until the preset number of cluster centers is selected.

[0021] Further, determining several cluster centers includes:

[0022] Use the hierarchical clustering algorithm or the Canopy algorithm to perform preliminary clustering on the dataset to obtain multiple preliminary cluster centers;

[0023] Select a data point from each preliminary cluster center to obtain several cluster centers.

[0024] Further, create a sub-neighborhood network within each cluster. The sub-neighborhood network calculates the relative distances between each vector and other vectors within the cluster and records the information of each vector and its neighbor vectors in the sub-neighborhood network, including:

[0025] Calculate the Euclidean distance between each vector within the cluster and other vectors;

[0026] Sort the Euclidean distances according to the distance values;

[0027] Determine the nearest neighbor of each vector according to the sorted Euclidean distance;

[0028] Save the identifier of the nearest neighbor of each vector and the corresponding distance.

[0029] Further, create a sub-neighborhood network within each cluster. The sub-neighborhood network calculates the relative distance between each vector and other vectors within the cluster, and records the information of each vector and its neighbor vectors in the sub-neighborhood network, including:

[0030] Calculate the cosine value of the angle between each vector within the cluster and other vectors;

[0031] Sort the cosine values of the angles according to their magnitudes;

[0032] Determine the nearest neighbor of each vector according to the sorted cosine values of the angles;

[0033] Save the identifier of the nearest neighbor of each vector by angle and the corresponding cosine value of the angle.

[0034] In a second aspect, an embodiment of the present invention provides an apparatus for constructing a vector index based on a sub-neighborhood network. The apparatus includes:

[0035] A clustering module for clustering the entire data set, dividing the data set into multiple clusters, and each cluster contains multiple vectors with similar distances;

[0036] A sub-neighborhood network creation module for creating a sub-neighborhood network within each cluster. The sub-neighborhood network calculates the relative distance between each vector and other vectors within the cluster, and records the information of each vector and its neighbor vectors in the sub-neighborhood network;

[0037] A retrieval module for retrieving a vector to be queried according to the cluster and the sub-neighbor network.

[0038] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the method described in any one of the first aspects when running.

[0039] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of the first aspects.

[0040] The technical solution provided by the embodiments of the present invention divides the data set into multiple clusters through clustering, effectively reducing the search space. When querying, it is possible to first locate the cluster to which the query vector is most likely to belong, thus avoiding aimless searching in the entire data set. At the same time, a sub-nearest neighbor network (SubNN) is constructed within each cluster to further refine the index structure of the vectors within the cluster. This refinement enables more rapid finding of potential nearest neighbor vectors when querying within the cluster, reducing unnecessary distance calculations. When querying, by preferentially accessing neighboring vectors, the unnecessary computational amount can be significantly reduced. This optimization strategy makes the query process more efficient, especially when dealing with large-scale data sets, and can significantly improve the retrieval speed and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flowchart of a method for constructing a vector index based on a sub-nearest neighbor network provided by Embodiment 1 of the present invention;

[0042] Figure 2 is a schematic structural diagram of a device for constructing a vector index based on a sub-nearest neighbor network provided by Embodiment 2 of the present invention;

[0043] Figure 3 is a schematic structural diagram of an electronic device provided by Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0045] Embodiment 1

[0046] Refer to Figure 1 , Figure 1 which is a flowchart of a method for constructing a vector index based on a sub-nearest neighbor network provided by an embodiment of the present invention. The method includes the following steps:

[0047] Step 11: Cluster the entire data set, divide the data set into multiple clusters, and each cluster contains multiple vectors with similar distances.

[0048] In this step, data set clustering is the cornerstone of the entire method. The purpose of this step is to divide the entire data set into multiple clusters to ensure that the vectors within each cluster are relatively close under a certain distance metric. Through clustering, similar vectors can be grouped together, thereby narrowing the search scope and improving the query efficiency in the subsequent query process.

[0049] After clustering is completed, the dataset is divided into multiple compact clusters. These clusters serve as the basis for constructing subsequent sub-neighborhood networks, helping to reduce unnecessary computational effort and improve the accuracy and speed of queries.

[0050] Step 12: Create a sub-neighborhood network within each cluster. The sub-neighborhood network calculates the relative distance between each vector and other vectors within the cluster and records the information of each vector and its neighbor vectors in the sub-neighborhood network.

[0051] In this step, creating a sub-neighborhood network (SubNN) within each cluster is to further refine the index structure of the vectors within the cluster. This step constructs a fine-grained local index structure by calculating the relative distance between each vector and other vectors within the cluster and recording the information of each vector and its neighbor vectors.

[0052] The creation of the sub-neighborhood network enables more rapid finding of potential nearest neighbor vectors when querying within the cluster. By leveraging neighbor relationships, unnecessary distance calculations can be significantly reduced, further improving query efficiency. At the same time, this local index structure also helps to enhance the accuracy of query results.

[0053] Step 13: Retrieve the vectors to be queried based on the cluster and the sub-neighbor network.

[0054] In this step, when a vector to be queried is given, it is first necessary to determine the cluster to which it belongs, and then use the sub-neighborhood network within that cluster for precise retrieval. By accessing the neighbor vectors of the vector to be queried and their neighbors' neighbors (performing depth or breadth search as needed), potential nearest neighbor vectors can be quickly filtered out.

[0055] Using the cluster and the sub-neighborhood network for retrieval can significantly improve query efficiency. By narrowing the search scope and reducing unnecessary computational effort, the vector closest to the vector to be queried can be found more quickly. In addition, this retrieval method also helps to enhance the stability and reliability of query results.

[0056] The implementation process of each step will be introduced in detail below.

[0057] In some embodiments, the implementation of step 11 may include the following steps:

[0058] Step 11: Preprocess the data in the entire dataset and convert the data into vector form.

[0059] In this step, preprocessing the data in the entire dataset mainly aims to improve data quality. Converting the dataset into vector form is to meet the input requirements of subsequent algorithms (such as clustering algorithms).

[0060] Specifically, step 11 may be implemented in the following ways:

[0061] Data cleaning: Remove noise, missing values, and outliers from the data. This is an important step to ensure data quality and helps improve the accuracy of subsequent clustering analysis. Noisy data may interfere with the accuracy of clustering; missing values need to be filled to avoid errors in subsequent calculations; outliers may need to be handled according to specific situations, such as deletion or replacement.

[0062] Feature selection: Select features from the original dataset that have an important impact on the clustering results. This helps reduce the dimensionality of the dataset, lower the computational complexity, and may improve the accuracy of the clustering results.

[0063] Data transformation: Transform the selected features into vector form. This usually involves normalizing or standardizing the data to ensure that different features are numerically comparable. Standardization is to transform the data into a form with zero mean and unit variance, while normalization is to scale the data to a specific range (such as [0,1]).

[0064] Step 12: Determine several cluster centers, calculate and save the distance from each vector to the cluster centers, and assign the vectors to the clusters corresponding to the nearest cluster centers according to the distances, thereby dividing the dataset into multiple clusters.

[0065] In this step, it can be achieved in the following ways:

[0066] Step 111a: Randomly select K data points as the initial cluster centers.

[0067] In this step, randomly select K data points from the dataset, and these points will be used as the initial cluster centers. The selection of K is usually based on the characteristics of the data and the goals of the clustering analysis.

[0068] Step 112a: Assign the vectors to the nearest clusters according to the distance from each vector to the cluster centers;

[0069] In this step, according to the positions of the current cluster centers, assign each data point (or vector) to the nearest cluster. Use a distance metric (such as Euclidean distance) to calculate the distances from each data point to all cluster centers. Assign each data point to the cluster to which the nearest cluster center belongs.

[0070] Step 113a: Take the average value of all points within the cluster as the new cluster center.

[0071] In this step, calculate the average value (i.e., the mean vector) of all data points within each cluster. Take the calculated mean vector as the new cluster center.

[0072] Step 114a: Repeat the steps of allocating vectors to the nearest clusters and calculating new centers until the change in the cluster center values is within a preset range or the preset number of iterations is reached.

[0073] In this step, through the iterative process, the positions of the cluster centers are gradually optimized to make the clustering results more stable. Repeat Step 112a and Step 113a until one of the following conditions is met: The change in the position of the cluster center is within the preset range (i.e., the convergence condition), which indicates that the position of the cluster center has become relatively stable and the clustering results will no longer change significantly. The preset number of iterations is reached, which is to prevent the algorithm from falling into an infinite loop or over-computation.

[0074] In some other embodiments of the present invention, Step 12 can also be implemented in the following manner:

[0075] Step 111b: Randomly select a data point as the initial cluster center.

[0076] This step provides a starting point for the subsequent selection of cluster centers. Randomly pick a data point from the dataset as the first initial cluster center. This step is similar to the random selection in the standard K-means algorithm, but it is just the beginning of the selection process.

[0077] Step 112b: Select the data point farthest from the initial cluster center as the second cluster center.

[0078] This step ensures that the second cluster center is as dispersed as possible from the first cluster center in the data space, thereby increasing the diversity of subsequent clusters. Calculate the distances between all data points and the first cluster center, and select the data point with the farthest distance as the second cluster center.

[0079] Step 113b: Sequentially select the data points farthest from the previous several cluster centers as new cluster centers until the preset number of cluster centers is selected.

[0080] In this step, by maximizing the distances between the cluster centers, the overlap and redundancy between the initial cluster centers are reduced, and the quality and stability of the clustering are improved. After n cluster centers have been selected (n is less than the preset number of cluster centers K), calculate the distances between the remaining data points and these n cluster centers, and find the data point with the farthest distance from all the current cluster centers as the (n + 1)-th cluster center. Repeat this process until the preset number K of cluster centers is selected.

[0081] In some other embodiments of the present invention, Step 11 can also be implemented through the following steps:

[0082] Step 111c: Use the hierarchical clustering algorithm or the Canopy algorithm to perform a preliminary clustering on the dataset to obtain multiple preliminary cluster centers.

[0083] In this step, the hierarchical clustering or the clustering ability of the Canopy algorithm is utilized to preliminarily partition the dataset and identify the natural cluster structure in the data.

[0084] Hierarchical clustering: By recursively merging or splitting data points, a dendrogram is constructed, and the dendrogram is cut at an appropriate level to obtain preliminary clusters. The center of each cluster can be obtained by calculating the average value or other statistics of the data points within the cluster.

[0085] Canopy algorithm: First, two distance thresholds T1 and T2 (T1 > T2) are selected, and then the data points are iteratively processed. Points are assigned to the nearest Canopy (if the distance to a certain Canopy is less than T1), or a new Canopy is created (if the distance to all Canopies is greater than T1). During the iteration, if the distance between point P and a certain Canopy is less than T2, it is removed from the dataset to avoid repeated assignment. Finally, each Canopy can be regarded as a preliminary cluster, and its center can be the average value or other representative points of the data points within the Canopy.

[0086] Step 112c: Select a data point from each preliminary cluster center to obtain several cluster centers.

[0087] In this step, representative points are selected from the preliminary clusters as the initial cluster centers for subsequent clustering.

[0088] For each preliminary cluster, a data point within the cluster can be selected as the initial cluster center. This point can be the centroid of the cluster (i.e., the average value of all points within the cluster), or an actual data point within the cluster that is closest to the centroid, or a representative point selected according to other criteria (such as density, distance, etc.).

[0089] When selecting the initial cluster centers, it should be ensured that they can represent the characteristics and structures of their respective clusters so that stable and meaningful clusters can be formed during the subsequent clustering process.

[0090] Using the hierarchical clustering algorithm or the Canopy algorithm to select the initial cluster centers for subsequent algorithms such as K-means clustering is an effective method. This method identifies the natural cluster structure in the dataset through preliminary clustering and selects representative cluster centers from it, thereby improving the quality and efficiency of subsequent clustering. In practical applications, it is necessary to select appropriate preliminary clustering algorithms and initial cluster center selection methods according to the characteristics of the data and the goals of the clustering analysis.

[0091] In some embodiments of the present invention, step 12 creates a sub-proximity network within each cluster. The sub-proximity network can be implemented by calculating the relative distance between each vector and other vectors within the cluster and recording the information of each vector and its neighbor vectors in the sub-proximity network through the following steps:

[0092] Step 121: Calculate the Euclidean distance between each vector within the cluster and other vectors.

[0093] The purpose of this step is to quantify the similarity or difference between vectors within the cluster. For each vector within the cluster, calculate the Euclidean distance between it and all other vectors within the cluster.

[0094] Step 122: Sort the Euclidean distances in ascending order of the distance values.

[0095] In this step, the distances between vectors are arranged in ascending order to determine the nearest neighbors subsequently.

[0096] For each vector, sort the distance values between it and other vectors. Sorting algorithms (such as quicksort, mergesort, etc.) can be used to improve efficiency.

[0097] Step 123: Determine the nearest neighbor of each vector based on the sorted Euclidean distances.

[0098] In this step, find the nearest neighbor of each vector (i.e., the vector with the smallest distance).

[0099] For each vector, view the sorted distance list and select the vector with the smallest distance as the nearest neighbor. If multiple neighbors are needed, the first k vectors with the smallest distances can be selected as the k nearest neighbors.

[0100] Step 124: Save the identifier of the nearest neighbor of each vector and the corresponding distance.

[0101] In this step, record the information of each vector and its nearest neighbor for subsequent analysis and use.

[0102] Specifically, a data structure (such as a dictionary, list, etc.) can be created to store the nearest neighbor information of each vector. For each vector, save the identifier (such as index, ID, etc.) of its nearest neighbor and the corresponding distance value into the data structure.

[0103] In some other embodiments of the present invention, step 12 creates a sub-proximity network within each cluster. The sub-proximity network can also be implemented by calculating the relative distance between each vector and other vectors within the cluster and recording the information of each vector and its neighbor vectors in the sub-proximity network through the following steps:

[0104] Step 121: Calculate the cosine value of the angle between each vector within the cluster and other vectors.

[0105] The purpose of this step is to quantify the directional similarity between vectors within the cluster.

[0106] Specifically, for each vector within the cluster, calculate the cosine value of the angle between it and all other vectors within the cluster. The range of the cosine value of the angle is [-1, 1]. The closer the value is to 1, the more similar the directions of the vectors are. The closer the value is to -1, the more opposite the directions are. The value close to 0 indicates that the directions are almost perpendicular.

[0107] Step 122: Sort the cosine values of the angles in descending order.

[0108] In this step, sort the cosine values of the angles between vectors in descending order to determine the nearest neighbors later. For each vector, sort the cosine values of the angles between it and other vectors. Since the larger the cosine value of the angle, the more similar the directions are, the largest value should be placed in the front during sorting.

[0109] Step 123: Determine the nearest neighbor of each vector according to the sorted cosine values of the angles.

[0110] In this step, find the nearest neighbor (i.e., the vector with the most similar direction) of each vector. For each vector, view the sorted list of cosine values of the angles and select the vector with the largest value as the nearest neighbor. If multiple neighbors are needed, the first k vectors with the largest values can be selected as the k nearest neighbors.

[0111] Step 124: Save the identifier of the angular nearest neighbor of each vector and the corresponding cosine value of the angle.

[0112] In this step, record the information of each vector and its nearest neighbor for subsequent analysis and use.

[0113] Specifically, a data structure (such as a dictionary, a list, etc.) can be created to store the information of the nearest neighbor of each vector. For each vector, save the identifier (such as index, ID, etc.) of its nearest neighbor and the corresponding cosine value of the angle into the data structure.

[0114] Through the above steps, a sub-neighborhood network can be created within each cluster based on the cosine value of the angle, and the information of each vector and its neighbor vector can be recorded. This method is particularly suitable for analyzing the directional similarity in high-dimensional data, rather than just the proximity in terms of distance. In practical applications, the details in the steps can be adjusted according to needs, such as selecting other similarity measurement methods, considering multiple neighbors, etc.

[0115] In some embodiments of the present invention, step 13 retrieves according to the cluster and the query vector of the sub-neighbor network, and the retrieval result can be returned by the following method:

[0116] Step 131: Build an index for each cluster and the corresponding sub-neighbor network of the cluster. The index includes the identifier of the cluster, the index of the vectors within the cluster, and the list of the nearest neighbors of each vector, etc.

[0117] In this step, the purpose of building the index is to improve the retrieval efficiency and quickly locate the relevant clusters and vectors.

[0118] Specifically, a unique identifier (such as ID or name) can be assigned to each cluster, the index of each vector within the cluster can be recorded for quick access, the list of the nearest neighbors of each vector can be built and stored in the index, and a database, hash table or other efficient data structures can be selected to store the index.

[0119] Step 132: Preprocess the query vector so that the preprocessed query vector can be compared with the vectors within the cluster in the same feature space.

[0120] The purpose of this step is to ensure that the query vector is compared with the vectors within the cluster in the same feature space to improve the retrieval accuracy. Specifically, appropriate preprocessing steps can be selected according to the characteristics of the data, such as normalization, dimensionality reduction (such as PCA, LDA, etc.), feature selection, etc. Apply the preprocessing steps to the query vector to make it have the same feature dimension and distribution as the vectors within the cluster.

[0121] Step 133: Select the cluster most similar to the preprocessed query vector as the initial search cluster according to the characteristics of the preprocessed query vector.

[0122] The purpose of this step is to narrow the search scope and improve the retrieval efficiency. Specifically, the similarity metric (such as cosine value of the angle, Euclidean distance, etc.) between the query vector and the center (or representative vector) of each cluster can be calculated, and the cluster with the highest similarity is selected as the initial search cluster. A similarity threshold can be set, and only when the similarity exceeds this threshold, the cluster is selected as the initial search cluster.

[0123] Step 134: Use the sub-neighbor network to perform an exact retrieval of the query vector within the initial search cluster.

[0124] In this step, the vector most similar to the query vector is found within the initial search cluster. Specifically, starting from the nearest neighbor of the query vector, it can be gradually extended to farther neighbors, and the neighbor relationship recorded in the sub-neighbor network can be used to quickly locate the relevant vector. A retrieval depth or neighbor number limit can be set to control the breadth and depth of the retrieval.

[0125] In some other embodiments of the present invention, if no satisfactory result is found within the initial search cluster, it is possible to consider expanding to other clusters for retrieval. This can be achieved by calculating the similarity between the query vector to be processed and the centers of other clusters, and selecting the clusters with higher similarity for further retrieval.

[0126] Step 135: According to the retrieval conditions, select the vectors that meet the conditions from the retrieved vectors as the final result and return them.

[0127] In this step, according to the retrieval conditions, the vectors that meet the requirements are filtered out. Specifically, the retrieved vectors can be filtered according to specific retrieval conditions (such as similarity threshold, quantity limit, category label, etc.). The vectors that meet the conditions are returned to the user or subsequent processing steps as the final result.

[0128] Through the above steps, the query vector can be retrieved efficiently by using the cluster and sub-neighbor network. In practical applications, it may be necessary to adjust and optimize the steps according to the specific data characteristics and retrieval requirements. For example, more complex similarity measurement methods can be introduced, more efficient index structures can be used, or parallel processing and other technologies can be adopted to improve the efficiency and accuracy of retrieval.

[0129] The technical solution provided by the embodiments of the present invention divides the data set into multiple clusters through initial clustering, effectively reducing the search space. During querying, it is possible to first locate the cluster to which the query vector most likely belongs, thus avoiding aimless search in the entire data set. At the same time, a sub-neighbor network (SubNN) is constructed within each cluster to further refine the index structure of the vectors within the cluster. This refinement enables potential nearest neighbor vectors to be found more quickly within the cluster during querying, reducing unnecessary distance calculations. During querying, by preferentially accessing neighboring vectors, a large amount of unnecessary computational effort can be significantly reduced. This optimization strategy makes the query process more efficient, especially when dealing with large-scale data sets, and can significantly improve the retrieval speed and user experience.

[0130] Embodiment 2

[0131] See Figure 2 , Figure 2 which is a schematic structural diagram of a vector index device based on a sub-neighbor network provided by Embodiment 2 of the present invention. The device includes:

[0132] A clustering module 21, configured to perform clustering on the entire data set, divide the data set into multiple clusters, and each cluster contains multiple vectors with close distances;

[0133] A sub - proximity network creation module 22, configured to create a sub - proximity network within each cluster. The sub - proximity network calculates the relative distance between each vector and other vectors within the cluster, and records the information of each vector and its neighbor vectors in the sub - proximity network;

[0134] A retrieval module 23, configured to perform retrieval according to the cluster and the vector to be queried in the sub - neighbor network, and return the retrieval result.

[0135] In some embodiments of the present invention, the clustering module 21 may include:

[0136] A pre - processing unit 211, configured to pre - process the data in the entire dataset and convert the data into a vector form;

[0137] A clustering unit 212, configured to determine several cluster centers, calculate and save the distance from each vector to the cluster centers, and allocate the vectors to the clusters corresponding to the nearest cluster centers according to the distances, thereby dividing the dataset into multiple clusters.

[0138] In some embodiments of the present invention, the clustering unit 212 may include:

[0139] A random selection subunit 2121a, configured to randomly select K data points as initial cluster centers;

[0140] An assignment subunit 2122a, configured to allocate vectors to the nearest clusters according to the distance between each vector and the cluster centers;

[0141] A cluster center update subunit 2123a, configured to use the average value of all points within the cluster as the new cluster center;

[0142] An iteration subunit 2124a, configured to repeat the steps of allocating vectors to the nearest clusters and calculating new centers until the change in the numerical value of the cluster centers is within a preset range or reaches a preset number of iterations.

[0143] In other embodiments of the present invention, the clustering unit 212 may include:

[0144] A second random selection subunit 2121b, configured to randomly select a data point as the initial cluster center;

[0145] A second cluster center selection unit 2122b, configured to select the data point farthest from the initial cluster center as the second cluster center;

[0146] A new cluster center selection unit 2123b, configured to sequentially select the data points farthest from the previous cluster centers as new cluster centers until a preset number of cluster centers are selected.

[0147] In some embodiments of the present invention, the sub - proximity network creation module 22 may include:

[0148] A distance calculation unit 221a for calculating the Euclidean distance between each vector in the cluster and other vectors;

[0149] A first sorting unit 222a for sorting the Euclidean distances according to the magnitude of the distance values;

[0150] A first determination unit 223a for determining the nearest neighbor of each vector according to the sorted Euclidean distances;

[0151] A first storage unit 224a for storing the identifier of the nearest neighbor of each vector and the corresponding distance.

[0152] In some other embodiments of the present invention, the sub-neighborhood network creation module 22 may include:

[0153] A cosine value calculation unit 221b for calculating the cosine value of the angle between each vector in the cluster and other vectors;

[0154] A second sorting unit 222b for sorting the cosine values of the angles according to the magnitude;

[0155] A second determination unit 223b for determining the nearest neighbor of each vector according to the sorted cosine values of the angles;

[0156] A second storage unit 224b for storing the identifier of the nearest neighbor of each vector in terms of angle and the corresponding cosine value of the angle.

[0157] In some embodiments of the present invention, the retrieval module 23 may include:

[0158] An index construction unit 231 for constructing an index for each cluster and the corresponding sub-neighborhood network thereof, the index including the identifier of the cluster, the index of the vectors in the cluster, and the list of the nearest neighbors of each vector;

[0159] A preprocessing unit 232 for preprocessing the vector to be queried so that the preprocessed vector to be queried can be compared with the vectors in the cluster in the same feature space;

[0160] A selection unit 233 for selecting the cluster most similar to it as the initial search cluster according to the features of the preprocessed vector to be queried;

[0161] An exact retrieval unit 234 for performing an exact retrieval of the vector to be queried in the initial search cluster by using the sub-neighborhood network;

[0162] A result return unit 235 for selecting the vectors that meet the conditions from the retrieved vectors as the final result and returning them according to the retrieval conditions.

[0163] Thus, the technical solution provided by the embodiments of the present invention divides the data set into multiple clusters through initial clustering, effectively reducing the search space. During querying, it is possible to first locate the cluster to which the query vector is most likely to belong, thus avoiding aimless searching in the entire data set. At the same time, a SubNN (Sub Nearest Neighbor) is constructed within each cluster to further refine the index structure of the vectors within the cluster. This refinement enables potential nearest neighbor vectors to be found more quickly within the cluster, reducing unnecessary distance calculations. During querying, by preferentially accessing neighboring vectors, a large amount of unnecessary computational effort can be significantly reduced. This optimization strategy makes the query process more efficient, especially when dealing with large-scale data sets, and can significantly improve the retrieval speed and user experience.

[0164] It should be noted that the vector index construction device based on the SubNN in the embodiments of the present invention and the vector index construction method in the above embodiments belong to the same inventive concept. Technical details not elaborated in this device can be referred to the relevant descriptions of the method above and will not be repeated here.

[0165] In addition, the embodiments of the present invention further provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the method described above when running.

[0166] Figure 3 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0167] As Figure 3As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0168] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0169] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the idle detection method.

[0170] In some embodiments, the idle detection method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the idle detection method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the idle detection method in any other appropriate manner (e.g., by means of firmware).

[0171] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0172] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0173] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0174] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0175] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0176] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0177] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0178] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing a vector index based on a sub-neighborhood network, characterized in that: The method comprises: Clustering the entire data set, dividing the data set into multiple clusters, each cluster containing multiple vectors with similar distances; Creating a sub-neighborhood network in each cluster, wherein the sub-neighborhood network calculates the relative distance between each vector and other vectors in the cluster, and records information about each vector and its neighbor vectors in the sub-neighborhood network; A search is performed according to the cluster and the sub-neighborhood network query vector, and a search result is returned.

2. The method according to claim 1, characterized in that The entire data set is clustered and divided into multiple clusters. Each cluster contains multiple vectors with similar distances, including: Preprocessing the data in the entire data set to convert the data into a vector form; A number of cluster centers are determined, the distance from each vector to the cluster center is calculated and saved, and the vector is assigned to a cluster corresponding to the cluster center closest to the vector according to the distance, thereby dividing the data set into a plurality of clusters.

3. The method according to claim 2, characterized in that Determining several cluster centers includes: Randomly select K data points as the initial cluster centers; Assign each vector to the closest cluster based on its distance from the center of the cluster; The average value of all points in the cluster is used as the new cluster center; The steps of assigning the vector to the nearest cluster and calculating the new center are repeated until the cluster center value changes within a preset range or the preset number of iterations is reached.

4. The method according to claim 2, characterized in that: Determining several cluster centers includes: Randomly select a data point as the initial cluster center; Select the data point farthest from the initial cluster center as the second cluster center; The data points with the farthest distance from the previous cluster centers are selected in sequence as new cluster centers until a preset number of cluster centers are selected.

5. The method according to claim 1, characterized in that A sub-neighborhood network is created in each cluster. The sub-neighborhood network calculates the relative distance between each vector and other vectors in the cluster, and records the information of each vector and its neighbor vectors in the sub-neighborhood network, including: Calculate the Euclidean distance between each vector in the cluster and other vectors; Sort the Euclidean distances according to the distance values; Determine the nearest neighbor of each vector based on the sorted Euclidean distance; Save the identity of each vector's nearest neighbor and the corresponding distance.

6. The method according to claim 1, characterized in that A sub-neighborhood network is created in each cluster. The sub-neighborhood network calculates the relative distance between each vector and other vectors in the cluster, and records the information of each vector and its neighbor vectors in the sub-neighborhood network, including: Calculate the cosine of the angle between each vector in the cluster and every other vector; Sort the angle cosine values ​​by size; Determine the nearest neighbors of each vector based on the sorted angle cosines; Save the identities of each vector's angular nearest neighbors and the corresponding angular cosine values.

7. The method according to claim 1, characterized in that Search is performed according to the cluster and the sub-neighborhood network query vector, and the search results returned include: Constructing an index for each cluster and the sub-neighborhood network corresponding to the cluster, wherein the index includes the identifier of the cluster, the index of the vector within the cluster, and the nearest neighbor list of each vector; Preprocessing the query vector so that the preprocessed query vector and the cluster vector can be compared in the same feature space; According to the characteristics of the preprocessed query vector, the most similar cluster is selected as the initial search cluster; Using a sub-neighborhood network, accurately searching the query vector within the initial search cluster; According to the search conditions, the vectors that meet the conditions are selected from the searched vectors and returned as the final result.

8. A device for constructing a vector index based on a sub-neighborhood network, characterized in that: The device comprises: A clustering module is used to cluster the entire data set and divide the data set into multiple clusters, each of which contains multiple vectors with similar distances; A sub-neighborhood network creation module is used to create a sub-neighborhood network in each cluster, wherein the sub-neighborhood network calculates the relative distance between each vector and other vectors in the cluster, and records the information of each vector and its neighbor vectors in the sub-neighborhood network; The retrieval module is used to perform retrieval according to the cluster and the sub-neighborhood network query vector, and return the retrieval result.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.