Similarity search method and system based on vector database
By clustering vector data in polar coordinates and determining sub-vectors and centroids using improved PQ encoding, the problem of difficult number of sub-vectors in traditional technology is solved, and efficient and high-precision vector database similarity search is achieved.
Patent Information
- Application Number
- CN202510277979.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-10
AI Technical Summary
When traditional PQ encoding determines the number of subvectors, it is difficult to balance the search efficiency and search accuracy, resulting in low efficiency and accuracy of search results.
By clustering vector data in polar coordinates, multiple global clusters are obtained, and the subvectors and centroids of the target cluster are determined using improved PQ encoding, thereby optimizing the number of subvectors to improve search accuracy and efficiency.
It realizes the efficiency of search efficiency while improving the search accuracy. By optimizing the number and centroid error of subvectors, the similarity search performance of vector database is significantly improved.
Smart Images

Figure CN120179872A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data retrieval, and particularly to a similarity search method and system based on a vector database. Background Art
[0002] With the rapid development of big data technology, the quantity of unstructured data such as images, videos, and audios has increased explosively, and this unstructured data is stored in a vector database in the form of vectors. When performing data recognition processing such as face recognition or speech recognition on unstructured data, it is often necessary to perform similarity search on the vector database to obtain multiple vector data with the highest similarity, and finally obtain the recognition result according to the data categories of these vector data.
[0003] Currently, Product Quantization (PQ) coding is a fast similarity search technology for vector databases. Traditional PQ coding first decomposes a high-dimensional vector into M sub-vectors, and uses a clustering algorithm to divide the vector space of each sub-vector into K clustering clusters. The clustering centers of the K clustering clusters are the centroids of the sub-vectors; for any one sub-vector, find the index of the clustering cluster to which the sub-vector belongs in the vector space, and use this index to represent this sub-vector. Finally, the high-dimensional vector can be encoded as a combination of M indexes, and subsequent vector search is implemented according to the combination of M indexes; for example, if the high-dimensional vector is 128-dimensional, the high-dimensional vector is decomposed into 8 sub-vectors, and the dimension of each sub-vector is 16, then the 128-dimensional high-dimensional vector is encoded as an 8-dimensional index combination.
[0004] However, traditional PQ coding directly specifies the number of sub-vectors and equally divides the high-dimensional vector to obtain multiple sub-vectors. The number of sub-vectors directly affects the search efficiency and search accuracy of vector search. The larger the number of sub-vectors, the search accuracy will increase to some extent, but the search efficiency will decrease. Therefore, how to determine the number of sub-vectors to ensure the search efficiency and search accuracy during vector search is an urgent problem to be solved. Summary of the Invention
[0005] In order to solve the technical problems of low search efficiency and search accuracy, the present application provides a similarity search method and system based on a vector database, which can ensure the search efficiency while improving the search accuracy.
[0006] In the first aspect of the present application, a similarity search method based on a vector database is provided. The search method includes: using the Euclidean distance between the vector data of the original information and the central data as the radius, and the angle with a preset direction as the angle to obtain the data points of the vector data in polar coordinates, where the original information is text or image, and the central data is the average value of all vector data; clustering the data points in polar coordinates to obtain multiple global clusters; taking the global cluster where the data point of the query vector is located as the target cluster; using the improved PQ coding to determine the sub-vectors of the target cluster and the centroid of each sub-vector, and then obtaining the search result of the query vector; the improved PQ coding includes: clustering the dimensions according to the correlation to obtain multiple sub-vectors; performing a clustering operation in the sub-vector space of each sub-vector to obtain the centroid of each sub-vector; constructing an objective function with the minimum centroid error and the maximum search efficiency to determine the optimal number of sub-vectors, where the centroid error is positively correlated with the Euclidean distance between the sub-vector and the corresponding centroid, and the search efficiency is negatively correlated with the number of data points in the target cluster and the number of sub-vectors.
[0007] Using the Euclidean distance between the vector data of the original information and the central data as the radius, and the angle with a preset direction as the angle to obtain the data points of each vector data in polar coordinates, and mapping all vector data to polar coordinates; dividing each vector data into multiple global clusters according to the position information of the data points of each vector data in polar coordinates, and the global cluster to which the query vector belongs can be directly located according to the position information of the data point of the query vector in polar coordinates, realizing the fast positioning of the target cluster and improving the search efficiency; further, calculating the similarity between each vector data and the query vector within the target cluster, avoiding traversing all vector data, and further improving the search efficiency. Further, in the process of calculating the similarity between each vector data and the query vector within the target cluster, an objective function is constructed with the minimum centroid error and the maximum search efficiency to determine the optimal number of sub-vectors, using the improved PQ coding to split the vector data into the optimal number of sub-vectors, encoding the query vector and each vector data in the target cluster into an index combination according to the sub-vector and the centroid of each sub-vector, and obtaining the search result of the query vector by calculating the Hamming distance between the index combinations, ensuring the search efficiency while improving the search accuracy.
[0008] Preferably, the obtaining of multiple global clusters includes: setting a first clustering number, and dividing all data points into multiple initial clusters according to the spatial distance of the data points in polar coordinates; taking the sum of the silhouette coefficient and the variance of the number of data points in each initial cluster as the evaluation value; adjusting the first clustering number multiple times to draw an evaluation value curve, and taking the first clustering number corresponding to the inflection point in the evaluation value curve as the first target number, and the clustering result of the first target number corresponds to multiple global clusters.
[0009] When performing vector search within the global cluster it is necessary to traverse the global cluster For all data points within, calculate the global cluster Calculate the similarity between all data points within and the query vector, and obtain the search results based on the similarity. Therefore, use the variance of the number of data points in the initial cluster as part of the evaluation value to measure the consistency of the number of data points in each initial cluster and ensure the search efficiency of each query vector.
[0010] Preferably, clustering the dimensions according to the correlation to obtain multiple sub-vectors includes: arranging the dimension values of any dimension of the vector data in each target cluster in a preset order to obtain the dimension value sequence of the dimension; defining the clustering distance, and dividing all dimensions into an initial number of dimension clusters, where the dimension clusters correspond to the sub-vectors; the clustering distance is negatively correlated with the correlation between the dimension value sequences and positively correlated with the difference between the dimension value distribution sequences.
[0011] Divide the dimensions with larger correlations into the same sub-vector to reduce information loss caused by dimension division. At the same time, divide the dimensions with smaller differences between the dimension value distribution sequences into the same sub-vector to effectively reduce the centroid error of the sub-vector and improve the search accuracy.
[0012] Preferably, the correlation is the absolute value of the Pearson correlation coefficient; within the dimension value sequence, count the number of occurrences of dimension values within a preset interval of each dimension value to obtain the dimension value distribution sequence.
[0013] Preferably, performing a clustering operation in the sub-vector space of each sub-vector to obtain the centroid of each sub-vector includes: determining the number of centroids of each sub-vector using the elbow method.
[0014] Preferably, the centroid error is: , is the number of sub-vectors, is the sub-vector the number of sub-vector data within, is the sub-vector the sub-vector data within , is the sub-vector the sub-vector data within the centroid of.
[0015] Preferably, the search efficiency is: , is the number of sub-vectors, is the number of data points in the target cluster.
[0016] Preferably, subtract the product of the centroid error and the adjustment weight from the search efficiency to obtain the objective function.
[0017] Give the specific calculation formula of the objective function. When the search efficiency is relatively large and the centroid error is relatively small, the value of the objective function is a relatively large value. By using an optimization algorithm to obtain the number of sub-vectors corresponding to the maximum value of the objective function, the optimal number can be accurately obtained.
[0018] Preferably, adjust the weight as: , is the number of data points in the target cluster, is the average number of data points in all global clusters, is function.
[0019] Since there are differences in the number of vector data in different global clusters, the adjustment weight is determined according to the number of data points in the target cluster to ensure that better search efficiency can be obtained in each global cluster.
[0020] In the second aspect of the present application, a similarity search system based on a vector database is further provided, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the similarity search method based on a vector database described in the first aspect of the present application is implemented.
[0021] The technical solution of the present application has the following beneficial technical effects: First, taking the Euclidean distance between the vector data of the original information and the central data as the radius and the angle with the preset direction as the angle, the data points of each vector data in polar coordinates are obtained, and all vector data are mapped into polar coordinates; according to the position information of the data points of each vector data in polar coordinates, each vector data is divided into multiple global clusters, and according to the position information of the data point of the query vector in polar coordinates, the global cluster to which the query vector belongs can be directly located, realizing the rapid positioning of the target cluster and improving the search efficiency; further, the similarity between each vector data and the query vector is calculated within the target cluster, avoiding traversing all vector data and further improving the search efficiency.
[0022] Further, in the process of calculating the similarity between each vector data and the query vector within the target cluster, an objective function is constructed with the minimum centroid error and the maximum search efficiency to determine the optimal number of sub-vectors. The vector data is split into the optimal number of sub-vectors by using improved PQ coding. According to the sub-vectors and the centroid of each sub-vector, the query vector and each vector data in the target cluster are encoded into index combinations, and the search result of the query vector is obtained by calculating the Hamming distance between the index combinations, ensuring the search efficiency while improving the search accuracy. Description of the Drawings
[0023] Figure 1It is a flowchart of a similarity search method based on a vector database according to an embodiment of the present application.
[0024] Figure 2 It is a structural block diagram of a similarity search system based on a vector database according to an embodiment of the present application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0026] According to the first aspect of the present application, the present application provides a similarity search method based on a vector database. Figure 1 It is a flowchart of a similarity search method based on a vector database according to an embodiment of the present application. As Figure 1 shown, the similarity search method based on the vector database includes steps S101 to S104, which are described in detail below.
[0027] S101, taking the Euclidean distance between the vector data of the original information and the central data as the radius and the included angle with the preset direction as the angle to obtain the data point of the vector data in polar coordinates, where the original information is text or image, and the central data is the average value of all vector data.
[0028] In one embodiment, the original information is text or image and is related to a specific application scenario. In the scenario of face recognition, the original information is a face image; in the scenario of text recommendation, the original information is a text sample. Among them, the vector data of the original information is obtained by using an autoencoder network or a PCA algorithm, and the present application does not make any restrictions.
[0029] After obtaining the vector data of each original information, calculate the average value of all vector data to obtain the central data; for the vector data of any original information, take the Euclidean distance between the vector data and the central data as the radius, and take the included angle between the vector data and the preset direction as the angle to obtain the data point of the vector data in polar coordinates; where the preset direction is the horizontal right direction or any other arbitrary direction, and the present application does not make any restrictions.
[0030] In this way, the data points of all vector data in polar coordinates are obtained in the same method.
[0031] S102, clustering the data points in polar coordinates to obtain a plurality of global clusters.
[0032] In one embodiment, obtaining multiple global clusters includes: setting a first number of clusters, and dividing all data points into multiple initial clusters according to the spatial distance of the data points in polar coordinates; taking the sum of the silhouette coefficient and the variance of the number of data points in each initial cluster as an evaluation value; adjusting the first number of clusters multiple times to draw an evaluation value curve, and taking the first number of clusters corresponding to the inflection point in the evaluation value curve as the first target number, and the clustering result of the first target number corresponds to multiple global clusters.
[0033] Wherein, the first number of clusters is 2, and the first number of clusters is increased by 1 each time of adjustment. The silhouette coefficient is a commonly used index for evaluating the clustering effect. The silhouette coefficient quantifies the clustering effect by comprehensively considering the intra-cluster compactness and inter-cluster separation, which is a common technical means for those skilled in the art and will not be elaborated here.
[0034] Wherein, the spatial distance of the data points in polar coordinates is the sum of the absolute value of the radius difference and the absolute value of the angle difference. Specifically, the data point and the data point The spatial distance in polar coordinates is: , wherein, and are the radii of the data points and the data point respectively, and are the angles of the data points and the data point respectively.
[0035] It can be understood that when performing vector search within the global cluster , it is necessary to traverse all data points within the global cluster , calculate the similarity between all data points within the global cluster and the query vector, and obtain the search result according to the similarity. Therefore, taking the variance of the number of data points in the initial cluster as part of the evaluation value can measure the consistency of the number of data points in each initial cluster, making the number of data points in each initial cluster basically the same. In this way, it can be ensured that when performing vector search within any global cluster, the number of vector data traversed is approximately the same, that is, the search efficiency of each query vector can be guaranteed.
[0036] S103, taking the global cluster where the data point of the query vector is located as the target cluster.
[0037] In one embodiment, in the scenario of face recognition, the query vector is the vector data of the face image to be queried; in the scenario of text recommendation, the query vector is the vector data of the text to be queried.
[0038] Quickly determine the global cluster where the query vector is located according to the position of the data point corresponding to the query vector in polar coordinates, use this global cluster as the target cluster, and search for the query vector in the target cluster, avoiding traversing all data points in polar coordinates and improving the search efficiency.
[0039] S104. Use the improved PQ coding to determine the sub-vectors of the target cluster and the centroid of each sub-vector, and then obtain the search result of the query vector.
[0040] In one embodiment, a data point in the target cluster corresponds to a vector data. Use the improved PQ coding to determine the sub-vectors of the target cluster and the centroid of each sub-vector, encode the query vector and each vector data in the target cluster into an index combination according to the sub-vectors and the centroid of each sub-vector, and obtain the search result of the query vector by calculating the Hamming distance between the index combinations.
[0041] The improved PQ coding includes: clustering the dimensions according to the correlation to obtain multiple sub-vectors, and performing a clustering operation in the sub-vector space of each sub-vector to obtain the centroid of each sub-vector; constructing an objective function to minimize the centroid error and maximize the search efficiency to determine the optimal number of sub-vectors; the centroid error is positively correlated with the Euclidean distance between the sub-vector and the corresponding centroid, and the search efficiency is negatively correlated with both the number of data points in the target cluster and the number of sub-vectors.
[0042] Among them, clustering the dimensions according to the correlation to obtain multiple sub-vectors includes: arranging the dimension values of any dimension of each vector data in the target cluster in a preset order to obtain the dimension value sequence of the dimension; defining a clustering distance, and dividing all dimensions into an initial number of dimension clusters, and the dimension clusters correspond to the sub-vectors; the clustering distance is negatively correlated with the correlation between the dimension value sequences and positively correlated with the difference between the dimension value distribution sequences.
[0043] Among them, the initial number is randomly set. In the embodiments of the present application, the value of the initial number is 2. After setting the initial number, the Kmeans algorithm or the Kmeans++ algorithm can be used to divide all dimensions into an initial number of dimension clusters. A dimension cluster contains at least one dimension. If a dimension cluster includes 3 dimensions, then this clustering cluster corresponds to a sub-vector composed of these 3 dimensions.
[0044] The correlation can be represented by the absolute value of the Pearson correlation coefficient.
[0045] In the dimension value sequence, count the number of occurrences of the dimension values within a preset interval of each dimension value, and the dimension value distribution sequence can be obtained. For example, for the dimension For example, the value range of the dimension value is equally divided into 5 preset dimension value intervals. There are 50 vector data in the target cluster, and the occurrence times in the 5 preset dimension value intervals are 25, 0, 0, 10, and 15 in sequence. Then the dimension value distribution sequence is If the difference between the dimension value distribution sequences of any two dimensions is small, it means that the dimension values between the two dimensions are roughly the same. Dividing the two dimensions into the same sub-vector can effectively reduce the centroid error of the sub-vector. In this way, dividing the dimensions with high correlation into the same sub-vector can reduce the information loss caused by dimension division. At the same time, dividing the dimensions with small differences in dimension value distribution sequences into the same sub-vector can effectively reduce the centroid error of the sub-vector and improve the search accuracy.
[0046] After obtaining the initial number of sub-vectors, one sub-vector corresponds to one vector space, and the sub-vector data contained in each vector space is the same. For example, there are 50 vector data in the target cluster, and dimension 1 and dimension 3 form a sub-vector. Then this sub-vector space includes 50 sub-vector data, and each sub-vector data contains dimension 1 and dimension 3. Further, in the vector space, use the Kmeans algorithm or the Kmeans++ algorithm to divide the 50 sub-vector data into K sub-vector clusters, and calculate the clustering center of each sub-vector cluster as the K centroids of the sub-vector formed by dimension 1 and dimension 3. Among them, the number of centroids K of the sub-vector is determined by the elbow method, and the number of centroids corresponding to different sub-vectors is different. In this way, the vector data is split into the initial number of sub-vectors, and the centroids of each sub-vector are obtained.
[0047] In one embodiment, adjust the number of sub-vectors, that is, continuously adjust the value of the initial number. Construct an objective function with the minimum centroid error and the maximum search efficiency to determine the optimal number of sub-vectors. The optimal number can balance the centroid error and the search efficiency, and while improving the retrieval efficiency, ensure the retrieval accuracy.
[0048] Among them, the centroid error is positively correlated with the Euclidean distance between the sub-vector and the corresponding centroid. Because the centroid will be used to replace the sub-vector data in the sub-vector cluster to which the centroid belongs for constructing the index combination, the centroid error directly affects the search accuracy. Specifically, the centroid error is: , is the number of sub-vectors, is the number of sub-vector data in the sub-vector is the sub-vector in the sub-vector data , is the sub-vector in the sub-vector data centroid.
[0049] Among them, the search efficiency is negatively correlated with both the number of data points in the target cluster and the number of sub-vectors. The larger the number of data points in the target cluster, the more vector data needs to be traversed during vector search, and the lower the search efficiency. The larger the number of sub-vectors, the longer the length of the index combination. When obtaining the search result based on the Hamming distance between index combinations, the required computational amount is also larger, and the search efficiency is lower. Specifically, the search efficiency is: , where is the number of sub-vectors, and is the number of data points in the target cluster.
[0050] Furthermore, a target function is constructed by minimizing the centroid error and maximizing the search efficiency, that is, the target function is negatively correlated with the centroid error and positively correlated with the search efficiency. Therefore, by subtracting the product of the centroid error and the adjustment weight from the search efficiency, the target function is obtained. The optimal number is obtained by using an optimization algorithm to find the number of sub-vectors corresponding to the maximum value of the target function. Specifically, the target function satisfies: , where is the search efficiency, is the centroid error, and is the adjustment weight. Among them, the adjustment weight is used to balance the importance of the search efficiency and the centroid error. The larger the adjustment weight, the higher the importance attached to the search efficiency. In the embodiments of the present application, the adjustment weight takes the value of 1.
[0051] In another embodiment, since the number of vector data in different global clusters may vary, in order to ensure good search efficiency in each global cluster, the adjustment weight is determined according to the number of data points in the target cluster. The adjustment weight is positively correlated with the number of data points in the target cluster. Specifically, in the process of using the improved PQ coding to determine the sub-vectors of the target cluster and the centroid of each sub-vector, the adjustment weight is: , where is the number of data points in the target cluster, is the average number of data points in all global clusters, and is a function.
[0052] In one embodiment, after obtaining the optimal number of the target cluster, the query vector and all vector data in the target cluster can be split into multiple sub-vectors, and the number of sub-vectors is equal to the optimal number; at the same time, the centroid of each sub-vector can be obtained.
[0053] Exemplarily, if the optimal number is recorded as 5, the query vector can be split into 5 sub-vectors. Sub-vector 1 is in the sub-vector cluster where the centroid 2 is located in the corresponding sub-vector space, and sub-vector 2 is in the sub-vector cluster where the centroid 1 is located in the corresponding sub-vector space. Similarly, it can be obtained that sub-vector 3, sub-vector 4, and sub-vector 5 are respectively in the sub-vector clusters where the centroid 2, centroid 5, and centroid 3 are located in their respective sub-vector spaces. Then, the query vector can be encoded as the index combination; obtain the index combinations of all vector data within the target cluster in the same way, calculate the similarity between the query vector and each vector data within the target cluster according to the Hamming distance between the index combinations, and use one or more vector data with the maximum similarity as the search result of the query vector.
[0054] According to the second aspect of the present application, the present application also provides a similarity search system based on a vector database. Figure 2 FIG. is a structural block diagram of a similarity search system based on a vector database according to an embodiment of the present application. As Figure 2 shown, the system 50 includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a similarity search method based on a vector database according to the first aspect of the present application is implemented. The system also includes other components well-known to those skilled in the art, such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be elaborated here.
[0055] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A similarity search method based on a vector database, characterized in that: The search method comprises: taking the Euclidean distance between the vector data of the original information and the center data as the radius and the angle with the preset direction as the angle, obtaining the data point of the vector data in polar coordinates, wherein the original information is text or image, and the center data is the average value of all vector data; Cluster the data points in polar coordinates to obtain multiple global clusters; The global cluster where the data points of the query vector are located is taken as the target cluster; The improved PQ coding is used to determine the sub-vectors of the target cluster and the centroid of each sub-vector, and then the search results of the query vector are obtained; The improved PQ coding includes: clustering dimensions according to correlation to obtain multiple sub-vectors; performing clustering operations in the sub-vector space of each sub-vector to obtain the centroid of each sub-vector; constructing an objective function to determine the optimal number of sub-vectors by minimizing the centroid error and maximizing the search efficiency, wherein the centroid error is positively correlated with the Euclidean distance between the sub-vector and the corresponding centroid, and the search efficiency is negatively correlated with both the number of data points in the target cluster and the number of sub-vectors.
2. A similarity search method based on a vector database according to claim 1, characterized in that: The obtaining of multiple global clusters comprises: Set the first clustering number and divide all data points into multiple initial clusters according to the spatial distance of the data points in polar coordinates; The sum of the silhouette coefficient and the variance of the number of data points in each initial cluster is taken as the evaluation value; The first clustering number is adjusted multiple times to draw an evaluation value curve, and the first clustering number corresponding to the inflection point in the evaluation value curve is used as the first target number, and the clustering result of the first target number corresponds to multiple global clusters.
3. A similarity search method based on a vector database according to claim 1, characterized in that: The dimensions are clustered according to the correlation to obtain multiple sub-vectors, including: Arrange the dimension values of any dimension of each vector data in the target cluster in a preset order to obtain a dimension value sequence of the dimension; A clustering distance is defined, and all dimensions are divided into an initial number of dimension clusters, where the dimension clusters correspond to the subvectors; the clustering distance is negatively correlated with the correlation between dimension value sequences, and is positively correlated with the difference between dimension value distribution sequences.
4. A similarity search method based on a vector database according to claim 3, characterized in that: The correlation is the absolute value of the Pearson correlation coefficient; In the dimension value sequence, the number of occurrences of the dimension value in the preset interval of each dimension value is counted to obtain the dimension value distribution sequence.
5. A similarity search method based on a vector database according to claim 1, characterized in that: The performing a clustering operation in the sub-vector space of each sub-vector to obtain the centroid of each sub-vector includes: determining the number of centroids of each sub-vector using an elbow method.
6. A similarity search method based on a vector database according to claim 1, characterized in that: The centroid error for: , is the number of sub-vectors, For subvector The number of inner sub-vector data, For subvector Sub-vector data within , For subvector Inner vector data The center of mass.
7. A similarity search method based on a vector database according to claim 1, characterized in that: The search efficiency for: , is the number of sub-vectors, is the number of data points in the target cluster.
8. A similarity search method based on a vector database according to claim 1, characterized in that: The objective function is obtained by subtracting the product of the centroid error and the adjustment weight from the search efficiency.
9. A similarity search method based on a vector database according to claim 8, characterized in that: Adjust weight for: , is the number of data points in the target cluster, is the average number of data points in all global clusters, for function.
10. A similarity search system based on a vector database, characterized in that: The invention comprises a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a similarity search method based on a vector database according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Vector retrieval method based on residual quantization
CN118132679A
Large-scale high-dimensional vector nearest neighbor data retrieval method and device
CN119089005A
Data processing method and related equipment
CN119441212A
Cited By
Method and Apparatus for Accelerating Vector Similarity Computation Using Identifier-Based Autonomous Loop Control
KR103000047B1