Vector retrieval acceleration method and system
By dividing the query vector and the center vector into blocks for matrix multiplication, the central vector database is loaded only once during the rough search process, which solves the performance bottleneck problem caused by repeated loading of the database during the vector search process in the prior art, and achieves more efficient vector retrieval.
Patent Information
- Application Number
- CN202311789155.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In the vector search process, the existing technology requires repeated loading of the entire central vector database, resulting in the query time increasing with the database scale and becoming a performance bottleneck.
By dividing the query vector and the center vector into blocks, matrix multiplication and distance calculation are performed separately, the center vector database is loaded only once during the rough search process, reducing the number of repeated loads.
It greatly reduces the time required for rough search, improves the processing speed of the hardware computing system, and improves the efficiency of vector retrieval.
Smart Images

Figure CN120196770A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and system for accelerating vector retrieval. Background Art
[0002] In business scenarios such as image search by image, video recommendation, text retrieval, etc., it is often necessary to obtain corresponding target samples from a large number of data samples stored in a database. Currently, vector retrieval algorithms are usually used to meet the requirements of these business scenarios through the similarity between vector data.
[0003] The retrieval algorithm based on Inverted file system product quantization (IVFPQ) includes two parts: coarse search and fine search. In the coarse search, several clusters closest to the query vector are selected from multiple clusters formed by clustering in the vector database. In the fine search, the topk vectors closest to the query vector are retrieved from all the vectors in the several clusters selected by the coarse search. The coarse search essentially calculates the distance between each query vector and the vector corresponding to each cluster center (also called the "center vector"), and selects the nprobe center vectors closest to the query vector. Figure 1 is a calculation process of an IVFPQ coarse search design, as Figure 1 shown, each query vector is sequentially serially input into a RISC-V VPU (a video processing unit with an open-source instruction set architecture). For each input query vector, the RISC-V VPU will load the entire center vector database (XC database) from the DDR to calculate the distance between the query vector and all the center vectors in the database, and perform distance sorting through a sorting acceleration module.
[0004] It can be seen that in the case of multiple query vectors, the distance calculation needs to repeatedly load the XC database. As the scale of the XC database increases, the time required to load the XC database from the DDR to the local will become longer and the time required for the coarse search will also increase rapidly. Therefore, repeatedly loading the entire XC database from the DDR for each query has become a performance bottleneck, and at the same time, it has also reduced the processing speed of the entire hardware computing system.
[0005] This section aims to provide background or context for the embodiments of the present application stated in the claims. The description herein is not admitted to be prior art that has been publicly disclosed just because it is included in this section. Summary of the Invention
[0006] The purpose of this application is to provide a method and system for accelerating vector retrieval, which only needs to load the center vector database once during the coarse search process, thereby reducing the time required for the coarse search and improving the processing speed of the hardware computing system.
[0007] This application discloses a vector retrieval acceleration system, including:
[0008] At least one computing module, the computing module includes a first receiving end and a second receiving end respectively used for receiving a query vector block and a central vector block, and performs operations on the received query vector block and central vector block to obtain a distance calculation result between each query vector and all central vectors, where nq query vectors are divided into nq / n query vector blocks The nlist central vectors in the central vector database are divided into nlist / m central vector blocks Each query vector block includes n query vectors, each central vector block includes m central vectors, where nq, nlist, m, and n are all integers greater than or equal to 1; and
[0009] At least one sorting acceleration module, the at least one computing module and the at least one sorting acceleration module correspond one by one, each sorting acceleration module respectively receives the distance calculation result output by the corresponding computing module, sorts the distance calculation result, and saves the obtained sorting intermediate state to the system memory.
[0010] In a preferred example, the operation of the computing module on the central vector block and the query vector block includes: performing a matrix multiplication operation on the query vector block and the central vector block, and adding the square of each element in the matrix multiplication result to the corresponding central vector respectively to obtain the distance calculation result.
[0011] In a preferred example, the second receiving end of each computing module receives its corresponding central vector block from the central vector database, corresponding to the corresponding central vector block, the first receiving end of each computing module sequentially receives each query vector block, so that each computing module can perform operations on its corresponding central vector block and each query vector block;
[0012] Among them, different computing modules respectively correspond to different central vector blocks
[0013] In a preferred example, the first receiving ends of each computing module are connected together and receive the same query vector block at the same time.
[0014] In a preferred example, the system further includes a final sorting module, the final sorting module obtains the sorting intermediate states of all sorting acceleration modules, and sorts according to the sorting intermediate states to obtain a sorting result corresponding to each query vector.
[0015] In a preferred example, the first receiving end of each computing module receives its corresponding query vector block. Corresponding to the corresponding query vector block, the second receiving end of each computing module sequentially receives each center vector block, so that each computing module can perform operations on its corresponding query vector block and each center vector block;
[0016] Among them, different computing modules respectively correspond to different query vector blocks.
[0017] In a preferred example, the computing modules are connected through a bus;
[0018] Each computing module respectively receives different center vector blocks from the center vector database through its second receiving end, and sequentially transfers the center vector blocks received by each through the bus among the computing modules, so that each computing module can sequentially receive different center vector blocks.
[0019] In a preferred example, the sorting acceleration module loads the sorting intermediate state from the system memory, and sorts the sorting intermediate state together with the distance calculation result from the computing module or the local memory to obtain a new sorting intermediate state, and updates the sorting intermediate state stored in the system memory with the new sorting intermediate state
[0020] In a preferred example, the sorting intermediate state obtained by each sorting acceleration module in the last sorting is the sorting result of the query vectors in the query vector block received by the corresponding computing module.
[0021] In a preferred example, the system further includes a local memory, and the distance calculation result calculated by the computing module is stored in the local memory and output to the corresponding sorting acceleration module through the local memory.
[0022] This application also discloses a vector retrieval acceleration method, including:
[0023] Dividing nq query vectors into nq / n query vector blocks Dividing nlist center vectors in the center vector database into nlist / m center vector blocks Each query vector block includes n query vectors, and each center vector block includes m center vectors, where nq, nlist, m, and n are all integers greater than or equal to 1;
[0024] Receiving center vector blocks C k 0, center vector block C k 1,..., center vector block And the center vector block Ck 0. Central vector block C k 1, …, central vector blocks Each central vector block in q 0. Query vector block x q 1, …, query vector blocks performs an operation with the query vector block x to obtain the distance calculation results between each query vector and all central vectors; and
[0025] The distance calculation results of the central vector block and the query vector block are obtained from the corresponding calculation module by at least one sorting acceleration module, and sorted according to the distance calculation results to obtain the sorting results corresponding to each query vector.
[0026] In a preferred example, each central vector block is respectively operated with the query vector block x q 0. Query vector block x q 1, …, query vector blocks The operation includes:
[0027] Each central vector block is respectively operated with the query vector block x q 0. Query vector block x q 1, …, query vector blocks Performs a matrix multiplication operation, and adds the square of the corresponding central vector to each element in the obtained matrix multiplication result to obtain the distance calculation result.
[0028] In a preferred example, each central vector block is respectively operated with the query vector block x q 0. Query vector block x q 1, …, query vector blocks The matrix multiplication operation includes:
[0029] Input the central vector block C k 0. Central vector block C k 1, …, central vector blocks into the corresponding calculation module in the at least one calculation module respectively; and
[0030] For each calculation module, input the query vector block x q 0. Query vector block x q 1, …, query vector blocks in sequence so that each calculation module calculates the matrix product of the corresponding central vector block and all query vector blocks.
[0031] In a preferred example, each central vector block is respectively operated with the query vector block x q 0. Query vector block x q 1, …, query vector blocks Performing matrix multiplication operations includes:
[0032] Inputting the query vector blocks x q 0, the query vector block x q 1, …, the query vector blocks into the corresponding computing modules among the at least one computing module respectively; and
[0033] For each computing module, inputting the central vector blocks C k 0, the central vector block C k 1, …, the central vector blocks in sequence, so that each computing module calculates the matrix product of the corresponding query vector block and all central vector blocks.
[0034] In a preferred example, obtaining the distance calculation results of the central vector blocks and the query vector blocks from the corresponding computing module through at least one sorting acceleration module, and sorting according to the distance calculation results includes:
[0035] The sorting acceleration module sorts based on the distance calculation results received from the corresponding computing module or the local memory and the sorting intermediate state obtained from the system memory to generate a new sorting intermediate state, and updates the sorting intermediate state stored in the system memory with the new sorting intermediate state.
[0036] In a preferred example, sorting according to the distance calculation results to obtain the sorting results corresponding to each query vector includes:
[0037] Taking the sorting intermediate state obtained by each sorting acceleration module in the last sorting as the sorting result of the query vector in the query vector block received by the corresponding computing module.
[0038] In a preferred example, sorting according to the distance calculation results to obtain the sorting results corresponding to each query vector includes: obtaining the sorting intermediate states of all sorting acceleration modules through the final sorting module, and sorting according to the sorting intermediate states to obtain the sorting results corresponding to each query vector.
[0039] This application also discloses a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, the steps in the method described above are implemented.
[0040] In the embodiments of this application, during the coarse search process, the central vector database only needs to be loaded once, so that the time required for the coarse search can be greatly reduced, and the processing speed of the hardware computing system is improved.
[0041] The description of this application records a large number of technical features, which are distributed in various technical solutions. If all possible combinations of technical features (i.e., technical solutions) of this application are listed, the description will be too long. To avoid this problem, each technical feature disclosed in the above-mentioned invention content of this application, each technical feature disclosed in the following embodiments and examples, and each technical feature disclosed in the drawings can be freely combined with each other to form various new technical solutions (all these technical solutions should be regarded as having been recorded in this specification), unless the combination of such technical features is technically infeasible. For example, in one example, features A + B + C are disclosed, and in another example, features A + B + D + E are disclosed. Features C and D are equivalent technical means that play the same role, and only one of them can be used technically and it is impossible to use both at the same time. Feature E can be combined with feature C technically. Then, the solution of A + B + C + D should not be regarded as having been recorded because it is technically infeasible, while the solution of A + B + C + E should be regarded as having been recorded. Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the calculation process of the traditional IVFPQ rough search design.
[0043] Figure 2 It is a schematic diagram of a vector retrieval acceleration system according to an embodiment of the present application.
[0044] Figure 3 It is a schematic diagram of the matrix multiplication and sorting process between each of a plurality of query vector blocks and each of a central vector block according to an embodiment of the present application.
[0045] Figure 4 It is a schematic diagram of stage 1 of the matrix multiplication and sorting process between each of a plurality of query vector blocks and each of a central vector block when both the query vector and the central vector are relatively large according to an embodiment of the present application.
[0046] Figure 5 It is a schematic flowchart of a vector retrieval acceleration method according to an embodiment of the present application. Detailed Embodiments
[0047] In the following description, many technical details are provided to help the reader better understand the present application. However, those of ordinary skill in the art can understand that even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0048] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0049] In the existing vector retrieval scheme, during the coarse search stage, every time a query vector is input, the entire central vector database needs to be loaded once. This repeated database loading approach causes the retrieval time to increase rapidly as the database scale grows, thereby affecting the processing speed of the hardware computing system.
[0050] Based on this problem, the present application proposes a vector retrieval acceleration method and system, which only needs to load the XC database once during the coarse search process, thus greatly reducing the time required for coarse search and improving the processing speed of the hardware computing system.
[0051] During the coarse search stage, calculate the query vector for each query vector x q in the distance from all central vectors q and select nprobe central vectors that are closest to it for each query vector x q , where nq is the number of query vectors, nlist is the number of central vectors, and both nq and nlist are integers greater than or equal to 1. The distance calculation formula between the query vector x k and the central vector C
[0052]
[0053] There are various types, such as the inner product distance calculation formula, the Euclidean distance calculation formula, etc. Here, the Euclidean distance calculation formula is taken as an example for illustration, as shown in Equation (1):
[0054] It can be seen from Equation (1) that for a specific query vector x q , the difference in the distance magnitudes between it and each central vector depends on the second term and the third term And the third term is the square of the central vector C k , and its magnitude depends on the central vector itself. Therefore, the process of calculating the distance between the query vector x q and the central vector C k mainly involves calculating the product of the query vector x q and the central vector C k and then adding the square of the corresponding central vector to the obtained product.
[0055] In one embodiment, the process of calculating the product of each query vector x q and all central vectors is transformed into calculating the product of all central vectors The formed nlist×d first matrix and the second matrix of nq×d formed by all query vectors The product of ,
[0056]
[0057] where each row vector of the first matrix is a center vector, each column vector of the second matrix is a query vector, and d is the vector dimension. After calculating the product of the above two matrices, the product of any query vector and any center vector can be obtained, and the matrix multiplication algorithm can be implemented through various matrix multiplication acceleration libraries.
[0058] After obtaining the product of the two matrices, the square of the corresponding center vector is added to each element of the matrix multiplication result of the query vector and the center vector, and then the distance calculation result is output to the sorting acceleration module for sorting. The sorting acceleration module can process the distance sorting of k query vectors in parallel, so all query vectors can be sorted by ceil(nq / k) (ceil() represents rounding up) times of sorting.
[0059] Considering that nlist and nq are relatively large, a large local memory is required to complete the product operation of the two matrices at one time. Therefore, in some embodiments, the two matrices can be divided first. For example, the matrix of query vectors is divided into multiple query vector blocks (xq block), and each query vector block can include n query vectors, then the nq query vectors are divided into nq / n query vector blocks Similarly, the matrix of center vectors can also be divided into multiple center vector blocks, and each center vector block can include m center vectors, that is, the nlist center vectors are divided into nlist / m center vector blocks where both m and n are integers greater than 1.
[0060] This application provides a vector retrieval acceleration system. The vector retrieval acceleration system may include: at least one computing module and at least one sorting acceleration module, and the computing module and the sorting acceleration module are connected in one-to-one correspondence. For the sake of simplicity, Figure 2 a vector retrieval acceleration system including only one computing module and one sorting acceleration module is shown. As Figure 2 shown, one end of the sorting acceleration module is coupled to the computing module, and is used to receive the distance calculation results output by the computing module and sort these results.
[0061] In some embodiments, each computing module has a first receiving end and a second receiving end, which are respectively used to receive a query vector block and a central vector block and perform operations on the received query vector block and central vector block. The sorting acceleration module respectively receives the distance calculation results output by the corresponding computing module, sorts these results, and then saves the obtained sorted intermediate state to the system memory. The sorted intermediate states of each sorting acceleration module are stored in different storage locations of the system memory. The system memory can be, for example, a DDR memory.
[0062] After the two matrices are divided, the central vector block and the query vector block are respectively input into at least one computing module. The computing module calculates the input central vector block and query vector block. In one embodiment, m central vectors in the central vector block are input into the computing module in parallel, and n query vectors in the query vector block are input into the computing module in parallel.
[0063] In one embodiment, calculating the central vector block and the query vector block includes: performing a matrix multiplication operation on the query vector block and the central vector block, and adding each element in the matrix multiplication result to the square of the corresponding central vector to obtain a distance calculation result.
[0064] In some embodiments, the vector retrieval acceleration system may include a local memory, such as Figure 2 As shown, the distance calculation results obtained by the computing module can be stored in the local memory and output to the sorting acceleration module via the local memory. Figure 2 The local memory shown is a separate module, but the present invention is not limited thereto, and the local memory can be integrated into the computing module.
[0065] The computing module can respectively load the central vector block from the system memory For the central vector block input to the computing module will be respectively calculated with the query vector block x q 0, query vector block x q 1,..., query vector block for calculation.
[0066] For example, calculating the central vector block with the query vector block x q 0 includes: performing a matrix multiplication operation on the central vector block C k j and the query vector block x q 0 (formula (3)) to obtain the product of each query vector in the query vector block x q 0 and each central vector in the central vector block Cj respectively, and then adding each obtained product (each element in the matrix multiplication result) to the square of the corresponding central vector, then the query vector block x can be obtainedq The n query vectors in 0 and the central vector block C k The distance calculation results between the m central vectors in j. These distance calculation results can be cached in the local memory.
[0067]
[0068] After obtaining the distance calculation results of a query vector block and a central vector block, the sorting acceleration module can be used for sorting. The sorting acceleration module can process the sorting of k ( Figure 2 illustrated with k = 16 as an example) query vectors in parallel. Therefore, when the number n of query vectors in the query vector block is greater than k, the distance calculation results of a query vector block and a central vector block need to be sorted ceil(n / k) times. And the sorting acceleration module can store the intermediate sorting state of each time in the system memory. So that in the next sorting, the sorting acceleration module can load the intermediate sorting state from the system memory and sort the intermediate sorting state together with the distance calculation results from the calculation module or the local memory to obtain a new intermediate sorting state, and at the same time update the corresponding storage location in the system memory with the new intermediate sorting state. The intermediate sorting state output after the sorting acceleration module sorts the distances of the query vectors from all the central vectors is used as the sorting result.
[0069] In one embodiment, only after the sorting acceleration module finishes sorting the distance calculation results corresponding to the n query vectors in a vector query block, will the calculation module calculate the next query vector block and the corresponding central vector block, and repeat the above process until all query vector blocks are calculated. In some embodiments, the vector retrieval acceleration system further includes a final sorting module, and the final sorting module can be one of the sorting acceleration modules among multiple sorting acceleration modules itself, or an additionally provided sorting module. The final sorting module can obtain the intermediate sorting states of all sorting acceleration modules and perform sorting according to the intermediate sorting states to obtain the sorting results corresponding to each query vector.
[0070] However, the embodiments of the present application are not limited thereto. In practical applications, the calculation of the calculation module and the sorting of the sorting acceleration module are processed in a pipeline manner, that is, while the sorting acceleration module sorts the received distance calculation results, the calculation module can perform the next round of calculations. This way can hide the processing time of the sorting acceleration module and is beneficial to improving efficiency.
[0071] In one embodiment, all query vectors can be divided into query vector blocks, and all central vectors in the central vector database can be sliced into a central vector block. It should be understood that can be an integer greater than 1, where for example, it can be 32, 16, 8, 4, etc., for example, it can be 32, 16, 8, 4, etc.
[0072] For the convenience of description, the following will take as an example for illustration, but those skilled in the art can understand that the embodiments of the present application do not impose any restrictions on the size of, and whether they are equal.
[0073] Figure 3 Exemplarily shown is a vector retrieval acceleration system including multiple computing modules and a sorting acceleration module according to another embodiment of the present application. In the vector retrieval acceleration system, the multiple computing modules can be respectively computing modules CU0 to CU15, and each computing module corresponds to one of the sorting acceleration modules SORT0 to SORT15. The first receiving end of each computing module receives a query vector block, and the second receiving end receives a central vector block. Each computing module calculates the received query vector block and the central vector block to obtain the distance calculation results between each query vector in the query vector block and all the central vectors in the central vector block. The computing modules can be connected through a bus for vector block transmission, and the calculation results of the computing modules can be stored in a local memory (not shown in the figure), and the local memory can be an independent component or integrated in the computing module. The sorting acceleration module can support parallel sorting of k query vectors, that is, it can simultaneously process the sorting of the distance calculation results corresponding to k query vectors and store the sorting intermediate state into the system memory.
[0074] Initially, the first receiving ends of the computing modules CU0 to CU15 respectively receive the query vector blocks x q 0, x q 1, …, x q 15, and the second receiving ends of the computing modules CU0 to CU15 respectively receive the central vector blocks C k 0, C k 1, …, C k 15 from the central vector database (XC database). In the first round of calculation, the computing modules CU0 to CU15 respectively calculate x q 0 and C k 0, x q 1 and C k 1, …, x q 15 and C k 15.
[0075] In the subsequent calculation, the central vector blocks C k 0, Ck 1, …, C k 15 flows sequentially between the computing modules CU0 to CU15 through the bus between the computing modules, so that for each query vector block, it is respectively matrix-multiplied with the central vector block C k 0, the central vector block C k 1, …, the central vector block C k 15.
[0076] Specifically, in the second round of calculation, the computing module CU0 transmits its central vector block C k 0 to the computing module CU1, the computing module CU1 transmits its central vector block C k 1 to the computing module CU2, …, the computing module CU15 transmits its central vector block C k 15 to the computing module CU0. The computing modules CU0 to CU15 respectively calculate x q 0 and C k 15, x q 1 and C k 0, …, x q 15 and C k 14. In the third round of calculation, the computing module CU0 transmits its central vector block C k 15 to the computing module CU1, the computing module CU1 transmits its central vector block C k 0 to the computing module CU2, …, the computing module CU15 transmits its central vector block C k 14 to the computing module CU0, and the computing modules CU0 to CU15 respectively calculate x q 0 and C k 14, x q 1 and C k 15, …, x q 15 and C k 13. And so on. In the last round of calculation, the computing module CU0 transmits its central vector block C k 2 to the computing module CU1, the computing module CU1 transmits its central vector block C k 3 to the computing module CU2, …, the computing module CU15 transmits its central vector block C k 1 to the computing module CU0. The computing modules CU0 to CU15 respectively calculate x q 0 and C k 1, x q 1 and C k 2, …, x q 15 and C k 0. Thus, each query vector block has been calculated with all the central vector blocks or each central vector block has been calculated with all the query vector blocks.
[0077] The sorting acceleration module sorts the distance calculation results, including sorting based on the distance calculation results received from the corresponding calculation module or the local memory and the sorting intermediate state obtained from the system memory, generating a new sorting intermediate state, and using the new sorting result to update the stored sorting intermediate state in the system memory. It should be noted that initially, the sorting intermediate state is empty.
[0078] The sorting intermediate state obtained by the sorting acceleration module sorting the distance calculation results output by the last round of calculation of the corresponding calculation module and the sorting intermediate state obtained from the system memory (the sorting intermediate state output after sorting the distances between the corresponding query vector and all center vectors) is the sorting result.
[0079] In this embodiment, the query vector blocks received by the first receiving end of each calculation module remain unchanged, and the second receiving end receives different center vector blocks.
[0080] Figure 4 Exemplarily shown is a vector retrieval acceleration system including multiple calculation modules and a sorting acceleration module according to another embodiment of the present application. Figure 4 Differing from Figure 3 is that the vector retrieval acceleration system may further include a final sorting module (not shown in the figure), and the first receiving ends of the calculation modules can be connected together to receive the same query vector block.
[0081] Initially, the first receiving ends of calculation modules CU0 to CU15 all receive the query vector block x q 0, and the second receiving ends of calculation modules CU0 to CU15 respectively receive the center vector blocks C k 0, C k 1, …, C k 15 from the center vector database (XC database). In the first round of calculation, calculation modules CU0 to CU15 respectively calculate x q 0 and C k 0, x q 0 and C k 1, …, x q 0 and C k 15 ( Figure 4 is a schematic diagram of the first round). In the second round of calculation, the first receiving ends of calculation modules CU0 to CU15 all receive the query vector block x q 1, and calculation modules CU0 to CU15 respectively calculate x q 1 and C k 0, x q 1 and C k 1, …, x q 1 and C kPerform calculations for 15, and so on. In the last round of calculation, the first receiving ends of calculation modules CU0 to CU15 all receive the query vector block x q 15, and calculation modules CU0 to CU15 respectively perform calculations on x q 15 and C k 0, x q 15 and C k 1, …, x q 15 and C k 15. Thus, each query vector block has been calculated with all center vector blocks or each center vector block has been calculated with all query vector blocks.
[0082] Finally, the sorting module obtains the sorting intermediate states of all sorting acceleration modules and performs sorting according to the sorting intermediate states to obtain the sorting results corresponding to each query vector.
[0083] In this embodiment, the first receiving end of each calculation module receives different query vector blocks, and the center vector blocks received by the second receiving end remain unchanged.
[0084] In another embodiment, if the center vector is sliced into 32 or more center vector blocks, then the calculation module needs 2 rounds or more to receive the center vector blocks. For example, in combination with Figure 4 , taking the center vector sliced into 32 center vector blocks C k 0 to C k 31 as an example, in the first round, 16 calculation modules CU0 to CU15 respectively receive the center vector blocks C k 0 to C k 15 from the XC database, and sequentially input t (t is greater than or equal to 1) query vector blocks x q 0 to x q (t - 1) into calculation modules CU0 to CU15 for calculation. Then, in the second round, calculation modules CU0 to CU15 respectively receive the center vector blocks C k 16 to C k 31 from the XC database, and sequentially input t query vector blocks x q 0 to x q (t - 1) into calculation modules CU0 to CU15 for calculation. After 2 rounds, the calculations between each query vector block and all center vector blocks C k 0 to C k 31 can be completed and the corresponding sorting results can be obtained.
[0085] In other embodiments, if the center vector If it is divided into 2 or 4 or 8 central vector blocks, then 16 computing modules can support 8, 4, or 2 coarse search tasks in parallel. Or, if the central vector is not divided, 1 computing module can complete one coarse search task, and 16 computing modules can support 16 coarse search tasks in parallel.
[0086] Next, the process of the vector retrieval acceleration method according to the embodiments of the present application will be described. As Figure 5 shown, this method can be applied to the above vector retrieval acceleration system, including:
[0087] Step 101: Divide nq query vectors into nq / n query vector blocks Divide nlist central vectors in the central vector database into nlist / m central vector blocks C k 0, Each query vector block includes n query vectors, and each central vector block includes m central vectors, where nq, nlist, m, and n are all integers greater than or equal to 1.
[0088] Step 102: Receive the central vector blocks C k 0, central vector block C k 1, …, central vector block from the central vector library through at least one computing module k 0, central vector block C k 1, …, central vector block and calculate each central vector block in C q 0, query vector block x q 1, …, query vector block respectively with the query vector blocks x
[0089] to obtain the distance calculation results between each query vector and all central vectors.
[0090] Among them, in step 102, calculating each central vector block respectively with the query vector blocks x q 0, query vector block x q 1, …, query vector block includes: calculating each central vector block respectively with the query vector blocks x q 0, query vector block x q 1, …, query vector block Perform matrix multiplication and add the square of the corresponding central vector to each element in the obtained matrix multiplication result to obtain the distance calculation result.
[0091] In some embodiments, each central vector block is respectively matrix-multiplied with query vector blocks x q 0, query vector blocks x q 1, …, query vector blocks The matrix multiplication operation can be implemented through the following steps in some embodiments:
[0092] Step 1021, input central vector blocks C k 0, central vector blocks C k 1, …, central vector blocks into the corresponding calculation modules in at least one calculation module respectively.
[0093] Step 1022, for each calculation module, sequentially input query vector blocks x q 0, query vector blocks x q 1, …, query vector blocks so that each calculation module calculates the matrix product of the corresponding central vector block and all query vector blocks.
[0094] In other embodiments, each central vector block is respectively matrix-multiplied with query vector blocks x q 0, query vector blocks x q 1, …, query vector blocks The matrix multiplication operation can be implemented through the following steps:
[0095] Step 1021′, input query vector blocks x q 0, query vector blocks x q 1, …, query vector blocks into the corresponding calculation modules in at least one calculation module respectively.
[0096] Step 1022′, for each calculation module, sequentially input central vector blocks C k 0, central vector blocks C k 1, …, central vector blocks so that each calculation module calculates the matrix product of the corresponding query vector block and all central vector blocks.
[0097] It can be understood that in step 1022, after each calculation module finishes calculating the matrix multiplication of a query vector block and a central vector block, the square of the corresponding central vector can be added to the product of any query vector in the query vector block and any central vector in the central vector block (any element in the matrix multiplication result) to obtain the distance calculation result.
[0098] After the computing module finishes computing a query vector block and a center vector block, the sorting acceleration module can perform sorting according to the distance calculation result, including:
[0099] Step 1031: Based on the distance calculation result received from the corresponding computing module or the local memory and the sorting intermediate state obtained from the system memory, perform sorting to generate a new sorting intermediate state, and update the corresponding storage location in the system memory with the new sorting intermediate state. Initially, the sorting intermediate state is empty.
[0100] It can be understood that each query vector block includes n query vectors. Therefore, after the computing module finishes computing a query vector block and a center vector block, distance calculation results of n query vectors will be obtained. The sorting acceleration module can process the distance calculation results of k query vectors in parallel each time. If n is less than k, or the number of query vectors corresponding to the remaining distance calculation results in the current local memory is less than k, the sorting acceleration module can obtain all the currently stored distance calculation results of the query vectors from the local memory; if n is greater than or equal to k, or the number of query vectors corresponding to the remaining distance calculation results in the current local memory is greater than or equal to k, the sorting acceleration module can obtain the distance calculation results of k query vectors from the local memory.
[0101] In some embodiments, the sorting acceleration module performing sorting according to the distance calculation result further includes:
[0102] Step 1032: The final sorting module obtains the sorting intermediate states of all sorting acceleration modules and performs sorting according to the sorting intermediate states to obtain the sorting results corresponding to each query vector.
[0103] In other embodiments, the sorting acceleration module performing sorting according to the distance calculation result further includes:
[0104] Step 1032': The sorting intermediate state output after each sorting acceleration module sorts the distances between the corresponding query vectors and all center vectors is used as the sorting result corresponding to the query vector.
[0105] In this application, all query vectors are divided into multiple query vector blocks, and all center vectors in the center vector database are also divided into multiple center vector blocks. The calculation between each query vector block and each center vector block is performed, and the entire center vector library only needs to be loaded once during this process, without repeated loading, which greatly reduces the coarse search time.
[0106] Table 1 gives the coarse search time and total time of this embodiment using two different logics of VCORE and TCORE respectively and Figure 1Comparison of the coarse search time and total time of the traditional scheme shown. It can be seen that in this embodiment, both the coarse search time and the total time have decreased to varying degrees. Among them, SIFT10M is the name of the test database, nb represents the number of database vectors, which is 10M here, d represents the dimension of each vector, M represents the parameter of how many segments a vector is split into during PQ quantization, that is, the number of PQ quantizers, code_num represents the number of quantization vectors of the PQ quantizer, nq represents the number of query vectors, topk represents the number of vectors finally selected by search and screening, and nprobe represents the number of clusters selected for coarse search.
[0107]
[0108] Table 1 Comparison of the coarse search time and total time in this embodiment with Figure 1 the comparison of the coarse search time and total time of the traditional scheme shown
[0109] The vector retrieval acceleration method of this application has greatly reduced the coarse search time and improved the processing efficiency of the hardware computing system, which allows users to select a larger number of central vectors nlist. We have conducted a performance theoretical analysis on the overall system using the vector retrieval acceleration method. As shown in Tables 2 to 4 below, the search times of different systems are given for SIFT1M vectors, SIFT10M vectors, and SIFT100M vectors respectively. Among them, ABU2.0 in the table is the system applying the vector retrieval acceleration method of this application, and when the system bandwidth is only half of that of the dual-Die H100 NVL, the system performance can be comparable to it.
[0110]
[0111] Table 2 Search times of different systems for 1M vectors
[0112]
[0113] Table 3 Search times of different systems for 10M vectors
[0114]
[0115] Table 4: Search time of different systems for 100M vectors Correspondingly, the embodiments of the present application also provide a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the various method embodiments of the present application are implemented. The computer-readable storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, CD-ROM, digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media do not include transitory media, such as modulated data signals and carrier waves.
[0116] In addition, the embodiments of the present application also provide a coarse search optimization system for IVFPQ, which includes a memory for storing computer-executable instructions and a processor. The processor is configured to implement the steps in the above-mentioned method embodiments when executing the computer-executable instructions in the memory. Among them, the processor can be a central processing module (Central Processing Unit, abbreviated as "CPU"), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as "DSP"), application-specific integrated circuits (Application Specific Integrated Circuit, abbreviated as "ASIC"), etc. The aforementioned memory can be a read-only memory (abbreviated as "ROM"), random access memory (abbreviated as "RAM"), flash memory, hard disk, or solid-state drive, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0117] It should be noted that in the application documents of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising said element. In the application documents of this patent, if it is mentioned that an act is performed according to a certain element, it means that the act is performed at least according to that element, including two cases: the act is performed only according to that element, and the act is performed according to that element and other elements. Expressions such as multiple, many times, and various include 2, 2 times, 2 kinds, and more than 2, more than 2 times, and more than 2 kinds.
[0118] All documents mentioned in this specification are considered to be integrally included in the disclosure of this application so that they can be used as a basis for modification if necessary. In addition, it should be understood that the above are only preferred embodiments of this specification and are not used to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the protection scope of one or more embodiments of this specification.
[0119] In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A vector retrieval acceleration system, characterized in that, including: At least one computing module, the computing module including a first receiving end and a second receiving end respectively for receiving a query vector block and a central vector block, and performing operations on the received query vector block and central vector block to obtain a distance calculation result between each query vector and all central vectors, wherein nq query vectors are divided into nq / n query vector blocks The nlist central vectors in the central vector database are divided into nlist / m central vector blocks Each query vector block includes n query vectors, and each central vector block includes m central vectors, where nq, nlist, m, and n are all integers greater than or equal to 1; and at least one sorting acceleration module, where the at least one computing module and the at least one sorting acceleration module correspond one by one. Each sorting acceleration module respectively receives the distance calculation results output by the corresponding computing module, sorts the distance calculation results, and saves the obtained sorting intermediate state to the system memory.
2. The system according to claim 1, wherein The operation of the computing module on the center vector block and the query vector block includes: performing a matrix multiplication operation on the query vector block and the center vector block, and adding each element in the matrix multiplication result to the square of the corresponding center vector respectively to obtain the distance calculation result.
3. The system according to claim 1, wherein the second receiving end of each computing module receives its corresponding center vector block from the center vector database. Corresponding to the corresponding center vector block, the first receiving end of each computing module sequentially receives each query vector block, so that each computing module can perform an operation on its corresponding center vector block and each query vector block; wherein, different computing modules respectively correspond to different center vector blocks.
4. The system according to claim 3, wherein The first receiving ends of the computing modules are connected together and receive the same query vector block at the same time.
5. The system according to claim 3, characterized in that, The system further includes a final sorting module, which obtains the sorting intermediate states of all sorting acceleration modules and sorts according to the sorting intermediate states to obtain the sorting results corresponding to each query vector.
6. The system according to claim 1, wherein the first receiving end of each computing module receives its corresponding query vector block. Corresponding to the corresponding query vector block, the second receiving end of each computing module sequentially receives each center vector block, so that each computing module can perform an operation on its corresponding query vector block and each center vector block; wherein, different computing modules respectively correspond to different query vector blocks.
7. The system according to claim 6, wherein The computing modules are connected by a bus; each computing module respectively receives different center vector blocks from the center vector database through its second receiving end, and sequentially transfers the center vector blocks received by each of them among the computing modules through the bus, so that each computing module can sequentially receive different center vector blocks.
8. The system according to claim 6, wherein The sorting acceleration module loads the sorting intermediate state from the system memory, and sorts the sorting intermediate state together with the distance calculation results from the computing module or the local memory to obtain a new sorting intermediate state, and updates the sorting intermediate state stored in the system memory with the new sorting intermediate state.
9. The system according to claim 8, characterized in that, The sorting intermediate state obtained by each sorting acceleration module in the last sorting is the sorting result of the query vectors in the query vector block received by the corresponding computing module.
10. The system according to claim 1, characterized in that, The system further includes a local memory, and the distance calculation results calculated by the computing module are stored in the local memory and output to the corresponding sorting acceleration module through the local memory.
11. A method for accelerating vector retrieval, characterized in that, including: Divide the nq query vectors into nq / n query vector blocks Divide the nlist center vectors in the center vector database into nlist / m center vector blocks Each query vector block contains n query vectors, and each center vector block contains m center vectors, where nq, nlist, m, and n are all integers greater than or equal to 1; Receiving, by at least one computing module, center vector blocks C k 0, center vector block C k 1, …, center vector blocks from the center vector library respectively, and k taking center vector block C k 0, center vector block C 1, …, center vector blocks q 0, query vector block x q 1, …, query vector blocks in each of the center vector blocks to perform an operation with the query vector block x respectively, so as to obtain distance calculation results between each query vector and all center vectors; and The distance calculation results of the central vector block and the query vector block are obtained from the corresponding computing module by at least one sorting acceleration module, and sorting is performed according to the distance calculation results to obtain the sorting results corresponding to each query vector.
12. The method according to claim 11, wherein Each central vector block is respectively operated on with the query vector block x q 0, the query vector block x q 1, …, the query vector block The operations include: Respectively perform matrix multiplication on each central vector block with the query vector block x q 0, the query vector block x q 1, …, the query vector block and add the square of the corresponding central vector to each element in the obtained matrix multiplication result to obtain the distance calculation result.
13. The method according to claim 12, characterized in that, Each central vector block is respectively matrix-multiplied with the query vector block x q 0, the query vector block x q 1, …, the query vector block The matrix multiplication operations include: Input the central vector block C k 0, the central vector block C k 1, …, the central vector block into the corresponding computing modules among the at least one computing module respectively; and For each computing module, input the query vector block x in sequence q 0, the query vector block x q 1, …, the query vector block so that each computing module calculates the matrix product of the corresponding central vector block and all query vector blocks.
14. The method according to claim 12, wherein Each central vector block is respectively matrix-multiplied with the query vector block x q 0, the query vector block x q 1, …, the query vector block The matrix multiplication operations include: Input the query vector block x q 0, the query vector block x q 1, …, the query vector block into the corresponding computing modules among the at least one computing module respectively; and For each computing module, the central vector block C is input sequentially k 0, the central vector block C k 1, …, the central vector block so that each computing module calculates the matrix product of the corresponding query vector block and all central vector blocks.
15. The method according to claim 14, wherein Obtaining the distance calculation results of the central vector block and the query vector block from the corresponding computing module by at least one sorting acceleration module and performing sorting according to the distance calculation results includes: The sorting acceleration module performs sorting based on the distance calculation results received from the corresponding computing module or the local memory and the sorting intermediate state obtained from the system memory to generate a new sorting intermediate state, and updates the sorting intermediate state stored in the system memory with the new sorting intermediate state.
16. The method according to claim 15, characterized in that, Performing sorting according to the distance calculation results to obtain the sorting results corresponding to each query vector includes: Taking the sorting intermediate state obtained by each sorting acceleration module in the last sorting as the sorting result of the query vector in the query vector block received by the corresponding computing module.
17. The method according to claim 13, wherein Performing sorting according to the distance calculation results to obtain the sorting results corresponding to each query vector includes: obtaining the sorting intermediate states of all sorting acceleration modules by the final sorting module and performing sorting according to the sorting intermediate states to obtain the sorting results corresponding to each query vector.
18. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, the steps in the method according to any one of claims 11 to 17 are implemented.