Vector search system for vector database

By setting up cluster search, vector capture, and global retrieval modules in the FPGA hardware processor, and optimizing data access in conjunction with the CPU's software runtime module, the low-latency search problem of large-scale vector databases is solved, and efficient vector approximate nearest neighbor search is achieved.

CN121070993APending Publication Date: 2025-12-05SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511200680.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing CPU-centric vector search systems cannot effectively solve the low-latency vector approximate nearest neighbor search problem for large-scale vector databases, mainly due to the limited computing resources of server CPUs and the high overhead of vector transmission across PCIe.

Method used

An architecture combining CPU and FPGA hardware acceleration card is adopted. By setting up a cluster search module, vector grabbing module, vector search module and global retrieval module in the FPGA hardware processor, hardware-accelerated approximate nearest neighbor search of large-scale vector database is realized. The software runtime module assists the hardware acceleration card to access NVMe SSD end-to-end, optimizing the computing and data transmission process.

Benefits of technology

It significantly improves the speed of approximate nearest neighbor search for large-scale vector databases, increasing the speed by 1.41 times to 171.28 times, while saving more than 90% of CPU computing resources, meeting the search requirements of low latency and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070993A_ABST
    Figure CN121070993A_ABST
Patent Text Reader

Abstract

The invention provides a vector search system for a vector database. A cluster search module and a vector search module receive query vectors input by an application; data transmission is carried out between the cluster search module and the vector capture module, the cluster search module is used for realizing cluster sampling calculation of the FPGA hardware processor, and the vector capture module is used for realizing end-to-end access of the FPGA hardware processor; data transmission is carried out between the vector search module and the global retrieval module, the vector search module is used for realizing vector search calculation of the FPGA hardware processor, and the global retrieval module is used for realizing sorting calculation of the FPGA hardware processor; and the software runtime module is used for assisting the vector capturing module to realize end-to-end access. The end-to-end approximate nearest neighbor search speed of the system is 1.41-171.28 times faster than that of an existing system, and meanwhile, more than 90% of CPU computing resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large-scale vector database technology, and also to the fields of vector approximate nearest neighbor search, computer hardware and software architecture, NVMe SSD storage system, etc. Specifically, it relates to a vector search system for vector databases, and more particularly to a vector search acceleration system for large-scale vector databases. Background Technology

[0002] Deep learning models abstract unstructured data (such as images, videos, and speech) into high-dimensional vectors. Vector databases persistently store these vectors and provide vector search services. Vector search, as a core service of vector databases, is deployed in multiple application domains, such as recommender systems, commercial search engines, large language models (LLMs), and bioinformatics retrieval. Its goal is to search the vector database for the top-k nearest neighbor vectors similar to a given vector. Specifically, vector search measures similarity by calculating distance metrics (commonly Euclidean and cosine distances) between a given vector and every vector in the database, and outputs the top-k vectors with the smallest distances (most similar). Therefore, vector search requires significant computational and memory resources to achieve high-precision and low-latency searches.

[0003] In recent years, the explosive growth of unstructured data has expanded the size of vector databases from GB to TB, and even PB. The sheer volume of data and the curse of high dimensionality have prompted vector search to abandon traditional exact nearest neighbor search (KNNS) and embrace approximate nearest neighbor search (ANNS). Currently, many ANNS algorithms have been proposed, such as KDTree, LSH, HNSW, and IVF. However, for large-scale (hundreds of millions) vector databases, achieving high-accuracy search with ANNs still requires massive computational and memory resources and exhibits high search latency. Specifically, since large-scale vector data far exceeds memory capacity, existing solutions typically implement disk-based search. Vectors are stored on high-performance disks, and the vector approximate nearest neighbor search engine reads vectors from the disk in batches and performs ANNS.

[0004] However, existing vector search systems are all implemented in CPU-centric architectures. CPU-centric architectures exhibit the following characteristics: (1) limited CPU computing resources; and (2) the CPU and host memory act as bridges for data interaction between accelerators (e.g., GPUs) and storage devices. These characteristics prevent existing CPU-centric vector search systems from providing low-latency vector approximate nearest neighbor search services for large-scale vector databases.

[0005] Existing CPU-centric vector search systems face the following challenges: (1) Computational challenges: server CPU computing resources cannot meet the computational requirements of large-scale vector searches, resulting in high latency for high-precision searches; (2) Vector movement challenges: CPU-centric vector reading causes existing vector approximate nearest neighbor search systems to incur overhead when reading vector features from storage devices in batches, involving vector movement across the Peripheral Component Interconnect Express (PCIe). Because current mainstream accelerators (such as GPUs) are connected to servers via PCIe, when using accelerators to accelerate search computation, large-scale vectors stored in storage devices must be read by the CPU into CPU memory and then transferred to accelerator memory via PCIe. Therefore, due to limited server CPU computing resources and severe vector cross-PCIe transmission overhead, existing vector approximate nearest neighbor search systems cannot provide low-latency vector approximate nearest neighbor search services for large-scale vector databases.

[0006] Patent document CN119669525A discloses a heterogeneous system for large-scale vector approximate nearest neighbor search based on a GPU-NVMe architecture. Its key feature is that the system is based on a direct GPU-NVMe architecture and designs a hybrid index combining clustering and graph algorithms. This hybrid index utilizes the parallel processing capabilities of the GPU to efficiently process well-clustered vectors, while simultaneously using a graph index as a compensation mechanism to improve the search accuracy for noisy vectors. This invention fully considers the collaborative optimization of hardware and software, implementing vector index data deployment at the NVMe driver level, and collaboratively utilizing CPU and GPU hardware to perform billion-level vector retrieval calculations, thus constructing a scalable large-scale retrieval system on a direct GPU-NVMe architecture. However, this patent document still has the drawback of not being able to provide low-latency vector approximate nearest neighbor search services for large-scale vector databases. Summary of the Invention

[0007] In view of the shortcomings of the prior art, the purpose of this invention is to provide a vector search system for vector databases.

[0008] A vector search system for a vector database according to the present invention includes: a server and a hardware acceleration card;

[0009] The server includes a CPU software processor, which has a software runtime module; the hardware acceleration card includes an FPGA hardware processor, a vector capture module, a vector search module, and a global retrieval module, which has a cluster search module.

[0010] The cluster search module and the vector search module receive query vectors input by the application;

[0011] Data is transmitted between the cluster search module and the vector capture module. The cluster search module is used to implement cluster sampling calculation of the FPGA hardware processor, and the vector capture module is used to implement end-to-end access of the FPGA hardware processor.

[0012] The vector search module and the global retrieval module transmit data. The vector search module is used to perform vector search calculations on the FPGA hardware processor, and the global retrieval module is used to perform sorting calculations on the FPGA hardware processor.

[0013] The software runtime module is used to assist the vector crawling module in achieving end-to-end access.

[0014] Preferably, the software runtime module polls the storage request submission queue in the local memory of the CPU software processor through memory read / write control.

[0015] When the storage request submission queue is not empty, the software runtime module notifies the doorbell register of the NVMe SSD controller through memory-mapped I / O to trigger the NVMe SSD controller to process the storage request. At the same time, it polls the doorbell register of the NVMe SSD controller through memory-mapped I / O to obtain the interrupt signal of the NVMe SSD controller to complete the storage request processing.

[0016] When an interrupt signal is received, the software runtime module informs the NVMe SSD controller of the pointer to the storage request completion queue in the local memory of the CPU software processor, and assists the NVMe SSD controller in submitting the storage completion request to the storage request completion queue in the local memory of the CPU software processor.

[0017] Preferably, the cluster search module receives the query vector input by the application, performs sampling calculations on the vector clusters in the vector database, and transmits the sampled cluster information to the vector capture module; the cluster information includes the index of the sampled vector database vectors;

[0018] The vector capture module reads vector features from the NVMe SSD controller into the local global memory of the hardware acceleration card based on the vector index in the sampled cluster output by the cluster search module; when the vector search system starts, the vector capture module prefetches the vector cluster features and the vector index of the vector database carried by the vector cluster into the local programmable memory and local global memory of the hardware acceleration card, respectively.

[0019] Preferably, the vector search module receives the query vector input by the application and performs concurrent and pipelined similarity calculation between the query vector and the sampled vector database vector at the granularity of the sampled vector cluster. After the similarity calculation is completed, the vector search module outputs 2×top-k locally most similar vectors to the global retrieval module to obtain the top-k vector database vectors that are most similar to the query vector.

[0020] Preferably, the global retrieval module aggregates all vector database vectors that are locally most similar to the query vector, output by the vector search module, and performs a fully parallel sorting operation accelerated by the FPGA hardware processor based on the similarity between the vector database vectors and the query vector; based on the sorting result, the global retrieval module selects the top-k vector database vectors that are globally most similar to the query vector and returns them to the application.

[0021] Preferably, the software runtime module includes: a doorbell register submodule, a storage request submission processing submodule, a storage request completion processing submodule, a queue address translation submodule, an accelerator card global memory address translation submodule, multiple storage request submission queues, and multiple storage request completion queues;

[0022] Multiple storage request submission queues and multiple storage request completion queues are used to cache storage access requests for the NVMe SSD controller generated by the vector fetching module; the multiple storage request submission queues and multiple storage request completion queues are allocated in the local memory of the CPU software processor;

[0023] The doorbell register submodule exposes the NVMe SSD controller's doorbell register on the PCIe interface and maps it to the CPU software processor's local memory via memory-mapped I / O;

[0024] The storage request submission processing submodule polls the storage request submission queue and manipulates the doorbell register submodule when a new storage request is generated to notify the NVMe SSD controller to process the storage request; the storage request completion processing submodule polls the doorbell register submodule to obtain the storage request completion interrupt generated by the NVMe SSD controller, and informs the doorbell register submodule of the pointer to the storage request completion queue to assist the NVMe SSD controller in writing the generated storage completion request into the storage request completion queue;

[0025] The queue address translation submodule translates the memory addresses of the storage request submission queue and the storage request completion queue into pointers that can be accessed by the hardware acceleration card. The vector capture module uses these pointers to insert storage access requests into the storage request queue or to retrieve storage completion requests from the storage request completion queue.

[0026] The global memory address translation submodule of the accelerator card converts the vector buffer and cluster buffer allocated in the global memory of the hardware accelerator card into pointers that can be accessed by the CPU software processor. These pointers are stored in the local global memory of the hardware accelerator card and are encoded into a storage request by the vector capture module to assist the NVMe controller in using direct memory access to directly transfer vector and cluster features to the local global memory of the hardware accelerator card without going through the local memory of the CPU software processor.

[0027] Preferably, the cluster search module includes: multiple processing units;

[0028] Each of the processing units independently calculates the query vector and the Euclidean similarity of a cluster center;

[0029] Within the processing unit, the Euclidean similarity calculation is expanded along the vector dimension; the subtraction and squaring operations between query vector elements and cluster center feature elements located in the same dimension are implemented as multiple independent first calculation units, which concurrently process each vector element and output the results to the summation unit; the summation unit outputs the summation result to the square root unit to complete the Euclidean similarity calculation.

[0030] The similarity metric output by the processing unit is used by a priority queue for cluster sampling. The priority queue only retains the indices of the top P clusters closest to the query vector, and the parameter P is set by the application.

[0031] Preferably, the vector capture module includes: a storage request processing submodule and a data processing submodule;

[0032] When the vector search system is initialized, the storage request processing submodule constructs a read request for the cluster and cluster features. The read request is used to transfer the cluster and cluster features directly from the NVMe SSD controller to the local global memory of the hardware acceleration card.

[0033] The data processing submodule transmits the cluster features to the local programmable memory of the hardware acceleration card and transmits the vector index carried by the cluster to the local global memory of the hardware acceleration card.

[0034] The storage request processing submodule and the data processing submodule complete the following vector feature reading process:

[0035] The data processing submodule uses the cluster index output by the cluster search module to retrieve the vector cluster in the local global memory of the hardware acceleration card in order to obtain the sampled vector index.

[0036] The storage request processing submodule calculates the address of the vector feature in the NVMe SSD controller based on the vector index, encodes the address into the storage request, and encodes the vector buffer pointer stored in the local global memory of the hardware acceleration card into the storage request to indicate the destination address of the vector feature read operation.

[0037] The storage request processing submodule inserts the storage request into the storage request submission queue in the local memory of the CPU software processor through the pointer of the storage request submission queue to trigger the reading of vector features. At the same time, the storage request processing submodule polls the storage request completion queue in the local memory of the CPU software processor through the pointer of the storage request completion queue to process the storage completion request.

[0038] The data processing submodule transfers vector features from the vector buffer in the local global memory of the hardware accelerator card to the vector feature buffer in the local programmable memory of the hardware accelerator card, so that the vector search module can perform similarity calculation.

[0039] Preferably, the vector search module includes: a plurality of second computing units;

[0040] The processing units in the vector grabbing module and the cluster search module are expanded to construct multiple second computing units, each of which performs an end-to-end similarity calculation workflow on vectors in a sampled cluster.

[0041] The similarity calculation workflow includes three steps: the vector grabbing module reads vector features from the NVMe SSD controller into the vector feature buffer in the local programmable memory of the hardware acceleration card; multiple processing units calculate the Euclidean similarity between the query vector and the vectors in the vector feature buffer in parallel; and the indexes and Euclidean similarities of the 2 top-k vectors most similar to the query vector are output to the global retrieval module.

[0042] Within the second computing unit, the three steps of the similarity calculation workflow are constructed into a pipeline; the sampled clusters are evenly distributed among multiple second computing units, and each second computing unit pipelines the calculation of the Euclidean similarity between the query vector and the vectors in the sampled clusters assigned to it.

[0043] Preferably, the global retrieval module includes: a comparator array, an accumulator array, and a selector;

[0044] The Euclidean similarity of the local most similar vectors input to the global retrieval module is copied and stored horizontally and vertically in the local programmable memory of the hardware acceleration card, respectively.

[0045] The horizontal and vertical elements are two elements, and comparators are built between each pair of elements to form a comparator array in matrix form; multiple comparators in the comparator array concurrently compare the size of the horizontal and vertical elements, and if the vertical element is smaller than the horizontal element, the comparator outputs 1, otherwise the comparator outputs 0.

[0046] An accumulator is implemented vertically after the comparator array to construct the accumulator array; multiple accumulators in the accumulator array concurrently accumulate the output of the comparator in a horizontal manner;

[0047] The selector is built after the accumulator array. Based on the values ​​in the accumulator, it selects the vector indices corresponding to the top-k values ​​and outputs them to the application in descending order.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. This invention sets up a software runtime module in the CPU to enable control flow processing for end-to-end access to NVMe SSD by the hardware acceleration card, and sets up a cluster search module, a vector grabbing module, a vector search module and a global retrieval module in the FPGA to realize hardware-accelerated approximate nearest neighbor search for large-scale vector databases.

[0050] 2. According to tests, the end-to-end approximate nearest neighbor search speed of the system proposed in this invention is 1.41 times to 171.28 times faster than the existing system, while saving more than 90% of CPU computing resources.

[0051] 3. This invention solves the problem of high search latency when searching large-scale vector databases in existing systems due to the reliance on the CPU for approximate nearest neighbor search calculations and large-scale vector data transmission. It better meets the application's need for low-latency approximate nearest neighbor search in large-scale vector databases and has certain innovative and practical application value.

[0052] 4. This invention solves the problem that existing CPU-centric vector approximate nearest neighbor search systems cannot provide low-latency vector approximate nearest neighbor search services for large-scale vector databases due to limited server CPU computing resources and severe vector cross-PCIe transmission overhead.

[0053] 5. The system of the present invention overcomes the computational challenges and vector cross-PCIe movement challenges existing in the existing CPU-centric vector approximate nearest neighbor search system, and can provide low-latency and high-accuracy vector approximate nearest neighbor search services for large-scale vector databases. Attached Figure Description

[0054] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0055] Figure 1 This is a system block diagram of the vector search system for a vector database according to the present invention;

[0056] Figure 2 This is a system block diagram of the vector search system for a vector database in Embodiment 3;

[0057] Figure 3 This is a flowchart of the software runtime module in Example 3;

[0058] Figure 4 This is a flowchart of the cluster search module in Example 3;

[0059] Figure 5 This is a flowchart of the vector capture module in Example 3;

[0060] Figure 6 This is a flowchart of the vector search module in Example 3;

[0061] Figure 7 This is a flowchart of the global retrieval module in Example 3. Detailed Implementation

[0062] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0063] Example 1

[0064] like Figure 1As shown, this embodiment provides a vector search system for a vector database, including: a server and a hardware accelerator card; the server includes a CPU software processor, and the CPU software processor has a software runtime module; the hardware accelerator card includes an FPGA hardware processor, a vector capture module, a vector search module, and a global retrieval module, and the FPGA hardware processor has a cluster search module; the cluster search module and the vector search module receive query vectors input by the application; data is transmitted between the cluster search module and the vector capture module, the cluster search module is used to implement cluster sampling calculation of the FPGA hardware processor, and the vector capture module is used to implement end-to-end access of the FPGA hardware processor; data is transmitted between the vector search module and the global retrieval module, the vector search module is used to implement vector search calculation of the FPGA hardware processor, and the global retrieval module is used to implement sorting calculation of the FPGA hardware processor; the software runtime module is used to assist the vector capture module in implementing end-to-end access.

[0065] The software runtime module polls the storage request submission queue in the CPU software processor's local memory via memory read / write control. When the storage request submission queue is not empty, the software runtime module notifies the NVMe SSD controller's doorbell register via memory-mapped I / O, triggering the NVMe SSD controller to process the storage request. Simultaneously, it polls the NVMe SSD controller's doorbell register via memory-mapped I / O to obtain the NVMe SSD controller's interrupt signal indicating that the storage request processing is complete. When the interrupt signal is obtained, the software runtime module informs the NVMe SSD controller of the pointer to the storage request completion queue in the CPU software processor's local memory, assisting the NVMe SSD controller in submitting the storage completion request to the storage request completion queue in the CPU software processor's local memory.

[0066] The software runtime modules include: a doorbell register submodule, a storage request submission processing submodule, a storage request completion processing submodule, a queue address translation submodule, an accelerator card global memory address translation submodule, multiple storage request submission queues, and multiple storage request completion queues.

[0067] Multiple storage request submission queues and multiple storage request completion queues are used to cache storage access requests for the NVMe SSD controller generated by the vector fetching module; multiple storage request submission queues and multiple storage request completion queues are allocated in the local memory of the CPU software processor; the doorbell register submodule exposes the doorbell register of the NVMe SSD controller on the PCIe interface, and maps it to the local memory of the CPU software processor through memory-mapped I / O; the storage request submission processing submodule polls the storage request submission queues and manipulates the doorbell register submodule when there are newly generated storage requests to notify the NVMe SSD controller to process the storage request.

[0068] The storage request completion processing submodule polls the doorbell register submodule to obtain storage request completion interrupts generated by the NVMe SSD controller and informs the doorbell register submodule of the pointer to the storage request completion queue, so as to assist the NVMe SSD controller in writing the generated storage completion request into the storage request completion queue. The queue address translation submodule translates the memory addresses of the storage request submission queue and the storage request completion queue into pointers that can be accessed by the hardware accelerator card. The vector capture module uses these pointers to insert storage access requests into the storage request queue or to retrieve storage completion requests from the storage request completion queue. The accelerator card global memory address translation submodule translates the vector buffer and cluster buffer allocated in the global memory of the hardware accelerator card into pointers that can be accessed by the CPU software processor. These pointers are stored in the local global memory of the hardware accelerator card. These pointers are encoded into the storage request by the vector capture module to assist the NVMe controller in using direct memory access to directly transfer vector and cluster features to the local global memory of the hardware accelerator card without going through the local memory of the CPU software processor.

[0069] The cluster search module receives the query vector input by the application and performs sampling calculations on the vector clusters in the vector database. The sampled cluster information is then transmitted to the vector capture module. The cluster information includes the indexes of the vectors in the sampled vector database. Based on the vector indexes in the sampled clusters output by the cluster search module, the vector capture module reads the vector features from the NVMe SSD controller into the local global memory of the hardware accelerator card. When the vector search system starts, the vector capture module prefetches the vector cluster features and the vector indexes of the vector database carried by the vector clusters into the local programmable memory and local global memory of the hardware accelerator card, respectively.

[0070] The specific calculation process is as follows: The cluster search module calculates the Euclidean similarity score between the query vector and the vector cluster centers in the vector database. It then uses the information from the top-k most similar vector clusters as the sampled cluster information. The top-k most similar vector clusters are those ordered from lowest to highest Euclidean similarity score; the top k clusters are the desired top-k vector clusters. A lower Euclidean similarity score indicates a greater similarity between the query vector and the vector cluster centers. k is a parameter that can be configured by the user through the software runtime module.

[0071] Top-k refers to sorting the sampled vectors in ascending order from low to high based on their Euclidean similarity score (the similarity between the query vector and the sampled vectors in the database). The top k (k is a parameter that can be configured by the user through the software runtime module) sampled vectors are the desired locally most similar vectors. 2×top-k means taking the top 2×k sampled vectors as the locally most similar vectors.

[0072] The cluster search module includes: multiple processing units; each processing unit independently calculates the Euclidean similarity between the query vector and a cluster center; within each processing unit, the Euclidean similarity calculation is expanded along the vector dimension; subtraction and squaring operations between query vector elements and cluster center feature elements located in the same dimension are implemented as multiple independent first calculation units, which concurrently process each vector element and output the results to a summing unit; the summing unit outputs the summation result to a square root unit to complete the Euclidean similarity calculation; the similarity metric output by the processing units is used by a priority queue for cluster sampling, and the priority queue only retains the indices of the top P clusters closest to the query vector, where parameter P is set by the application.

[0073] The vector crawling module includes a storage request processing submodule and a data processing submodule.

[0074] When the vector search system is initialized, the storage request processing submodule constructs read requests for the cluster and cluster features. These read requests are used to transfer the cluster and cluster features directly from the NVMe SSD controller to the local global memory of the hardware accelerator card. The data processing submodule transfers the cluster features to the local programmable memory of the hardware accelerator card and transfers the vector index carried by the cluster to the local global memory of the hardware accelerator card.

[0075] The storage request processing submodule and the data processing submodule complete the following vector feature reading process:

[0076] The data processing submodule uses the cluster index output by the cluster search module to retrieve the vector cluster in the local global memory of the hardware acceleration card in order to obtain the sampled vector index.

[0077] The storage request processing submodule calculates the address of the vector feature in the NVMe SSD controller based on the vector index, encodes the address into the storage request, and encodes the vector buffer pointer stored in the local global memory of the hardware acceleration card into the storage request to indicate the destination address of the vector feature read operation.

[0078] The storage request processing submodule inserts the storage request into the storage request submission queue in the local memory of the CPU software processor through the pointer of the storage request submission queue to trigger the reading of vector features. At the same time, the storage request processing submodule polls the storage request completion queue in the local memory of the CPU software processor through the pointer of the storage request completion queue to process the storage completion request.

[0079] The data processing submodule transfers vector features from the vector buffer in the local global memory of the hardware accelerator card to the vector feature buffer in the local programmable memory of the hardware accelerator card, so that the vector search module can perform similarity calculations.

[0080] The vector search module receives the query vector input by the application and performs concurrent and pipelined similarity calculation between the query vector and the sampled vector database vectors at the granularity of the sampled vector cluster. After the similarity calculation is completed, the vector search module outputs the top-k locally most similar vectors to the global retrieval module to obtain the top-k vector database vectors that are most similar to the query vector.

[0081] The vector search module includes: multiple second computing units; the processing units in the vector grabbing module and the cluster search module are expanded to construct multiple second computing units, each of which performs an end-to-end similarity calculation workflow on vectors in a sampled cluster; the similarity calculation workflow includes three steps: the vector grabbing module reads vector features from the NVMe SSD controller into a vector feature buffer in the local programmable memory of the hardware acceleration card; multiple processing units calculate the Euclidean similarity between the query vector and the vectors in the vector feature buffer in parallel; the indexes and Euclidean similarities of the 2 top-k vectors most similar to the query vector are output to the global retrieval module; within the second computing unit, the three steps of the similarity calculation workflow are constructed into a pipeline; the sampled cluster is evenly distributed among multiple second computing units, and each second computing unit pipelines the calculation of the Euclidean similarity between the query vector and the vectors in the sampled cluster assigned to it.

[0082] The first computing unit is different from the second computing unit. The first computing unit can be modified into multiple hardware subtraction units and multiple hardware squaring units.

[0083] The global retrieval module aggregates all vectors in the vector database that are most locally similar to the query vector, as output by the vector search module. Based on the similarity between the vectors in the vector database and the query vector, it performs a fully parallel sorting operation accelerated by the FPGA hardware processor. According to the sorting results, the global retrieval module selects the top-k vectors in the vector database that are globally most similar to the query vector and returns them to the application.

[0084] According to the vector search module, the global retrieval module calculates the similarity between the query vector and the sampled vector database vectors in parallel at the vector cluster granularity, i.e., the Euclidean similarity score, and outputs 2×top-k locally most similar vectors. The locally most similar vectors indicate that the vector is most similar to the query vector in its own vector cluster, without involving comparison with sampled vectors in other vector clusters. The global retrieval module performs cross-vector cluster comparisons on the sampled vectors that are most similar to the query vector in all vector clusters, thereby obtaining the final top-k sampled vectors that are most similar to the query vector, i.e., the top-k vector database vectors that are globally most similar to the query vector.

[0085] The global retrieval module includes a comparator array, an accumulator array, and a selector. The Euclidean similarity of the locally most similar vectors input to the global retrieval module is copied and stored horizontally and vertically in the local programmable memory of the hardware accelerator card. Each horizontal and vertical element consists of two elements, with comparators built between each pair to form a matrix-like comparator array. Multiple comparators in the comparator array concurrently compare the horizontal and vertical elements; if the vertical element is smaller than the horizontal element, the comparator outputs 1, otherwise it outputs 0. An accumulator is implemented vertically after the comparator array to construct an accumulator array. Multiple accumulators in the accumulator array concurrently accumulate the comparator outputs horizontally. A selector is built after the accumulator array, selecting the vector indices corresponding to the top-k values ​​based on the values ​​in the accumulators and outputting them to the application in descending order.

[0086] This embodiment discloses a vector search acceleration system for large-scale vector databases. The system includes a server equipped with a CPU software processor and a hardware acceleration card equipped with an FPGA hardware processor. This embodiment incorporates a software runtime module within the CPU to enable control flow processing for end-to-end access to NVMe SSDs by the hardware acceleration card. The FPGA includes a cluster search module, a vector fetching module, a vector search module, and a global retrieval module, enabling hardware-accelerated approximate nearest neighbor search for large-scale vector databases.

[0087] Tests have shown that the end-to-end approximate nearest neighbor search speed of the system proposed in this embodiment is 1.41 times to 171.28 times faster than existing systems, while saving more than 90% of CPU computing resources.

[0088] This embodiment solves the problem of high search latency when searching large-scale vector databases in existing systems due to the reliance on the CPU for approximate nearest neighbor search calculations and large-scale vector data transmission. It better meets the application's need for low-latency approximate nearest neighbor search in large-scale vector databases and has certain innovativeness and practical application value.

[0089] Example 2

[0090] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1.

[0091] This embodiment provides a vector search acceleration system for large-scale vector databases. The system includes a server equipped with a CPU software processor and a hardware acceleration card equipped with a Field-Programmable Gate Array (FPGA) hardware processor. The hardware acceleration card is connected to the server via a high-speed serial computer expansion bus (Peripheral Component Interconnect Express, PCIe).

[0092] Generally, a vector database can be considered a large-scale vector database if it contains 1 billion or more vectors and requires hundreds of GB of disk storage.

[0093] The server equipped with a CPU software processor includes:

[0094] A software runtime system, i.e., a CPU software processor, contains a software runtime module. This module assists the vector fetching module in the FPGA hardware processor in implementing end-to-end access to Non-Volatile Memory Express Solid State Drives (NVMe SSDs) according to the Non-Volatile Memory Express Solid State Drive (NVMe SSD) interface specification. The software runtime module polls the storage request submission queue in local memory via memory read / write control. When the queue is not empty, it notifies the NVMe SSD controller's doorbell register via Memory Mapped I / O (MMIO) to trigger NVMe SSD processing of the storage request. Simultaneously, the software runtime module polls the NVMe SSD controller's doorbell register via MMIO to obtain an interrupt signal indicating that the NVMe SSD has completed processing the storage request. When an interrupt signal is received, it informs the NVMe SSD controller of the pointer to the storage request completion queue in local memory, assisting the NVMe SSD controller in submitting the storage completion request to the local memory storage request completion queue.

[0095] The hardware acceleration card includes:

[0096] An FPGA hardware processor, used for approximate nearest neighbor vector search hardware acceleration, incorporates a cluster search module for FPGA-accelerated cluster sampling computation. The cluster search module receives query vectors from the application and performs sampling computation on vector clusters in a vector database. The sampled cluster information (including indices of the sampled vector database vectors) is transmitted to a vector fetching module to enable end-to-end reading of large-scale sampled vector database vectors from an NVMe SSD.

[0097] A vector fetching module is provided to implement end-to-end fast path NVMe SSD access accelerated by FPGA hardware, thereby optimizing the reading of vectors from a large-scale sampled vector database. The vector fetching module reads large-scale high-dimensional vector features from the NVMe SSD into the local global memory of the FPGA based on the vector indexes in the sampled clusters output by the cluster search module. Furthermore, the vector fetching module prefetches vector cluster features and the vector database vector indexes carried by the vector clusters into the FPGA's local programmable memory and local global memory respectively during system startup, enabling the reuse of vector cluster information across different query requests, thus accelerating the end-to-end processing flow of large-scale vector approximate nearest neighbor search.

[0098] A vector search module is provided to implement FPGA hardware-accelerated vector search computation. The module receives a query vector from the application and performs concurrent and pipelined similarity calculations between the query vector and the sampled vector database vectors at the granularity of sampled vector clusters. After similarity calculation, the module outputs the top-k locally most similar vectors to the global retrieval module to obtain the top-k vectors in the database that are most similar to the query vector. Here, top-k can be 1, 10, or 100, or can be set by the application.

[0099] A global retrieval module is used to implement FPGA hardware-accelerated sorting computation. The global retrieval module aggregates all vectors from the vector search module that are locally most similar to the query vector in the vector database, and performs a fully parallel sorting operation accelerated by FPGA hardware based on the similarity between the vectors in the vector database and the query vector. Finally, based on the sorting results, the global retrieval module selects the top-k vectors from the vector database that are globally most similar to the query vector and returns them to the application.

[0100] Furthermore, the software runtime module in the CPU software processor consists of multiple storage request submission queues / storage request completion queues, a doorbell register submodule, a storage request submission processing submodule, a storage request completion processing submodule, a queue address translation submodule, and an accelerator card global memory address translation submodule.

[0101] The storage request submission queue / storage request completion queue handles NVMe SSD storage access requests generated by the cache vector fetch module. These requests are allocated in CPU memory to conserve global memory on the hardware accelerator card. The doorbell register submodule maps the doorbell registers exposed on the PCIe interface of the NVMe SSD controller to CPU memory via MMIO. The software runtime module can then manipulate the doorbell registers and the NVMe SSD controller via memory read / write control. The storage request submission processing submodule polls the storage request submission queue and manipulates the doorbell register submodule to notify the NVMe SSD controller to process new storage requests. The storage request completion processing submodule polls the doorbell register submodule to receive storage request completion interrupts generated by the NVMe SSD controller. It then informs the doorbell register submodule of the pointer to the storage request completion queue to assist the NVMe SSD controller in writing the generated storage completion request to the storage request completion queue. The queue address translation submodule translates the memory addresses of the storage request submission queue and the storage request completion queue into pointers accessible to the hardware accelerator card. The vector fetching module uses this pointer to insert storage access requests into the storage request queue or to retrieve storage completion requests from the storage request completion queue. The accelerator card global memory address translation submodule translates the vector buffers and cluster buffers allocated in the hardware accelerator card's global memory into CPU-accessible pointers. These pointers are stored in the hardware accelerator card's global memory. They are encoded into storage requests by the vector fetching module to assist the NVMe controller in using Direct Memory Access (DMA) to directly transfer vector and cluster features to the hardware accelerator card's global memory without going through CPU memory.

[0102] Furthermore, the cluster search module in the FPGA hardware processor consists of multiple processing elements (PEs), each PE independently calculates the query vector and the Euclidean similarity of a cluster center.

[0103] Within the PE (Predicted Similarity Measure), Euclidean similarity calculation is unfolded along the vector dimension. Subtraction and squaring operations between query vector elements and cluster center feature elements located in the same dimension are implemented as independent computational units. These computational units process each vector element concurrently and output the results to a summation unit. The summation result is then output to a square root unit to complete the final step of the Euclidean similarity calculation. Finally, the similarity metric output by the PE is used for cluster sampling by a priority queue. The priority queue retains only the indices of the top P clusters closest to (most similar to) the query vector. The parameter P is set by the application.

[0104] Furthermore, the vector capture module in the FPGA hardware processor consists of a storage request processing submodule and a data processing submodule.

[0105] The storage request processing submodule constructs read requests for the cluster and its features when the system is initialized. These read requests are used to transfer the cluster and its features directly from the NVMe SSD to the hardware accelerator card's local global memory. The data processing submodule further transfers the cluster features to the hardware accelerator card's local programmable memory, while the vector index carried by the cluster is transferred to the hardware accelerator card's local global memory. Furthermore, the storage request processing submodule and the data processing submodule complete the large-scale vector feature reading process as described below. The data processing submodule uses the cluster index output by the cluster search module to retrieve the vector cluster in the hardware accelerator card's local global memory to obtain the sampled vector index. Subsequently, the storage request processing submodule calculates the address of the vector feature in the NVMe SSD based on the vector index and encodes it into the storage request. The vector buffer pointer stored in the hardware accelerator card's local global memory is also encoded into the storage request to indicate the destination address of the vector feature read operation. The storage request processing submodule inserts the storage request into the storage request submission queue in CPU memory via the storage request submission queue pointer to trigger the reading of the vector feature. Simultaneously, the storage request processing submodule polls the storage request completion queue in CPU memory via the storage request completion queue pointer to process the storage completion request. Finally, the data processing submodule transfers the vector features from the vector buffer in the local global memory of the hardware accelerator card to the vector feature buffer in the local programmable memory of the hardware accelerator card for the vector search module to perform similarity calculations.

[0106] Furthermore, the vector search module in the FPGA hardware processor consists of eight computing units (CUs).

[0107] The PEs in the vector grabbing module and the cluster search module are expanded to construct 8 compute units (CUs). Each CU performs an end-to-end similarity calculation workflow for vectors in a sampled cluster. This includes the vector grabbing module reading vector features from the NVMe SSD into a vector feature buffer in the local programmable memory of the hardware accelerator card; multiple PEs computing in parallel the Euclidean similarity between the query vector and the vectors in the vector feature buffer; and outputting the indexes and Euclidean similarities of the 2 top-k vectors most similar to the query vector to the global search module. Internally, these three steps are pipelined to further optimize latency and improve throughput. Furthermore, the sampled clusters are evenly distributed across the 8 CUs. Each CU pipelines the computation of the Euclidean similarity between the query vector and the vectors in the sampled cluster assigned to it.

[0108] Furthermore, the global retrieval module in the FPGA hardware processor consists of a comparator array, an accumulator array, and a selector.

[0109] The Euclidean similarity of the locally most similar vectors input to the global retrieval module is copied and stored horizontally and vertically in the local programmable memory of the hardware accelerator card. Comparators are built between every two elements to form a matrix array of comparators. The comparators in the comparator array concurrently compare the size of the horizontal and vertical elements. If the vertical element is smaller than the horizontal element, the comparator outputs 1; otherwise, it outputs 0. Accumulators are implemented vertically after the comparator array to build an accumulator array. The accumulators in the accumulator array concurrently accumulate the outputs of the comparators horizontally. Finally, a selector is built after the accumulator array. Based on the values ​​in the accumulators, it selects the vector indices corresponding to the top-k values ​​and outputs them to the application in descending order.

[0110] The purpose of this embodiment is to address the problem that existing CPU-centric vector approximate nearest neighbor search systems cannot provide low-latency vector approximate nearest neighbor search services for large-scale vector databases due to limited server CPU computing resources and severe vector cross-PCIe transmission overhead. This embodiment provides a vector search acceleration system for large-scale vector databases that overcomes the computational and cross-PCIe vector movement challenges of existing CPU-centric vector approximate nearest neighbor search systems, enabling it to provide low-latency, high-accuracy vector approximate nearest neighbor search services for large-scale vector databases.

[0111] This embodiment provides an end-to-end approximate nearest neighbor search system for large-scale vector databases that is 1.41 to 171.28 times faster than existing systems, while saving more than 90% of CPU computing resources. Compared with existing CPU-centric vector approximate nearest neighbor search systems, this embodiment successfully overcomes the computational challenges and cross-PCIe vector movement challenges in large-scale vector database approximate nearest neighbor search. It solves the problem of high search latency in large-scale vector databases caused by the reliance on the CPU for approximate nearest neighbor search computation and large-scale vector data transmission in existing systems, thus better meeting the application's need for low-latency approximate nearest neighbor search in large-scale vector databases.

[0112] Example 3

[0113] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1.

[0114] like Figure 2As shown, the server in this embodiment uses a commercial server equipped with an Intel(R) Xeon(R) Gold 6330 CPU software processor, and the hardware accelerator card uses a commercial hardware accelerator card equipped with an AMD Xilinx XCUX35-3VSVA1365E FPGA hardware processor. In addition, the hardware accelerator card also includes some commonly used hardware interfaces and peripheral devices. In this embodiment, the hardware interface used is a PCIe Gen4×8 high-speed serial computer expansion bus interface, and the peripheral devices used are local global memory DDR4 and local programmable memory BRAM. The hardware accelerator card connects to the server via the PCIe Gen4×8 high-speed serial computer expansion bus interface.

[0115] The CPU software processor corresponds to a 28-core Intel(R) Xeon(R) Gold 6330 CPU software processor, which has a software runtime module; the FPGA hardware processor corresponds to an AMD Xilinx XCUX35-3VSVA1365E Field Programmable Gate Array (FPGA) hardware processor, which has a cluster search module, a vector capture module, a vector search module and a global retrieval module.

[0116] First, the vector fetching module located in the FPGA hardware processor constructs a storage read request for the vector cluster features and vector database indexes carried by the vector clusters stored in the NVMe SSD during system startup in this embodiment. This is to prefetch the vector cluster features and vector database indexes carried by the vector clusters into the local programmable memory BRAM and local global memory DDR4 of the hardware accelerator card (FPGA), respectively. The process of the vector fetching module reading the vector cluster features and vector database indexes carried by the vector clusters requires real-time assistance from the software runtime module located in the CPU software processor.

[0117] Specifically, the software runtime module located in the CPU software processor polls the storage request submission queue in the server's local memory via memory read / write control. When the queue is not empty, it notifies the NVMe SSD controller's doorbell register via Memory-Mapped I / O (MMIO) to trigger NVMe SSD processing of the storage request. Simultaneously, the software runtime module polls the NVMe SSD controller's doorbell register via MMIO to obtain an interrupt signal indicating that the NVMe SSD has completed processing the storage request. When an interrupt signal is received, it informs the NVMe SSD controller of the pointer to the storage request completion queue in the server's local memory, assisting the NVMe SSD controller in submitting the storage completion request to the queue. Subsequently, the cluster search module located in the FPGA hardware processor receives the query vector input by the application and performs sampling calculations using the vector cluster features prefetched into the accelerator card's local programmable memory. The sampled cluster information (including the index of the sampled vector database vectors) is transmitted to the vector grabbing module to achieve end-to-end reading of large-scale sampled vector database vector features from the NVMe SSD. The vector grabbing module, located in the FPGA hardware processor, reads large-scale, high-dimensional vector features from the NVMe SSD into the local global memory of the hardware accelerator card based on the vector indices in the sampled cluster output by the cluster search module, and then further transfers them in batches to the accelerator card's local programmable memory. The process of the vector grabbing module reading large-scale vector features also requires the software runtime module located in the CPU software processor to implement the aforementioned assistance process.

[0118] Furthermore, the vector search module located in the FPGA hardware processor receives the query vector input by the application and performs concurrent and pipelined similarity calculations between the query vector and the sampled vector database vectors at the granularity of sampled vector clusters. The sampled vector database vectors refer to the large-scale, high-dimensional vector features that have been read in batches into the accelerator card's local programmable memory by the vector grabbing module. After the similarity calculation is completed, the vector search module outputs 2 × top-k locally most similar vectors to the global retrieval module to obtain the top-k vector database vectors most similar to the query vector. Here, top-k can be 1, 10, or 100, or can be set by the application. Finally, the global retrieval module in the FPGA hardware processor aggregates all locally most similar vector database vectors output by the vector search module and performs a fully parallel sorting operation accelerated by the FPGA hardware based on the similarity between the vector database vectors and the query vector. Based on the sorting results, the global retrieval module selects the top-k globally most similar vector database vectors to the query vector and returns them to the application. The software runtime module of this embodiment is implemented in one core of the CPU software processor, namely an Intel(R) Xeon(R) Gold 6330 CPU software processor core, using the C / C++ software programming language. The cluster search module, vector capture module, vector search module, and global retrieval module of this embodiment are implemented in the FPGA hardware processor of the hardware accelerator card, namely the AMD Xilinx XCUX35-3VSVA1365E FPGA hardware processor, using the High-Level Synthesis (HLS) language.

[0119] like Figure 3 The diagram shown is a flowchart of the software runtime module in this embodiment. In this embodiment, the software runtime module is implemented in a single core of a 28-core Intel(R) Xeon(R) Gold6330 CPU software processor in C / C++. The software runtime module consists of multiple storage request submission queues / storage request completion queues, a doorbell register submodule, a storage request submission processing submodule, a storage request completion processing submodule, a queue address translation submodule, and an accelerator card global memory address translation submodule.

[0120] The storage request submission queue / storage request completion queue handles NVMe SSD storage access requests generated by the cache vector fetch module. These requests are allocated in CPU (server) memory to conserve global DDR4 memory on the hardware accelerator card. The doorbell register submodule maps the doorbell registers exposed on the PCIe interface of the NVMe SSD controller to CPU memory via MMIO. The software runtime module can then manipulate the doorbell registers and the NVMe SSD controller via memory read / write control. The storage request submission processing submodule polls the storage request submission queue and manipulates the doorbell register submodule to notify the NVMe SSD controller to process new storage requests. The storage request completion processing submodule polls the doorbell register submodule to receive storage request completion interrupts generated by the NVMe SSD controller. It then informs the doorbell register submodule of the pointer to the storage request completion queue to assist the NVMe SSD controller in writing the generated storage completion request to the storage request completion queue. The queue address translation submodule translates the memory addresses of the storage request submission queue and the storage request completion queue into pointers accessible to the hardware accelerator card. The vector grabbing module located in the FPGA hardware processor uses this pointer to insert storage access requests into the storage request queue or to retrieve storage completion requests from the storage request completion queue. The accelerator card global memory address translation submodule translates the vector buffers and cluster buffers allocated in the hardware accelerator card's global memory DDR4 into CPU-accessible pointers. These pointers are stored in the hardware accelerator card's global memory DDR4. They are encoded into storage requests by the vector grabbing module in the FPGA hardware processor to assist the NVMe controller in using Direct Memory Access (DMA) to directly transfer vector and cluster features to the hardware accelerator card's global memory DDR4 without going through CPU memory.

[0121] like Figure 4 The diagram shown is a flowchart of the cluster search module in this embodiment. In this embodiment, the cluster search module is implemented in the FPGA hardware processor of the hardware acceleration card using HLS language encoding. The cluster search module consists of multiple processing elements (PEs), each PE independently calculating the Euclidean similarity between the query vector and a cluster center.

[0122] Within the PE (Predicted Similarity Measure), Euclidean similarity calculation is unfolded along the vector dimension. Subtraction and squaring operations between query vector elements and cluster center feature elements located in the same dimension are implemented as independent computational units. These computational units process each vector element concurrently and output the results to a summation unit. The summation result is then output to a square root unit to complete the final step of the Euclidean similarity calculation. Finally, the similarity metric output by the PE is used for cluster sampling by a priority queue. The priority queue retains only the indices of the top P clusters closest to (most similar to) the query vector. The parameter P is set by the application.

[0123] like Figure 5 The diagram shown is a flowchart of the vector capture module in this embodiment. In this embodiment, the vector capture module is implemented in the FPGA hardware processor of the hardware accelerator card using HLS language encoding. The vector capture module consists of a storage request processing submodule and a data processing submodule.

[0124] In this embodiment, the storage request processing submodule constructs a read request for the cluster and its features during system initialization. The read request is used to transfer the cluster and cluster features directly from the NVMe SSD to the hardware accelerator card's local global memory (DDR4). The data processing submodule further transfers the cluster features to the hardware accelerator card's local programmable memory (BRAM), while the vector index carried by the cluster is transferred to the hardware accelerator card's local global memory (DDR4). Furthermore, the storage request processing submodule and the data processing submodule complete the large-scale vector feature reading process as described below. The data processing submodule uses the cluster index output by the cluster search module to retrieve the vector cluster in the hardware accelerator card's local global memory (DDR4) to obtain the sampled vector index. Subsequently, the storage request processing submodule calculates the address of the vector feature in the NVMe SSD based on the vector index and encodes it into the storage request. The vector buffer pointer stored in the hardware accelerator card's local global memory (DDR4) is also encoded into the storage request to indicate the destination address of the vector feature read operation. The storage request processing submodule inserts the storage request into the storage request submission queue in CPU memory via the storage request submission queue pointer to trigger the reading of the vector features. Simultaneously, the storage request processing submodule polls the storage request completion queue in CPU memory via the storage request completion queue pointer to process storage completion requests. Finally, the data processing submodule transfers vector features from the vector buffer in the hardware accelerator card's local global memory DDR4 to the vector feature buffer in the hardware accelerator card's local programmable memory BRAM for the vector search module to perform similarity calculations.

[0125] like Figure 6The diagram shown is a flowchart of the vector search module in this embodiment. In this embodiment, the vector search module is implemented in the FPGA hardware processor of the hardware accelerator card using HLS language encoding. The vector search module consists of 8 computing units (CUs).

[0126] The vector fetching module and the physical search (PE) modules within the FPGA hardware processor of the accelerator card are scaled in parallel to construct 8 compute units (CUs). Each CU performs an end-to-end similarity calculation workflow for vectors in a sampled cluster. This includes the vector fetching module reading vector features from the NVMe SSD into a vector feature buffer in the accelerator card's local programmable memory (BRAM); multiple PEs computing the Euclidean similarity between the query vector and the vectors in the feature buffer in parallel; and outputting the indices and Euclidean similarities of the 2×top-k vectors most similar to the query vector to the global search module in the FPGA hardware processor. Within each CU, these three steps are pipelined to further optimize latency and improve throughput. Furthermore, the sampled clusters are evenly distributed across the 8 CUs. Each CU pipelines the computation of the Euclidean similarity between the query vector and the vectors in its assigned sampled cluster.

[0127] like Figure 7 The diagram shown is a flowchart of the global search module in this embodiment. In this embodiment, the global search module is implemented in the FPGA hardware processor of the hardware acceleration card using HLS language encoding. The global search module consists of a comparator array, an accumulator array, and a selector.

[0128] The Euclidean similarity of the locally most similar vectors input to the global retrieval module is copied and stored horizontally and vertically in the local programmable memory (BRAM) of the hardware accelerator card. Comparators are built between every two elements to form a matrix array of comparators. The comparators in the comparator array concurrently compare the size of the horizontal and vertical elements. If the vertical element is smaller than the horizontal element, the comparator outputs 1; otherwise, it outputs 0. Accumulators are implemented vertically after the comparator array to build an accumulator array. The accumulators in the accumulator array concurrently accumulate the outputs of the comparators horizontally. Finally, a selector is built after the accumulator array. Based on the values ​​in the accumulators, it selects the vector indices corresponding to the top-k values ​​and outputs them to the application in descending order.

[0129] Finally, the software (software runtime module) of this embodiment was built using the standard C / C++ development tool Visual Studio Code on a server equipped with a 28-core Intel(R) Xeon(R) Gold 6330 CPU software processor. The hardware (cluster search module, vector crawling module, vector search module, and global retrieval module) of this embodiment was built using the AMD Xilinx Vivado HLS development tool on a hardware acceleration card equipped with an AMD Xilinx XCUX35-3VSVA1365E FPGA hardware processor.

[0130] The purpose of this embodiment is to address the problem that existing CPU-centric vector approximate nearest neighbor search systems cannot provide low-latency vector approximate nearest neighbor search services for large-scale vector databases due to limited server CPU computing resources and severe vector cross-PCIe transmission overhead, and to provide a vector search acceleration system for large-scale vector databases.

[0131] Experiments were conducted on the vector search acceleration system for large-scale vector databases provided in this embodiment, and the experimental results were analyzed. The system prototype of this embodiment was implemented on a commercial hardware accelerator card and a commercial server. The commercial hardware accelerator card is a mainstream FPGA hardware accelerator card equipped with an AMD Xilinx XCUX35-3VSVA1365E FPGA hardware processor, a PCIe Gen4×8 high-speed serial computer expansion bus interface, local global memory DDR4, and local programmable memory BRAM, among other hardware devices. The commercial server consists of a 28-core Intel(R) Xeon(R) Gold 6330 CPU software processor, 32 x 16GB DDR4 memory modules, and eight 480GiB Intel Optane 900P NVMe SSDs. The hardware accelerator card is connected to the server via the PCIe Gen4×8 high-speed serial computer expansion bus interface. In this embodiment, the software module (software runtime module) is housed within a core of the server's Intel(R) Xeon(R) Gold 6330 CPU software processor, while the hardware modules (cluster search module, vector capture module, vector search module, and global retrieval module) are housed within the AMD Xilinx XCUX35-3VSVA1365E FPGA hardware processor. An Ubuntu 20.04.6LTS operating system with a Linux 5.15.0-72-generic kernel is installed on the server. Furthermore, mainstream CPU-centric vector approximate nearest neighbor search systems such as Faiss and Milvus are also deployed on the server to enable performance comparison and evaluation between the system of this invention and existing systems.

[0132] The test datasets used in the experiments of this embodiment are open-source standard large-scale vector datasets widely used in academia and industry, including SIFT1M, GIST1M, DEEP1B, TEXT1B, and SIFT1B. They are stored on eight 480GiB Intel Optane 900P NVMe SSDs on the server. The parameters of the vector datasets used are shown in Table 1.

[0133] To ensure a fair comparison between the system in this embodiment and existing systems, the relevant experiments in this embodiment used the SIFT1M, GIST1M, DEEP1B, TEXT1B, and SIFT1B datasets to evaluate the end-to-end search latency of the system in this embodiment and existing systems (Faiss and Milvus) when the top-1, top-10, and top-100 accuracies are greater than 85%, 90%, and 95%, respectively. Furthermore, the relevant experiments in this embodiment also used the GIST1M dataset to compare the CPU resource consumption of the system in this embodiment and the Faiss system when performing vector approximate nearest neighbor search, to demonstrate the advantage of the system in saving CPU computing resources.

[0134] Tables 2 and 3 show the experimental verification results of the vector search acceleration system for large-scale vector databases provided in this embodiment. Table 2 shows the comparison results of this embodiment and the existing CPU-centric vector approximate nearest neighbor search system in terms of end-to-end search speed, while Table 3 shows the comparison results of this embodiment and the Faiss system in terms of CPU computing resource consumption.

[0135] Table 1. Vector datasets used in the relevant experiments of this invention.

[0136] Vector data name Vector database vector size Query vector size Vector Dimension SIFT1M 1 trillion 10000 128 GIST1M 1 trillion 1000 960 DEEP1B 1 billion 10000 96 TEXT1B 1 billion 100000 200 SIFT1B 1 billion 10000 128

[0137] Table 2. End-to-end search speed test results of the vector search acceleration system (this invention) for large-scale vector databases.

[0138]

[0139]

[0140] Table 3. CPU resource consumption test results of the vector search acceleration system (this invention) for large-scale vector databases.

[0141]

[0142] Experimental results show that the vector search acceleration system for large-scale vector databases provided in this embodiment achieves an end-to-end (including the step of reading large-scale vector data from storage disk) approximate nearest neighbor search latency between 6.69ms and 2262.90ms when searching vector datasets of different sizes, while ensuring a reasonable approximate nearest neighbor search accuracy. Its end-to-end approximate nearest neighbor search speed for large-scale vector databases is 1.41 times to 171.28 times faster than existing CPU-centric approximate nearest neighbor search systems (Faiss and Milvus). Furthermore, the system in this embodiment consumes only 3.8% of CPU computing resources when performing end-to-end approximate nearest neighbor search, which is 94.5% lower than existing systems. Therefore, the vector search acceleration system for large-scale vector databases provided in this embodiment successfully solves the problem that existing CPU-centric vector approximate nearest neighbor search systems cannot provide low-latency vector approximate nearest neighbor search services for large-scale vector databases due to limited server CPU computing resources and significant vector cross-PCIe transmission overhead. This embodiment can provide a low-latency, high-accuracy vector approximate nearest neighbor search service for large-scale vector databases, and has certain innovative and practical value.

[0143] The end-to-end near nearest neighbor search speed of the system of the present invention is 1.41 times to 171.28 times faster than the existing system, while saving more than 90% of CPU computing resources.

[0144] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A vector search system for a vector database, characterized by, The application relates to a server and a hardware acceleration card. The server comprises a CPU software processor, and the CPU software processor is internally provided with a software runtime module; the hardware acceleration card comprises an FPGA hardware processor, a vector grabbing module, a vector search module and a global retrieval module, and the FPGA hardware processor is internally provided with a cluster search module. The cluster search module and the vector search module receive an input query vector. The cluster search module and the vector grabbing module are connected for data transmission, the cluster search module is used for realizing cluster sampling calculation of the FPGA hardware processor, and the vector grabbing module is used for realizing end-to-end access of the FPGA hardware processor. The vector search module and the global retrieval module are connected for data transmission, the vector search module is used for realizing vector search calculation of the FPGA hardware processor, and the global retrieval module is used for realizing sorting calculation of the FPGA hardware processor. The software runtime module is used for assisting the vector grabbing module to realize end-to-end access. The software runtime module polls a storage request submission queue in a local memory of the CPU software processor in a memory reading and writing control mode. 2.The vector search system for a vector database of claim 1, wherein, When the storage request submission queue is not empty, the software runtime module notifies a doorbell register of an NVMe SSD controller through memory mapping I / O, triggers the NVMe SSD controller to process a storage request, and simultaneously polls the doorbell register of the NVMe SSD controller through memory mapping I / O to acquire an interrupt signal of the NVMe SSD controller. When the interrupt signal is acquired, the software runtime module informs the NVMe SSD controller of a pointer of a storage request completion queue in the local memory of the CPU software processor, and assists the NVMe SSD controller to submit a storage completion request to the storage request completion queue in the local memory of the CPU software processor. The cluster search module receives an input query vector, and performs sampling calculation on a vector cluster in a vector database, and transmits cluster information obtained through sampling to the vector grabbing module; the cluster information comprises indexes of the sampled vector database vectors. 3.The vector search system for a vector database of claim 2, wherein, The vector grabbing module reads vector features from the NVMe SSD controller to a local global memory of the hardware acceleration card according to the vector indexes in the sampled cluster output by the cluster search module; the vector grabbing module pre-fetches vector cluster features and vector database vector indexes carried by the vector cluster to a local programmable memory and a local global memory of the hardware acceleration card respectively when the vector search system is started. The vector search module receives an input query vector, and realizes concurrent and pipelined similarity calculation between the query vector and the sampled vector database vectors in a sampled vector cluster granularity, and outputs 2*top-k local most similar vectors to the global retrieval module after the similarity calculation is completed, so as to acquire top-k vector database vectors most similar to the query vector. 4.The vector search system for a vector database of claim 3, wherein, ​ 5.The vector search system for a vector database of claim 4, wherein, The global search module aggregates all the locally most similar vector database vectors output by the vector search module to the query vector, and implements a full-parallel sorting operation accelerated by the FPGA hardware processor according to the similarity between the vector database vectors and the query vector; according to the sorting result, the global search module selects the top-k globally most similar vector database vectors to the query vector and returns them to the application. 6.The vector search system for a vector database of claim 2, wherein, The software runtime module comprises a doorbell register submodule, a storage request submission processing submodule, a storage request completion processing submodule, a queue address conversion submodule, an acceleration card global memory address conversion submodule, a plurality of storage request submission queues, and a plurality of storage request completion queues; The plurality of storage request submission queues and the plurality of storage request completion queues are used to cache the storage access requests of the NVMe SSD controller generated by the vector grabbing module; the plurality of storage request submission queues and the plurality of storage request completion queues are allocated in the local memory of the CPU software processor; The doorbell register submodule exposes the doorbell register of the NVMe SSD controller on the PCIe interface and maps it to the local memory of the CPU software processor through memory-mapped I / O; The storage request submission processing submodule polls the storage request submission queue and manipulates the doorbell register submodule when a new storage request is generated to inform the NVMe SSD controller to process the storage request; the storage request completion processing submodule polls the doorbell register submodule to obtain the storage request completion interrupt generated by the NVMe SSD controller, and informs the doorbell register submodule of the pointer of the storage request completion queue to assist the NVMe SSD controller to write the generated storage completion request to the storage request completion queue; The queue address conversion submodule converts the memory addresses of the storage request submission queue and the storage request completion queue into pointers that can be accessed by the hardware acceleration card, and the vector grabbing module inserts storage access requests into the storage request queue through the pointer or obtains storage completion requests from the storage request completion queue through the pointer; The acceleration card global memory address conversion submodule converts the vector buffer and cluster buffer allocated in the global memory of the hardware acceleration card into pointers that can be accessed by the CPU software processor, and the pointers are stored in the local global memory of the hardware acceleration card, which are encoded into storage requests by the vector grabbing module to assist the NVMe controller to use direct memory access to directly transmit the vector and cluster features to the local global memory of the hardware acceleration card without passing through the local memory of the CPU software processor. 7.The vector search system for a vector database of claim 3, wherein, The cluster search module comprises a plurality of processing units; Each processing unit independently calculates the Euclidean similarity between the query vector and a cluster center; Within the processing unit, the Euclidean similarity computation is unfolded over the vector dimension; the subtraction and squaring operations between the query vector elements and the cluster center feature elements in the same dimension are implemented as a plurality of first computing units independent of each other, the plurality of first computing units concurrently process each vector element and output the result to a summation unit; the summation unit outputs the summation result to a square root unit to complete the Euclidean similarity computation; The similarity metric output by the processing unit is used by a priority queue for cluster sampling, the priority queue only retains the indexes of the top P clusters closest to the query vector, the parameter P is set by the application. 8.The vector search system for a vector database of claim 3, wherein, The vector grabbing module comprises a storage request processing submodule and a data processing submodule; The storage request processing submodule constructs a read request for the cluster and the cluster feature when the vector search system is initialized, the read request is used to transmit the cluster and the cluster feature from the NVMe SSD controller to the local global memory of the hardware acceleration card directly; The data processing submodule transmits the cluster feature to the local programmable memory of the hardware acceleration card and transmits the vector index carried by the cluster to the local global memory of the hardware acceleration card; The storage request processing submodule and the data processing submodule complete the following vector feature reading process: The data processing submodule uses the cluster index output by the cluster search module to retrieve the vector cluster in the local global memory of the hardware acceleration card to obtain the sampled vector index; The storage request processing submodule calculates the address of the vector feature in the NVMe SSD controller according to the vector index, encodes the address into the storage request, and encodes the vector buffer pointer stored in the local global memory of the hardware acceleration card into the storage request to indicate the destination address of the vector feature reading operation; The storage request processing submodule inserts the storage request into the storage request submission queue in the local memory of the CPU software processor through the pointer of the storage request submission queue to trigger the reading of the vector feature, and at the same time, the storage request processing submodule polls the storage request completion queue in the local memory of the CPU software processor through the pointer of the storage request completion queue to process the storage completion request; The data processing submodule transmits the vector feature from the vector buffer in the local global memory of the hardware acceleration card to the vector feature buffer in the local programmable memory of the hardware acceleration card for the vector search module to implement the similarity computation. 9.The vector search system for a vector database of claim 4, wherein, The vector search module comprises a plurality of second computing units; The processing units in the vector grabbing module and the cluster search module are expanded to construct a plurality of second computing units, each second computing unit implements an end-to-end similarity computation workflow for the vectors in a sampled cluster; The similarity calculation workflow includes three steps: the vector grabbing module reads the vector features from the NVMe SSD controller to the vector feature buffer in the local programmable memory of the hardware acceleration card; multiple processing units calculate the Euclidean similarity between the query vector and the vectors in the vector feature buffer in parallel; the indexes and Euclidean similarities of the top-k most similar vectors to the query vector are output to the global retrieval module; Inside the second computing unit, the three steps of the similarity calculation workflow are constructed into a pipeline; the sampled clusters are evenly distributed to multiple second computing units, and each second computing unit pipeline calculates the Euclidean similarity between the query vector and the vectors in the sampled cluster assigned to it. 10.The vector search system for a vector database of claim 5, wherein, The global retrieval module includes a comparator array, an accumulator array, and a selector; The Euclidean similarities of the locally most similar vectors input into the global retrieval module are copied once and stored in the local programmable memory of the hardware acceleration card in the horizontal and vertical directions respectively; The horizontal elements and the vertical elements are two elements, and the comparator is constructed between each two elements to form the comparator array in the form of a matrix; multiple comparators in the comparator array concurrently compare the sizes of the horizontal elements and the vertical elements, and if the vertical element is smaller than the horizontal element, the comparator outputs 1, otherwise, the comparator outputs 0; The accumulator is implemented vertically after the comparator array to construct the accumulator array; multiple accumulators in the accumulator array concurrently accumulate the output results of the comparators in the horizontal direction; The selector is constructed after the accumulator array, and according to the values in the accumulator, the vector indexes corresponding to the top-k values are selected and output in descending order to the application.

Citation Information

Patent Citations

  • Heterogeneous system for large-scale vector approximate nearest neighbor search based on GPU-NVMe architecture

    CN119669525A