Vector query method and device

Through the coordinated work of the computing core and multiple target read and write cores, the problem of low vector data reading efficiency is solved, and the efficiency of high concurrent query is improved.

CN120492693APending Publication Date: 2025-08-15LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510472517.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the reading efficiency of vector data is limited by the processing capability of a single read and write core in the micro file system, which leads to a bottleneck when concurrent filtering query requests, affecting query efficiency.

Method used

The calculation core determines the set of data blocks that match the query data, and allocates multiple target read and write cores based on the preset strategy, respectively process the read request, and realizes parallel reading of vector data.

Benefits of technology

It improves the read and write throughput of vector queries, reduces the overhead of the file system layer, saves the reading time of vector data, and improves the reading efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492693A_ABST
    Figure CN120492693A_ABST
Patent Text Reader

Abstract

The invention discloses a vector query method and device which are applied to a processor, the processor at least comprises a calculation core and a read-write core, the method comprises the steps that the calculation core determines a data block set matched with query data, the data block set comprises N data blocks, and N is an integer larger than or equal to 1; the calculation core determines N target read-write cores in one-to-one correspondence with the N data blocks based on a preset strategy, and sends N read requests to the corresponding N target read-write cores respectively; and the Nth target read-write core reads the Nth vector data stored in the Nth disk block number, and returns the Nth vector data to the calculation core.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing technology, and relate to, but are not limited to, a vector query method and device. Background Art

[0002] Query scenarios, such as product retrieval and Retrieval-Augmented Generation (RAG) in e-commerce, involve vector queries, which are becoming increasingly important. In related technologies, a single read / write core (IO core) provides file read and write services for dataset shards (corresponding to micro-file systems). As concurrent filtering query requests increase, read requests become a bottleneck for improving query efficiency. Improving the read efficiency of vector data has become a pressing technical challenge. Summary of the Invention

[0003] In view of this, embodiments of the present application provide a vector query method and apparatus.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] In a first aspect, an embodiment of the present application provides a vector query method, which is applied to a processor, wherein the processor includes at least a computing core and a read / write core, including:

[0006] The computing core determines a set of data blocks that match the query data, where the set of data blocks includes N data blocks, where N is an integer greater than or equal to 1;

[0007] The computing core determines N target read-write cores corresponding to the N data blocks based on a preset strategy, and sends the N read requests to the corresponding N target read-write cores respectively;

[0008] The Nth target read-write core reads the Nth vector data stored in the Nth disk block number and returns the Nth vector data to the computing core.

[0009] In a second aspect, an embodiment of the present application provides a vector query device, including:

[0010] A first determining module is configured to determine a data block set matching the query data using a computing core, wherein the data block set includes N data blocks, where N is an integer greater than or equal to 1;

[0011] A second determination module is configured to use the computing core to determine N target read / write cores corresponding to the N data blocks based on a preset strategy, and send the N read requests to the corresponding N target read / write cores respectively;

[0012] The reading module is used to use the Nth target read-write core to read the Nth vector data stored in the Nth disk block number and return the Nth vector data to the computing core.

[0013] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, it uses a computing core to determine a set of data blocks that match the query data, wherein the data block set includes N data blocks, and N is an integer greater than or equal to 1; the computing core determines N target read-write cores corresponding one-to-one to the N data blocks based on a preset strategy, and sends N read requests to the corresponding N target read-write cores respectively; the Nth target read-write core reads the Nth vector data stored in the Nth disk block number, and returns the Nth vector data to the computing core.

[0014] In a fourth aspect, an embodiment of the present application provides a storage medium storing executable instructions, which, when executed by a processor, enables a computing core to determine a data block set that matches the query data, wherein the data block set includes N data blocks, and N is an integer greater than or equal to 1; the computing core determines N target read-write cores corresponding one-to-one to the N data blocks based on a preset strategy, and sends N read requests to the corresponding N target read-write cores respectively; the Nth target read-write core reads the Nth vector data stored in the Nth disk block number, and returns the Nth vector data to the computing core.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, a computing core determines a data block set that matches the query data, wherein the data block set includes N data blocks, and N is an integer greater than or equal to 1; the computing core determines N target read-write cores corresponding one-to-one to the N data blocks based on a preset strategy, and sends N read requests to the corresponding N target read-write cores respectively; the Nth target read-write core reads the Nth vector data stored in the Nth disk block number, and returns the Nth vector data to the computing core. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of an implementation flow of a vector query method provided in an embodiment of the present application;

[0017] Figure 2 A schematic diagram of an implementation flow for determining a data block set provided in an embodiment of the present application;

[0018] Figure 3 A schematic diagram of an implementation flow of constructing a read request provided in an embodiment of the present application;

[0019] Figure 4A A schematic diagram of a process for index construction provided in an embodiment of the present application;

[0020] Figure 4B A schematic diagram of the implementation process of loading an index is provided for an embodiment of the present application;

[0021] Figure 4C A schematic diagram of a filtering query main process provided in an embodiment of the present application;

[0022] Figure 4D A schematic diagram of an intra-cluster filtering query sub-process provided in an embodiment of the present application;

[0023] Figure 5 A schematic diagram of the structure of a vector query device provided in an embodiment of the present application;

[0024] Figure 6 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] To make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the specific technical solutions of the embodiments of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0026] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0027] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0029] The embodiment of the present application provides a vector query method, which is applied to a processor, wherein the processor includes at least a computing core and a read-write core, such as Figure 1 As shown, the method includes:

[0030] Step S110: The computing core determines a set of data blocks that match the query data, wherein the set of data blocks includes N data blocks, where N is an integer greater than or equal to 1;

[0031] Here, the CPU's compute core, also known as a job core, refers to the CPU's general-purpose processing core, used to execute compute-intensive tasks such as applications, algorithms, and logical operations. The read / write core, also known as the IO core, is a hardware module designed to efficiently handle input / output (I / O) tasks. In practice, the compute core runs the operating system and applications. Frequent processing of I / O interrupts (such as disk reads and writes) incurs context switching overhead. Dedicated I / O cores can reduce CPU utilization. Data reading tasks are simple but time-consuming, and dedicated read / write cores are more efficient than general-purpose cores.

[0032] Query data refers to the input data used by users or systems when performing query operations. Query operations are typically used to retrieve specific information from a database, data warehouse, or other data store. Query data can include structured scalar fields such as numeric (age, price), categorical (gender, city), and time (creation time). It can also include unstructured data (text, images, audio) embedded as high-dimensional vectors through deep learning models (such as BERT and CLIP), i.e., vector fields.

[0033] In databases and storage systems, data blocks are the basic unit of disk storage and the core carrier for query data matching. A data block is a group of data stored contiguously on disk, meaning that a data block can store entity data that matches the query data.

[0034] During implementation, the computing core may be used to determine N data blocks matching the query data in the database based on the query data, that is, the entity data stored in the N data blocks is found to match the query data.

[0035] Step S120: The computing core determines N target read / write cores corresponding to the N data blocks based on a preset strategy, and sends N read requests to the corresponding N target read / write cores respectively.

[0036] Here, the design of the preset strategy can be based on load balancing, or based on the mapping relationship between the read-write core and the data block, or can be determined based on the location information stored in the database.

[0037] During the implementation process, the computing core can allocate a specific computing core to each data block based on a preset strategy, which is responsible for the read and write operations of the data block.

[0038] A read request is constructed by the compute core based on different data blocks. This read request includes the disk location information (disk block number) of the data block to be read. During implementation, after the compute core determines the Nth data block to be read by the Nth target read / write core, it sends the read request constructed for this Nth data block to the Nth target read / write core.

[0039] Step S130: The Nth target read / write core reads the Nth vector data stored in the Nth disk block number, and returns the Nth vector data to the computing core.

[0040] During implementation, after receiving a read request from the computing core including the disk block number corresponding to the data block to be read, the Nth target read-write core reads the Nth vector data stored in the Nth disk block number based on the read request.

[0041] In this way, N target read-write cores can obtain N sets of vector data and return the N sets of vector data to the computing core, and the computing core can obtain N sets of vector data that match the query data.

[0042] In an embodiment of the present application, the computing core first determines a set of data blocks that match the query data, and then determines N target read-write cores corresponding one-to-one to the N data blocks based on a preset strategy; finally, the Nth target read-write core reads the Nth vector data from the Nth disk block number based on the corresponding read request. In this way, the computing core can be used to determine that multiple target read-write cores serve a data set shard at the same time, supporting high-concurrency filtering queries, breaking the limitation that a micro file system can only be served by one target read-write core, improving read and write throughput, and effectively improving vector query efficiency. Multiple target read-write cores issue requests for disk block numbers, which can effectively eliminate file system layer overhead, that is, directly read vector data from the disk, effectively saving vector data reading time and improving reading efficiency.

[0043] In some embodiments, in step S120 above, "the computing core determines, based on a preset strategy, N target read / write cores corresponding one-to-one to the N data blocks" can be implemented by one of the following steps:

[0044] Step 121: The computing core determines the N target read-write cores based on the idle computing power of the read-write cores.

[0045] or,

[0046] During implementation, hardware performance counters can be used to collect read / write core computing power indicators, such as utilization, which is the occupancy rate of the current read / write core; queue depth, which is the number of pending I / O requests; and bandwidth saturation, which is the actual / theoretical bandwidth of the storage interface of the read / write core.

[0047] In some embodiments, based on a greedy algorithm, the read and write cores with the lowest utilization can be preferentially selected to minimize the overall latency.

[0048] In some embodiments, based on load balancing, the load variance of all read and write cores may be determined, and a core set that minimizes the variance may be selected.

[0049] In some embodiments, affinity constraints can also be used to prioritize read / write cores that are physically close to the computing cores, thereby reducing communication delays.

[0050] Step 122: The computing core determines the N target read-write cores based on the mapping relationship between the read-write cores and the disk block numbers.

[0051] Here, a mapping relationship between the disk block number and the read-write core may be preset, that is, the Nth target read-write core may be determined to correspond to the Nth disk block number.

[0052] During implementation, the Nth target read / write core may be determined based on the mapping relationship and the Nth disk block number to be read.

[0053] In an embodiment of the present application, the computing core determines N target read / write cores based on the idle computing power of the read / write cores. Thus, by evaluating the idle computing power, the computing core can efficiently utilize the idle computing power of the read / write cores and optimize vector data reading performance. The computing core determines the N target read / write cores based on the mapping relationship between the read / write cores and the disk block numbers. Thus, the computing core can efficiently determine the target read / write cores and reduce data access latency.

[0054] In some embodiments, the query data includes a query scalar and a query vector. In the above step S110, "the computing core determines a set of data blocks matching the query data" is as follows: Figure 2 As shown, this can be achieved by following the steps below:

[0055] Step S210: The computing core determines the target number of target clusters to be acquired based on a filtering rate, wherein the filtering rate is determined based on the query scalar;

[0056] The records or entities stored in the vector database contain both structured scalar fields and vector fields generated by embedding unstructured data. When performing a vector-scalar hybrid query, the user provides a scalar filter condition (query scalar) and a query vector. The vector database system then finds the top k records closest to the query vector from all records that meet the filter condition. This vector-scalar hybrid query is also known as a metadata filtering query, or simply a filtering query.

[0057] The filter ratio is an indicator that measures the efficiency of the query condition in filtering the data set. The ratio of data that meets the query condition to the total data is expressed as follows:

[0058] Filter Ratio = total data volume / data volume that meets the conditions (1);

[0059] When performing a filtering query, the scalar domain filtering can be performed on all entities in the database using calculations, and then the filtering rate can be obtained based on the above formula (1).

[0060] During implementation, a computing core may be used to perform scalar domain filtering based on a scalar index or raw scalar data to generate a filtered entity bitmap (filtered_entities_bitmap).

[0061] While generating the filtered entity bitmap, the filtering rate s of this query is calculated, that is, the proportion of entities that meet the conditions.

[0062] The optimal head batch (head_batch), i.e. the target number of target clusters to be acquired, is calculated using the following formula (2):

[0063] head_batch = (topk / (a*pn*s)) (2);

[0064] Where Pn is the average number of vectors / entities contained in each cluster, a is the iteration coefficient, which can be set to 3 for example, s is the filtering rate of this query, and topk represents the size of the candidate set.

[0065] Step S220: determining a first cluster center that meets the target number from a vector index file based on the query vector, and determining the target cluster using the first cluster center, wherein the vector index file is stored in a memory;

[0066] Here, you can pre-build a cluster tree for all vectors in the database and select appropriate cluster center nodes from the tree according to specified rules. For example, you can select the vector with the smallest average distance to other members in each cluster as the center node, or select the vector at a specific location in the cluster (such as the vector closest to the cluster centroid). Save the cluster center node vector and its identity document (ID) in a vector index file (header index file). Build an in-memory index (header index file) for all cluster center nodes to quickly find the appropriate cluster center for the query variable. Save this index in the in-memory index.

[0067] During implementation, the computing kernel can be used to find the best head_batch first cluster centers from the vector index file according to the approximate nearest neighbor rule, and sort these first cluster center vectors from small to large according to the distance.

[0068] The first cluster center found above can be filtered again based on the filtering conditions to obtain the target cluster that matches the query vector. For example, the job core can be used to process the head_batch cluster center vectors in a loop: if the entity of this cluster center vector satisfies the filtering expression expr, then this center vector is added to the result set (resultset), where the center vector is stored in memory, and the center vectors that meet the filtering expression can be added to the result set. At the same time, the computing core can be used to process the head_batch clusters in a loop, obtain the value range of each scalar domain of the current cluster from the cluster scalar range file, and perform scalar filtering based on the range. If the filtering conditions are not met, skip this cluster and take the next cluster to continue processing. If the filtering conditions are met, put it in the inverted list to obtain the target cluster. The above two screening steps are not limited to the order and can also be executed simultaneously.

[0069] Step S230: Determine the data block set based on the target cluster.

[0070] During implementation, the vectors in the target cluster may be filtered based on the filtering entity bitmap to obtain vectors that meet the conditions, and then a data block set corresponding to the vectors that meet the conditions is determined.

[0071] In this embodiment, the computation core first determines the target number of target clusters to be acquired based on the filter rate. Then, based on the query vector, the center of the first cluster that meets the target number is determined from the vector index file, using this first cluster center to determine the target cluster. Finally, the data block set is determined based on the target cluster. This dynamically determines the number of clusters acquired in each batch based on the filter rate, and further concurrently filters queries within each cluster, resulting in fewer iterations and lower query latency. Because the vector index file is stored in memory, determining the center of the first cluster that meets the target number in the vector index file improves query efficiency.

[0072] In some embodiments, the above step S230 “determining the data block set based on the target cluster” may be implemented by the following steps:

[0073] Step 231: Identify a target entity matching the query scalar based on the vector list corresponding to the target cluster;

[0074] During implementation, the cluster's vector list (idlist) can be obtained from the cluster mapping (postinglist_vectorids_map). Then, based on the filtered entity bitmap (filtered_entities_bitmap), each entity in the vector list is determined to meet the filtering criteria, generating a cluster bitmap (plbitmap). The cluster bitmap is a bitmap used to quickly determine whether an entity has passed the filtering criteria. If an entity has been previously determined to meet the criteria, its corresponding bit in the bitmap is set to 1; otherwise, it is set to 0.

[0075] Step 232: Based on the identifier of the target entity, merge the data that meets the threshold distance into a data block to obtain the data block set.

[0076] During implementation, blocks with bits set to 1 can be merged based on the intra-cluster bitmap. If the distance between two adjacent blocks is not far, whether to merge them can be determined based on a threshold, where the threshold can be set in relation to the vector dimension (dim) and the vector type (type). For example, the threshold can be set using the following formula (3):

[0077] threshold = 4K / dim*sizeof(type)(3);

[0078] Where dim is the vector dimension and sizeof(type) is the size of the vector type. This threshold can be used to determine whether to merge two interval blocks that are not far apart into one interval block (data block).

[0079] In this embodiment, the target entity that matches the query scalar is first identified based on the vector list corresponding to the target cluster. Then, based on the identification of the target entity, the data that meets the threshold distance is merged into a data block to obtain a data block set. In this way, through matching with the query scalar and merging the data, a data block set that meets the scalar matching requirements can be obtained.

[0080] In some embodiments, the above step S220 of “determining the target cluster using the first cluster center” can be implemented by the following steps:

[0081] Step 221: Obtain a cluster scalar range file for each of the first cluster centers, wherein the cluster scalar range file is used to record the scalar value range corresponding to the first cluster center;

[0082] Here, the cluster scalar range file (postinglist_field_range_map) is pre-stored. In real-time, for all cluster vectors in the database, the following steps are performed: The values of each scalar field of the cluster vector entity are obtained; the maximum and minimum values of each scalar field within the cluster are recorded to obtain the cluster scalar range. The cluster scalar range file is loaded into memory, and based on the cluster scalar range, the cluster scalar range file, i.e., the vector inverted index (postinglist_field_range_map), is constructed.

[0083] During implementation, a pre-stored cluster scalar range file that records the scalar value range corresponding to the first cluster center may be obtained.

[0084] Step 222: Determine the target cluster based on the scalar value range corresponding to the first cluster center.

[0085] During implementation, the value ranges for each scalar field in the current cluster are obtained from the cluster scalar range file and scalar filtering is performed based on the ranges. If the filtering conditions are not met, the cluster is skipped and the next cluster is selected for processing. If the filtering conditions are met, the cluster is stored in the filtered posting list (filtered_postinglist).

[0086] In an embodiment of the present application, the cluster scalar range file of each first cluster center is first obtained; then the target cluster is determined based on the scalar value range corresponding to the first cluster center. In this way, since the cluster scalar range file (in the vector inverted index) records the value range (maximum value, minimum value) of the scalar field (domain) of the entity to which the cluster belongs, the clusters that do not meet the filter conditions are screened out according to the value range before reading the inverted list, thereby eliminating unnecessary disk reads. For the vector inverted index, the vectors in the cluster are rearranged based on the scalar field (domain) data information. After rearrangement, the vectors of the entities that meet the filter conditions can usually be clustered together. When reading the inverted list, only the vectors of the entities that meet the filter conditions are read from the disk, further reducing the amount of disk data read.

[0087] In some embodiments, the present application provides a method for constructing a read request, such as Figure 3 As shown, this can be achieved by following the steps below:

[0088] Step S310: The computing core obtains a pre-stored file disk block mapping relationship, wherein the file disk block mapping relationship represents a mapping relationship between file index information and disk block numbers;

[0089] Here, the pre-stored file disk block mapping relationship (file_offset_plog_map) is constructed based on the mapping information between the offset (offset) of the disk index file and the underlying block (plog).

[0090] Step S320: Determine the Nth disk block number corresponding to the Nth data block based on the Nth file index information corresponding to the Nth data block using the file-disk block mapping relationship;

[0091] During implementation, based on the file-disk block mapping relationship, the offset of the index file corresponding to the Nth data block can be converted into the underlying disk block number (plog_id), that is, the Nth disk block number corresponding to the Nth data block can be determined.

[0092] Step S330: Construct the Nth read request using the Nth disk block number.

[0093] During implementation, the constructed Nth read request includes the corresponding Nth disk block number, so that the vector data stored in the Nth disk block number can be read based on the Nth read request.

[0094] In this embodiment of the present application, a pre-stored file-disk block mapping relationship is first obtained; then, based on the Nth file index information corresponding to the Nth data block, the file-disk block mapping relationship is used to determine the Nth disk block number corresponding to the Nth data block; finally, the Nth disk block number is used to construct the Nth read request. Thus, the resulting Nth read request includes the corresponding Nth disk block number, and the vector data stored at the Nth disk block number can be read based on the Nth read request, eliminating file system layer overhead and effectively improving vector data reading efficiency.

[0095] In some embodiments, the above step S120 in which "the computing core sends N read requests to the corresponding N target read / write cores" can be implemented by the following steps:

[0096] Step 123: Obtain N target message objects from the message cache of the computing core;

[0097] Here, the computing core can maintain a message cache to store pending messages. The message cache can be a queue, hash table, or other data structure. A message object in the cache refers to a message entity stored in a cache system (such as a memory cache or message queue) and represented by a specific data structure. It is an entity used in programs to encapsulate data and metadata and can be represented in a structure, class, or JSON format.

[0098] The message object is used to generate a read request for reading the database. During implementation, since N read requests are to be generated, N target message objects can be obtained from the message cache.

[0099] Step 124: Use N target message objects to send the N read requests to the corresponding N target read-write cores respectively.

[0100] During the implementation process, N message objects can be obtained from the message cache of the computing core, and a read request can be sent to the IO core (selected_iocore) providing the service, that is, the corresponding N target read and write cores.

[0101] In this embodiment of the present application, N target message objects are first obtained from the message cache of the computing core; then, N read requests are sent to the corresponding N target read / write cores using the N target message objects. In this way, the read request can be sent to the N target read / write cores providing services using the N target message objects, so that the target read / write cores can execute the request to read the data block.

[0102] In some embodiments, the present application also provides a method for retrieving a message object, which can be implemented by the following steps:

[0103] Step 125: Determine that there are no idle message objects in the message cache of the computing core;

[0104] During the implementation process, if there is no available message object in the idle message cache of the computing core, the message object cannot be directly obtained, and step 126 is executed.

[0105] Step 126: Recover the message object that completes the read request from the preset message pump;

[0106] During implementation, several message objects that complete the read request can be recovered from the lightweight Storage Performance Development Kit (SPDK) message pump.

[0107] Step 127: Store the message object that completes the read request into the message cache of the computing core.

[0108] During the implementation process, the message object that completes the read request can be stored in the message cache of the computing core to achieve recycling and reuse of the message object.

[0109] In the embodiment of the present application, it is first determined that there are no idle message objects in the message cache of the computing core; then the message object that completes the read request is recovered from the preset message pump; finally, the message object that completes the read request is stored in the message cache of the computing core. In this way, a lightweight SPDK message engine is built for the computing core (which is only responsible for recovering messages and does not need to perform other processing). After the read-write core processes the read request of the computing core, it can directly return the SPDK message (which contains the identifier of the source computing core) to the message engine of the source computing core. Messages can be transmitted in both directions, providing an effective message object recovery and management method, thereby improving the performance of filtering queries.

[0110] In some embodiments, the vector query method further includes the following steps:

[0111] Step S140: The computing core sorts the vector data and the first cluster center vector data based on the distance information between the Nth vector data and the M first cluster center vector data and the query data to obtain a sorting result, wherein the M first cluster center vector data are obtained by filtering the first cluster centers based on the filter expression, where M is an integer greater than or equal to 1.

[0112] Here, the computation kernel can be used to loop through step S220 to obtain the target number (head_batch) of first cluster center vectors. If the entity of this cluster center vector satisfies the filter expression expr, this center vector is added to the result set (resultset). In this way, the result set includes the N vector data obtained from disk and the M first cluster center vector data obtained from memory.

[0113] During the implementation process, the distance between each vector data and the query data, as well as the distance between each first cluster center vector data and the query data are calculated respectively, and the vector data in the result set (including N vector data and M first cluster center vector data) are sorted based on the obtained distance information.

[0114] Step S150: Based on the sorting result, sequentially output the entities corresponding to the first K vector data as query results, where K is a preset query entity threshold and K is less than or equal to N.

[0115] During implementation, if the number of elements in the result set does not meet the specified threshold (related to recall, which can be set to topk*24), a new search will be performed to find cluster centers that meet the target number. If the number of elements in the result set meets the specified threshold, the top k (topk) vectors with the smallest distances will be found from the result set and these top k results will be returned.

[0116] In this embodiment of the present application, the computing core sorts the vector data and the first cluster center vector data based on the distance information between the Nth vector data and the M first cluster center vector data and the query data, obtaining a sorted result. Based on the sorted result, the computing core then sequentially outputs the entities corresponding to the first K vector data as the query result. In this way, a query result that meets the output requirements can be obtained based on the vector data corresponding to the cluster center and other vector data not belonging to the cluster center. Furthermore, since the vector data corresponding to the cluster center is stored in memory, query efficiency can be effectively improved.

[0117] Records or entities stored in vector databases typically contain both structured scalar fields and vector fields generated by embedding unstructured data. When performing a mixed vector-scalar query, the user provides scalar filtering conditions and a query vector. The vector database system then finds multiple records closest to the query vector from all records that meet the filtering conditions. This mixed vector-scalar query is also known as a metadata filtering query, or simply a filtering query. The filtering query method has the following problems:

[0118] A. When the filtering rate is low, most of the inverted list data read from the disk is invalid data (entities do not meet the filtering conditions);

[0119] B. To achieve a specified recall rate, different filtering conditions (different filtering rates) require different numbers of iterations, resulting in large fluctuations in query latency.

[0120] C. The query process is executed sequentially (first find a batch of clusters, then read the cluster data, and then perform vector calculations). This has low concurrency and causes poor query latency.

[0121] D. Each dataset shard (corresponding to a micro file system) is provided with file read and write services by a read / write core. When the number of concurrent filtering query requests increases, I / O reads can easily become a bottleneck.

[0122] E. When the computing core calls the read-write core service, it needs to construct and send a message. Since the message is only transmitted in one direction, it cannot be effectively recycled and managed, which affects the filtering query performance. To solve the above problems, the embodiment of the present application proposes an efficient filtering query method that supports complex expressions.

[0123] Figure 4A A schematic diagram of an implementation flow of index construction provided in an embodiment of the present application is shown as follows: Figure 4A As shown, this can be achieved by following the steps below:

[0124] Step S401: construct a cluster tree and select a cluster center node;

[0125] During implementation, a cluster tree is constructed for all vectors in the database, and appropriate cluster centers are selected from the tree according to specified rules. For example, the center node can be the vector with the smallest average distance to other members in each cluster, or the center node can be selected from a vector at a specific position in the cluster (such as the vector closest to the cluster centroid).

[0126] Step S402: Build a memory index for the cluster center node;

[0127] During implementation, the cluster center node vectors and their identifiers are saved in a vector index file (header index file). A memory index (header index file) is constructed for all cluster center nodes to quickly find the appropriate cluster center for the query variable. This index is saved in the memory index.

[0128] Step S403: adding other vector entities to the adapted cluster;

[0129] In this implementation, all non-cluster center vectors (xq) are indexed using the header to find the cluster centers closest to the vectors. Some cluster centers are then pruned using the relative neighborhood graph (RNG) rule. The vectors are then added to the clusters represented by the remaining cluster centers. Using the RNG rule to prune cluster centers is a geometrically-based approach that selects more representative cluster centers while removing redundant centers that may be caused by noise or uneven data distribution.

[0130] Step S404: obtaining the priority order of each scalar domain;

[0131] If a dataset contains multiple scalar fields, check the dataset metadata to see if there is a scalar field sort hint (for example, the scalar field's sort order). If there is no sort hint, calculate the cardinality of each scalar field (based on sampled data or directly on the original data) and sort the scalar fields in ascending order of cardinality.

[0132] Step S405: sort the vector entities in the cluster according to the size of the scalar domain value;

[0133] During implementation, all vectors within all clusters are sorted based on the scalar field order obtained in step S404. First, the current scalar field value of each vector entity within the cluster is obtained. Then, the vectors are sorted by the magnitude of the current scalar field value, and vectors with the same current scalar field value are recorded. For vectors with the same current scalar value, the value of the next scalar field is obtained, and the order is repeated, recording vectors with the same value. This continues until no vectors with the same scalar field value remain or all scalar fields provided in step S404 have been sorted.

[0134] Step S406: Obtain the value range of each scalar field of the vector entity in the cluster;

[0135] During the implementation process, for all clusters' intra-cluster vectors: obtain the values of each scalar field of the intra-cluster vector entity; and record the maximum and minimum values of each scalar field in the cluster.

[0136] Step S407: output the memory index file;

[0137] During implementation, the mapping relationship between the cluster center vector identifier and the maximum or minimum value of each scalar domain of the entities in the cluster is saved in the cluster scalar range file or the header index file (memory index file).

[0138] Step S408: Output the entity scalar value range within the cluster and the vector identification file within the cluster;

[0139] During the implementation process, the mapping relationship between the cluster center vector identifier and the cluster inner vector identifier is saved in the cluster mapping file (this step is optional and can speed up the loading process).

[0140] Step S409: Output the disk index file.

[0141] During the implementation process, the cluster center vector identifier, cluster intra-vector identifier, and cluster intra-vector data are saved to the disk index file.

[0142] Figure 4B A schematic diagram of the implementation process of loading an index is provided for the embodiment of the present application, such as Figure 4B As shown, this can be achieved by following the steps below:

[0143] Step S411: Load the header index file into the memory;

[0144] Here, the header index file is an index file of cluster center node vectors and their identities (Identity Document, ID). Header index files can be constructed for all cluster center nodes so as to quickly find a suitable cluster center for the query variable.

[0145] Step S412: Load the cluster scalar range into the memory;

[0146] Here, for all clusters' intra-cluster vectors: obtain the values of each scalar domain of the intra-cluster vector entity; record the maximum and minimum values of each scalar domain in the cluster to obtain the cluster scalar range.

[0147] During implementation, the cluster scalar range file can be loaded into memory, and a vector inverted index (postinglist_field_range_map) can be constructed based on the cluster scalar range.

[0148] Step S413: Load the cluster mapping relationship into the memory;

[0149] Here, the mapping relationship between the cluster center vector identifier and the maximum or minimum value of each scalar domain of the entities in the cluster is saved in the cluster scalar range file.

[0150] During implementation, the cluster mapping relationship may be loaded into memory, and postinglist_vectorids_map may be constructed based on the cluster mapping relationship.

[0151] Step S414: Construct a mapping relationship between the disk index file and the underlying blocks.

[0152] During the implementation process, the mapping information from the offset of the disk index file to the underlying block (plog) is obtained, and file_offset_plog_map is constructed.

[0153] Figure 4C A schematic diagram of a filtering query main process provided by an embodiment of the present application, such as Figure 4C As shown, this can be achieved by following the steps below:

[0154] Step S421: Filter all entities based on the scalar domain and calculate a reasonable head_batch according to the filtering rate;

[0155] During implementation, the job core may be used to perform scalar domain filtering based on scalar indexes or raw scalar data to generate a filtered entity bitmap (filtered_entities_bitmap).

[0156] While generating the filtered entity bitmap, the filtering rate s of this query is calculated, that is, the proportion of entities that meet the conditions.

[0157] The optimal head batch (head_batch) is calculated using the following formula (2):

[0158] head_batch = (topk / (a*pn*s)) (2);

[0159] Where Pn is the average number of vectors / entities contained in each cluster, a is the iteration coefficient, which can be set to 3 for example, s is the filtering rate of this query, and topk represents the size of the candidate set.

[0160] Step S422, find the head_batch clusters with the closest distance;

[0161] During implementation, the job core can be used to find the best head_batch clusters from the memory index according to the approximate nearest neighbor rule, and sort the cluster center vectors from small to large by distance.

[0162] Step S423: adding the cluster center vectors that meet the conditions to the result set;

[0163] During implementation, the job core loop can be used to process head_batch cluster center vectors: if the entity of this cluster center vector satisfies the filter expression expr, then this center vector is added to the result set (resultset).

[0164] Step S424: Perform range filtering (RangeCheck) on the head_batch clusters;

[0165] During implementation, the job core is used to loop over head_batch clusters:

[0166] Obtain the value ranges for each scalar field in the current cluster from the cluster scalar range file (postinglist_field_range_map) and perform scalar filtering based on the ranges. If the filtering conditions are not met, skip the cluster and continue processing with the next one. If the filtering conditions are met, store the results in the filtered posting list (filtered_postinglist).

[0167] In some embodiments, steps S423 and S424 are not limited to the order in which they are performed. Step S423 can be performed first, followed by step S424; step S424 can be performed first, followed by step S423; or steps S423 and S424 can be performed simultaneously. Step S423 is performed to obtain a central vector that satisfies the filtering criteria and is stored in memory. Step S424 is performed to obtain a posting list, which is then further processed in step S425.

[0168] Step S425: The cluster construction coroutine that meets the RangeCheck requirement triggers the execution of the intra-cluster filtering query sub-process;

[0169] On the job core, a task / thread is constructed for each cluster in the filtered inverted list, i.e., the filtering query sub-process. Figure 4D A schematic diagram of an intra-cluster filtering query sub-process provided in an embodiment of the present application is shown as follows: Figure 4D As shown, this can be achieved by following the steps below:

[0170] Step S4251: Get a list of vectors within the cluster and find vectors that meet the filtering conditions;

[0171] During implementation, the job core can first retrieve the cluster's vector list (idlist) from the cluster mapping (postinglist_vectorids_map). The job core then uses the filtered_entities_bitmap to determine whether each entity in the vector list meets the filtering criteria, generating a cluster bitmap (plbitmap). The cluster bitmap is a bitmap used to quickly determine whether an entity has passed the filtering criteria. If an entity has been previously determined to meet the criteria, its corresponding bit in the bitmap is set to 1; otherwise, it is set to 0.

[0172] Step S4252: Generate a file read request for the vector that meets the filtering conditions, and merge adjacent read requests;

[0173] During implementation, the job core can be used to merge blocks with the bitmap within the cluster. If the distance between two adjacent blocks is not far, the decision of whether to merge can be made based on a threshold. The threshold can be set in relation to the vector dimension (dim) and the vector type (type). For example, the threshold can be set using the following formula (3):

[0174] threshold = 4K / dim*sizeof(type)(3);

[0175] Where dim is the vector dimension and sizeof(type) is the size of the vector type. This threshold can be used to determine whether to merge two interval blocks that are not far apart into one interval block.

[0176] Step S4253: Convert the file read request into an underlying block read request and select a suitable IO core;

[0177] During implementation, the job core constructs a disk read request for each interval block. Based on the file_offset_plog_map, the offset of the interval block's corresponding index file is converted to the underlying disk block number (plog_id). A service IO core (selected_iocore) is selected according to a preset load balancing strategy. The preset load balancing strategy can be based on the idle computing power of the IO core or the mapping between the preset IO core and the disk block number.

[0178] Step S4254: Get the message from the lightweight message pump and send a read request message to the IO core;

[0179] During implementation, a message object is retrieved from the job core's message cache and a read request is sent to the IO core (selected_iocore) providing the service. If there are no available message objects in the job core's free message cache, several message objects are retrieved from the lightweight Storage Performance Development Kit (SPDK) message pump and placed in the job core's message cache.

[0180] Step S4255: The IO core reads the data block and returns a message to the job core;

[0181] During the implementation process, the IO core executes the request to read the data block (plog_id) and sends the message back to the message pump of the job core according to the source of the message (that is, the job core that obtains the message object).

[0182] Step S4256: Calculate the vector distance and add it to the result set.

[0183] During implementation, the job core can be used to calculate the distance between the query vector and the vectors that meet the filter conditions based on the read vector data, and add it to the result set.

[0184] Step S426: Wait until all intra-cluster filtering query sub-processes are completed;

[0185] Step S427: Determine whether the number of vector entities in the result set meets the conditions;

[0186] Here, if it is determined that the number of entities in the vector set meets the condition, step S428 is executed; if it is determined that the number of entities in the vector set does not meet the condition, step S422 is executed.

[0187] During implementation, if the number of elements in the result set does not reach a specified threshold (related to recall, which can be set to topk*24), step S422 is executed. If the number of elements in the result set does meet the specified threshold, the top k (topk) vectors with the smallest distances are found from the result set and these top k results are returned.

[0188] In an embodiment of the present application, the vector inverted index records the value range (maximum value, minimum value) of the scalar field (domain) of the entity to which the cluster belongs. Before reading the vector inverted index, clusters that do not meet the filter conditions are screened out according to the value range, eliminating unnecessary disk reads (this solution is referred to as RangeCheck). For the vector inverted index, the vectors in the cluster are rearranged based on the scalar field (domain) data information. After rearrangement, the vectors of the entities that meet the filter conditions are usually clustered together. When reading the inverted list, only the vectors of the entities that meet the filter conditions are read from the disk, which further reduces the amount of disk data read and solves the above problem A.

[0189] When loading the disk index file, the mapping relationship between the disk index file offset and the underlying disk blocks is obtained and recorded. During filtering queries, requests for underlying blocks are directly sent to multiple IO cores. This eliminates file system layer overhead and breaks the limitation that a micro-file system can only be serviced by a single IO core, improving IO throughput and solving the aforementioned problem D. A lightweight SPDK message engine is built for the job core (which is only responsible for message recycling and does not require any other processing). After the IO core processes the job core's IO read request, it directly returns the SPDK message (which contains the source job core's ID) to the source job core's message engine, solving the aforementioned problem E.

[0190] After performing scalar domain filtering on all entities, the number of clusters obtained in each batch is dynamically calculated based on the filtering rate, reducing the number of iterations and solving problem B above. After finding a batch of clusters closest to the query vector from the in-memory index, a RangeCheck is first performed. For each qualifying cluster, a pipeline (thread / coroutine) is launched to concurrently process the following process: finding a valid data block => reading the data block => vector calculation and adding it to the result set, solving problem C above.

[0191] The embodiments of the present application can achieve the following beneficial technical effects:

[0192] (1) Eliminate unnecessary cluster inverted list data reading, reduce disk read volume, and improve filtering query throughput.

[0193] (2) Multiple IO cores serve a dataset shard simultaneously, supporting high-concurrency filtering queries.

[0194] (3) A lightweight job core message engine that reduces message management overhead, supports high-throughput IO reads, and improves filtering query performance.

[0195] (4) Dynamically determine the number of clusters obtained in each batch, and concurrently filter queries in each cluster, with fewer iterations and lower query latency.

[0196] (5) Independent scalar filtering steps do not limit the content of the filter expression and support complex expressions, such as equal to / greater than / less than / range for numeric scalar fields, equal to / prefix / greater than / less than / range for string scalar fields, etc.

[0197] Based on the aforementioned embodiments, an embodiment of the present application provides a vector query device, which includes modules, each module includes sub-modules, each sub-module includes a unit, and can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; during implementation, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0198] Figure 5 A schematic diagram of the structure of the vector query device provided in the embodiment of the present application is shown in FIG. Figure 5 As shown, the apparatus 500 includes:

[0199] A first determining module 510 is configured to determine a data block set matching the query data using the computing core, wherein the data block set includes N data blocks, where N is an integer greater than or equal to 1;

[0200] A second determining module 520 is configured to use the computing core to determine N target read / write cores corresponding to the N data blocks based on a preset strategy, and send N read requests to the corresponding N target read / write cores respectively;

[0201] The reading module 530 is configured to use the Nth target read / write core to read the Nth vector data stored in the Nth disk block number, and return the Nth vector data to the computing core.

[0202] In some embodiments, the second determination module 520 includes a first determination submodule or a second determination submodule, wherein the first determination submodule is used to use the computing core to determine the N target read-write cores based on the idle computing power of the read-write core; the second determination submodule is used to use the computing core to determine the N target read-write cores based on the mapping relationship between the read-write core and the disk block number.

[0203] In some embodiments, the query data includes a query scalar and a query vector, and the first determination module 510 includes a third determination submodule, a fourth determination submodule and a fifth determination submodule, wherein the third determination submodule is used to determine the target number of target clusters to be acquired based on the filtering rate using the computing core, wherein the filtering rate is determined based on the query scalar; the fourth determination submodule is used to determine the first cluster center that meets the target number from the vector index file based on the query vector, so as to determine the target cluster using the first cluster center; the fifth determination submodule is used to determine the data block set based on the target cluster.

[0204] In some embodiments, the fifth determination submodule includes an identification unit and a merging unit, wherein the identification unit is used to identify the target entity matching the query scalar based on the vector list corresponding to the target cluster; the merging unit is used to merge the data that meets the threshold distance into a data block based on the identification of the target entity to obtain the data block set.

[0205] In some embodiments, the fourth determination submodule includes an acquisition unit and a determination unit, wherein the acquisition unit is used to obtain a cluster scalar range file of each first cluster center, wherein the cluster scalar range file is used to record the scalar value range corresponding to the first cluster center; and the determination unit is used to determine the target cluster based on the scalar value range corresponding to the first cluster center.

[0206] In some embodiments, the vector query device further includes an acquisition module, a third determination module and a construction module, wherein the acquisition module is used to use the computing core to obtain a pre-stored file disk block mapping relationship, wherein the file disk block mapping relationship represents a mapping relationship between file index information and disk block numbers; the third determination module is used to determine the Nth disk block number stored corresponding to the Nth data block based on the Nth file index information corresponding to the Nth data block using the file disk block mapping relationship; and the construction module is used to use the Nth disk block number to construct the Nth read request.

[0207] In some embodiments, the second determination module 520 includes an acquisition submodule and a sending submodule, wherein the acquisition submodule is used to obtain N target message objects from the message cache of the computing core; the sending submodule is used to use the N target message objects to send the N read requests to the corresponding N target read-write cores respectively.

[0208] In some embodiments, the second determination module 520 also includes a fifth determination submodule, a recycling submodule and a storage submodule, wherein the fifth determination submodule is used to determine that there are no idle message objects in the message cache of the computing core; the recycling submodule is used to recycle the message object that completes the read request from a preset message pump; and the storage submodule is used to store the message object that completes the read request into the message cache of the computing core.

[0209] In some embodiments, the vector query device also includes a sorting module and an output module, wherein the sorting module is used to use the computing core to sort the N vector data and the M first cluster center vector data based on the distance information between the Nth vector data and the M first cluster center vector data and the query data to obtain a sorting result, and the M first cluster center vector data are obtained by filtering the first cluster center based on a filter expression, and M is an integer greater than or equal to 1; the output module is used to sequentially output the entities corresponding to the first K vector data as query results based on the sorting result, wherein K is a preset query entity threshold, and K is less than or equal to N.

[0210] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.

[0211] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0212] Correspondingly, an embodiment of the present application provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the vector query method provided in the above embodiment.

[0213] Correspondingly, an embodiment of the present application provides an electronic device, Figure 6 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the hardware entity of the device 600 includes: a memory 601 and a processor 602, wherein the memory 601 stores a computer program that can be run on the processor 602, and when the processor 602 executes the program, the steps in the vector query method provided in the above embodiment are implemented.

[0214] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or processed by the processor 602 and various modules in the electronic device 600 (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (RAM).

[0215] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0216] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0217] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0218] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0219] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0220] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0221] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0222] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words be embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0223] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0224] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0225] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0226] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A vector query method, applied to a processor, wherein the processor includes at least a computing core and a read / write core, the method comprising: The computing core determines a set of data blocks that match the query data, wherein the set of data blocks includes N data blocks, where N is an integer greater than or equal to 1; The computing core determines N target read-write cores corresponding to the N data blocks based on a preset strategy, and sends N read requests to the corresponding N target read-write cores respectively; The Nth target read-write core reads the Nth vector data stored in the Nth disk block number, and returns the Nth vector data to the computing core.

2. The method according to claim 1, wherein the computing core determines N target read / write cores corresponding to the N data blocks based on a preset strategy, comprising: The computing core determines the N target read-write cores based on the idle computing power of the read-write core; or, The computing core determines the N target read-write cores based on a mapping relationship between the read-write cores and the disk block numbers.

3. The method of claim 1 , wherein the query data includes a query scalar and a query vector, and the computing core determines a set of data blocks matching the query data, comprising: The computing core determines a target number of target clusters to be acquired based on a filtering rate, wherein the filtering rate is determined based on the query scalar; determining a first cluster center that meets the target number from a vector index file based on the query vector, and determining the target cluster using the first cluster center, wherein the vector index file is stored in a memory; The data block set is determined based on the target cluster.

4. The method according to claim 3, wherein determining the data block set based on the target cluster comprises: Identifying a target entity matching the query scalar based on the vector list corresponding to the target cluster; Based on the identifier of the target entity, the data meeting the threshold distance are merged into a data block to obtain the data block set.

5. The method according to claim 3, wherein determining the target cluster using the first cluster center comprises: Obtaining a cluster scalar range file for each of the first cluster centers, wherein the cluster scalar range file is used to record a scalar value range corresponding to the first cluster center; The target cluster is determined based on a scalar value range corresponding to the first cluster center.

6. The method according to any one of claims 1 to 5, further comprising: before the computing core sends the N read requests to the corresponding N target read / write cores respectively; The computing core obtains a pre-stored file disk block mapping relationship, wherein the file disk block mapping relationship represents a mapping relationship between file index information and disk block numbers; Determine the Nth disk block number corresponding to the Nth data block based on the Nth file index information corresponding to the Nth data block using the file disk block mapping relationship; The Nth read request is constructed using the Nth disk block number.

7. The method according to any one of claims 1 to 5, wherein the computing core sends N read requests to the corresponding N target read / write cores, respectively, comprising: Obtain N target message objects from the message cache of the computing core; The N target message objects are used to send the N read requests to the corresponding N target read-write cores respectively.

8. The method of claim 7, further comprising: Determining that there are no idle message objects in the message cache of the computing core; Recover the message object that completes the read request from the preset message pump; The message object of the completed read request is stored in the message cache of the computing core.

9. The method according to any one of claims 1 to 5, further comprising: The computing core sorts the N vector data and the M first cluster center vector data based on distance information between the Nth vector data and the M first cluster center vector data and the query data to obtain a sorting result, wherein the M first cluster center vector data are obtained by filtering the first cluster centers based on a filter expression, where M is an integer greater than or equal to 1; Based on the sorting result, the entities corresponding to the first K vector data are sequentially output as query results, where K is a preset query entity threshold and K is less than or equal to N.

10. A vector query device, comprising: A first determining module, configured to use a computing core of a processor to determine a set of data blocks matching the query data, wherein the set of data blocks includes N data blocks, where N is an integer greater than or equal to 1; A second determining module is configured to use the computing core to determine N target read-write cores corresponding to the N data blocks based on a preset strategy, and send N read requests to the corresponding N target read-write cores respectively; A reading module is used to use the Nth target read-write core to read the Nth vector data stored in the Nth disk block number, and return the Nth vector data to the computing core.