Data processing method, data processing system, and computer-readable storage medium
Patent Information
- Application Number
- PCT/CN2025/148206
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-12-31
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025148206_01102026_PF_FP_ABST
Abstract
Description
Data processing methods, data processing systems, and computer-readable storage media
[0001] This application claims priority to Chinese Patent Application No. 202510380060.9, filed on March 27, 2025, entitled "Data Processing Method, Data Processing System and Computer-Readable Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computing, and in particular to data processing methods, data processing systems and computer-readable storage media. Background Technology
[0003] Data quantization is a common data processing method in computing, which converts high-bit data into low-bit data. High-bit data refers to data containing a large number of bits, while low-bit data refers to data containing a small number of bits. The resources required for storing and processing low-bit data are less than those required for storing and processing high-bit data. Therefore, in computing, high-bit data is typically quantized into low-bit data before further processing, thus reducing resource consumption.
[0004] For example, data quantization can be applied in vector data retrieval scenarios. Vector data retrieval is the process of finding vector data that meets specific requirements from a large amount of vector data. The vector data that meets the specific requirements can be those with a high degree of similarity to the retrieved vector data. During vector data retrieval, the processor quantizes the high-bit vector data and the high-bit retrieved vector data to obtain low-bit vector data and low-bit retrieved vector data, and calculates the distance between the two. The distance between the low-bit vector data and the low-bit retrieved vector data is negatively correlated with their similarity. Therefore, by calculating the distance between the low-bit vector data and the low-bit retrieved vector data, vector data with a high degree of similarity to the retrieved vector data can be identified, thus achieving data retrieval.
[0005] In the process of quantizing high-bit vector data and high-bit retrieval vector data, related technologies first construct a low-dimensional Hamming space based on the high-bit vector data. Then, they quantize the high-bit vector data and high-bit retrieval vector data using the Hamming space to obtain low-bit vector data and low-bit retrieval vector data. The low-bit vector data or low-bit retrieval vector data can be regarded as points in the Hamming space. By calculating the Hamming distance between the low-bit vector data and low-bit retrieval vector data in the Hamming space, the similarity between the low-bit vector data and low-bit retrieval vector data can be determined, thereby realizing data retrieval.
[0006] However, different vector data are suited to different Hamming spaces. Therefore, when the vector data changes, the relevant technologies also need to reconstruct a low-dimensional Hamming space. Since constructing a low-dimensional Hamming space is computationally expensive, and data quantization through Hamming space is difficult, the data quantization process provided by the relevant technologies is not only computationally costly but also computationally complex. Summary of the Invention
[0007] This application provides a data processing method, a data processing system, and a computer-readable storage medium to reduce the computational load and complexity of the quantization process. The technical solution is as follows:
[0008] Firstly, a data processing method is provided, comprising: acquiring N quantization reference values, the N quantization reference values being arranged in a preset order to form N+1 value domains, where N is a positive integer greater than 1; acquiring vector data to be quantized, comparing the vector data with the N quantization reference values in a preset order to obtain N-bit quantized data of the vector data, wherein any bit of the N-bit quantized data is determined by the comparison result between the vector data and the quantization reference value corresponding to the order of any bit of the N quantization reference values, and the N-bit data indicates that the vector data belongs to one of the value domains in the N+1 value domains.
[0009] In this application, N quantization reference values form N+1 value ranges. By comparing the vector data with the N quantization reference values, N comparison results are obtained. Based on these N comparison results, the N bits of data included in the quantization data are determined, thus obtaining N-bit quantized data and realizing the quantization of vector data. The quantization method of this application is simple, reducing the computational load and complexity of the quantization process, resulting in lower computational overhead.
[0010] In one possible implementation, the vector data is vector data from a vector database. All vector data in the database are quantized into quantized data belonging to N+1 value ranges, each value range being indicated by an N-bit quantized data. During the quantization process for each vector data in the vector database, N comparison results obtained by comparing the magnitude of the vector data with the magnitudes of N quantization reference values can respectively indicate the magnitude relationship between the vector data and the N quantization reference values, thereby accurately indicating the value range to which the vector data belongs and achieving uniform quantization of the vector data.
[0011] In one possible implementation, the vector data is compared with N quantization reference values in a preset order to obtain N-bit quantized data after quantization. This includes: comparing the vector data with the i-th quantization reference value, and determining the i-th bit of data included in the quantized data based on the comparison result; wherein, when the comparison result is that the vector data is less than or equal to the i-th quantization reference value, the i-th bit of data included in the quantized data is the first data; when the comparison result is that the vector data is greater than the i-th quantization reference value, the i-th bit of data included in the quantized data is the second data, and the second data is different from the first data.
[0012] The data for each bit determined based on different comparison results varies. Each bit indicates the relationship between the vector data and each quantization reference value, ensuring that the data in the quantized data accurately indicates the range of the vector data before quantization. Determining the data for each bit in the quantized data based on the comparison results reduces the computational load of the quantization process and improves quantization efficiency.
[0013] In one possible implementation, obtaining the vector data to be quantized includes: retrieving X vector data points from a vector database, where X is a positive integer greater than 1; comparing the vector data points with N quantization reference values in a preset order to obtain N-bit quantized data, which includes: comparing the X vector data points with the N quantization reference values in a preset order to obtain N X-bit data points; and transposing the N X-bit data points to obtain X N-bit quantized data points. By comparing the X vector data points from the vector database with the N quantization reference values in a preset order to finally obtain X quantized data points, parallel quantization of the X vector data points is achieved, improving quantization efficiency.
[0014] In one possible implementation, X vector data are compared with N quantization reference values in a preset order to obtain N X-bit data, including: comparing the size of the X vector data with the i-th judgment data, determining X third data based on the comparison result, where each third data includes X identical bits, and X is a positive integer greater than 1; performing bitwise AND operations on the X third data with the reference data of the same order among the X reference data to obtain X fourth data, where each reference data includes X bits, the j-th bit of the j-th reference data is the fifth data, and the bits other than the j-th bit are the sixth data, where the fifth and sixth data are different, and j is a positive integer less than or equal to X; and performing a bitwise sum operation on the X fourth data to obtain the i-th X-bit data corresponding to the X vector data.
[0015] Since the j-th bit of the j-th reference data is different from the other bits, while the j-th third data contains the same bits, when the j-th reference data and the j-th third data are bitwise ANDed, the j-th bit of the resulting fourth data can reflect the data of the bits included in the j-th third data. After performing a bitwise summation operation on X fourth data, the j-th bit of the resulting X-bit data can reflect the data of the bits included in the j-th third data. This reflects the magnitude relationship between the j-th vector data before quantization and the i-th quantization reference value, thus obtaining accurate X-bit data. This ensures the accuracy of the N-bit quantized data obtained subsequently based on the transpose of the X-bit data and improves the quantization accuracy.
[0016] In one possible implementation, when the comparison result indicates that any vector data is less than or equal to the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all first data; when the comparison result indicates that any vector data is greater than the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all second data, which are different from the first data. When the comparison results are different, the data included in the quantized third data are different, ensuring that the data of each bit in the third data accurately reflects the relationship between the vector data and the i-th quantization reference value, thus accurately reflecting the value range of the vector data.
[0017] In one possible implementation, the method further includes: acquiring a vector retrieval request, which includes retrieval vector data; comparing the retrieval vector data with N quantization benchmark values in a preset order to obtain N-bit quantized retrieval data after quantization; determining the target value range corresponding to the retrieval quantization data based on the distance between the retrieval quantization data and N+1 types of quantization data in the vector database; and comparing the retrieval vector data with vector data in the target value range to determine the target data corresponding to the retrieval vector data. Based on the same quantization method as the vector data, the retrieval vector data is quantized to accurately determine the distance between the retrieval quantization data and N+1 types of quantization data in the vector database, thereby accurately determining the target value range corresponding to the retrieval quantization data. This allows for further accurate determination of the target data corresponding to the retrieval vector data among the vector data belonging to the target value range.
[0018] In one possible implementation, the retrieval vector data is compared with N quantization reference values in a preset order to obtain N-bit quantized retrieval data after quantization of the retrieval vector data. This includes: comparing the retrieval vector data with the i-th quantization reference value, and determining the i-th bit of data included in the retrieval quantization data based on the comparison result; wherein, when the retrieval vector data is less than or equal to the i-th quantization reference value, the i-th bit of data in the retrieval quantization data is the first data; when the retrieval vector data is greater than the i-th quantization reference value, the i-th bit of data in the retrieval quantization data is the second data, and the second data is different from the first data.
[0019] The data for each bit determined based on different comparison results varies. Each bit indicates the relationship between the unquantized retrieval vector data and each quantization reference value, ensuring that the data for each bit in the quantized retrieval data accurately indicates the value range of the unquantized retrieval vector data. Determining the data for each bit in the quantized retrieval data based on the comparison results reduces the computational load of the quantization process and improves quantization efficiency.
[0020] In one possible implementation, before determining the target value range of the retrieved quantized data based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database, the method further includes: for any type of quantized data, performing a bitwise AND operation on the N bits of data included in the retrieved quantized data and the N bits of data included in the retrieved quantized data to obtain N AND operation results; summing the N AND operation results to obtain the distance between the arbitrary type of quantized data and the retrieved quantized data.
[0021] In the quantization process of this application, the comparison results are determined based on N quantization benchmark values used for value range division, vector data, and retrieved vector data. The data of each bit included in the quantized data or retrieved quantized data obtained based on the comparison results can reflect the value range to which the vector data or retrieved vector data belongs. Therefore, the Euclidean distance between the quantized data and the retrieved quantized data can reflect the similarity between the retrieved quantized data and the quantized data. Thus, by calculating the Euclidean distance, the distance between the quantized data and the retrieved quantized data can be accurately determined.
[0022] In one possible implementation, before determining the target value range of the retrieved quantized data based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database, the method further includes: for any type of quantized data, performing a bitwise XOR operation on the N bits of data included in the retrieved quantized data and the N bits of data included in the retrieved quantized data to obtain N XOR operation results; summing the N XOR operation results to obtain the distance between the quantized data and the retrieved quantized data.
[0023] In the quantization process, after obtaining the comparison results, this application determines the data of each bit of the quantized data or retrieved quantized data according to the Hamming encoding method. Therefore, the Hamming distance between the quantized data and the retrieved quantized data can also reflect the similarity between the two data. Thus, by calculating the Hamming distance, the distance between the quantized data and the retrieved quantized data can be accurately determined.
[0024] In one possible implementation, the target value range corresponding to the retrieved quantized data is determined based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database. This includes: based on the distance between the N+1 types of quantized data and the retrieved quantized data, determining the quantized data with the smallest distance from the retrieved quantized data among the N+1 types of quantized data; and determining the value range of the quantized data with the smallest distance from the retrieved quantized data as the target value range.
[0025] Based on the calculated distance between the quantized data and the retrieved quantized data, the similarity between the two can be accurately reflected. Furthermore, the similarity between the quantized data and the retrieved quantized data is inversely proportional to the distance between them. Therefore, the quantized data with the smallest distance to the retrieved quantized data has the highest similarity. The value range of the quantized data with the smallest distance to the retrieved quantized data is closest to the value range of the retrieved quantized data. Consequently, the distance between vector data in the value range of the quantized data with the smallest distance to the retrieved vector data is smaller and the similarity is higher compared to vector data in other value ranges. Therefore, using the value range of the quantized data with the smallest distance to the retrieved quantized data as the target value range allows for the efficient identification of target data similar to the retrieved vector data within the vector data included in the target value range.
[0026] In one possible implementation, obtaining N quantization reference values includes: determining the value range of the vector data; dividing the value range into N+1 value domains, where the difference between the two data values of vector data belonging to any two value domains is less than a difference threshold; and determining N quantization reference values based on the values of N intersection points of the N+1 value domains. The fact that the difference between the two data values of vector data belonging to any two value domains is less than the difference threshold ensures a relatively uniform distribution of vector data in the N+1 value domains in the vector database, enabling uniform quantization of the vector data based on the intersection points of the N+1 value domains.
[0027] Secondly, a data processing system is provided, comprising: an acquisition module for acquiring N quantization reference values, the N quantization reference values being arranged in a preset order to form N+1 value domains, where N is a positive integer greater than 1; and a quantization module for acquiring vector data to be quantized, comparing the vector data with the N quantization reference values in a preset order to obtain N-bit quantized data of the vector data, wherein any bit of the N-bit quantized data is determined by the comparison result between the vector data and the quantization reference value corresponding to any bit of the N quantization reference values, and the N-bit data indicates that the vector data belongs to one of the N+1 value domains.
[0028] In one possible implementation, the vector data is vector data in a vector database, and all vector data in the vector database is quantized into quantized data belonging to N+1 value ranges, each value range being indicated by an N-bit quantized data.
[0029] In one possible implementation, a quantization module is used to compare the vector data with the i-th quantization reference value, and determine the data of the i-th bit included in the quantization data based on the comparison result; wherein, when the comparison result is that the vector data is less than or equal to the i-th quantization reference value, the data of the i-th bit included in the quantization data is the first data; when the comparison result is that the vector data is greater than the i-th quantization reference value, the data of the i-th bit included in the quantization data is the second data, and the second data is different from the first data.
[0030] In one possible implementation, an acquisition module is used to acquire X vector data from a vector database, where X is a positive integer greater than 1; a quantization module is used to compare the X vector data with N quantization reference values in a preset order to obtain N X-bit data; and to transpose the N X-bit data to obtain X N-bit quantized data after quantizing the X vector data.
[0031] In one possible implementation, a quantization module is used to compare the size of X vector data with the i-th judgment data, and determine X third data based on the comparison result. Each third data consists of X identical bits, where X is a positive integer greater than 1. The X third data are then bitwise ANDed with the X reference data in the same order to obtain X fourth data. Each reference data consists of X bits. The j-th bit of the j-th reference data is the fifth data, and the bits other than the j-th bit are the sixth data. The fifth and sixth data are different, where j is a positive integer less than or equal to X. Finally, the X fourth data are bitwise summed to obtain the i-th X-bit data corresponding to the X vector data.
[0032] In one possible implementation, when the comparison result indicates that any vector data is less than or equal to the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all the first data; when the comparison result indicates that any vector data is greater than the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all the second data, and the second data is different from the first data.
[0033] In one possible implementation, the data processing system further includes a determination module; an acquisition module, which is further used to acquire a vector retrieval request, the vector retrieval request including retrieval vector data; a quantization module, which is further used to compare the retrieval vector data with N quantization benchmark values in a preset order to obtain N-bit quantized retrieval data after quantization of the retrieval vector data; and a determination module, which is used to determine the target value range corresponding to the retrieval quantization data based on the distance between the retrieval quantization data and N+1 types of quantization data in the vector database; and to compare the retrieval vector data with the vector data in the target value range to determine the target data corresponding to the retrieval vector data.
[0034] In one possible implementation, a quantization module is used to compare the retrieved vector data with the i-th quantization reference value, and based on the comparison result, determine the data of the i-th bit included in the retrieved quantized data; wherein, when the retrieved vector data is less than or equal to the i-th quantization reference value, the data of the i-th bit of the retrieved quantized data is the first data; when the retrieved vector data is greater than the i-th quantization reference value, the data of the i-th bit of the retrieved quantized data is the second data, and the second data is different from the first data.
[0035] In one possible implementation, the data processing system further includes a calculation module, which is used to perform a bitwise AND operation on N bits of any quantized data and N bits of the retrieved quantized data to obtain N AND operation results; and to sum the N AND operation results to obtain the distance between the arbitrary quantized data and the retrieved quantized data.
[0036] In one possible implementation, the data processing system further includes a calculation module, which is used to perform a bitwise XOR operation on N bits of any quantized data and N bits of the retrieved quantized data to obtain N XOR operation results; and to sum the N XOR operation results to obtain the distance between the arbitrary quantized data and the retrieved quantized data.
[0037] In one possible implementation, a determination module is used to determine, based on the distance between N+1 types of quantized data and the retrieved quantized data, the quantized data with the smallest distance from the retrieved quantized data among the N+1 types of quantized data; and to determine the value range of the quantized data with the smallest distance from the retrieved quantized data as the target value range.
[0038] In one possible implementation, an acquisition module is used to determine the value range of the vector data; the value range is divided into N+1 value domains, and the difference between the two data quantities of the vector data belonging to any two value domains is less than the difference threshold; N quantization benchmark values are determined based on the values of the N intersection points of the N+1 value domains.
[0039] Thirdly, a computer program (product) is provided, comprising: computer program code, which, when executed by a computer, causes the computer to perform the data processing method described in the first aspect and any possible implementation thereof.
[0040] Fourthly, a computer-readable storage medium is provided that stores a program or instructions, wherein when the program or instructions are run on a computer, the data processing method described in the first aspect and any possible implementation thereof is executed.
[0041] It should be understood that the beneficial effects of the technical solutions and corresponding possible implementations of the second to fourth aspects of this application can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0042] Figure 1 is a schematic diagram of a data retrieval process provided by a related technology;
[0043] Figure 2 is a schematic diagram of a data retrieval process provided by related technology 2;
[0044] Figure 3 is a schematic diagram of an implementation scenario provided by an embodiment of this application;
[0045] Figure 4 is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0046] Figure 5 is a schematic diagram of the distribution of vector data values in a vector database provided in an embodiment of this application;
[0047] Figure 6 is a schematic diagram illustrating the principle of a quantization method provided in an embodiment of this application;
[0048] Figure 7 is a schematic diagram of quantization and storage provided in an embodiment of this application;
[0049] Figure 8 is a schematic diagram of a quantization process provided in an embodiment of this application;
[0050] Figure 9 is a schematic diagram of another quantization and storage method provided in an embodiment of this application;
[0051] Figure 10 is a schematic diagram of a data processing process provided in an embodiment of this application;
[0052] Figure 11 is a schematic diagram of the structure of a data processing system provided in an embodiment of this application;
[0053] Figure 12 is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Detailed Implementation
[0054] The terminology used in the implementation section of this application is for the purpose of explaining specific embodiments of this application only, and is not intended to limit this application.
[0055] Data quantization is a common data processing technique in computing, which converts high-bit data into low-bit data. High-bit data refers to data containing a large number of bits, while low-bit data refers to data containing a small number of bits. Data quantization can be applied in various data processing scenarios; for example, it can be used in vector data retrieval.
[0056] Vector data retrieval is a search method that finds the most similar data among the retrieved vector data by calculating the similarity or distance between them. Vector data refers to the data that needs to be compared with the retrieved vector data during the retrieval process. It is typically stored in a vector database, which contains target data similar to the retrieved vector data among multiple vector data sets. During vector data retrieval, quantization can modify the number of bits occupied by each vector data set and the number of bits occupied by the retrieved vector data. This allows high-bit vector data (or vector data containing multiple bits, high-precision vector data) to be quantized into low-bit vector data (or vector data containing fewer bits, low-precision vector data) with a relatively low loss of precision. This reduces the space required for storing and retrieving vector data and increases the efficiency of data retrieval. The number of bits occupied and the range of data represented by different data types in a computer vary; the fewer bits used for quantization, the higher the compression ratio of the quantized data.
[0057] The industry standard for quantization bits is 8 bits; quantization with fewer than 8 bits is called low-bit quantization. 1 bit is the limit for quantization compression. Common quantization techniques in vector data retrieval scenarios focus on 8-bit integer quantization and 4-bit integer quantization. Int8 quantization converts vector data containing more than 8 bits into 8-bit integer data, while int4 quantization further reduces the number of bits to 4, reducing data storage requirements and improving computational efficiency. For example, converting 32-bit single-precision floating-point format (FP32) data and FP32 query data in a vector into int4 data through a linear transformation transforms the similarity calculation during data retrieval into a calculation of the similarity between int4 data, reducing computational complexity. Furthermore, int4 data occupies only 1 / 8 the number of bits of FP32 data, thus reducing memory usage.
[0058] Related technologies provide a uniform linear quantization method, in which the distance between two adjacent quantized data points is fixed. For example, related technologies perform uniform linear quantization using the following formula (1): Q = round(r / s) (1)
[0059] In formula (1), r represents the high-bit vector data before quantization, i.e., the original vector data; Q represents the low-bit vector data after quantization; and s represents the scaling factor, which is a parameter related to quantization. The quantization process involves dividing the high-bit vector data by the scaling factor and then rounding it to obtain the quantized result, which is the low-bit vector data.
[0060] For example, see Figure 1, which illustrates a vector data retrieval process provided by a related technology. This related technology determines the similarity between vector data in the vector database and the queried vector data (i.e., the retrieved vector data) by calculating the distance between them, thereby identifying one or more vector data points in the vector database that have the highest similarity to the queried vector data.
[0061] In the process of calculating the distance between vector data in the vector database (hereinafter referred to as vector data) and the queried vector data, as shown in Figure 1, the vector data in the vector database in fp32, int8 or int16 format is encoded, that is, int4 quantization is performed to obtain 4-bit encoded data.
[0062] After receiving the data retrieval request, the 4-bit encoded data is decoded, which is also called dequantization, to obtain vector data in fp32, int8, or int16 format. Related techniques use the following formula (2) for dequantization: r'=s*Q (2)
[0063] In formula (2), r' is the high-bit data obtained by dequantization, s is the scaling factor used in the quantization process, and Q is the low-bit data after quantization. The dequantization process is the inverse process of multiplying the low-bit data after quantization by the scaling factor.
[0064] Accordingly, the related technology also performs quantization on the queried vector data to obtain 4-bit compressed vector encoded data, and then performs dequantization on the 4-bit compressed vector encoded data to obtain dequantized vector data. The dequantized vector data is in fp32, int8, or int16 format. Distance is calculated between the dequantized vector data corresponding to the vector data in the vector database and the dequantized data corresponding to the queried vector data to obtain the distance result, thereby enabling the retrieval of data similar to the queried vector data.
[0065] The existing technology involves quantization and dequantization of data, which is time-consuming. Furthermore, because the quantization process introduces rounding operations, the dequantized data may differ from the original data, whether it's vector data from a vector database or queried vector data, leading to lower accuracy in distance calculations. Additionally, distance calculations involve numerous dot product operations between vector data, which are less compatible with central processing units (CPUs) and result in slow processing speeds.
[0066] Related Technique 2 provides a method to accelerate vector distance calculation using a pre-trained Hamming space. Hamming space is a method for projecting high-dimensional vectors onto a low-dimensional space. In Hamming space, vectors are XORed bitwise to obtain a shorter binary string. This binary string can be considered as points in Hamming space. As shown in Figure 2, Related Technique 2 pre-trains N-dimensional Hamming spaces based on fp32, int8, or int16 format vector data from a vector database. The generated N-dimensional Hamming space is then used to encode the fp32, int8, or int16 format vector data from the vector database, obtaining N-bit Hamming encoded data corresponding to the vector data in the database. Similarly, the N-dimensional Hamming space is used to encode the query vector data in fp32, int8, or int16 format, obtaining N-bit Hamming encoded data corresponding to the query vector data. The Hamming distance between the N-bit Hamming encoded data generated from the query vector data and the N-bit Hamming encoded data generated from the vector data in the vector database is calculated, yielding the distance result. The distance result is then used to determine the data most similar to the query data, thus achieving data retrieval.
[0067] Related technique two requires pre-training to generate an N-dimensional Hamming space before calculating the distance result. Furthermore, whenever data is added or deleted from the vector database, the matched N-dimensional Hamming space changes, necessitating retraining and incurring significant computational overhead. Additionally, the encoding process using the N-dimensional Hamming space is cumbersome, data retrieval requires substantial computational resources, and demands significant computing power.
[0068] This application provides a data processing method that reduces the computational load during quantization, improves quantization efficiency, and increases the efficiency of calculating the distance between quantized data. Referring to Figure 3, a schematic diagram of an implementation scenario provided by this application is shown. This scenario includes one or more processors 31, each processor 31 connected to its own memory 32. The memory 32 connected to the processor 31 stores the data processed by the processor 31. Different processors 31 and connected memory 32 can belong to the same data processing device (e.g., a computer, server, etc.), or they can belong to different data processing devices. Different data processing devices can belong to the same device cluster, such as a device cluster in a database scenario, an artificial intelligence (AI) server cluster, a device cluster in a cloud service scenario, or a device cluster in a large model development scenario. For example, the implementation scenario includes two processors 31. One processor 31 and its connected memory 32 belong to data processing device A, and the other processor 31 and its connected memory 32 can belong to either data processing device A or data processing device B. Data processing device A and data processing device B can be two independent data processing devices or two data processing devices belonging to the same device cluster.
[0069] This application does not limit the type of processor 31. For example, processor 31 can be a graphics processing unit (GPU), a neural processing unit (NPU), a data processing unit (DPU), or a central processing unit (CPU). The types of different processors 31 in the implementation scenario can be the same or different. If there are multiple processors 31 in the implementation scenario, wireless communication connections or wired communication connections can be established between the multiple processors 31.
[0070] The data processing method provided in this application embodiment can be executed by one processor 31 or by multiple processors 31 working together. This data processing method can be used to implement data retrieval, or in other words, for data retrieval scenarios. During the data retrieval process, vector data in the vector database and retrieval vector data are quantized to obtain quantized data and retrieval quantized data. Based on the quantized data and retrieval quantized data, operations such as distance calculation are performed to achieve the retrieval of data similar to the retrieved data.
[0071] Referring to Figure 4, a flowchart of a data processing method provided in an embodiment of this application is shown. The data processing method includes, but is not limited to, the following steps S401 to S402.
[0072] S401, acquire N quantization reference values, arrange the N quantization reference values in a preset order, and form N+1 value ranges, where N is a positive integer greater than 1.
[0073] A quantization reference value is data used to determine the value range of the data before quantization during the quantization process. The values of the N quantization reference values are different. This application does not limit the quantization reference values, the number of bits included in the data representing the N quantization reference values, or the format of the data representing the N quantization reference values. For example, the quantization reference value can be 0, -1, -3.2, etc.; the number of bits included in the data representing the quantization reference value can be 4, 8, 16, or 32, etc., and the number of bits included in the data representing different quantization reference values can be the same or different; the data representing the quantization reference value can be in single-precision floating-point, double-precision floating-point, integer, etc., and the format of the data representing different quantization reference values can be the same or different. N is the number of bits included in the data after quantization based on the quantization reference value. The value of N can be set based on experience, set by the user, determined based on the data characteristics of the data to be quantized, or determined based on the data storage resources for storing the quantized data.
[0074] This application does not limit the method of obtaining N quantization benchmark values. For example, the N quantization benchmark values can be set based on experience, set by the user, or determined based on the data to be quantized. In this application embodiment, the data to be quantized includes vector data, which is vector data in a vector database. The vector database can also be called a base database. The vector data in the vector database can be data uploaded by the user or data generated during the processor's execution of tasks.
[0075] Taking the determination of N quantization benchmark values based on vectors as an example, the process of obtaining N quantization benchmark values includes: determining the value range of vector data; dividing the value range into N+1 value domains, where the difference between the two data values of vector data belonging to any two value domains is less than the difference threshold; and determining N quantization benchmark values based on the values of N intersection points of the N+1 value domains.
[0076] The vector data used to determine the N quantization benchmark values can be all vector data in the vector database. There are usually multiple vector data points. The minimum value of these multiple vector data points can be used as the minimum value of the range, and the maximum value can be used as the maximum value of the range. Then, based on the values of these multiple vector data points, the range is divided into N+1 value domains. During this division, the number of vector data points belonging to each value domain is compared, and the range of each value domain is adjusted until the number of vector data points belonging to each sub-value domain is similar. That is, the difference between the two types of data points belonging to any two value domains is less than a difference threshold. This difference threshold can be set based on experience or by the user.
[0077] For example, referring to Figure 5, a schematic diagram of the distribution of vector data values in a vector database provided in an embodiment of this application is shown. The values of the vector data in the vector database follow a normal distribution. The horizontal axis represents the values of the vector data, and the range of values formed by multiple vector data in the vector database is from -0.32 to 0.32. The vertical axis represents the amount of vector data in the corresponding range of the horizontal axis in the vector database. For example, the amount of vector data with values between -0.02 and 0.02 is approximately 2.8 * 1e. 8 Based on the distribution of multiple vector data values and the number of N (taking N=2 as an example), the value range is divided into three value ranges: value range 1, value range 2, and value range 3. The amount of vector data belonging to these three value ranges is similar. In this embodiment, the difference between the two data values of vector data belonging to any two value ranges is less than the difference threshold, making the distribution of multiple vector data in the N+1 sub-domains relatively uniform, thereby achieving uniform quantization of multiple vector data.
[0078] S402, acquire the vector data to be quantized, compare the vector data with N quantization reference values in a preset order to obtain N-bit quantized data after quantization of the vector data. Any bit of the N-bit quantized data is determined by the comparison result between the vector data and the quantization reference value corresponding to any bit of the N quantization reference values in the order of the data. The N-bit data indicates that the vector data belongs to one of the N+1 value ranges.
[0079] Vector data refers to the data that makes up a vector. For example, data in a vector structure includes multiple elements, or in other words, data in a vector structure can be viewed as a one-dimensional array, where each element can be considered a vector data. For instance, vector V can be represented in array form as V = [v1, v2, v3], where v1, v2, and v3 are the three elements of vector V, or the three data points in vector V. A single vector data point can be any one of v1, v2, or v3; for example, a single vector data point could be v1. The vector data to be quantized includes M bits, where M is a positive integer greater than N, such as 32, 16, or 8.
[0080] Optionally, the processor can acquire the vector data to be quantized by reading vector data from a vector database. The processor can read one or more vector data at a time.
[0081] The process of quantizing vector data involves comparing it with N quantization reference values in a preset order to obtain quantized data. Figure 6 illustrates the principle of the quantization method provided in this application. In Figure 6, the minimum value is the minimum value of each vector data to be quantized, and the maximum value is the maximum value of each vector data to be quantized. P1, P2, P3, ..., PN represent N quantization reference values. These N quantization reference values serve as N quantization judgment points, dividing the value range of the vector data to be quantized into N+1 value ranges. For example, the first value range is the value between the minimum value and the quantization reference value P1, the second value range is the value between the quantization reference value P1 and the quantization reference value P2, and so on, thus dividing the value range of the vector data to be quantized. The quantization process involves comparing the vector data to be quantized with N quantization reference values to determine the value range to which the vector data belongs. Based on the comparison results indicating the value range of the vector data to be quantized, the data of each bit of the quantized data is determined. One comparison result determines one bit of data. Thus, a vector data to be quantized, consisting of M bits, is quantized into quantized data consisting of N bits. In other words, the vector data is the vector data in the vector database, and all vector data in the vector database are quantized into quantized data belonging to N+1 value ranges, with each value range indicated by an N-bit quantized data.
[0082] See Table 1 below, which shows the mapping relationship between the number of bits included in the quantized data, the number of value ranges to which the vector data to be quantized belongs, and the binary values of the quantized data.
[0083] Table 1
[0084] The quantization storage space in Table 1 refers to the number of bits required to store one quantized data point, i.e., the number of bits included in one quantized data point, which is N, for example, 2, 3, or 4. The mapping value in Table 1 refers to the possible values of the quantized data. The number of mapping values refers to the number of possible values of the quantized data obtained after quantizing different vector data to be quantized, i.e., the number of value ranges to which the vector data to be quantized belongs, which is N+1, for example, 3, 4, or 5. The binary values in Table 1 refer to the possible values of the quantized data. Binary values can also be understood as mapping values. Since different vector data to be quantized belonging to the same value range result in the same quantized data, or in other words, different data to be quantized belonging to the same value range result in the same mapping value, the number of binary values is the same as the number of mapping values.
[0085] As can be seen from the foregoing description, there are multiple vector data to be quantized. Therefore, in this embodiment, one vector data to be quantized can be quantized in one quantization process, or multiple vector data to be quantized can be quantized in parallel in one quantization process. The following will use Example 1 and Example 2 as examples to illustrate the two quantization methods: quantizing one vector data to be quantized at a time and quantizing multiple vector data to be quantized in parallel at a time.
[0086] Example 1: Quantize one vector of data at a time.
[0087] The vector data is compared with N quantization reference values in a preset order to obtain N-bit quantized data after quantization. This includes: for the i-th quantization reference value, comparing the vector data with the i-th quantization reference value, and determining the i-th bit of data included in the quantized data based on the comparison result; wherein, when the comparison result is that the vector data is less than or equal to the i-th quantization reference value, the i-th bit of data included in the quantized data is the first data; when the comparison result is that the vector data is greater than the i-th quantization reference value, the i-th bit of data included in the quantized data is the second data, and the second data is different from the first data.
[0088] For example, the first data is 1, the second data is 0, the vector data consists of 16 bits (M is 16), the value of the vector data is 3, N is 2, the first quantization reference value is 0, and the second quantization reference value is 1. Comparing the vector data 3 with the first quantization reference value 0, the vector data 3 is greater than the first quantization reference value 0, so the first bit of the quantized data is the second data 0; comparing the vector data 3 with the second quantization reference value 1, the vector data 3 is greater than the second quantization reference value 1, so the second bit of the quantized data is the second data 0, that is, the obtained quantized data is 00.
[0089] In the example above, different quantization reference values correspond to the same first data, and different quantization reference values correspond to the same second data. This ensures that the XOR distance between the quantized data of two vector data belonging to adjacent value ranges is 1. For example, if vector data 3 belongs to the value range greater than the first quantization reference value and greater than the second quantization reference value, then the value of vector data belonging to the value range adjacent to vector data 3 should belong to the value range greater than the first quantization reference value and less than the second quantization reference value, that is, the value range greater than 0 and less than 1. Taking the vector data 0.5, which belongs to the value range adjacent to vector data 3, as an example, following the example above, quantizing vector data 0.5, comparing the size of vector data 0.5 with the first quantization reference value 0, since vector data 0.5 is greater than the first quantization reference value 0, the first bit of the quantized data is the second data 0; comparing the size of vector data 0.5 with the second quantization reference value 1, since vector data 0.5 is less than the second quantization reference value 1, the second bit of the quantized data is the first data 1, that is, the obtained quantized data is 01. Therefore, the XOR distance between quantized data 01 and quantized data 00 is 1, which means that the number of different bits between quantized data 01 and quantized data 00 is 1.
[0090] Optionally, the first and second data corresponding to different quantization reference values can also be different. For example, the first data corresponding to the first quantization reference value is 1, and the second data is 0; the first data corresponding to the second quantization reference value is 0, and the second data is 1. Taking a vector data value of -3, N as 2, a first quantization reference value of 4, and a second quantization reference value of -4 as an example, during the quantization process, the vector data -3 is compared with the first quantization reference value 4. The comparison result is that the vector data -3 is less than the first quantization reference value 4, so the data of the first bit included in the quantized data is the first data 1 corresponding to the first quantization reference value; the vector data -3 is compared with the second quantization reference value -4. The comparison result is that the vector data -3 is greater than the second quantization reference value -4, so the data of the second bit included in the quantized data is the second data 1 corresponding to the second quantization reference value, that is, the obtained quantized data is 11.
[0091] The data for each bit determined based on different comparison results varies. Each bit indicates the relationship between the vector data and each quantization reference value, ensuring that the data for each bit in the quantized data accurately indicates the range of the vector data. Determining the data for each bit in the quantized data based on the comparison results reduces the computational load of the quantization process and improves quantization efficiency.
[0092] After quantizing to obtain the data of each bit in the quantized data, the data of each bit in the quantized data can be stored. For example, the storage method can be sequential storage, that is, storing the data of the first to Nth bits of the first quantized data, storing the data of the first to Nth bits of the second quantized data, and so on. The data of each bit in the quantized data can be stored in memory connected to the processor.
[0093] Referring to Figure 7, a schematic diagram of quantization and storage provided in an embodiment of this application is shown. The vector data includes four data points: A1, A2, A3, and A4. The process of storing the data of each bit in the quantized data can also be called the process of encoding the data of each bit in the quantized data. The storage result can also be regarded as the encoding result. Taking N as 2 as an example, the vector data A1 is quantized to obtain quantized data A1'. The first bit of quantized data A1' is A1[1], and the second bit of quantized data A1' is A1[2]. A1[1] and A1[2] are stored adjacently. The vector data A2 is quantized to obtain quantized data A2'. The first bit of quantized data A2' is A2[1], and the second bit of quantized data A2' is A2[2]. A2[1] and A2[2] are stored adjacently. Adjacent storage means that the physical address or logical address of the stored two bits of data are adjacent. The vector data A3 is quantized to obtain quantized data A3'. The first bit of quantized data A3' is A3[1], and the second bit of quantized data A3' is A3[2]. A3[1] and A3[2] are stored adjacently. The vector data A4 is quantized to obtain quantized data A4'. The first bit of quantized data A4' is A4[1], and the second bit of quantized data A4' is A4[2]. A4[1] and A4[2] are stored adjacently.
[0094] Storing quantized data, which includes individual bits of data, can reduce the resource consumption of storing data.
[0095] Example 2: Parallel quantization of multiple vector data at once.
[0096] The vector data is compared with N quantization reference values in a preset order to obtain N-bit quantized data after quantization, including: comparing X vector data with N quantization reference values in a preset order to obtain N X-bit data; and transposing the N X-bit data to obtain X N-bit quantized data after quantization of the X vector data.
[0097] In Example 2, obtaining the vector data to be quantized includes: retrieving X vector data points from a vector database, where X is a positive integer greater than 1. The value of X can be set based on experience, user requirements, or the characteristics of the vector data, processor performance, or memory performance. For example, X can be the number of bits that a single memory storage unit can store; if a storage unit can store a maximum of 8 bits, then X can be set to 8. Alternatively, X can be the number of bits that the processor can process; for example, if the minimum number of bits a processor can process is 4, then X can be 4 or a multiple of 4, such as 8, 16, etc. Vector data consists of data that makes up a vector. A vector includes 4 dimensions, meaning it includes 4 elements, with each element representing a vector data point. Therefore, setting X to 4 ensures that vector data points belonging to the same vector are quantized in the same batch, preserving the concept of vector dimensions.
[0098] X vector data are compared with N quantization reference values in a preset order to obtain N X-bit data, including: comparing the size of the X vector data with the i-th judgment data, and determining X third data based on the comparison results. Each third data consists of X identical bits, where X is a positive integer greater than 1; performing bitwise AND operations on the X third data with the reference data of the same order among the X reference data to obtain X fourth data. Each reference data consists of X bits. The j-th bit of the j-th reference data is the fifth data, and the bits other than the j-th bit are the sixth data. The fifth and sixth data are different, where j is a positive integer less than or equal to X; performing a bitwise sum operation on the X fourth data to obtain the i-th X-bit data corresponding to the X vector data.
[0099] For example, in the process of quantizing X vector data at one time, the quantization of X vector data can be achieved through the following steps 21 to 24 to obtain the data of each bit of the X quantized data, that is, N X-bit data.
[0100] Step 21: Compare the magnitudes of the X vector data with the i-th quantization reference value to obtain the comparison results.
[0101] For example, referring to Figure 8, which shows a partial quantization process according to an embodiment of this application, X is 8, and the values of the 8 vector data to be quantized in the same batch are -3, 2, 0, -2, -2, 2, 0, and -2, respectively. Each of the 8 vector data includes 16 bits, that is, M is 16. N is 2, that is, each quantized data includes 2 bits, and the number of quantization reference values is 2. The first quantization reference value is -1, and the second quantization reference value is 1.
[0102] Comparing the first vector data -3 with the first quantization reference value -1, the result is that the first vector data -3 is less than the first quantization reference value -1. Comparing the first vector data -3 with the second quantization reference value 1, the result is also that the first vector data -3 is less than the second quantization reference value 1. Comparing the second vector data 2 with the first quantization reference value -1, the result is that the second vector data 2 is greater than the first quantization reference value -1. Comparing the second vector data 2 with the second quantization reference value 1, the result is also that the second vector data 2 is greater than the second quantization reference value 1. This process continues, resulting in 16 comparison results for each of the 8 vector data points compared to the two quantization reference values. These 16 comparison results are divided into two types: the first type, where the vector data is less than or equal to the quantization reference value, can be set to true; the second type, where the vector data is greater than the quantization reference value, can be set to false.
[0103] Step 22: Based on the comparison results, determine the third data corresponding to each vector data.
[0104] Specifically, a third data point is determined based on a comparison result. X comparison results between X vector data points and the i-th quantization reference value can determine X third data points. Therefore, for X bits of data before quantization and N quantization reference values, a total of X*N third data points can be determined.
[0105] For example, when any vector data is less than or equal to the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all the first data; when any vector data is greater than the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all the second data, and the second data is different from the first data.
[0106] Referring to Figure 8, the comparison result between the first vector data -3 and the first quantization reference value -1 is the first type of comparison result, meaning the first vector data -3 is less than the first quantization reference value -1. Therefore, the X bits in the third data determined based on the first comparison result of the first vector data are all first data. Taking the first data as 1 as an example, the first third data corresponding to the first vector data -3 is 11111111, a total of 8 bits of 1, which is 255 in decimal. The comparison result between the first vector data -3 and the second quantization reference value -1 is also the first type of comparison result, meaning the first vector data -3 is less than the second quantization reference value 1. Therefore, the X bits in the third data determined based on the second comparison result of the first vector data are all first data 1. The second third data corresponding to the first vector data -3 is 11111111, a total of 8 bits of 1, which is 255 in decimal. The comparison result between the second vector data 2 and the first quantization reference value -1 is the second comparison result, meaning the second vector data 2 is greater than the first quantization reference value -1. Therefore, the X bits in the third data determined based on the first comparison result of the second vector data are all the second data. Taking the second data as 0 as an example, the first third data corresponding to the second vector data 2 is 00000000, a total of 8 bits of 0, which is 0 in decimal. The comparison result between the second vector data 2 and the second quantization reference value 1 is also the second comparison result, meaning the second vector data 2 is less than the second quantization reference value 1. Therefore, the X bits in the third data determined based on the second comparison result of the second vector data are all the second data 0. The second third data corresponding to the second vector data 2 is 00000000, a total of 8 bits of 0, which is 0 in decimal. Following this pattern, a total of 16 third data points are obtained. The third data point corresponding to the first comparison result is 11111111, and the third data point corresponding to the second comparison result is 00000000. The decimal representation of the 16 third data points is shown in Figure 8. The 8 third data points on the left side of Figure 8 are the third data points determined based on the comparison result of the 8 vector data points and the first quantization reference value, and the 8 third data points on the right side are the third data points determined based on the comparison result of the 8 vector data points and the second quantization reference value.
[0107] Step 23: Perform a bitwise AND operation between each of the X third data and the X reference data that are in the same order to obtain X fourth data.
[0108] Each of the X reference data consists of X bits. For example, if X is 8, each reference data consists of 8 bits. Optionally, the j-th bit of the j-th reference data is the fifth data, and the bits excluding the j-th bit are the sixth data. The fifth and sixth data are different, and j is a positive integer less than or equal to X. For example, if the fifth data is 1 and the sixth data is 0, when j is 1, the first bit of the first reference data is 1, and the second to eighth bits are 0, meaning the first reference data is 10000000, which is 128 in decimal. When j is 2, the second bit of the second reference data is 1, and the first, third, and eighth bits are all 0, meaning the second reference data is 01000000, which is 64 in decimal; and so on.
[0109] In this embodiment, if the processor stores and reads the data bit by bit in reverse order, then in another optional method, the data of the (X-j+1)th bit of the j-th reference data is the fifth data, and the data of the remaining bits is the sixth data. The fifth data and the sixth data are different, and j is a positive integer less than or equal to X. Taking the fifth data as 1 and the sixth data as 0 as an example, when j is 1, the data of the 8th bit of the 1st reference data is 1, and the data of the 1st to 7th bits are 0, that is, the 1st reference data is 00000001, which is 1 in decimal representation; when j is 2, the data of the 7th bit of the 2nd reference data is 1, and the data of the 1st to 6th bits and the 8th bit are 0, that is, the 2nd reference data is 00000010, which is 2 in decimal representation; and so on.
[0110] Here, we will temporarily take the j-th bit of the j-th reference data as the fifth data 1, and the bits other than the j-th bit as the sixth data 0, to explain the content of step 23. X third data obtained by comparing X vector data with the i-th quantization reference value are respectively bitwise ANDed with the reference data of the same order among the X reference data to obtain X fourth data.
[0111] Referring to Figure 8, taking i as 1 as an example, the eight vector data points compared with the first quantization reference value result in eight third data points, which are represented in decimal as 255, 0, 0, 255, 255, 0, 0, 255. A bitwise AND operation is performed between the first third data point (255) and the first reference data. The first third data point has all eight bits set to 1, resulting in 11111111. The first reference data point has the first bit set to 1 and the remaining bits set to 0, resulting in 10000000. The resulting fourth data point (the first of the eight fourth data points) is 10000000. Perform a bitwise AND operation between the second third data (0) and the second reference data. The data of the second third data is 0 for all 8 bits, so the second third data is 00000000. The data of the second reference data is 1 for the second bit and 0 for the remaining bits, so the second reference data is 01000000. After performing the bitwise AND operation between the second third data and the second reference data, the resulting fourth data (the second fourth data among the 8 fourth data) is 00000000. Perform a bitwise AND operation between the third third data (0) and the third reference data. The data of the third third data is 0 for all 8 bits, so the third third data is 00000000. The data of the third reference data is 1 for the third bit and 0 for the remaining bits, so the third reference data is 00100000. After performing the bitwise AND operation between the third third data and the third reference data, the resulting fourth data (the third fourth data among the eight fourth data) is 00000000. Perform a bitwise AND operation between the fourth third data (255) and the fourth reference data. The fourth third data will have all 8 bits set to 1, resulting in 11111111. The fourth reference data will have 1 bit set to 1 and the remaining bits set to 0, resulting in 00010000. The resulting fourth data (the fourth of the eight fourth data) will then be 00010000. This process is repeated to obtain the eight fourth data corresponding to the eight vector data and the first quantization reference value.
[0112] Step 24: Perform a bitwise summation operation on the X fourth data to obtain N X-bit data corresponding to the X vector data.
[0113] The i-th X-bit data (the X-bit data obtained by bitwise summation of the X fourth data corresponding to the X vector data and the i-th quantization reference value) includes X bits, and the j-th bit of the i-th X-bit data is the i-th bit of the j-th quantization data.
[0114] Referring to Figure 8, the eight fourth data points corresponding to the eight vector data points and the first quantization reference value are 10000000, 00000000, 00000000, 00010000, 00001000, 00000000, 00000000, and 00000001. Performing a bitwise summation on these eight fourth data points yields an X-bit data point of 10011001, which is 153 in decimal. This X-bit data point is determined based on the comparison between the eight vector data points and the first quantization reference value; therefore, this X-bit data point is the first X-bit data point corresponding to the eight vector data points. According to steps 23 and 24, the X-bit data determined by the comparison result of the three vector data and the second quantization reference value is 10111011, that is, the second X-bit data corresponding to the eight vector data is 10111011, which is 187 in decimal representation.
[0115] The first bit of the first X-bit data is the first bit of the first quantized data; the second bit of the first X-bit data is the first bit of the second quantized data; and so on, until the eighth bit of the first X-bit data is the first bit of the eighth quantized data. Similarly, the first bit of the second X-bit data is the second bit of the first quantized data; the second bit of the second X-bit data is the second bit of the second quantized data; and so on, until the eighth bit of the second X-bit data is the second bit of the eighth quantized data. Thus, the data for each bit of each quantized data is obtained.
[0116] For example, the code for steps 23 and 24 is as follows:
[0117] uint8 t mask[]={1,2,4,8,16,32,64,128};
[0118] for p in pn{
[0119] / / 1. Compare vector elements with the decision point p.
[0120] uint8x8 t compareA=vcgt s8(inputVector,p);
[0121] / / 2. After performing a bitwise AND operation on the two vectors, sum all the values to obtain the compressed value stored int8.
[0122] dstData[j]=vaddv u8(vand u8(compareA,vld1 u8(mask)));
[0123] }
[0124] In this embodiment of the application, when the comparison results are different, the data included in the third data obtained by quantization are different, so that the data of each bit included in the third data can accurately reflect the relationship between the vector data and the i-th quantization reference value, thereby accurately reflecting the value range of the vector data.
[0125] Since the j-th bit of the j-th reference data differs from the other bits, while the j-th third data contains the same bits, when the j-th reference data and the j-th third data are bitwise ANDed, the j-th bit of the resulting fourth data reflects the data of all the bits in the j-th third data. Furthermore, after performing a bitwise summation on X fourth data data, the j-th bit of the resulting X data reflects the data of all the bits in the j-th third data, thus reflecting the magnitude relationship between the j-th vector data before quantization and the i-th quantization reference value. Parallel quantization of the X vector data improves quantization efficiency.
[0126] After obtaining the data of each bit in the quantized data, the data of each bit in the quantized data can be stored. For example, the storage method can be to store the X bits of data sequentially, that is, to store the data of the first bit of the first quantized data to the first bit of the Xth quantized data, to store the data of the second bit of the first quantized data to the second bit of the Xth quantized data, and so on. The data of each bit in the quantized data can be stored in memory connected to the processor.
[0127] Referring to Figure 9, a schematic diagram of quantization and storage provided in an embodiment of this application is shown. The vector data includes four data A1, A2, A3 and A4, where X is 4. The process of storing the data of each bit included in the quantized data can also be called the process of encoding the data of each bit included in the quantized data. The storage result can also be regarded as the encoding result. Taking N as 2 as an example, the vector data A1 to vector data A4 are quantized in the same batch. The first X-bit data obtained includes the data of the first bit of the first quantized data A1[1], the data of the first bit of the second quantized data A2[1], the data of the first bit of the third quantized data A3[1], and the data of the first bit of the fourth quantized data A4[1]. A1[1], A2[1], A3[1] and A4[1] are stored adjacently. Adjacent storage means that the physical address or logical address of the data of two bits are adjacent. The obtained second X-bit data includes the second bit of the first quantized data A1[2], the second bit of the second quantized data A2[2], the second bit of the third quantized data A3[2], and the second bit of the fourth quantized data A4[2]. A1[2], A2[2], A3[2], and A4[2] are stored adjacently. Adjacent storage means that the physical or logical addresses of the two bits of data are adjacent.
[0128] Storing quantized data, which includes individual bits of data, can reduce the resource consumption of storing data.
[0129] After obtaining N X-bit data points, these N X-bit data points can be transposed to obtain X N-bit data points, which are the X quantized data points corresponding to the X vector data points. These quantized data points can then be used for further processing. Such processing could include distance calculations in data retrieval scenarios.
[0130] In a data retrieval scenario, the data processing method provided in this application embodiment further includes: obtaining a vector retrieval request, the vector retrieval request including retrieval vector data; comparing the retrieval vector data with N quantization benchmark values in a preset order to obtain N-bit quantized retrieval data after quantization of the retrieval vector data; determining the target value range corresponding to the retrieval quantization data based on the distance between the retrieval quantization data and N+1 types of quantization data in the vector database; and comparing the retrieval vector data with vector data in the target value range to determine the target data corresponding to the retrieval vector data.
[0131] In this context, retrieval vector data can be data generated during business execution or user input, used to find the data needed to achieve the business objectives within the vector data. Retrieval vector data can be the data that makes up a vector, or a retrieval vector can be a portion of data within a vector structure.
[0132] This application does not limit the method of obtaining the vector retrieval request. For example, the vector retrieval request can be input by the user, sent to the processor by other devices, or automatically generated by the processor when performing a task.
[0133] The process of comparing the retrieval vector data with N quantization reference values in a preset order to obtain N-bit quantized retrieval vector data is the process of quantizing the retrieval vector data. The principle of this process can be found in Figure 6 and the explanation of Figure 6, which will not be repeated here.
[0134] Since there are multiple retrieval vector data, in this embodiment of the application, quantization can be performed on one retrieval vector data in one quantization process, or quantization can be performed on multiple retrieval vector data in parallel in one quantization process. The following will use Example 3 and Example 4 as examples to illustrate the two quantization methods of quantizing one retrieval vector data at a time and quantizing multiple retrieval vector data in parallel at a time.
[0135] Example 3: Quantize one retrieval vector at a time.
[0136] The retrieval vector data is compared with N quantization reference values in a preset order to obtain N-bit quantized retrieval data after quantization of the retrieval vector data. This includes: comparing the retrieval vector data with the i-th quantization reference value, and determining the i-th bit of data included in the quantized retrieval data based on the comparison result; wherein, when the retrieval vector data is less than or equal to the i-th quantization reference value, the i-th bit of data in the quantized retrieval data is the first data; when the retrieval vector data is greater than the i-th quantization reference value, the i-th bit of data in the quantized retrieval data is the second data, and the second data is different from the first data.
[0137] For example, the first data is 1, the second data is 0, the retrieval vector data includes 16 bits (i.e., M is 16), the value of the retrieval vector data is 3, N is 2, the first quantization reference value is 0, and the second quantization reference value is 1; comparing the size of the retrieval vector data 3 with the first quantization reference value 0, the retrieval vector data 3 is greater than the first quantization reference value 0, so the first bit of the retrieval quantization data is the second data 0; comparing the size of the retrieval vector data 3 with the second quantization reference value 1, the retrieval vector data 3 is greater than the second quantization reference value 1, so the second bit of the retrieval quantization data is the second data 0, that is, the obtained retrieval quantization data is 00.
[0138] In the example above, different quantization reference values correspond to the same first data, and different quantization reference values correspond to the same second data. This ensures that the XOR distance between two retrieval vector data belonging to adjacent subdomains is 1 after quantization. For example, if retrieval vector data 3 belongs to a subdomain greater than the first quantization reference value and greater than the second quantization reference value, then the value of retrieval vector data belonging to an adjacent subdomain of retrieval vector data 3 should belong to a subdomain greater than the first quantization reference value and less than the second quantization reference value, that is, a subdomain with a value greater than 0 and less than 1. Taking retrieval vector data 0.5, which belongs to an adjacent subdomain of retrieval vector data 3, as an example, following the example above, quantize retrieval vector data 0.5. Compare retrieval vector data 0.5 with the first quantization reference value 0. Since retrieval vector data 0.5 is greater than the first quantization reference value 0, the first bit of the quantized data is the second data 0. Compare retrieval vector data 0.5 with the second quantization reference value 1. Since retrieval vector data 0.5 is less than the second quantization reference value 1, the second bit of the quantized data is the first data 1, that is, the obtained quantized data is 01. Therefore, the XOR distance between quantized data 01 and quantized data 00 is 1, that is, the number of different bits between quantized data 01 and quantized data 00 is 1.
[0139] Optionally, the first and second data corresponding to different quantization reference values can also be different. For example, the first data corresponding to the first quantization reference value is 1, and the second data is 0; the first data corresponding to the second quantization reference value is 0, and the second data is 1. Taking the value of the retrieval vector data as -3, N as 2, the first quantization reference value as 4, and the second quantization reference value as -4 as an example, during the quantization process, the retrieval vector data -3 is compared with the first quantization reference value 4. The comparison result is that the retrieval vector data -3 is less than the first quantization reference value 4. Therefore, the data of the first bit included in the retrieval quantization data is the first data 1 corresponding to the first quantization reference value. The retrieval vector data -3 is compared with the second quantization reference value -4. The comparison result is that the retrieval vector data -3 is greater than the second quantization reference value -4. Therefore, the data of the second bit included in the retrieval quantization data is the second data 1 corresponding to the second quantization reference value. That is, the obtained retrieval quantization data is 11.
[0140] The data for each bit determined based on different comparison results varies. Each bit indicates the relationship between the retrieval vector data and each quantization reference value, ensuring that the data for each bit included in the retrieval quantization data accurately indicates the value range of the retrieval vector data. Determining the data for each bit included in the retrieval quantization data based on the comparison results reduces the computational load of the quantization process and improves quantization efficiency.
[0141] After quantizing the retrieval quantized data to obtain the data of each bit, the data of each bit can be stored. For example, the storage method can be sequential storage, that is, storing the data of the first to Nth bits of the first retrieval quantized data, the data of the first to Nth bits of the second retrieval quantized data, and so on. The data of each bit of the retrieval quantized data can be stored in memory connected to the processor. The process of quantizing the retrieval vector data and storing the retrieval quantized data is shown in Figure 7 and its description, and will not be repeated here. Storing the data of each bit of the retrieval quantized data can reduce the resource consumption of storing query data.
[0142] Example 4: Parallel quantization of multiple retrieval vector data at once.
[0143] The process involves comparing X retrieval vector data with N quantization benchmark values in a preset order to obtain N-bit quantized retrieval data after quantization of the retrieval vector data. This includes: comparing X retrieval vector data with N quantization benchmark values in a preset order to obtain N X-bit retrieval data; and transposing the N X-bit retrieval data to obtain X N-bit quantized retrieval data after quantization of the X retrieval vector data.
[0144] X retrieval vector data are compared with N quantization reference values in a preset order to obtain N retrieval X-bit data. This includes: comparing the size of the X retrieval vector data with the i-th quantization reference value, and determining X third data based on the comparison results. Each third data consists of X identical bits, where X is a positive integer greater than 1. The X third data are then bitwise ANDed with the X reference data in the same order to obtain X fourth data. Each reference data consists of X bits. The j-th bit of the j-th reference data is the fifth data, and the bits other than the j-th bit are the sixth data. The fifth and sixth data are different, where j is a positive integer less than or equal to X. Finally, the X fourth data are bitwise summed to obtain the i-th retrieval X-bit data corresponding to the X retrieval vector data.
[0145] The value of X can be set based on experience, user needs, or the characteristics of the retrieval vector data, processor performance, or memory performance. For example, X can be the number of bits that a single memory cell can store; if a cell can store a maximum of 8 bits, X can be set to 8. Alternatively, X can be the number of bits that the processor can process; if the minimum number of bits a processor can process is 4, X can be 4 or a multiple of 4, such as 8, 16, etc. Furthermore, if the retrieval vector data consists of vectors, each with four dimensions (four elements), and each element representing a retrieval vector data point, then X can be set to 4. This ensures that retrieval vector data points belonging to the same vector are quantized in the same batch, preserving the concept of vector dimensions.
[0146] In the process of quantizing X retrieval vector data, the quantization of X retrieval vector data can be achieved through the following steps 31 to 34, thereby obtaining the data of each bit included in the X quantized retrieval data.
[0147] Step 31: Compare the size of each of the X retrieval vector data with the i-th quantization benchmark value to obtain the comparison results.
[0148] For example, referring to Figure 8, which illustrates a quantization process according to an embodiment of this application, X is 8, and the values of the 8 retrieval vector data quantized in the same batch are -3, 2, 0, -2, -2, 2, 0, and -2, respectively. Each of the 8 retrieval vector data includes 16 bits, that is, M is 16. N is 2, that is, each quantized retrieval data includes 2 bits, and the number of quantization reference values is 2. The first quantization reference value is -1, and the second quantization reference value is 1.
[0149] Comparing the first search vector data -3 with the first quantization reference value -1, the result is that the first search vector data -3 is less than the first quantization reference value -1. Comparing the first search vector data -3 with the second quantization reference value 1, the result is also that the first search vector data -3 is less than the second quantization reference value 1. Comparing the second search vector data 2 with the first quantization reference value -1, the result is that the second search vector data 2 is greater than the first quantization reference value -1. Comparing the second search vector data 2 with the second quantization reference value 1, the result is also that the second search vector data 2 is greater than the second quantization reference value 1. This process continues, resulting in comparisons of 8 search vector data with 2 quantization reference values, for a total of 16 comparison results. These 16 comparison results are divided into two types: the first type, where the search vector data is less than or equal to the quantization reference value, can be set to true; the second type, where the search vector data is greater than the quantization reference value, can be set to false.
[0150] Step 32: Based on the comparison results, determine the third data corresponding to each retrieval vector data.
[0151] Specifically, a third data point is determined based on a comparison result. X comparison results between X retrieval vector data and the i-th quantization reference value can determine X third data points. Therefore, for X bits of data before quantization and N quantization reference values, a total of X*N third data points can be determined.
[0152] For example, when any query data before quantization is less than or equal to the i-th quantization reference value, the X bits of the third data corresponding to any query data before quantization are all the first data; when any query data before quantization is greater than the i-th quantization reference value, the X bits of the third data corresponding to any query data before quantization are all the second data, and the second data is different from the first data.
[0153] Referring to Figure 8, the comparison result between the first retrieval vector data -3 and the first quantization reference value -1 is the first type of comparison result, meaning the first retrieval vector data -3 is less than the first quantization reference value -1. Therefore, the X bits in the third data determined based on the first comparison result of the first retrieval vector data are all first data. Taking the first data as 1 as an example, the first third data corresponding to the first retrieval vector data -3 is 11111111, a total of 8 bits of 1, which is 255 in decimal. The comparison result between the first retrieval vector data -3 and the second quantization reference value -1 is also the first type of comparison result, meaning the first retrieval vector data -3 is less than the second quantization reference value 1. Therefore, the X bits in the third data determined based on the second comparison result of the first retrieval vector data are all first data 1. The second third data corresponding to the first retrieval vector data -3 is 11111111, a total of 8 bits of 1, which is 255 in decimal. The comparison result between the second search vector data 2 and the first quantization reference value -1 is the second comparison result, meaning the second search vector data 2 is greater than the first quantization reference value -1. Therefore, the X bits in the third data determined based on the first comparison result of the second search vector data are all the second data. Taking the second data as 0 as an example, the first third data corresponding to the second search vector data 2 is 00000000, a total of 8 bits of 0, which is 0 in decimal. The comparison result between the second search vector data 2 and the second quantization reference value 1 is also the second comparison result, meaning the second search vector data 2 is less than the second quantization reference value 1. Therefore, the X bits in the third data determined based on the second comparison result of the second search vector data are all the second data 0. The second third data corresponding to the second search vector data 2 is 00000000, a total of 8 bits of 0, which is 0 in decimal. Similarly, a total of 16 third data points are obtained. The third data point corresponding to the first comparison result is 11111111, and the third data point corresponding to the second comparison result is 00000000. The decimal representation of the 16 third data points is shown in Figure 8. The 8 third data points on the left side of Figure 8 are the third data points determined based on the comparison results of the 8 retrieval vector data points and the first quantization benchmark value, and the 8 third data points on the right side are the third data points determined based on the comparison results of the 8 retrieval vector data points and the second quantization benchmark value.
[0154] Step 33: Perform a bitwise AND operation between each of the X third data and the X reference data that are in the same order to obtain X fourth data.
[0155] Each of the X reference data consists of X bits. For example, if X is 8, each reference data consists of 8 bits. Optionally, the j-th bit of the j-th reference data is the fifth data, and the bits excluding the j-th bit are the sixth data. The fifth and sixth data are different, and j is a positive integer less than or equal to X. For example, if the fifth data is 1 and the sixth data is 0, when j is 1, the first bit of the first reference data is 1, and the second to eighth bits are 0, meaning the first reference data is 10000000, which is 128 in decimal. When j is 2, the second bit of the second reference data is 1, and the first, third, and eighth bits are all 0, meaning the second reference data is 01000000, which is 64 in decimal; and so on.
[0156] In this embodiment, if the processor stores and reads the data bit by bit in reverse order, then in another optional method, the data of the (X-j+1)th bit of the j-th reference data is the fifth data, and the data of the remaining bits is the sixth data. The fifth data and the sixth data are different, and j is a positive integer less than or equal to X. Taking the fifth data as 1 and the sixth data as 0 as an example, when j is 1, the data of the 8th bit of the 1st reference data is 1, and the data of the 1st to 7th bits are 0, that is, the 1st reference data is 00000001, which is 1 in decimal representation; when j is 2, the data of the 7th bit of the 2nd reference data is 1, and the data of the 1st to 6th bits and the 8th bit are 0, that is, the 2nd reference data is 00000010, which is 2 in decimal representation; and so on.
[0157] Here, we will temporarily take the j-th bit of the j-th reference data as the fifth data 1, and the bits other than the j-th bit as the sixth data 0, to explain the content of step 23. The X third data obtained by comparing the X retrieval vector data with the i-th quantization reference value are respectively subjected to bitwise AND operations with the reference data of the same order among the X reference data to obtain X fourth data.
[0158] Referring to Figure 8, taking i as 1 as an example, the eight third data obtained by comparing the eight retrieval vector data with the first quantization reference value are represented in decimal as 255, 0, 0, 255, 255, 0, 0, 255. A bitwise AND operation is performed between the first third data (255) and the first reference data. The first third data has 8 bits set to 1, resulting in 11111111. The first reference data has 1 bit set to 1 and the remaining bits set to 0, resulting in 10000000. After performing the bitwise AND operation between the first third data and the first reference data, the resulting fourth data (the first of the eight fourth data) is 10000000. Perform a bitwise AND operation between the second third data (0) and the second reference data. The data of the second third data is 0 for all 8 bits, so the second third data is 00000000. The data of the second reference data is 1 for the second bit and 0 for the remaining bits, so the second reference data is 01000000. After performing the bitwise AND operation between the second third data and the second reference data, the resulting fourth data (the second fourth data among the 8 fourth data) is 00000000. Perform a bitwise AND operation between the third third data (0) and the third reference data. The data of the third third data is 0 for all 8 bits, so the third third data is 00000000. The data of the third reference data is 1 for the third bit and 0 for the remaining bits, so the third reference data is 00100000. After performing the bitwise AND operation between the third third data and the third reference data, the resulting fourth data (the third fourth data among the eight fourth data) is 00000000. Perform a bitwise AND operation between the fourth third data (255) and the fourth reference data. The fourth third data will have all 8 bits set to 1, resulting in 11111111. The fourth reference data will have 1 bit set to 1 and the remaining bits set to 0, resulting in 00010000. The resulting fourth data (the fourth of the eight fourth data) is 00010000. This process is repeated to obtain the eight fourth data corresponding to the eight retrieval vector data and the first quantization reference value.
[0159] Step 34: Perform a bitwise summation operation on the X fourth data to obtain the X retrieval data corresponding to the X retrieval vector data.
[0160] The i-th X-bit data (the X-bit data obtained by bitwise summation of the X fourth data corresponding to the X retrieval vector data and the i-th quantization reference value) includes X bits, and the j-th bit of the i-th X-bit data is the i-th bit of the j-th retrieval quantization data.
[0161] In this embodiment of the application, when the comparison results are different, the data included in the third data obtained by quantization are different, so that the data of each bit included in the third data can accurately reflect the relationship between the size of the retrieval vector data and the i-th quantization reference value, thereby accurately reflecting the value range of the retrieval vector data.
[0162] Since the j-th bit of the j-th reference data differs from the other bits, while the j-th third data contains the same bits, when the j-th reference data and the j-th third data are bitwise ANDed, the j-th bit of the resulting fourth data reflects the data of all the bits in the j-th third data. Furthermore, after bitwise summing X fourth data data, the j-th bit of the resulting X data reflects the data of all the bits in the j-th third data, thus reflecting the relationship between the j-th query data and the i-th quantization reference value before quantization. Parallel quantization of the X retrieval vector data improves quantization efficiency.
[0163] After determining the quantized data to be retrieved, the processor determines the target value range corresponding to the quantized data based on the distance between the quantized data and the N+1 quantized data in the vector database. The method by which the processor determines the distance between the quantized data and the N+1 quantized data in the vector database is not limited. Below, examples five and six illustrate two methods for determining the distance between the quantized data and the N+1 quantized data; the different methods result in different types of distances.
[0164] Example 5: Optionally, the distance between the quantized data and the retrieved quantized data can be a Euclidean-like distance. Euclidean distance, also known as Euclidean distance, measures the distance between two points in two-dimensional, three-dimensional, or higher-dimensional space. The Euclidean-like distance in this embodiment measures the distance between two points in one-dimensional space; that is, the distance between two data points is the distance between their values on a one-dimensional number axis. Since the quantization process in this application is based on the comparison results determined by N quantization benchmark values used for value range division, vector data, and retrieved vector data, the data of each bit included in the quantized data or retrieved quantized data obtained based on the comparison results can reflect the subdomain to which the vector data or retrieved vector data belongs, that is, the position of the vector data on the number axis and the position of the retrieved vector data on the number axis. Therefore, the Euclidean-like distance between the quantized data and the retrieved quantized data can reflect the similarity between the retrieved vector data and the vector data.
[0165] When the distance between the quantized data and the retrieved quantized data is a Euclidean distance, before determining the target value range of the retrieved quantized data based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database, the following steps are also included: for any type of quantized data, perform a bitwise AND operation on the N bits of data included in the quantized data and the N bits of data included in the retrieved quantized data to obtain N AND operation results; sum the N AND operation results to obtain the distance between the quantized data and the retrieved quantized data.
[0166] For example, if the quantized data is 101 and the retrieved quantized data is 110, the first bit of the quantized data (1) is ANDed with the first bit of the retrieved quantized data (1), resulting in a first AND operation of 1. The second bit of the quantized data (0) is ANDed with the second bit of the retrieved quantized data (1), resulting in a second AND operation of 0. The third bit of the quantized data (1) is ANDed with the third bit of the retrieved quantized data (0), resulting in a third AND operation of 0. Therefore, the distance between the quantized data and the retrieved quantized data is 1 + 0 + 0 = 1.
[0167] Example 6: The distance between quantized data and retrieved quantized data is the Hamming distance. The Hamming distance represents the number of different bits in the same data segments of two datasets. In this embodiment, after obtaining the comparison result during quantization, the data for each bit of the quantized data or retrieved quantized data is determined using Hamming encoding. The closer the values of the two unquantized data are, the smaller the difference in corresponding bits between the quantized data obtained from these two unquantized data. Therefore, the Hamming distance between the two quantized data segments can also reflect the similarity between the two datasets. Thus, calculating the Hamming distance can accurately determine the distance between the quantized data and the retrieved quantized data.
[0168] When the distance between the quantized data and the retrieved quantized data is the Hamming distance, before determining the target value range of the retrieved quantized data based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database, the following steps are also included: for any type of quantized data, perform a bitwise XOR operation on the N bits of data included in the quantized data and the N bits of data included in the retrieved quantized data to obtain N XOR operation results; sum the N XOR operation results to obtain the distance between the quantized data and the retrieved quantized data.
[0169] For example, if the quantized data is 101 and the retrieved quantized data is 110, the first bit of the quantized data (1) is XORed with the first bit of the retrieved quantized data (1), resulting in a first XOR result of 0. The second bit of the quantized data (0) is XORed with the second bit of the retrieved quantized data (1), resulting in a second XOR result of 1. The third bit of the quantized data (1) is XORed with the third bit of the retrieved quantized data (0), resulting in a third XOR result of 1. Therefore, the Hamming distance between the quantized data and the retrieved quantized data is 0 + 1 + 1 = 2.
[0170] Referring to Figure 10, a schematic diagram of a data processing procedure provided in an embodiment of this application is shown, where the data type is an X-dimensional vector. The vector data of type fp32 / int8 / int16 stored in the vector database is quantized to obtain quantized (compressed) quantized data (vector encoded data) including N bits, totaling N*X bits; the query vector data of type fp32 / int8 / int16 in the query vector is quantized to obtain quantized (compressed) retrieval quantized data (vector encoded data) including N bits, totaling N*X bits; the distance between the retrieval quantized data and the quantized data is calculated to obtain the distance result.
[0171] After determining the distance between the retrieved quantized data and N+1 other quantized data, the processor determines the target value range corresponding to the retrieved quantized data based on the determined distance. Optionally, the target value range can be the value range of one or more quantized data that has the smallest distance to the retrieved quantized data, or it can be the value range of one or more quantized data that has a distance to the retrieved quantized data that is less than a distance threshold.
[0172] For example, for the same distance type, N+1 distances can be determined between a retrieved quantized data and N+1 other quantized data. In this case, determining the target value range corresponding to the retrieved quantized data based on the distances between the retrieved quantized data and the N+1 other quantized data in the vector database includes: determining the quantized data with the smallest distance from the retrieved quantized data among the N+1 quantized data based on the distances between the retrieved quantized data and the N+1 other quantized data; and determining the value range of the quantized data with the smallest distance from the retrieved quantized data as the target value range. For example, the N+1 distances between the retrieved quantized data and the N+1 other quantized data can be sorted to determine the smallest distance among the N+1 distances, and the quantized data corresponding to the smallest distance can be determined. The value range of the quantized data corresponding to the smallest distance can then be determined as the target value range.
[0173] Since the distance between the quantized data in the target value range and the retrieved quantized data is close in the scheme described in the above embodiments, but the target value range includes multiple values, and the distance between the vector data corresponding to the target value range and the retrieved vector data before quantization is also multiple, in this embodiment, the retrieved vector data can be compared with the vector data in the target value range to determine the target data corresponding to the retrieved vector data. That is, the vector data belonging to the target value range is further filtered to more accurately determine the vector data that is close to the retrieved vector data.
[0174] Optionally, comparing the retrieved vector data with the vector data in the target value domain can be done by determining the similarity or distance between the retrieved vector data and the vector data in the target value domain. Distance and similarity are negatively correlated; that is, the greater the distance between the retrieved vector data and the retrieved vector data, the lower the similarity, and the smaller the distance, the higher the similarity. The distance between the retrieved vector data and the vector data in the target value domain can be either Euclidean distance or Hamming distance.
[0175] This application does not limit the method for determining the similarity between the retrieved vector data and the vector data in the target threshold. For example, the reciprocal of the distance between any given vector data and the retrieved vector data can be used as the similarity between them, making distance and similarity negatively correlated. Alternatively, the minimum distance between each vector data in the target value range and the retrieved vector data can be used as the numerator, and the distance between any given vector data and the retrieved vector data as the denominator. The determined value is the similarity between that given vector data and the retrieved vector data. Thus, the larger the distance between any given vector data and the retrieved vector data, the smaller the determined similarity, making distance and similarity negatively correlated.
[0176] After determining the similarity or distance between each vector data and the retrieval vector data, vector data with a similarity greater than a similarity threshold or a distance less than a specified distance threshold can be identified as target data, thereby achieving accurate data retrieval.
[0177] In summary, the N quantization reference values in this application form N+1 value ranges. By comparing the vector data with the N quantization reference values, N comparison results are obtained. Based on these N comparison results, the N bits of data included in the quantization data are determined, thus obtaining N-bit quantized data and achieving the quantization of vector data. The quantization method of this application is simple, reducing the computational load and complexity of the quantization process, resulting in lower computational overhead.
[0178] The data processing method provided by the embodiments of this application has been described above. Corresponding to the above method, the embodiments of this application also provide a data processing system. This data processing system is used to execute the method shown in Figure 4 through the various modules shown in Figure 11. As shown in Figure 11, the data processing system provided by the embodiments of this application includes the following modules.
[0179] The acquisition module 1101 is used to acquire N quantization reference values, which are arranged in a preset order to form N+1 value domains, where N is a positive integer greater than 1. The quantization module 1102 is used to acquire the vector data to be quantized, and compare the vector data with the N quantization reference values in a preset order to obtain N-bit quantized data after quantization. Any bit of the N-bit quantized data is determined by the comparison result between the vector data and the quantization reference value that corresponds to any bit of the N quantization reference values in the order of the data. The N-bit data indicates that the vector data belongs to one of the N+1 value domains.
[0180] In one possible implementation, the vector data is vector data in a vector database, and all vector data in the vector database is quantized into quantized data belonging to N+1 value ranges, each value range being indicated by an N-bit quantized data.
[0181] In one possible implementation, the quantization module 1102 is used to compare the size of the vector data with the i-th quantization reference value, and determine the data of the i-th bit included in the quantization data based on the comparison result; wherein, when the comparison result is that the vector data is less than or equal to the i-th quantization reference value, the data of the i-th bit included in the quantization data is the first data; when the comparison result is that the vector data is greater than the i-th quantization reference value, the data of the i-th bit included in the quantization data is the second data, and the second data is different from the first data.
[0182] In one possible implementation, the acquisition module 1101 is used to acquire X vector data from the vector database, where X is a positive integer greater than 1; the quantization module 1102 is used to compare the X vector data with N quantization reference values in a preset order to obtain N X-bit data; and to transpose the N X-bit data to obtain X N-bit quantized data after quantizing the X vector data.
[0183] In one possible implementation, the quantization module 1102 is used to compare the size of X vector data with the i-th judgment data, and determine X third data based on the comparison result. Each third data includes X identical bits, where X is a positive integer greater than 1. The X third data are then bitwise ANDed with the X reference data in the same order to obtain X fourth data. Each reference data includes X bits. The j-th bit of the j-th reference data is the fifth data, and the bits other than the j-th bit are the sixth data. The fifth and sixth data are different, where j is a positive integer less than or equal to X. The X fourth data are then bitwise summed to obtain the i-th X-bit data corresponding to the X vector data.
[0184] In one possible implementation, when the comparison result indicates that any vector data is less than or equal to the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all the first data; when the comparison result indicates that any vector data is greater than the i-th quantization reference value, the X bits of the third data corresponding to any vector data are all the second data, and the second data is different from the first data.
[0185] In one possible implementation, the data processing system further includes a determination module; an acquisition module 1101, which is further configured to acquire a vector retrieval request, the vector retrieval request including retrieval vector data; a quantization module 1102, which is further configured to compare the retrieval vector data with N quantization benchmark values in a preset order to obtain N-bit quantized retrieval data after quantization of the retrieval vector data; and a determination module, which is configured to determine the target value range corresponding to the retrieval quantization data based on the distance between the retrieval quantization data and N+1 types of quantization data in the vector database; and to compare the retrieval vector data with the vector data in the target value range to determine the target data corresponding to the retrieval vector data.
[0186] In one possible implementation, the quantization module 1102 is used to compare the size of the retrieved vector data with the i-th quantization reference value, and based on the comparison result, determine the data of the i-th bit included in the retrieved quantization data; wherein, when the retrieved vector data is less than or equal to the i-th quantization reference value, the data of the i-th bit of the retrieved quantization data is the first data; when the retrieved vector data is greater than the i-th quantization reference value, the data of the i-th bit of the retrieved quantization data is the second data, and the second data is different from the first data.
[0187] In one possible implementation, the data processing system further includes a calculation module, which is used to perform a bitwise AND operation on N bits of any quantized data and N bits of the retrieved quantized data to obtain N AND operation results; and to sum the N AND operation results to obtain the distance between the arbitrary quantized data and the retrieved quantized data.
[0188] In one possible implementation, the data processing system further includes a calculation module, which is used to perform a bitwise XOR operation on N bits of any quantized data and N bits of the retrieved quantized data to obtain N XOR operation results; and to sum the N XOR operation results to obtain the distance between the arbitrary quantized data and the retrieved quantized data.
[0189] In one possible implementation, a determination module is used to determine, based on the distance between N+1 types of quantized data and the retrieved quantized data, the quantized data with the smallest distance from the retrieved quantized data among the N+1 types of quantized data; and to determine the value range of the quantized data with the smallest distance from the retrieved quantized data as the target value range.
[0190] In one possible implementation, the acquisition module 1101 is used to determine the value range of the vector data; divide the value range into N+1 value domains, and the difference between the two data quantities of the vector data belonging to any two value domains is less than the difference threshold; and determine N quantization benchmark values based on the values of the N intersection points of the N+1 value domains.
[0191] It should be understood that the beneficial effects of the data processing system provided in Figure 11 are the same as those of the data processing method provided in Figure 4 when implementing its functions, and will not be repeated here. Furthermore, the data processing system provided in Figure 11 is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the data processing system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.
[0192] Referring to Figure 12, which shows a schematic diagram of the structure of an exemplary data processing device 1200 of this application, the data processing device 1200 includes at least one processor 1201, a memory 1203 and at least one network interface 1204.
[0193] The processor 1201 is, for example, a general-purpose central processing unit (CPU), a digital signal processor (DSP), a network processor (NP), a GPU, a neural-network processing unit (NPU), a data processing unit (DPU), a microprocessor, or one or more integrated circuits or application-specific integrated circuits (ASICs), programmable logic devices (PLDs), other general-purpose processors or other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components, or any combination thereof used to implement the scheme of this application. The PLD is, for example, a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor. It is worth noting that the processor can be a processor supporting an advanced reduced instruction set machine (RISC) machine (ARM) architecture. It can implement or execute the various logic blocks, modules, and circuits described in conjunction with the disclosure of this application. A processor can also be a combination of components that perform computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on.
[0194] Optionally, the data processing device 1200 also includes a bus 1202. The bus 1202 is used to transmit information between the components of the data processing device 1200. The bus 1202 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 1202 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one line is used in Figure 12, but this does not mean that there is only one bus or one type of bus.
[0195] The memory 1203 may be, for example, volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache.
[0196] By way of example, but not limitation, many forms of ROM and RAM are available. For example, ROM is a compact disc read-only memory (CD-ROM). RAM includes, but is not limited to, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0197] The memory 1203 can also be other types of storage devices capable of storing static information and instructions. Alternatively, it can be other types of dynamic storage devices capable of storing information and instructions. It can also be other optical disc storage, optical disk storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. The memory 1203 may exist independently, for example, and be connected to the processor 1201 via bus 1202. The memory 1203 may also be integrated with the processor 1201.
[0198] Network interface 1204 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), or wireless local area network (WLAN). Network interface 1204 may include wired network interfaces and wireless network interfaces. Specifically, network interface 1204 can be an Ethernet interface, such as Fast Ethernet (FE), Gigabit Ethernet (GE), asynchronous transfer mode (ATM), WLAN, cellular network, or combinations thereof. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. In some embodiments of this application, network interface 1204 can be used for data processing device 1200 to communicate with other devices.
[0199] In specific implementations, as some embodiments, processor 1201 may include one or more CPUs, such as CPU0 and CPU1 shown in FIG12. Each of these processors may be a single-core processor or a multi-core processor. Here, processor may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0200] In specific implementations, as some embodiments, the data processing device 1200 may include multiple processors, such as processor 1201 and processor 1205 shown in FIG12. Each of these processors may be a single-core processor or a multi-core processor. Here, a processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0201] In some embodiments, memory 1203 is used to store program instructions 1210 for executing the present application scheme, and processor 1201 can execute the program instructions 1210 stored in memory 1203. That is, data processing device 1200 can implement the method provided in the method embodiment, i.e., the method shown in FIG4, through processor 1201 and program instructions 1210 in memory 1203. Program instructions 1210 may include one or more software modules. Optionally, processor 1201 itself may also store program instructions for executing the present application scheme.
[0202] In specific implementation, the processor 1201 in the data processing device 1200 of this application can correspond to the processor used to execute the above method. The processor 1201 in the data processing device 1200 reads the instructions in the memory 1203 so that the data processing device 1200 shown in FIG12 can execute all or part of the steps in the method embodiment.
[0203] The data processing device 1200 can also correspond to the data processing system shown in Figure 11 above. Each functional module in the data processing system shown in Figure 11 is implemented using software from the data processing device 1200. In other words, the functional modules included in the data processing system shown in Figure 11 are generated by the processor 1201 of the data processing device 1200 reading the program instructions 1210 stored in the memory 1203.
[0204] In the method shown in Figure 4, each step is completed through integrated logic circuits in the hardware or instructions in the software form of the processor in the data processing device 1200. The steps of the method embodiments disclosed in this application can be directly implemented by the hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. Since the storage medium is located in memory, the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method embodiments. To avoid repetition, these steps will not be described in detail here.
[0205] In an exemplary embodiment, a computer program (product) is provided, comprising: computer program code, which, when executed by a computer, causes the computer to perform the method in FIG4.
[0206] In an exemplary embodiment, a computer-readable storage medium is provided that stores a program or instructions, which, when run on a computer, enable the computer to perform the method described in FIG4 above.
[0207] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk).
[0208] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with substantially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or order of execution. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another.
[0209] It should also be understood that, in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0210] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple second devices means two or more second devices. The terms "system" and "network" are often used interchangeably herein.
[0211] It should be understood that the terminology used in the description of the various examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0212] It should also be understood that the term "and / or" as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. The term "and / or" describes an association between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " in this application generally indicates that the preceding and following related objects are in an "or" relationship.
[0213] It should also be understood that the terms “if” and “if” can be interpreted as meaning “when” or “upon”, or “in response to determination” or “in response to detection”. Similarly, depending on the context, the phrases “if determination…” or “if detection [the stated condition or event]” can be interpreted as meaning “when determination…”, or “in response to determination…”, or “when detection [the stated condition or event]” or “in response to detection [the stated condition or event]”.
[0214] The above description is merely an embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
A data processing method, characterized in that, The method includes: Obtain N quantization reference values, which are arranged in a preset order to form N+1 value ranges, where N is a positive integer greater than 1; Obtain the vector data to be quantized, and compare the vector data with the N quantization reference values in the preset order to obtain N-bit quantized data after quantization of the vector data. Each bit of the N-bit quantized data is determined by the comparison result between the vector data and the quantization reference value that corresponds to the order of the N-bit quantized data. The N-bit data indicates that the vector data belongs to one of the N+1 value domains. The method according to claim 1, characterized in that, The vector data is vector data in a vector database. All vector data in the vector database are quantized into quantized data belonging to N+1 value ranges, and each value range is indicated by an N-bit quantized data. The method according to claim 1 or 2, characterized in that, The step of comparing the vector data with the N quantization reference values in the preset order to obtain N-bit quantized data after quantization of the vector data includes: Compare the vector data with the size of the i-th quantization reference value, and determine the data of the i-th bit included in the quantization data based on the comparison result; Wherein, when the comparison result is that the vector data is less than or equal to the i-th quantization reference value, the data of the i-th bit included in the quantization data is the first data; when the comparison result is that the vector data is greater than the i-th quantization reference value, the data of the i-th bit included in the quantization data is the second data, and the second data is different from the first data. The method according to claim 1 or 2, characterized in that, The process of obtaining the vector data to be quantized includes: Obtain X vector data from the vector database, where X is a positive integer greater than 1; The step of comparing the vector data with the N quantization reference values in the preset order to obtain N-bit quantized data after quantization of the vector data includes: The X vector data are compared with the N quantization reference values in the preset order to obtain N X-bit data. The N X-bit data are transposed to obtain X N-bit quantized data after the X vector data are quantized. The method according to claim 4, characterized in that, The step of comparing the X vector data with the N quantization reference values in the preset order to obtain N X-bit data includes: Compare the size of the X vector data with the i-th judgment data, and determine X third data based on the comparison result. Each third data includes X identical data bits, where X is a positive integer greater than 1. The X third data are each bitwise ANDed with the X reference data in the same order to obtain X fourth data. Each reference data includes X bits. The j-th bit of the j-th reference data is the fifth data, and the bits other than the j-th bit are the sixth data. The fifth data and the sixth data are different. j is a positive integer less than or equal to X. Perform a bitwise summation operation on the X fourth data to obtain the i-th X-bit data corresponding to the X vector data. The method according to any one of claims 1-5 is characterized in that, The method further includes: Obtain a vector retrieval request, the vector retrieval request including retrieving vector data; The retrieval vector data is compared with the N quantization benchmark values in the preset order to obtain N-bit retrieval quantization data after quantizing the retrieval vector data. The target value range corresponding to the retrieved quantized data is determined based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database. The retrieval vector data is compared with the vector data in the target value domain to determine the target data corresponding to the retrieval vector data. The method according to claim 6, characterized in that, The step of determining the target value range corresponding to the retrieved quantized data based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database includes: Based on the distance between N+1 types of quantized data and the retrieved quantized data, determine the quantized data with the smallest distance from the retrieved quantized data among the N+1 types of quantized data; The value range to which the quantized data with the smallest distance from the retrieved quantized data belongs is determined as the target value range. The method according to any one of claims 1-7, characterized in that, The acquisition of N quantization benchmark values includes: Determine the value range of the vector data; The value range is divided into N+1 value domains, and the difference between the two data quantities of vector data belonging to any two value domains is less than the difference threshold. The N quantization reference values are determined based on the values of the N intersection points of the N+1 value ranges. A data processing system, characterized in that, The data processing system includes: The acquisition module is used to acquire N quantization reference values, which are arranged in a preset order to form N+1 value ranges, where N is a positive integer greater than 1; The quantization module is used to acquire vector data to be quantized, compare the vector data with the N quantization reference values in the preset order, and obtain N-bit quantized data after quantization of the vector data. Each bit of the N-bit quantized data is determined by the comparison result between the vector data and the quantization reference value that corresponds to the order of the N-bit quantized data. The N-bit data indicates that the vector data belongs to one of the N+1 value domains. The data processing system according to claim 9 is characterized in that, The vector data is vector data in a vector database. All vector data in the vector database are quantized into quantized data belonging to N+1 value ranges, and each value range is indicated by an N-bit quantized data. The data processing system according to claim 9 or 10 is characterized in that, The quantization module is used to compare the vector data with the i-th quantization reference value, and determine the data of the i-th bit included in the quantization data based on the comparison result; wherein, when the comparison result is that the vector data is less than or equal to the i-th quantization reference value, the data of the i-th bit included in the quantization data is first data; when the comparison result is that the vector data is greater than the i-th quantization reference value, the data of the i-th bit included in the quantization data is second data, and the second data is different from the first data. The data processing system according to claim 9 or 10 is characterized in that, The acquisition module is used to acquire X vector data from the vector database, where X is a positive integer greater than 1; The quantization module is used to compare the X vector data with the N quantization reference values in the preset order to obtain N X-bit data. The N X-bit data are transposed to obtain X N-bit quantized data after the X vector data are quantized. The data processing system according to claim 12 is characterized in that, The quantization module is used to compare the size of the X vector data with the i-th judgment data, and determine X third data based on the comparison result. Each third data includes X identical bits, where X is a positive integer greater than 1. The X third data are then bitwise ANDed with the X reference data in the same order to obtain X fourth data. Each reference data includes X bits. The j-th bit of the j-th reference data is the fifth data, and the bits other than the j-th bit are the sixth data. The fifth data and the sixth data are different, where j is a positive integer less than or equal to X. The X fourth data are then bitwise summed to obtain the i-th X-bit data corresponding to the X vector data. The data processing system according to any one of claims 9-13 is characterized in that, The data processing system further includes a determination module; the acquisition module is also used to acquire a vector retrieval request, the vector retrieval request including retrieval vector data; the quantization module is also used to compare the retrieval vector data with the N quantization benchmark values respectively according to the preset order to obtain N-bit quantized retrieval data after quantization of the retrieval vector data. The determining module is used to determine the target value range corresponding to the retrieved quantized data based on the distance between the retrieved quantized data and N+1 types of quantized data in the vector database. The retrieval vector data is compared with the vector data in the target value domain to determine the target data corresponding to the retrieval vector data. The data processing system according to claim 14 is characterized in that, The determining module is used to determine, based on the distance between N+1 types of quantized data and the retrieved quantized data, the quantized data with the smallest distance from the retrieved quantized data among the N+1 types of quantized data; The value range to which the quantized data with the smallest distance from the retrieved quantized data belongs is determined as the target value range. The data processing system according to any one of claims 9-15 is characterized in that, The acquisition module is used to determine the value range of the vector data; divide the value range into N+1 value domains, and the difference between the two data quantities of the vector data belonging to any two value domains is less than the difference threshold; and determine the N quantization benchmark values based on the values of the N intersection points of the N+1 value domains. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement the data processing method as described in any one of claims 1-8. A computer program product, characterized in that, The computer program product includes a computer program or instructions that are executed by a processor to enable a computer to implement the data processing method according to any one of claims 1-8.