Data processing method and device

By filtering the bond feature vectors from the attention mechanism, using Hamming distance and similarity thresholds, redundant calculations are reduced, and the processing efficiency and resource utilization of the model are improved.

CN120387008APending Publication Date: 2025-07-29SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510526309.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When processing data, the model based on the self-attention mechanism has low computational efficiency and consumes a lot of computing resources, mainly due to a large number of multiplication operations and redundant key feature vector calculations.

Method used

By calculating the Hamming distance between the key feature vectors and the query feature vectors, we filter out the key feature vectors with a smaller degree of difference, and only internal product operations are performed on these vectors to reduce redundant calculations.

Benefits of technology

Without affecting the accuracy of the processing results, the calculation amount of the self-attention mechanism is reduced, processing efficiency is improved, and computing resource consumption is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387008A_ABST
    Figure CN120387008A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, and the method comprises the steps: obtaining a key feature vector and a query feature vector, and enabling the key feature vector to be a feature vector obtained through the key weight processing of target data based on a target model containing a self-attention mechanism, the query feature vector is a feature vector obtained by processing the target data based on the query weight of the target model; determining a Hamming distance between each key feature vector and the query feature vector; screening each key feature vector according to the Hamming distance to obtain a screened key feature vector; and processing the screened key feature vector and the query feature vector to obtain an attention weight of the target data so as to obtain a processing result of the target data output by the target model based on the attention weight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular to a data processing method and apparatus. Background Art

[0002] When a model based on the self-attention mechanism processes data, it first extracts the key feature vectors and query feature vectors of the data, and then calculates the key feature vectors and query feature vectors to obtain a processing result.

[0003] Calculating the key feature vectors and query feature vectors requires a large number of operations, resulting in low processing efficiency of the model and consuming more computing resources during processing. Summary of the Invention

[0004] For this reason, the present application discloses the following technical solutions:

[0005] The first aspect of this application provides a data processing method, including:

[0006] Obtain key feature vectors and query feature vectors, where the key feature vectors are feature vectors obtained by processing target data based on the key weights of a target model including a self-attention mechanism, and the query feature vectors are feature vectors obtained by processing target data based on the query weights of the target model;

[0007] Determine the Hamming distance between each of the key feature vectors and the query feature vectors;

[0008] Screen each of the key feature vectors according to the Hamming distance to obtain screened key feature vectors;

[0009] Process the screened key feature vectors and the query feature vectors to obtain the attention weights of the target data, so as to obtain the processing result of the target data output by the target model based on the attention weights.

[0010] Optionally, the screening each of the key feature vectors according to the Hamming distance to obtain screened key feature vectors includes:

[0011] Perform a first screening on each of the key feature vectors according to the Hamming distance;

[0012] Determine the similarity between the key feature vectors after the first screening and the query feature vectors according to the Hamming distance;

[0013] Perform a second screening on each of the key feature vectors after the first screening based on the similarity to obtain screened key feature vectors.

[0014] Optionally, determining the similarity between the key feature vector after the first screening and the query feature vector according to the Hamming distance includes:

[0015] Determining an angle value between the key feature vector after the first screening and the query feature vector according to the Hamming distance;

[0016] Determining the similarity between the key feature vector after the first screening and the query feature vector according to the angle value.

[0017] Optionally, determining the similarity between the key feature vector after the first screening and the query feature vector according to the angle value includes:

[0018] Performing multiple iterative operations on the angle value and a preset initial parameter by a target calculation unit to obtain a cosine value corresponding to the angle value, and using the cosine value as the similarity between the key feature vector after the first screening and the query feature vector;

[0019] Wherein, the input of the first iterative operation includes the angle value and the initial parameter, and the input of each iterative operation after the first iterative operation is the output of the previous iterative operation, and the target calculation unit includes a shifter and multiple adders.

[0020] Optionally, performing a second screening on each of the key feature vectors after the first screening based on the similarity to obtain screened key feature vectors includes:

[0021] Obtaining an inner product value of each of the key feature vectors based on the similarity and the modulus value of each of the key feature vectors after the first screening;

[0022] Performing a second screening according to the inner product values of each of the key feature vectors after the first screening to obtain screened key feature vectors.

[0023] Optionally, the process of obtaining the modulus value of the key feature vector includes:

[0024] Calculating a plurality of partial products corresponding to each component in the key feature vector, and the number of partial products corresponding to each component is less than the number of binary digits of the component;

[0025] Obtaining a squared value corresponding to each component according to the partial products of each component, so as to determine the modulus value of the key feature vector based on the squared values of each component.

[0026] Optionally, it further includes:

[0027] Obtaining a target key feature vector whose corresponding normalized attention weight is less than a preset weight threshold;

[0028] Determine the maximum value among the Hamming distances corresponding to each of the target key feature vectors as the Hamming distance threshold;

[0029] Wherein, the Hamming distance threshold is used for the first screening.

[0030] Optionally, it further includes:

[0031] Obtain the normalized attention weight corresponding to the key feature vector;

[0032] Determine the minimum value among multiple normalized attention weights greater than or equal to a preset weight threshold;

[0033] Determine the similarity threshold according to the minimum value;

[0034] Wherein, the similarity threshold is used for the second screening.

[0035] Optionally, the obtaining the processing result of the target data output by the target model based on the attention weight includes:

[0036] Obtain the value feature vector corresponding to the screened key feature vector, where the value feature vector is a feature vector obtained by processing the target data based on the value weight of the target model;

[0037] Calculate the value feature vector according to the attention weight and perform normalization processing on the calculation result to obtain the processing result of the target data output by the target model.

[0038] The second aspect of the present application provides a data processing device, including:

[0039] An obtaining module, configured to obtain a key feature vector and a query feature vector, where the key feature vector is a feature vector obtained by processing target data based on the key weight of a target model including a self-attention mechanism, and the query feature vector is a feature vector obtained by processing target data based on the query weight of the target model;

[0040] A screening module, configured to:

[0041] Determine the Hamming distance between each of the key feature vectors and the query feature vector;

[0042] Screen each of the key feature vectors according to the Hamming distance to obtain a screened key feature vector;

[0043] A processing module, configured to process the screened key feature vector and the query feature vector to obtain the attention weight of the target data, so as to obtain the processing result of the target data output by the target model based on the attention weight. Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided accompanying drawings.

[0045] Figure 1 is a flowchart of a data processing method provided by an embodiment of the present application;

[0046] Figure 2 is a schematic diagram of a bitwise multiplication and squaring calculation process provided by an embodiment of the present application;

[0047] Figure 3 is a schematic diagram of a first simplification of a bitwise multiplication and squaring calculation process provided by an embodiment of the present application;

[0048] Figure 4 is a schematic diagram of a second simplification of a bitwise multiplication and squaring calculation process provided by an embodiment of the present application;

[0049] Figure 5 is a schematic diagram of the structure of a target calculation unit provided by an embodiment of the present application;

[0050] Figure 6 is a schematic diagram of the structure of cascading multiple target calculation units provided by an embodiment of the present application;

[0051] Figure 7 is a schematic diagram of the RTL synthesis area corresponding to different calculation methods provided by an embodiment of the present application;

[0052] Figure 8 is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application;

[0053] Figure 9 is a schematic diagram of the structure of another data processing device provided by an embodiment of the present application. Detailed implementation manners

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0055] Data processing models based on self-attention mechanisms (such as Transformer neural network models) have currently been widely applied in various fields. However, such models have high computing power requirements, making it difficult to deploy such models on terminal devices (such as mobile phones and personal computers).

[0056] In related technologies, when applying the self-attention mechanism, a query feature vector, multiple key feature vectors (Key), and multiple value feature vectors (Value) can be obtained based on the target data. The key feature vectors and value feature vectors correspond one by one, and then these feature vectors are calculated according to the following formula (1) to obtain the processing result corresponding to the target data based on the calculation result. The target data can be any data that needs to be processed by the model, such as text data to be processed, image data, audio data, or other data.

[0057]

[0058] Among them, Query i represents the query feature vector, output i represents the calculation result of the self-attention mechanism, T represents transpose, Key k represents the k-th key feature vector, Value k represents the corresponding k-th value feature vector, d represents the dimension of the key feature vector, softmax() represents the exponential normalization function, M represents the total number of key feature vectors, and the dimensions of the key feature vector, value feature vector, and query feature vector are the same. Query i Key k T represents the dot product (also known as the inner product) operation on the transpose of the query feature vector and the key feature vector, and the obtained operation result is a real number.

[0059] Analyzing the processing process of the above self-attention mechanism, it can be found that the operations of the self-attention mechanism mainly revolve around the multiplication operations of the query feature vector (Query), key feature vector (Key), and value feature vector (Value) of the target data. In particular, the dot product operation of the transpose of the query feature vector and each key feature vector requires a large number of multiplication operations.

[0060] Furthermore, the above feature vectors are all determined based on the target data, and each of them contains partial information of the target data. Therefore, when performing the dot product operation on the transpose of the query feature vector and multiple key feature vectors, the operation results corresponding to some key feature vectors may be equal to 0 or approximately equal to 0 (such as equal to 0.000001), and only the operation results corresponding to some key feature vectors are greater than 0.

[0061] Obviously, the calculation result equal to 0 or approximately equal to 0 has an impact on the calculation result output of the self-attention mechanism. i There is almost no impact. Therefore, if we can pre-screen those key feature vectors whose operation results are equal to 0 or approximately equal to 0 before performing calculations based on formula (1), we can avoid performing dot product operations on these key feature vectors, thereby reducing the amount of multiplication operations and basically not affecting the accuracy of the calculation results of the self-attention mechanism.

[0062] Based on the above principles, this embodiment provides a data processing method, see Figure 1 , is a flowchart of the method, which may include the following steps.

[0063] S101, obtain a key feature vector and a query feature vector, the key feature vector is a feature vector obtained by processing the target data based on the key weight of the target model including the self-attention mechanism, and the query feature vector is a feature vector obtained by processing the target data based on the query weight of the target model.

[0064] The target model can be any data processing model that can process data based on the self-attention mechanism, for example, a Transformer neural network model or a Large Language Model (LLM) model.

[0065] The target data may be any type of data that can be processed by the target model. For example, if the target model is a large language model, the target data may be text data input into the large language model.

[0066] The target model based on the self-attention mechanism can have key weights, query weights, and value weights. These weights can be a weight matrix. The values of the elements in the matrix can be determined during the training process of the target model. The training process can be referred to related technologies and will not be described in detail.

[0067] In step S101, the target data can be processed with the key weight of the target model to obtain multiple key feature vectors, and the target data can be processed with the query weight of the target model to obtain a query feature vector. The target data can also be processed with the value weight of the target model to obtain multiple value feature vectors. The multiple key feature vectors and the multiple value feature vectors correspond one to one.

[0068] Taking the example that the target model is a large language model, the target data can be divided into multiple text units (each text unit is equivalent to a token), obtaining a sequence of text units. Each text unit is converted into a text unit vector representing the text unit through a word embedding method or other optional methods. Then, by calculating each text unit vector with key weights, multiple key feature vectors can be obtained. By calculating each text unit vector with value weights, multiple value feature vectors can be obtained. By calculating the text unit vector of the last text unit in the text unit sequence with query weights, a query feature vector can be obtained.

[0069] The value feature vectors can also be calculated at any stage before obtaining the processing result of the target data based on formula (1), not limited to being calculated in step S101.

[0070] S102, determine the Hamming distance between each key feature vector and the query feature vector.

[0071] In step S102, the query feature vector and a key feature vector can be calculated to obtain the Hamming distance between this key feature vector and the query feature vector. Executing this step for each key feature vector can obtain the Hamming distance corresponding to each key feature vector.

[0072] The Hamming distance between any two vectors can represent the degree of difference between these two vectors. The larger the Hamming distance, the greater the degree of difference between the two vectors, and the less similar the two vectors are. The smaller the Hamming distance, the smaller the degree of difference between the two vectors, and the more similar the two vectors are.

[0073] In this embodiment, the Hamming distance between the key feature vector and the query feature vector can be directly calculated based on the calculation method of the Hamming distance. The calculation method of the Hamming distance can refer to the related technology and will not be elaborated here.

[0074] Alternatively, the key feature vector and the query feature vector can be converted first, and then the converted vectors are calculated based on the calculation method of the Hamming distance, and the calculation result is used as the Hamming distance between the key feature vector and the query feature vector.

[0075] For example, each key feature vector can be processed based on a hash algorithm to obtain a key hash vector corresponding to each key feature vector. The query feature vector is processed based on the hash algorithm to obtain a query hash vector corresponding to the query feature vector. Then, the key hash vector and the query hash vector are calculated based on the calculation method of the Hamming distance, and the obtained calculation result is used as the Hamming distance between the key feature vector and the query feature vector. The implementation method of the hash algorithm can refer to the related technology and will not be elaborated here.

[0076] As an example, the query feature vector obtained in S101 can be (0.1, 0.0, -0.1, 0.7), and the query hash vector obtained by processing with the hash algorithm can be (1, 1, 1). The multiple key feature vectors obtained in S101 can include key feature vector 1 (-0.3, -0.2, 1.1, 0.7), key feature vector 2 (1.4, -0.7, 0.9, -0.8), key feature vector 3 (0.4, 0.3, 0.8, -0.2), and key feature vector 4 (0.9, 0.5, 0.1, -1.3). The key hash vectors obtained by processing with the hash algorithm can include key hash vector 1 (1, 0, 1) corresponding to key feature vector 1, key hash vector 2 (0, 1, 1) corresponding to key feature vector 2, key hash vector 3 (1, 0, 1) corresponding to key feature vector 3, and key hash vector 4 (0, 0, 0) corresponding to key feature vector 4.

[0077] When calculating the Hamming distance based on the hash vectors, the elements at the corresponding positions in the two hash vectors can be subjected to exclusive OR (XOR) processing to obtain the XOR operation result at that position, and then the XOR operation results at each position are summed to obtain the Hamming distance between the two hash vectors. Taking key hash vector 4 and the query hash vector as an example, the first elements of the two, that is, 0 and 1, are subjected to XOR operation to obtain the XOR operation result 1. Similarly, the second elements of the two are subjected to XOR operation, and the third elements of the two are subjected to XOR operation, and the obtained XOR operation results are all 1. The sum of the three XOR operation results gives the Hamming distance between key hash vector 4 and the query hash vector equal to 3.

[0078] Combined with the above example, the Hamming distance between key feature vector 1 and the query feature vector can be equal to 1, the Hamming distance between key feature vector 2 and the query feature vector can be equal to 2, the Hamming distance between key feature vector 3 and the query feature vector can be equal to 1, and the Hamming distance between key feature vector 4 and the query feature vector can be equal to 3.

[0079] Converting to hash vectors first and then calculating the Hamming distance can reduce the amount of calculation when calculating the Hamming distance, so as to improve the calculation efficiency and reduce the consumed computing resources.

[0080] S103, filter each key feature vector according to the Hamming distance to obtain the filtered key feature vector.

[0081] In step S103, based on the Hamming distances corresponding to each key feature vector, those key feature vectors with a large degree of difference from the query feature vector can be filtered out, and only those key feature vectors with a small degree of difference are retained as the filtered key feature vectors.

[0082] The screening method is not limited. For example, since the Hamming distance is positively correlated with the degree of difference between vectors, a Hamming distance threshold can be determined, and all key feature vectors corresponding to a Hamming distance greater than or equal to the Hamming distance threshold are filtered out, and only those key feature vectors corresponding to a Hamming distance less than the Hamming distance threshold are retained as the screened key feature vectors.

[0083] S104. Process the screened key feature vectors and the query feature vectors to obtain the attention weights of the target data, so as to obtain the processing result of the target data output by the target model based on the attention weights.

[0084] In step S104, based on the calculation method of the foregoing formula (1), the inner product of the screened key feature vectors and the query feature vectors can be calculated to obtain the attention weight corresponding to each key feature vector, and then according to the attention weight and the value feature vector corresponding to the screened key feature vector, the calculation result output of the self-attention mechanism is calculated. i Then, based on this calculation result, the processing result of the target data output by the target model is obtained.

[0085] Taking the target model as a large language model as an example, the calculation result output of the self-attention mechanism is obtained. i Based on the word embedding method or other optional methods, this calculation result can be converted into a text unit, and this text unit is the processing result of the target data output by the target model.

[0086] According to the difference of the target model, the method for obtaining the processing result of the target data based on the calculation result of the self-attention mechanism may be different, and the specific implementation manner can refer to the related technology and will not be elaborated.

[0087] The beneficial effect of this embodiment is as follows:

[0088] When performing the inner product operation on the key feature vectors and the query feature vectors, when the degree of difference between the key feature vectors and the query feature vectors is small, the calculated real number will be greater than 0, and when the degree of difference between the key feature vectors and the query feature vectors is large, the calculated real number may be equal to 0 or approximately equal to 0.

[0089] According to this principle, before performing the inner product operation on the key feature vectors and the query feature vectors, the method of this embodiment filters out those key feature vectors whose inner product operation results may be equal to or approximately equal to 0 according to the Hamming distance, and only retains those screened key feature vectors with a small degree of difference and a real number greater than 0 in the inner product operation for the inner product operation. In this way, the method of this embodiment can reduce the operation amount of the self-attention mechanism by filtering out some key feature vectors without affecting the accuracy of the processing result, achieving the effect of improving the processing efficiency and reducing the consumed computing resources.

[0090] In some alternative embodiments, screening each key feature vector according to the Hamming distance to obtain the screened key feature vector may include the following steps:

[0091] Perform a first screening on each key feature vector according to the Hamming distance;

[0092] Determine the similarity between the key feature vector after the first screening and the query feature vector according to the Hamming distance;

[0093] Perform a second screening on each key feature vector after the first screening based on the similarity to obtain the screened key feature vector.

[0094] During the first screening, all key feature vectors corresponding to a Hamming distance greater than or equal to a preset Hamming distance threshold can be filtered out, and only those key feature vectors corresponding to a Hamming distance less than the Hamming distance threshold are retained as the key feature vectors after the first screening.

[0095] Combined with the foregoing example, assuming that the Hamming distance threshold is equal to 3, the key feature vector 4 corresponding to a Hamming distance equal to 3 can be filtered out, and the key feature vectors 1, 2, and 3 are retained as the key feature vectors after the first screening.

[0096] Among them, the Hamming distance threshold can be manually configured by the relevant user, or it can also be dynamically determined according to the obtained key feature vectors and query feature vectors.

[0097] The similarity between the key feature vector and the query feature vector can be calculated by various algorithms for calculating vector similarity, which is not limited.

[0098] For example, based on the calculation method of cosine similarity, the key feature vector after the first screening and the query feature vector can be directly calculated, and the obtained calculation result is used as the similarity between the key feature vector after the first screening and the query feature vector.

[0099] During the second screening, all key feature vectors corresponding to a similarity less than a preset similarity threshold can be filtered out, and only those key feature vectors corresponding to a similarity greater than or equal to the similarity threshold are retained as the screened key feature vectors.

[0100] The similarity threshold can be manually configured by the relevant user, or it can also be dynamically determined according to the obtained key feature vectors and query feature vectors.

[0101] The beneficial effect of this embodiment is that the key feature vectors are screened twice based on the Hamming distance and similarity respectively to obtain the screened key feature vectors. Screening the key feature vectors from two different dimensions in this way is beneficial to improving the fineness and accuracy of the screening and fully filtering out the key feature vectors with a large degree of difference.

[0102] An optional method for determining the similarity between the key feature vector and the query feature vector can be as follows:

[0103] Determine the angular value between the key feature vector after the first screening and the query feature vector according to the Hamming distance;

[0104] Determine the similarity between the key feature vector after the first screening and the query feature vector according to the angular value.

[0105] When determining the angular value, the Hamming distance corresponding to the key feature vector can be directly determined as the angular value in radians. Combining the foregoing example, if the Hamming distance of the key feature vector 1 is equal to 1, then the angular value corresponding to the key feature vector 1 can be determined as 1 radian (rad).

[0106] When determining the angular value, the Hamming distance corresponding to the key feature vector can also be converted into the angular value in radians based on the following conversion formula (2).

[0107]

[0108] Where, π represents the pi, K represents the total number of key feature vectors after the first screening, Hamming represents the Hamming distance of the key feature vector, Bias represents a preset reference value, its value can be set as required without limitation, and Theta represents the angular value corresponding to the converted key feature vector.

[0109] Combining the foregoing example, the key feature vectors after the first screening include the key feature vectors 1 to 3, so K is equal to 3. For the key feature vector 1, its Hamming distance Hamming is equal to 1. Assuming that the reference value is set to 0.127, the angular value corresponding to the key feature vector 1 calculated according to the formula (2) can be 0.92, and the angular value corresponding to the key feature vector 2 can be 1.97.

[0110] After obtaining the angular value, the cosine value corresponding to the angular value can be determined, and this cosine value is used as the similarity between the key feature vector and the query feature vector.

[0111] Combining the foregoing example, if the angular value corresponding to the key feature vector 1 is 0.92, then the similarity between the key feature vector 1 and the query feature vector can be equal to the cosine value of 0.92, that is, cos(0.92). If the angular value corresponding to the key feature vector 2 is 1.97, then the similarity between the key feature vector 2 and the query feature vector can be equal to cos(1.97).

[0112] Optionally, based on the similarity, perform a secondary screening on each key feature vector after the first screening to obtain the screened key feature vector, including:

[0113] Obtain the inner product value of each key feature vector based on the similarity and the modulus value of each key feature vector after the first screening;

[0114] Perform a second screening based on the inner product values of each key feature vector after the first screening to obtain the screened key feature vectors.

[0115] For any key feature vector, the method for calculating its modulus value can be to calculate the square of each element contained in the key feature vector, sum the squares of all elements, and take the square root of the obtained sum, and the obtained result is used as the modulus value of the key feature vector.

[0116] For each key feature vector after the first screening, the similarity corresponding to this key feature vector can be multiplied by the modulus value corresponding to this key feature vector, and the obtained result is used as the inner product value of this key feature vector.

[0117] After obtaining the inner product value, the inner product value of each key feature vector after the first screening can be compared with a preset similarity threshold, and those key feature vectors with inner product values less than or equal to the similarity threshold are filtered out, and only the key feature vectors with inner product values greater than the similarity threshold are retained as the above-mentioned screened key feature vectors.

[0118] The advantage of performing the second screening according to the above method is that:

[0119] Compared with screening only based on similarity, the above method can perform a second screening by combining the similarity and modulus value of the key feature vector, thereby improving the accuracy of the second screening.

[0120] In some optional embodiments, the method of this embodiment may further include the following method for determining the Hamming distance threshold:

[0121] Obtain the target key feature vectors whose corresponding normalized attention weights are less than the preset weight threshold;

[0122] Determine the maximum value among the Hamming distances corresponding to each target key feature vector as the Hamming distance threshold;

[0123] Among them, the Hamming distance threshold is used for the first screening.

[0124] For each key feature vector Key k , the attention weight Sc corresponding to this key feature vector can be calculated according to the following formula (3) k , and the attention weight of this key feature vector is calculated based on the normalized exponential function softmax(), and the normalized attention weight NorSc of this key feature vector is obtained k .

[0125]

[0126] The weight threshold can be determined according to a pre-configured hyperparameter p. For example, the weight threshold can be equal to p divided by n, that is, the weight threshold = p / n, where n is the dimension of the key feature vector. Exemplarily, assuming the key feature vector is a 100-dimensional vector, then n is equal to 100. The specific value of the hyperparameter p can be manually configured by relevant users according to needs, without limitation.

[0127] After obtaining the normalized attention weights and the weight threshold, all key feature vectors corresponding to the normalized attention weights less than or equal to the weight threshold can be found as target key feature vectors, and then the Hamming distances of these target key feature vectors are obtained, and the maximum value among them is determined as the Hamming distance threshold.

[0128] As an example, p can be set to be equal to 1. Assuming the dimension n of the key feature vector is equal to 4, then the above weight threshold is equal to 0.25. Assuming the normalized attention weights of the aforementioned key feature vectors 1 to 4 are 0.48, 0.06, 0.30, and 0.16 in sequence, it can be determined that the normalized attention weights of key feature vector 2 and key feature vector 4 are less than the weight threshold. The Hamming distances of key feature vector 2 and key feature vector 4 are equal to 3 and 2 respectively. Then the maximum value among them can be determined as the Hamming distance threshold, that is, the Hamming distance threshold is equal to 3.

[0129] In some alternative embodiments, new key feature vectors and new value feature vectors will be continuously generated during the operation of the target model, and the query feature vector will also be continuously updated. Based on this, after each new key feature vector is generated, the above method for determining the Hamming distance threshold can be executed based on the new key feature vector and the original key feature vectors, and the final Hamming distance threshold is determined by combining the multiple maximum values obtained after multiple executions. For example, the average value of the multiple maximum values is used as the Hamming distance threshold.

[0130] In some alternative embodiments, the method of this embodiment may further include the following method for determining the similarity threshold:

[0131] Obtain the normalized attention weights corresponding to the key feature vectors;

[0132] Determine the minimum value among the multiple normalized attention weights greater than or equal to the preset weight threshold;

[0133] Determine the similarity threshold according to the minimum value;

[0134] Among them, the similarity threshold is used for secondary screening.

[0135] The method for obtaining the normalized attention weights and the weight threshold can refer to the above embodiments and will not be elaborated.

[0136] After screening out at least one normalized attention weight greater than or equal to the weight threshold based on the weight threshold, the minimum value can be determined therefrom, and the attention weight corresponding to the minimum value can be obtained.

[0137] Combined with the foregoing example, assume that the normalized attention weights of the foregoing key feature vectors 1 to 4 are 0.48, 0.06, 0.30, and 0.16 in sequence, and the normalized attention weights of key feature vectors 1 and 3 are greater than the weight threshold. Thus, the minimum value among the normalized attention weights corresponding to these two key feature vectors can be obtained, that is, 0.30, and then the unnormalized attention weight corresponding to the minimum value can be obtained, that is, 0.86.

[0138] After obtaining the attention weight corresponding to the minimum value, the modulus value of the query feature vector can be multiplied by the maximum value among the modulus values of multiple key feature vectors to obtain a modulus product, and then the attention weight can be divided by the modulus product to obtain a similarity threshold. That is to say, the similarity threshold Sim can be calculated by the following formula (4).

[0139]

[0140] Among them, Score represents the attention weight corresponding to the above minimum value, ||q|| represents the modulus value of the query feature vector, and ||Kmax|| represents the maximum value among the modulus values of multiple key feature vectors.

[0141] In some alternative embodiments, new key feature vectors and new value feature vectors will be continuously generated during the operation of the target model, and the query feature vector will also be continuously updated. Based on this, after each new key feature vector is generated, the above method for determining the similarity threshold can be executed based on the new key feature vector and the original key feature vectors, and the final similarity threshold can be determined by combining the similarity thresholds obtained after multiple executions. For example, the average value of the similarity thresholds determined multiple times can be used as the final similarity threshold.

[0142] The advantage of determining the Hamming distance threshold and the similarity threshold according to the above method is that the corresponding thresholds can be dynamically determined according to different key feature vectors and query feature vectors, so as to more accurately screen out the filtered key feature vectors whose operation results are greater than 0.

[0143] Optionally, obtaining the processing result of the target data output by the target model based on the attention weight may include:

[0144] 1. Obtain the value feature vectors corresponding to the filtered key feature vectors, where the value feature vectors are the feature vectors obtained by processing the target data based on the value weights of the target model;

[0145] 2. Calculate the value feature vector according to the attention weight, and normalize the calculation result to obtain the processing result of the target data output by the target model.

[0146] The method for calculating the value feature vector corresponding to the target data based on the value weight can be referred to the related technology and will not be elaborated here.

[0147] As described above, each key feature vector uniquely corresponds to a value feature vector. Taking the large language model as an example, the target data can be the text data input into the large language model, and this text data can be divided into multiple text units. Calculating the text unit vectors corresponding to the text units with the key weight and the value weight can obtain the key feature vector and the value feature vector of this text unit. If a key feature vector and a value feature vector are calculated from the same text unit vector, it can be considered that this key feature vector and this value feature vector correspond to each other.

[0148] For each key feature vector, the attention weight corresponding to this key feature vector can be calculated according to the aforementioned formula (3).

[0149] In step 2, the power exponent operation can be first performed on the attention weight of each key feature vector to obtain the corresponding exponential attention weight Se k , and the calculation process can be expressed by the following formula (5).

[0150]

[0151] e is the base of the natural logarithm, and Sc k represents the attention weight of the k-th screened key feature vector.

[0152] Then, the exponential attention weight corresponding to each key feature vector can be multiplied by the value feature vector corresponding to this key feature vector, and the sums of the obtained multiple products are calculated to obtain the calculation result Sum. This process can be expressed by the following formula (6).

[0153]

[0154] Among them, L represents the total number of screened key feature vectors, and Value k represents the value feature vector corresponding to the k-th screened key feature vector.

[0155] When normalizing the calculation result, the exponential attention weights of all screened key feature vectors can be summed, and then Sum is divided by the sum of these exponential attention weights. The obtained result is the calculation result output obtained by calculating the screened key feature vectors based on the self-attention mechanism i . This calculation process can be expressed by the following formula (7). Finally, according to outputi Obtain the processing result of the target data. For example, convert output i to obtain the corresponding text unit, which will not be elaborated here.

[0156]

[0157] Optionally, when calculating the modulus value of a vector, it is necessary to perform a square operation on each component (which can also be called an element of the vector) in the vector.

[0158] In the related art, during the process of squaring a binary number, a number of partial products equal to the number of bits of the binary number will be generated. By adding these partial products, the result of the square operation is obtained. Figure 2 Taking as an example, when squaring an 8-bit binary number, 8 partial products will be generated, which will result in an addition tree with a relatively long accumulation path, causing a large computational resource overhead. Figure 2 The combination of any two one-bit binary numbers in is equivalent to their multiplication. For example, x7x4 is equivalent to x7 * x4, and also equivalent to x7 multiplied by x4.

[0159] Figure 2 In, x0 to x7 are all one-bit binary numbers, that is, x0 is equal to 0 or 1, and the same is true for x1 to x7. As an example, the 8-bit binary number to be squared can be represented by Table 1.

[0160] Table 1

[0161] <![CDATA[x7]]> <![CDATA[x6]]> <![CDATA[x5]]> <![CDATA[x4]]> <![CDATA[x3]]> <![CDATA[x2]]> <![CDATA[x1]]> <![CDATA[x0]]> 1 1 0 0 0 1 1 0

[0162] In view of the above problems, this embodiment provides a method for calculating the square value of a binary number. Based on this method, the process of obtaining the modulus value of the key feature vector may include:

[0163] Calculate multiple partial products corresponding to each component in the key feature vector, and the number of partial products corresponding to each component is less than the number of bits of the component in binary;

[0164] Obtain the square value corresponding to each component according to the partial products of each component, so as to determine the modulus value of the key feature vector based on the square values of each component.

[0165] First, refer to Figure 2 , since 1 * 1 = 1 and 0 * 0 = 0, it can be seen that the product obtained by multiplying any two identical one-bit binary numbers can be simplified to the one-bit binary number itself. For example, x0 * x0 = x0.

[0166] Furthermore, since 1 * 1 = 10 and 0 * 0 = 0 = 00, therefore, adding any two identical one-bit binary numbers can be simplified to shifting the one-bit binary number one bit to the left. For example Figure 2Among them, the addition of the two x1x0 at the 1st position is equivalent to shifting x1x0 to the left to the 2nd position. Figure 2 The rightmost bit in Figure 2 is the 0th bit, and the bit numbers increase one by one from right to left.

[0167] According to the above simplification rules, Figure 2 Simplify the 8 partial products shown in Figure 2 to obtain Figure 3 the 5 partial products shown in Figure 3 .

[0168] Furthermore, for the case of adding the product of two binary numbers to one of the binary numbers, it can be converted according to formula (8) into multiplying the product of two binary numbers by 2 and then adding the product of the other binary number after taking the inverse and one binary number.

[0169]

[0170] Among them, represents taking the inverse of x0.

[0171] After converting to the form on the rightmost side of formula (8), the 2 times of x1x0 can be further shifted one bit to the left. Therefore, as Figure 3 shown, based on formula (8), the above 5 partial products can be simplified to obtain Figure 4 the 4 partial products shown in Figure 4 .

[0172] To sum up, when calculating the modulus value of a key feature vector, each component of the key feature vector can be converted into an 8-bit binary number. Then, based on this 8-bit binary number, calculate Figure 4 the four partial products shown in Figure 4 . Finally, sum these 4 partial products, and the result obtained is the square value of this component.

[0173] After obtaining the square value of each component, the square values of all components in the key feature vector can be summed, and then the square root operation is performed on the obtained sum of square values. The result obtained is the modulus value of this key feature vector.

[0174] It can be seen that through the above calculation method, during the process of performing the square operation on binary data, the number of partial products that need to be calculated can be effectively reduced, achieving the effect of shortening the path of the adder tree and reducing the overhead of computing resources.

[0175] When it is necessary to calculate the modulus value of the query feature vector and the modulus value of the query feature vector, it can also be calculated according to the above method, which will not be elaborated here.

[0176] When determining the similarity according to the angle value, the cosine value corresponding to this angle value can be obtained as the corresponding similarity. In this embodiment, the cosine value corresponding to the angle value can be obtained by looking up a table, or it can also be obtained through the following method:

[0177] Based on the target calculation unit, perform multiple iterative operations on the angle value and the preset initial parameters to obtain the cosine value corresponding to the angle value. The cosine value is used as the similarity between the key feature vector after the first screening and the query feature vector.

[0178] Among them, the input of the first iterative operation includes the angle value and the initial parameters. The input of each iterative operation after the first iterative operation is the output of the previous iterative operation. The target calculation unit includes a shifter and multiple adders.

[0179] Please refer to Figure 5 , which is the structural schematic diagram of the target calculation unit provided in this embodiment. The target calculation unit can at least include 3 adders and two shifters, denoted as adder 1, adder 2, and adder 3 in sequence, as well as shifter 1 and shifter 2.

[0180] In addition, it can also include a register for storing the parameter di of the i-th iteration, that is, the box corresponding to di in Figure 5 , and a register for storing the arctangent value of the i-th iteration, that is, the box corresponding to tan Figure 5 in -1 (2 -i ).

[0181] The connection relationships of the above devices can be referred to in Figure 5 .

[0182] Among them, adder 1 is used to calculate the output x[i + 1] of the i-th iteration. Shifter 1 is used to perform a right shift operation of i bits on the input y[i] of the i-th iteration. The register di is used to provide the parameter di of the i-th iteration to adder 1, adder 2, and adder 3 for calculation.

[0183] The calculation process of adder 1 can be represented by the following formula (9).

[0184]

[0185] Adder 2 is used to calculate the output y[i + 1] of the i-th iteration. Shifter 2 is used to perform a right shift operation of i bits on the input x[i] of the i-th iteration. The calculation process of adder 2 can be represented by the following formula (10).

[0186]

[0187] Adder 3 is used to calculate the output z[i + 1] of the i-th iteration. The register of the tangent value is used to provide the arctangent value of the i-th iteration to adder 3. The calculation process of adder 3 can be represented by the following formula (11).

[0188]

[0189] tan -1 (2 -i )represents the tangent value 2 -i corresponding angle value. It can be seen that for any integer i, tan -1 (2 -i )has a fixed value, so the corresponding tan -1 (2 -i )can be pre-stored in the corresponding register before the i-th iteration operation without calculation.

[0190] The parameter di for the i-th iteration can be determined according to formula (12) and pre-stored in the corresponding register.

[0191]

[0192] Whether z[i] is less than 0 can be determined according to the most significant bit (MSB) of z[i].

[0193] In the above iterative operation process, the initial value of i can be equal to 0, z[0] is equal to the angle value to be operated. For example, when calculating the cosine value cos(0.92) of 0.92, z[0] is equal to 0.92. x[0] and y[0] are preset initial parameters, and their values can be set as needed. For example, x[0] is set equal to 1 and y[0] is set equal to 0.

[0194] In the above embodiment, starting from i equal to 0, the input of the i-th iteration, that is, z[i], x[i] and y[i], can be input into Figure 5 the target calculation unit shown to obtain the output of the i-th iteration, that is, z[i + 1], x[i + 1] and y[i + 1]. If i + 1 is less than the preset upper limit of the number of iteration times at this time, the output of the i-th iteration can be used as the input of the (i + 1)-th iteration, and the input of the (i + 1)-th iteration is continuously sent to the target calculation unit for operation, while i is incremented by 1, and so on until i + 1 is equal to the upper limit of the number of iteration times.

[0195] In the case where i + 1 is equal to the upper limit of the number of iteration times, the output x[i + 1] of the i-th iteration can be determined as the cosine value corresponding to the input angle value z[0], and the output y[i + 1] of the i-th iteration can be determined as the sine value corresponding to the input angle value z[0].

[0196] In some embodiments, to improve the operation efficiency, please refer to Figure 6, the target computing units corresponding to the upper limit of the number of iterations can be connected in series. The i-th target computing unit corresponds to the (i - 1)-th iteration. The input of the first target computing unit is the aforementioned angle value and the initial parameters. Thereafter, the input of each target computing unit is the output of the previous target computing unit. The x[i + 1] and y[i + 1] in the output of the last target computing unit are used as the cosine value and sine value corresponding to the input angle value.

[0197] For example Figure 6 In, assume that the upper limit of the number of iterations is 8. Eight target computing units can be connected in series. The first target computing unit on the left corresponds to i = 0, that is, the 0-th iteration. The second target computing unit corresponds to the 1-st iteration, and so on. The eighth target computing unit corresponds to the 7-th iteration. When calculating the similarity, the angle value and the initial parameters are used as z[0], x[0], and y[0] and input into the first target computing unit on the left. After being calculated one by one through eight target computing units, x[i + 1] output by the eighth target computing unit, that is, x[8], is used as the cosine value corresponding to the angle value.

[0198] Compared with obtaining the cosine value corresponding to the angle value by looking up a table, using the above method to calculate the cosine value can reduce the number of computing elements (such as adders, shifters, registers, etc.) required in the calculation process, thereby reducing the occupied register transfer level (RTL) synthesis area, enabling the processing method of this embodiment to run on a chip with a smaller size.

[0199] As an example, please refer to Figure 7 , which is a schematic diagram of the variation law of the RTL synthesis area (in square millimeters) with the accuracy of the calculation result when using the table lookup method and the above method of this embodiment to calculate the cosine value. It can be seen that at any accuracy, the RTL synthesis area occupied by the table lookup method is greater than that of the calculation method provided in this embodiment. And as the accuracy increases, the RTL synthesis area occupied by the table lookup method increases significantly, while the RTL synthesis area occupied by the method of this embodiment remains basically unchanged. It can be seen that the method of this embodiment can reduce the occupied register transfer level (RTL) synthesis area.

[0200] This embodiment provides a data processing device. Please refer to Figure 8 , and the device may include the following modules.

[0201] An obtaining module 801, configured to obtain a key feature vector and a query feature vector. The key feature vector is a feature vector obtained by processing target data based on the key weights of a target model including a self-attention mechanism, and the query feature vector is a feature vector obtained by processing target data based on the query weights of the target model;

[0202] A screening module 802, configured to:

[0203] Determine the Hamming distance between each of the key feature vectors and the query feature vector;

[0204] Filter each of the key feature vectors according to the Hamming distance to obtain the filtered key feature vectors;

[0205] A processing module 803, configured to process the filtered key feature vectors and the query feature vector to obtain the attention weights of the target data, so as to obtain the processing result of the target data output by the target model based on the attention weights.

[0206] Optionally, when the screening module 802 filters each key feature vector according to the Hamming distance to obtain the filtered key feature vectors, it can be used for:

[0207] Perform a primary screening on each key feature vector according to the Hamming distance;

[0208] Determine the similarity between the key feature vectors after the primary screening and the query feature vector according to the Hamming distance;

[0209] Perform a secondary screening on each of the key feature vectors after the primary screening based on the similarity to obtain the filtered key feature vectors.

[0210] Optionally, when the screening module 802 determines the similarity between the key feature vectors after the primary screening and the query feature vector according to the Hamming distance, it can be used for:

[0211] Determine the angle value between the key feature vectors after the primary screening and the query feature vector according to the Hamming distance;

[0212] Determine the similarity between the key feature vectors after the primary screening and the query feature vector according to the angle value.

[0213] Optionally, when the screening module 802 determines the similarity between the key feature vectors after the primary screening and the query feature vector according to the angle value, it can be used for:

[0214] Perform multiple iterative operations on the angle value and a preset initial parameter by a target calculation unit to obtain the cosine value corresponding to the angle value, and the cosine value is used as the similarity between the key feature vectors after the primary screening and the query feature vector;

[0215] Wherein, the input of the first iterative operation includes the angle value and the initial parameter, the input of each iterative operation after the first iterative operation is the output of the previous iterative operation, and the target calculation unit includes a shifter and a plurality of adders.

[0216] Optionally, when the screening module 802 performs a secondary screening on each of the key feature vectors after the primary screening based on the similarity to obtain the filtered key feature vectors, it can be used for:

[0217] Obtain the inner product values of each key feature vector based on the similarity and the modulus values of each key feature vector after the first screening;

[0218] Perform a second screening based on the inner product values of each key feature vector after the first screening to obtain the screened key feature vectors.

[0219] Optionally, the process by which the screening module 802 obtains the modulus value of the key feature vector may include:

[0220] Calculate multiple partial products corresponding to each component in the key feature vector, and the number of partial products corresponding to each component is less than the number of binary digits of the component;

[0221] Obtain the squared value corresponding to each component based on the partial product of each component, and determine the modulus value of the key feature vector based on the squared values of each component.

[0222] Optionally, the screening module 802 can also be used for:

[0223] Obtain the target key feature vectors whose corresponding normalized attention weights are less than the preset weight threshold;

[0224] Determine the maximum value among the Hamming distances corresponding to each target key feature vector as the Hamming distance threshold;

[0225] Wherein, the Hamming distance threshold is used for the first screening.

[0226] Optionally, the screening module 802 can also be used for:

[0227] Obtain the first key feature vectors whose corresponding normalized attention weights are less than the preset weight threshold;

[0228] Determine the minimum value among the multiple normalized attention weights greater than or equal to the preset weight threshold;

[0229] Determine the similarity threshold according to the minimum value;

[0230] Wherein, the similarity threshold is used for the second screening.

[0231] Optionally, when the processing module 803 obtains the processing result of the target data output by the target model based on the attention weight, it can be used for:

[0232] Obtain the value feature vectors corresponding to the screened key feature vectors, where the value feature vectors are feature vectors obtained by processing the target data based on the value weights of the target model;

[0233] Calculate the value feature vectors according to the attention weight and perform normalization processing on the calculation result to obtain the processing result of the target data output by the target model.

[0234] Optionally, please refer toFigure 9 , which is a schematic diagram of another data processing device provided in this embodiment.

[0235] The query cache area and the key cache area therein are respectively used to cache the query feature vector and the key feature vector, and these two cache areas can be equivalent to the aforementioned obtaining module 801.

[0236] The hash operation module is used to calculate the query hash corresponding to the query feature vector and the key hash vector. The query hash vector is stored in the query hash storage area (i.e., Query Hash Buffer), and the key hash vector (i.e., key hash 1, key hash 2... key hash n in the figure) is stored in the key hash storage area (Key Hash Mem).

[0237] The first screening module reads the query hash vector and the key hash vector from the above storage area, and calculates the Hamming distance corresponding to each key feature vector according to the aforementioned method.

[0238] The Hamming distance threshold module determines the Hamming distance threshold and sends it to the first screening module.

[0239] The first screening module performs the first screening based on the calculated Hamming distance and the obtained Hamming distance threshold, and writes the key identifier (Key ID) of the key feature vector after the first screening, the query identifier (Query ID) of the query feature vector, and the Hamming distance corresponding to the key feature vector after the first screening into the Hamming distance buffer (Hamming DistanceBuffer). Thus, the first screening is completed.

[0240] The modulus operation module can calculate the modulus of each key feature vector and the modulus of the query feature vector, and store the calculated modulus in the modulus storage area (Norm Mem).

[0241] The modulus operation module can also use the calculated modulus to calculate the similarity threshold based on the aforementioned method and provide it to the second screening module.

[0242] The second screening module obtains the Hamming distance of the key feature vector after the first screening from the Hamming distance buffer, and obtains the modulus of these key feature vectors from the modulus storage area based on the key identifier of the key feature vector after the first screening. Then, according to the Hamming distance and the modulus, the second screening is performed according to the aforementioned method to finally determine the key feature vector after screening. Then, the key identifier of the key feature vector after screening and the query identifier of the query feature vector are written into the identifier buffer (ID Buffer). Thus, the second screening ends.

[0243] Figure 9 All the modules and storage areas / cache areas from the hash operation module to the identifier buffer are equivalent to the screening module 802 in the aforementioned embodiment.

[0244] Finally, the query identifier that identifies the buffer is sent to the query buffer (Query Mem) so that the corresponding query feature vector can be read out from the query buffer. The key identifier is sent to the key buffer (Key Mem) and the value buffer (Value Mem) to read out the filtered key feature vector corresponding to the key identifier from the key buffer, and read out the value feature vector corresponding to the filtered key feature vector from the value buffer according to the key identifier.

[0245] The query feature vector and the filtered key feature vector enter the inner product operation module. The inner product operation module performs an inner product operation on each filtered key feature vector according to formula (3) to obtain the attention weight Sc corresponding to each filtered key feature vector. k The attention weight is sent to the exponential operation module for calculation to obtain the exponential attention weight.

[0246] The exponential attention weight is sent to the multiplication accumulation module and the summation module; the summation module sums up the exponential attention weights corresponding to all the filtered key feature vectors, and the obtained sum is sent to the normalization operation module; the multiplication accumulation module multiplies the exponential attention weight corresponding to the filtered key feature vector and the value feature vector provided by the value buffer to obtain the corresponding product, and accumulates the multiplications corresponding to all the filtered key feature vectors to obtain Sum shown in formula (6). This result can be stored in the accumulation result buffer.

[0247] Finally, the normalization operation module reads out the sum Sum of the product of the exponential attention weight and the corresponding value feature vector from the accumulation result buffer, and divides Sum by the sum of the exponential attention weights corresponding to all the filtered key feature vectors to obtain the calculation result Output of the self-attention mechanism. i Output i After being processed by the target model, the processing result corresponding to the target data can be obtained.

[0248] Figure 9 All the modules and buffers after the identification buffer in to Figure 9 are equivalent to the processing module 803 in the foregoing embodiment.

[0249] For the data processing device provided in this embodiment, the working principle can refer to the relevant steps in the data processing method provided in the foregoing embodiment, which will not be elaborated here.

[0250] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0251] For the convenience of description, when describing the above system or device, it is divided into various modules or units according to functions for separate description. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0252] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0253] Finally, it should also be noted that in this article, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the said element.

[0254] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A data processing method, comprising: Obtaining a key feature vector and a query feature vector, where the key feature vector is a feature vector obtained by processing target data based on the key weights of a target model including a self-attention mechanism, and the query feature vector is a feature vector obtained by processing target data based on the query weights of the target model; Determining the Hamming distance between each of the key feature vectors and the query feature vector; Screening each of the key feature vectors according to the Hamming distance to obtain screened key feature vectors; Processing the screened key feature vectors and the query feature vector to obtain the attention weights of the target data, so as to obtain the processing result of the target data output by the target model based on the attention weights.

2. The method according to claim 1, wherein the screening each of the key feature vectors according to the Hamming distance to obtain screened key feature vectors comprises: Performing a first screening on each of the key feature vectors according to the Hamming distance; Determining the similarity between the key feature vectors after the first screening and the query feature vector according to the Hamming distance; Performing a second screening on each of the key feature vectors after the first screening based on the similarity to obtain screened key feature vectors.

3. The method according to claim 2, wherein the determining the similarity between the key feature vectors after the first screening and the query feature vector according to the Hamming distance comprises: Determining an angle value between the key feature vectors after the first screening and the query feature vector according to the Hamming distance; Determining the similarity between the key feature vectors after the first screening and the query feature vector according to the angle value.

4. The method according to claim 3, wherein the determining the similarity between the key feature vectors after the first screening and the query feature vector according to the angle value comprises: Performing multiple iterative operations on the angle value and a preset initial parameter by a target calculation unit to obtain a cosine value corresponding to the angle value, and using the cosine value as the similarity between the key feature vectors after the first screening and the query feature vector; Wherein, the input of the first iterative operation includes the angle value and the initial parameter, and the input of each iterative operation after the first iterative operation is the output of the previous iterative operation, and the target calculation unit includes a shifter and a plurality of adders.

5. The method according to claim 2, wherein the performing a second screening on each of the key feature vectors after the first screening based on the similarity to obtain screened key feature vectors comprises: Obtaining the inner product values of each of the key feature vectors based on the similarity and the modulus values of each of the key feature vectors after the first screening; Performing a second screening according to the inner product values of each of the key feature vectors after the first screening to obtain screened key feature vectors.

6. The method according to claim 5, wherein the process of obtaining the modulus value of the key feature vector comprises: Calculating a plurality of partial products corresponding to each component in the key feature vector, and the number of partial products corresponding to each component is less than the number of binary digits of the component; Obtain the square value corresponding to each component according to the partial product of each component, so as to determine the modulus value of the key feature vector based on the square values of each component.

7. The method according to claim 2, further comprising: Obtain a target key feature vector whose corresponding normalized attention weight is less than a preset weight threshold; Determine the maximum value among the Hamming distances corresponding to each of the target key feature vectors as the Hamming distance threshold; Wherein, the Hamming distance threshold is used for the first screening.

8. The method according to claim 2, further comprising: Obtain the normalized attention weight corresponding to the key feature vector; Determine the minimum value among multiple normalized attention weights greater than or equal to the preset weight threshold; Determine the similarity threshold according to the minimum value; Wherein, the similarity threshold is used for the second screening.

9. The method according to claim 1, wherein the process of obtaining the processing result of the target data output by the target model based on the attention weight comprises: Obtain the value feature vector corresponding to the screened key feature vector, where the value feature vector is a feature vector obtained by processing the target data based on the value weight of the target model; Calculate the value feature vector according to the attention weight and perform normalization processing on the calculation result to obtain the processing result of the target data output by the target model.

10. A data processing device, comprising: An obtaining module, configured to obtain a key feature vector and a query feature vector, where the key feature vector is a feature vector obtained by processing target data based on the key weight of a target model including a self-attention mechanism, and the query feature vector is a feature vector obtained by processing target data based on the query weight of the target model; A screening module, configured to: Determine the Hamming distance between each of the key feature vectors and the query feature vector; Screen each of the key feature vectors according to the Hamming distance to obtain a screened key feature vector; A processing module, configured to process the screened key feature vector and the query feature vector to obtain the attention weight of the target data, so as to obtain the processing result of the target data output by the target model based on the attention weight.