Data processing method and device and terminal equipment

By combining and interleaving the relevant weight matrix to form the target weight matrix, the problem of multiple read accesses of input data in the self-attention mechanism is solved, and the computing efficiency is improved.

CN120216850APending Publication Date: 2025-06-27NIO TECH ANHUI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510275039.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing self-attention mechanism reads and accesses to the same input data multiple times during the calculation process, resulting in high memory access and reducing the efficiency of calculating self-attention scores.

Method used

By obtaining the target weight matrix, the matrix is ​​combined by multiple weight matrices related to the self-attention mechanism, and each column of data in each weight matrix is ​​arranged in sequence with preset interleaving rules, and thus participates in the operation of all weight matrices in just one data load.

Benefits of technology

It reduces multiple read accesses to the same input data, reduces memory access time, and improves the efficiency of calculating self-attention scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216850A_ABST
    Figure CN120216850A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device and terminal equipment, and the method comprises the steps: obtaining a target weight matrix, enabling the target weight matrix to be obtained through the combination of a plurality of weight matrixes related to a self-attention mechanism, and enabling the target weight matrix to comprise each column of data in each weight matrix, each column of data in each weight matrix is sequentially arranged according to a preset interleaving rule, and the self-attention score of the input data is obtained by obtaining the input data and inputting the input data into the target weight matrix, so that the self-attention score of the input data can be obtained only by inputting the input data into the target weight matrix; the input data does not need to be accessed for multiple times and are respectively input into the query weight matrix, the key weight matrix and the value weight matrix, so that the condition that the same input data is read and accessed for multiple times is reduced, the time consumption of memory access is reduced, and the efficiency of calculating the self-attention score is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a method, device, and terminal device for data processing. Background Art

[0002] The self-attention mechanism is an attention mechanism based on human vision, and its core lies in being able to screen out relatively important information from a large amount of information while ignoring other relatively unimportant information. Currently, the self-attention mechanism has been widely applied to deep learning tasks such as natural language processing and text generation.

[0003] In the calculation process of the self-attention mechanism, the core operations include the calculation of the weight matrices of query (Query), key (Key), and value (Value) vectors. The input data needs to be calculated through the query weight matrix, key weight matrix, and value weight matrix respectively, resulting in the situation of multiple read accesses to the same input data, with high memory access time consumption and reduced efficiency of calculating the self-attention score. Summary of the Invention

[0004] Embodiments of this application provide a method, device, and terminal device for data processing, aiming to solve the problem that in the calculation process of the existing self-attention mechanism, there are multiple read accesses to the same input data, resulting in high memory access time consumption and reduced efficiency of calculating the self-attention score.

[0005] In a first aspect, embodiments of this application provide a method for data processing, and the method includes:

[0006] Obtain a target weight matrix; wherein, the target weight matrix is obtained by combining multiple weight matrices related to the self-attention mechanism; the target weight matrix includes each column of data in each of the weight matrices, and the data in each column of each weight matrix are arranged in sequence according to a preset interleaving rule;

[0007] Obtain input data;

[0008] Input the input data into the target weight matrix to obtain the self-attention score of the input data.

[0009] In a possible implementation manner of the above first aspect, the obtaining of the target weight matrix includes:

[0010] Obtain multiple weight matrices related to the self-attention mechanism;

[0011] Based on the self-attention mechanism, determine the arrangement order of each of the weight matrices;

[0012] Arrange and combine each column of data in each of the weight matrices according to the described arrangement order to obtain the target weight matrix.

[0013] In a possible implementation manner of the first aspect above, the arranging and combining each column of data in each of the weight matrices according to the described arrangement order to obtain the target weight matrix includes:

[0014] According to the arrangement order in the interleaving rule, sequentially store the elements in the i-th column of each weight matrix in the target array; where i is greater than or equal to 1, and i + 1 is less than or equal to N, and N represents the number of columns of any one of the weight matrices;

[0015] According to the arrangement order, sequentially store the elements in the (i + 1)-th column of each weight matrix in the target array until all the data in all the weight matrices are stored in the target array to obtain the target weight matrix.

[0016] In a possible implementation manner of the first aspect above, the weight matrix includes a query weight matrix, a key weight matrix, and a value weight matrix. The arrangement order indicates that the storage priority of the matrix data of the query weight matrix in the target array is higher than the storage priority of the matrix data of the key weight matrix in the target array, and the storage priority of the matrix data of the key weight matrix in the target array is higher than the storage priority of the matrix data of the value weight matrix in the target array.

[0017] In a possible implementation manner of the first aspect above, the method further includes:

[0018] Input the target weight matrix into a quantization algorithm to obtain a quantized target weight matrix;

[0019] Process the input data based on the quantized target weight matrix; where the quantization algorithm includes quantization parameters corresponding to each weight matrix.

[0020] In a possible implementation manner of the first aspect above, the inputting the target weight matrix into a quantization algorithm to obtain a quantized target weight matrix includes:

[0021] Traverse the target weight matrix and sequentially extract the matrix data of the target weight matrix;

[0022] Input each matrix data into the quantization algorithm respectively to obtain a quantized target weight matrix; where the quantization algorithm is used to quantize each matrix data through the quantization parameters, and the quantization parameters are determined based on the matrix type of each matrix data.

[0023] In a possible implementation of the above first aspect, before respectively inputting each of the matrix data into the quantization algorithm to obtain a quantized target weight matrix, the method further includes:

[0024] Determine the storage location of each matrix data in the target weight matrix according to the arrangement order;

[0025] Determine the traversal order of traversing the target weight matrix, and determine the matrix data currently traversed and extracted and its corresponding matrix type according to the traversal order and the storage location;

[0026] Determine the quantization parameter corresponding to the matrix type, so as to quantize each matrix data based on the quantization parameter.

[0027] In a second aspect, a data processing device includes:

[0028] An acquisition module, configured to acquire a target weight matrix; wherein, the target weight matrix is obtained by combining a plurality of weight matrices related to the self-attention mechanism; the target weight matrix includes each column data in each of the weight matrices, and the data in each column of each weight matrix are arranged in sequence according to a preset interleaving rule;

[0029] An input module, configured to acquire input data;

[0030] A determination module, configured to input the input data into the target weight matrix to obtain the self-attention score of the input data.

[0031] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the data processing method provided in the above first aspect or any possible implementation of the first aspect.

[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the data processing method provided in the above first aspect or any possible implementation of the first aspect.

[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program runs on a computer, it causes the computer to execute the data processing method provided in the above first aspect or any possible implementation of the first aspect.

[0034] It can be understood that the beneficial effects of the above second to fifth aspects can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here.

[0035] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0036] In the embodiments of the present application, by obtaining a target weight matrix, which is obtained by combining a plurality of weight matrices related to the self-attention mechanism, the target weight matrix includes each column of data in each weight matrix, and each column of data in each weight matrix is arranged in sequence according to a preset interleaving rule. Then, after obtaining the input data, only one data loading is required, and the input data can directly participate in the operations of all weight matrices to obtain the self-attention score of the input data, without repeatedly accessing the input data and separately inputting the input data into the query weight matrix, key weight matrix, and value weight matrix. This reduces the situation of repeatedly reading and accessing the same input data, reduces the memory access time consumption, and improves the efficiency of calculating the self-attention score. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart of the steps of a data processing method provided by an embodiment of the present application;

[0038] Figure 2 is a flowchart of the steps of another data processing method provided by an embodiment of the present application;

[0039] Figure 3 is a schematic diagram of extracting matrix data provided by an embodiment of the present application;

[0040] Figure 4 is a schematic diagram of storing matrix data provided by an embodiment of the present application;

[0041] Figure 5 is a schematic diagram of processing for determining a target weight matrix provided by an embodiment of the present application;

[0042] Figure 6 is a schematic flowchart of applying a target weight matrix provided by an embodiment of the present application;

[0043] Figure 7 is a schematic structural diagram of a data processing device provided by an embodiment of the present application;

[0044] Figure 8 is a schematic block diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0046] The self-attention mechanism is an attention mechanism based on human vision. Its core lies in being able to screen out more important information from a large amount of information while ignoring other less important information. At present, the self-attention mechanism has been widely applied to deep learning tasks such as natural language processing and text generation.

[0047] In the calculation process of the self-attention mechanism, the core operations include the calculation of the weight matrices of query (Query), key (Key) and value (Value) vectors. The specific operation is to obtain the input sequence in deep learning tasks such as natural language processing and text generation, and calculate the query vector of each input data in the input sequence through the query weight matrix, and, calculate the key vector of each input data in the input sequence through the key weight matrix, and calculate the value vector of each input data in the input sequence through the value weight matrix. And in this process, each input data in the input sequence is read by the query weight matrix, the key weight matrix and the value weight matrix respectively, that is, there is a situation of multiple read accesses to the same input data, and the memory access time consumption is relatively high, reducing the efficiency of calculating the self-attention score.

[0048] Based on this, the present application provides a data processing method. By obtaining a target weight matrix, the target weight matrix is composed of multiple weight matrices related to the self-attention mechanism. The target weight matrix includes each column of data in each weight matrix, and each column of data in each weight matrix is arranged in sequence according to a preset interleaving rule. By obtaining the input data and inputting the input data into the target weight matrix to obtain the self-attention score of the input data, it is only necessary to input the input data into the target weight matrix to obtain the self-attention score of the input data, without multiple accesses to the input data and inputting the input data into the query weight matrix, the key weight matrix and the value weight matrix respectively, reducing the situation of multiple read accesses to the same input data, reducing the memory access time consumption, and improving the efficiency of calculating the self-attention score.

[0049] See Figure 1 , Figure 1 shows the step flow chart of a data processing method provided by an embodiment of the present application, which may specifically include the following steps:

[0050] Step 101, obtain a target weight matrix.

[0051] Among them, the target weight matrix can be obtained by permuting and combining multiple weight matrices related to the self-attention mechanism. The target weight matrix includes the data in each column of each weight matrix, and the data in each column of each weight matrix is rearranged according to a preset interleaving rule. The interleaving rule can be a rule for arranging the data in each column according to a preset arrangement order, and the arrangement order can be the order of arranging the data in the weight matrix.

[0052] For example, the weight matrices can include a first weight matrix, a second weight matrix, and a third weight matrix. Taking the first weight matrix, the second weight matrix, and the third weight matrix as P rows × Q columns as an example, and taking the target weight matrix as an example that can be obtained by combining the column data of the first weight matrix, the second weight matrix, and the third weight matrix related to the self-attention mechanism, the target weight matrix can include: the first column of the first weight matrix, the first column of the second weight matrix, and the first column of the third weight matrix, and include the second column of the first weight matrix, the first column of the second weight matrix, and the second column of the third weight matrix,..., the Qth column of the first weight matrix, the Qth column of the second weight matrix, and the Qth column of the third weight matrix. Or it can also be considered that the target weight matrix includes Q sub-weight matrices, and any sub-weight matrix is obtained by arranging the column data of the first weight matrix, the second weight matrix, and the third weight matrix according to a preset interleaving rule, and different sub-weight matrices include the data in different columns of the first weight matrix, the second weight matrix, and the third weight matrix.

[0053] For example, the first weight matrix in the embodiments of the present application can be a query weight matrix. The second weight matrix can be a key weight matrix. The third weight matrix can be a value weight matrix. Of course, the first weight matrix, the second weight matrix, and the third weight matrix can also be other weight matrices, and the embodiments of the present application do not limit this.

[0054] Among them, the query weight matrix can be a matrix used to calculate the query vector in the self-attention mechanism, the key weight matrix can be a matrix used to calculate the key vector in the self-attention mechanism, and the value weight matrix can be a matrix used to calculate the value vector in the self-attention mechanism. The query vector can represent the features of the input data used for matching and correlation calculation with other input data in the self-attention mechanism. The key vector can represent the features of the input data used for matching and retrieval with other input data in the self-attention mechanism. The query vector can be used for dot product operation with the key vector to determine the correlation between the input data and other input data in the input sequence. The value vector can represent the specific content of the input data in this specific task.

[0055] In the embodiments of the present application, the query weight matrix, the key weight matrix, and the value weight matrix may be matrices with the same number of rows and columns. For example, the query weight matrix, the key weight matrix, and the value weight matrix are all matrices of 5 rows × 8 columns or matrices of 6 rows × 3 columns. The embodiments of the present application do not limit this.

[0056] When performing deep learning tasks such as natural language processing and text generation, a preset target weight matrix can be obtained.

[0057] In practical applications, different deep learning models can be pre-trained for different deep learning tasks such as natural language processing and text generation. Furthermore, the preset target weight matrix can be obtained from the pre-trained neural network model.

[0058] Specifically, the pre-trained neural network model may include a query weight matrix, a key weight matrix, and a value weight matrix related to the self-attention mechanism. Furthermore, the query weight matrix, the key weight matrix, and the value weight matrix related to the self-attention mechanism can be extracted from the pre-trained neural network model, and the extracted query weight matrix, key weight matrix, and value weight matrix can be combined to obtain the target weight matrix.

[0059] Among them, the combination method for combining the extracted query weight matrix, key weight matrix, and value weight matrix may include combining the query weight matrix, key weight matrix, and value weight matrix in a certain arrangement order, and dividing the query weight matrix, key weight matrix, and value weight matrix into multiple matrix blocks and combining each matrix block in a preset arrangement order, etc.

[0060] In an embodiment of the present application, step 101 may include steps 1011 to 1013:

[0061] Step 1011, obtaining multiple weight matrices related to the self-attention mechanism.

[0062] When performing deep learning tasks such as natural language processing and text generation, multiple weight matrices related to the self-attention mechanism can be extracted from the pre-trained neural network model. Specifically, the query weight matrix, the key weight matrix, and the value weight matrix can be extracted.

[0063] Step 1012, based on the self-attention mechanism, determining an interleaving rule for arranging each weight matrix.

[0064] Among them, the interleaving rule may include the arrangement order of each weight matrix.

[0065] After extracting multiple weight matrices, an interleaving rule for arranging each weight matrix can be determined based on the self-attention mechanism, and the arrangement order of each weight matrix can be determined from the interleaving rule.

[0066] In practical applications, based on the self-attention mechanism, the operation process of the input data during the actual calculation of the self-attention score can be determined, that is, the forward propagation process of the input data can be determined.

[0067] Specifically, a static computational graph related to the self-attention mechanism can be determined from a pre-trained neural network model, and the propagation direction of the input data can be determined through the edges in the static computational graph, and the processing operation of the input data at this node can be determined through the nodes in the static computational graph. Furthermore, based on the propagation direction of the input data and the processing operation of each node, the specific steps for each weight matrix to be applied to process the input data to calculate the self-attention score can be determined. For example, the specific steps when the query weight matrix, key weight matrix, and value weight matrix process the input data can be determined, and an interleaving rule for arranging each weight matrix can be determined based on the specific steps corresponding to each weight matrix, and the arrangement order of each weight matrix can be determined from the interleaving rule.

[0068] Step 1013, arrange and combine each column of data in each weight matrix according to the arrangement order in the interleaving rule to obtain a target weight matrix.

[0069] After determining the arrangement order, each column of data in each weight matrix can be arranged and combined according to the arrangement order in the interleaving rule to obtain a target weight matrix.

[0070] It should be understood that based on the self-attention mechanism, the specific steps of each weight matrix during the self-attention calculation can be determined. Furthermore, based on the specific steps of each weight matrix, the arrangement order of each weight matrix can be determined, and each column of data in each weight matrix can be arranged and combined according to the arrangement order to obtain a target weight matrix. And inputting the input data into the target weight matrix arranged according to the arrangement order can make the process of processing the input data in the target weight matrix the same as the operation process of the static computational graph related to the self-attention mechanism, that is, the error between the self-attention score calculated through the target weight matrix and the self-attention score calculated by respectively determining the query vector, key vector, and value vector through the query weight matrix, key weight matrix, and value weight matrix and then calculating through the query vector, key vector, and value vector is less than the preset range.

[0071] In an embodiment of the present application, the following steps may further be included:

[0072] Input the target weight matrix into the quantization algorithm to obtain the quantized target weight matrix, and process the input data based on the quantized target weight matrix.

[0073] Among them, the quantization algorithm can be used to quantize each matrix data through quantization parameters. The quantization algorithm can be used to convert the weight values of the target weight matrix into weight values with lower precision to reduce the storage space occupied by the target weight matrix. The quantization algorithm can include quantization parameters corresponding to each weight matrix. Specifically, it can include quantization parameters corresponding to the query weight matrix, quantization parameters corresponding to the key weight matrix, and quantization parameters corresponding to the value weight matrix. The quantization parameters can be used to determine the quantization precision and quantization range of the target weight matrix.

[0074] After obtaining the target weight matrix, the target weight matrix can be input into a preset quantization algorithm. Then, the quantization algorithm can use the quantization parameters corresponding to each weight matrix to quantize the target weight matrix to obtain the quantized target weight matrix, so as to process the input data based on the quantized target weight matrix.

[0075] In practical applications, since the precision of the weight values of the quantized target weight matrix is relatively low, the computing resources and bandwidth required for processing the input data are less, thereby improving the operation speed of processing the input data and the processing efficiency of the input data.

[0076] It should be understood that the process of self-attention calculation can include stages such as a preprocessing stage, a self-attention score calculation stage, an attention weight normalization stage, and a weighted summation stage. Before calculating the self-attention score, that is, in the preprocessing stage, the quantization parameters corresponding to each weight matrix can be obtained from a pre-trained neural network model. For example, the quantization parameters corresponding to the query weight matrix can be obtained, the quantization parameters corresponding to the key weight matrix can be obtained, and the quantization parameters corresponding to the value weight matrix can be obtained, and a quantization operator can be configured based on each quantization parameter to quantize the target weight matrix based on the same quantization operator to obtain the quantized target weight matrix.

[0077] Specifically, when quantizing the target weight matrix based on the same quantization operator, the quantization operator can select the corresponding quantization parameters from the configured multiple quantization parameters to quantize the target weight matrix.

[0078] In a specific implementation, if a deep learning task is executed by separately determining the query vector, key vector, and value vector of the input data and calculating the self-attention score using the query vector, key vector, and value vector, then in order to improve the processing efficiency, it is necessary to quantize the query weight matrix, key weight matrix, and value weight matrix respectively. Specifically, it is necessary to determine the quantization parameters of each weight matrix, and for the quantization parameters of different weight matrices, configure different quantization operators to quantize different weight matrices based on different quantization operators. Moreover, each quantization operator needs to be repeatedly called multiple times during the quantization process, which reduces the quantization efficiency.

[0079] However, by configuring a quantization operator based on the quantization parameters of each weight matrix and quantizing the target weight matrix based on this quantization operator, only one quantization operator needs to be called for processing, reducing the situation of repeatedly calling the quantization operator multiple times and improving the quantization efficiency.

[0080] Step 102, obtain the input data.

[0081] Among them, the input data can be the data input when performing deep learning tasks such as natural language processing and text generation. Specifically, when performing deep learning tasks such as natural language processing and text generation, the input data is usually one or more, and the one or more data are input in a sequence form.

[0082] When performing deep learning tasks such as natural language processing and text generation, the input related to the deep learning task can be obtained.

[0083] Step 103, input the input data into the target weight matrix to obtain the self-attention score of the input data.

[0084] After obtaining the input data, the obtained input data can be input into the target weight matrix. Then, the target weight matrix can process the input data to obtain the self-attention score of the input data, so as to perform the next operation in the deep learning task based on the obtained self-attention score.

[0085] In an embodiment of the present application, by obtaining a target weight matrix, which is obtained by combining multiple weight matrices related to the self-attention mechanism, the target weight matrix includes each column data in each weight matrix, and each column data in each weight matrix is arranged in sequence according to a preset interleaving rule. By obtaining input data and inputting the input data into the target weight matrix to obtain the self-attention score of the input data, it is only necessary to input the input data into the target weight matrix to obtain the self-attention score of the input data, without repeatedly accessing the input data and inputting the input data into the query weight matrix, key weight matrix, and value weight matrix respectively, reducing the situation of repeatedly reading and accessing the same input data, reducing the memory access time consumption, and improving the efficiency of calculating the self-attention score.

[0086] See Figure 2 , Figure 2 which shows a flowchart of steps of another data processing method provided by an embodiment of the present application, and specifically may include the following steps:

[0087] Step 201, obtain multiple weight matrices related to the self-attention mechanism.

[0088] Step 202, based on the self-attention mechanism, determine an interleaving rule for arranging each weight matrix.

[0089] Step 203, arrange and combine each column data in each weight matrix according to the arrangement order in the interleaving rule to obtain a target weight matrix.

[0090] In an embodiment of the present application, step 203 may include steps 2031 to 2032:

[0091] Step 2031, in accordance with the arrangement order in the interleaving rule, sequentially store the elements of the i-th column in each weight matrix in the target array.

[0092] Wherein, i is greater than or equal to 1, and i + 1 is less than or equal to N, N represents the number of columns of any weight matrix, the weight matrix may include a query weight matrix, a key weight matrix, and a value weight matrix, and the element may be matrix data in the weight matrix. The target array may be an array for storing data of all weight matrices, the target array may be a two-dimensional array, the number of rows of the target array may be the same as the maximum number of rows in all weight matrices, and the number of columns of the target array may be the sum of the number of columns of all weight matrices.

[0093] After determining the permutation order of each weight matrix, for each weight matrix, the elements of each column in the weight matrix can be extracted. For example, for the query weight matrix, the matrix data of each column in the query weight matrix can be extracted; for the key weight matrix, the matrix data of each column in the key weight matrix can be extracted; for the value weight matrix, the matrix data of each column in the value weight matrix can be extracted, and thus the matrix data of each column in each weight matrix can be obtained.

[0094] In a specific implementation, the matrix data in all weight matrices can be extracted sequentially. For example, first, the matrix data of the first column in all weight matrices can be extracted, obtaining the matrix data of the first column in the query weight matrix, the matrix data of the first column in the key weight matrix, and the matrix data of the first column in the value weight matrix. Then, the matrix data of the second column in all weight matrices can be extracted, and so on, and thus the matrix data of each column in all weight matrices can be obtained.

[0095] Participate Figure 3 , Figure 3 FIG. shows a schematic diagram of extracting matrix data provided by an embodiment of the present application. Taking the query weight matrix Q, the key weight matrix K, and the value weight matrix V as matrices with a two-column and three-row structure as an example for illustration.

[0096] In practical applications, the matrix data of each column in the query weight matrix Q, the key weight matrix K, and the value weight matrix V can be extracted, obtaining the matrix data Q1 of the first column and the matrix data Q2 of the second column in the query weight matrix Q, the matrix data K1 of the first column and the matrix data K2 of the second column in the key weight matrix K, and the matrix data V1 of the first column and the matrix data V2 of the second column in the value weight matrix V.

[0097] Among them, among all the matrix data, the matrix data of the first column can include the matrix data Q1, the matrix data K1, and the matrix data V1; the matrix data of the second column can include the matrix data Q2, the matrix data K2, and the matrix data V2.

[0098] After obtaining the matrix data of each column in each weight matrix, according to the permutation order in the interleaving rule, the elements of the i-th column in each weight matrix can be sequentially stored in the target array.

[0099] In practical applications, the elements of the i-th column in each weight matrix can be determined, and according to the permutation order in the interleaving rule, the elements of the i-th column can be sequentially stored in the target array.

[0100] Specifically, the permutation order can include the order of storing the matrix data of the i-th column in the target array. The permutation order can indicate the storage priority of each weight matrix, and the storage priority can be the priority of storing data in the target array.

[0101] When storing matrix data in a target array, generally, the matrix data with the highest priority is stored first. After storing the matrix data with the highest priority in the target array, the matrix data with the priority second only to the highest priority is then stored, and so on. Thus, all matrix data can be stored in the target array according to the storage priority of the matrix data.

[0102] In a specific implementation, according to the arrangement order, the storage priority of each matrix data in all the matrix data of the i-th column can be determined, and based on the storage priority of each matrix data, all the matrix data of the i-th column is stored in the target array.

[0103] Specifically, if the storage priority of the matrix data of the query weight matrix in the target array is higher than that of the matrix data of the key weight matrix in the target array, and the storage priority of the matrix data of the key weight matrix in the target array is higher than that of the matrix data of the value weight matrix in the target array, then the matrix data of the query weight matrix can be preferentially stored in the target array. After storing the matrix data of the query weight matrix, the matrix data of the key weight matrix is stored in the target array, and finally, the matrix data of the value weight matrix is stored in the target array.

[0104] Step 2032: Store the elements of the (i + 1)-th column in each weight matrix in the target array in sequence according to the arrangement order until all the data in all the weight matrices are stored in the target array, obtaining the target weight matrix.

[0105] After storing the elements of the i-th column in each weight matrix in the target array, the elements of the (i + 1)-th column in each weight matrix can be stored in the target array in sequence according to the arrangement order until all the data in all the weight matrices are stored in the target array, obtaining the target weight matrix.

[0106] In a specific implementation, the matrix data of each column can be stored in sequence, thereby determining the storage priority of the matrix data of each column, and based on the storage priority, the matrix data of each column is stored in the target array in sequence. For example, for the matrix data of the i-th column, the storage priority of the matrix data of the i-th column can be determined, that is, the storage priority of the matrix data of the i-th column in the query weight matrix, the key weight matrix, and the value weight matrix is determined, and according to the determined storage priority, the matrix data of the i-th column is stored in the target array in sequence. Similarly, the matrix data of the (i + 1)-th column can be stored in the target array until all the data in all the weight matrices are stored in the target array, obtaining the target weight matrix.

[0107] See Figure 4 , Figure 4 shows a schematic diagram of storing matrix data provided by an embodiment of the present application. As Figure 4As shown, the matrix data in the first column may include matrix data Q1, matrix data K1, and matrix data V1, and the matrix data in the second column may include matrix data Q2, matrix data K2, and matrix data V2. Furthermore, the storage priorities of the matrix data in each column can be determined.

[0108] Among them, the storage priority of matrix data Q1 is greater than that of matrix data K1, and the storage priority of matrix data K1 is greater than that of matrix data V1.

[0109] After obtaining the storage priorities of each matrix data, matrix data Q1 can be stored in the target array according to the storage priority, then matrix data K1 can be stored in target array 41, and finally matrix data V1 can be stored in target array 41. Similarly, the matrix data in the second column can be stored in target array 41 according to the storage priority to obtain target weight matrix 42.

[0110] Specifically, target array 41 can be a blank array with six columns and three rows, and target weight matrix 42 can be a matrix with six columns and three rows. It should be understood that the sizes of the above matrices and arrays are for illustrative purposes only.

[0111] It should be understood that the determination of the target weight matrix can be done in the preprocessing stage or when calculating the self-attention score is needed. Moreover, the target weight matrix can be used to calculate the self-attention score offline and online.

[0112] Step 204, obtain the target weight matrix.

[0113] Step 205, obtain the input data.

[0114] Step 206, input the target weight matrix into the quantization algorithm to obtain the quantized target weight matrix, and process the input data based on the quantized target weight matrix.

[0115] In an embodiment of the present application, the target weight matrix can also be quantized in the following manner:

[0116] Traverse the target weight matrix, and sequentially extract the matrix data of the target weight matrix. Each matrix data is respectively input into the quantization algorithm to obtain the quantized target weight matrix.

[0117] Among them, the quantization parameter is determined based on the matrix type of each matrix data. The matrix type can be the weight matrix to which the matrix data belongs, specifically including the matrix type of the query weight matrix, the matrix type of the key weight matrix, and the matrix type of the value weight matrix.

[0118] After obtaining the target weight matrix, the target weight matrix can be traversed, and the matrix data currently traversed can be extracted. Furthermore, the matrix data currently traversed can be input into a quantization algorithm to obtain quantized matrix data. After quantizing the matrix data currently traversed, the next matrix data is traversed, and similarly, the next matrix data is input into the quantization algorithm to obtain quantized matrix data. Thus, all matrix data can be quantized to obtain a quantized target weight matrix.

[0119] Specifically, the matrix data currently traversed can be input into a pre-configured quantization operator, and the quantization operator can determine the weight matrix to which the matrix data currently traversed belongs to obtain the matrix type of the matrix data. After determining the matrix type of the matrix data currently traversed, the quantization operator can call the quantization parameters corresponding to the matrix type of the matrix data to quantize the matrix data and obtain quantized matrix data.

[0120] In an embodiment of the present application, before respectively inputting each matrix data into the quantization algorithm to obtain a quantized target weight matrix, the following steps may further be included:

[0121] According to the arrangement order, determine the storage position of each matrix data in the target weight matrix, determine the traversal order of traversing the target weight matrix, and according to the traversal order and the storage position, determine the matrix data currently traversed and extracted and its corresponding matrix type, and determine the quantization parameters corresponding to the matrix type, so as to quantize each matrix data based on the quantization parameters.

[0122] Among them, the traversal order may be the order of traversing the target weight matrix.

[0123] After obtaining the target weight matrix, the storage position of each matrix data in the target weight matrix can be determined according to the arrangement order of each weight matrix.

[0124] Exemplarily, as Figure 4 shown, according to the arrangement order of the query weight matrix, the key weight matrix, and the value weight matrix, it can be determined that the matrix data Q1 in the first column of the query weight matrix is stored in the first column of the target weight matrix 42, the matrix data V1 in the first column of the value weight matrix is stored in the second column of the target weight matrix 42, and the matrix data Q2 in the second column of the query weight matrix is stored in the fourth column of the target weight matrix 42.

[0125] After determining the storage position of each matrix data in the target weight matrix, the traversal order of traversing the target weight matrix can be determined. Specifically, the traversal order of traversing the target weight matrix can be the same as the arrangement order.

[0126] After determining the traversal order, the target weight matrix can be traversed according to the traversal order, and the matrix data of the current traversal can be extracted. According to the storage location, the matrix type corresponding to the matrix data of the current traversal can be determined. Furthermore, the quantization parameter corresponding to the matrix data can be determined according to the matrix type of the matrix data, so as to quantize each matrix data based on the quantization parameter.

[0127] Exemplarily, the matrix type corresponding to the matrix data of the current traversal can be the matrix type of the query weight matrix, and then the quantization parameter corresponding to the query weight matrix can be determined as the quantization parameter for quantizing the matrix data.

[0128] It should be understood that by determining the matrix type of the matrix data of the current traversal, the corresponding quantization parameter can be called to quantize the matrix data, that is, for the matrix data of different matrix types, different quantization parameters can be called for quantization, which improves the accuracy of quantization based on the same quantization operator.

[0129] Step 207, input the input data into the quantized target weight matrix to obtain the self-attention score of the input data.

[0130] In the embodiments of the present application, multiple weight matrices related to the self-attention mechanism are obtained, and based on the self-attention mechanism, an interleaving rule for arranging each weight matrix is determined, and the arrangement order of each weight matrix in the interleaving rule is obtained, so as to perform permutation and combination on each column of data in each weight matrix according to the arrangement order to obtain the target weight matrix. After obtaining the input data, the target weight matrix is input into the quantization algorithm to obtain the quantized target weight matrix, so as to process the input data based on the quantized target weight matrix, and input the input data into the quantized target weight matrix to obtain the self-attention score of the input data. Then, it is only necessary to input the input data into the target weight matrix to obtain the self-attention score of the input data, without repeatedly accessing the input data and inputting the input data into the query weight matrix, the key weight matrix, and the value weight matrix respectively, reducing the situation of repeatedly reading and accessing the same input data, thereby reducing the number of memory accesses, reducing the memory access time consumption, and improving the efficiency of calculating the self-attention score.

[0131] See Figure 5 , Figure 5 shows a processing schematic diagram of determining a target weight matrix provided by an embodiment of the present application. As Figure 5 shown, when it is necessary to determine the target weight matrix, the query weight matrix Q, the key weight matrix K, and the value weight matrix V related to the self-attention mechanism can be extracted from the original model file, and the static computational graph related to the self-attention mechanism can be extracted from the original model file.

[0132] Among them, the original model file can be a pre-trained deep learning model. Specifically, the storage format of the deep learning model includes, but is not limited to, the ONNX model representation format, the Pytorch model file, and the TensorFlow Checkpoint format.

[0133] In practical applications, the original model file can be loaded through a parsing tool for the deep learning model, such as loading the original model file through parsing tools such as ONNX Runtime and TorchScript. Furthermore, the static computational graph of the original model file and the weight matrices related to the self-attention mechanism can be parsed, such as the query weight matrix Q, the key weight matrix K, and the value weight matrix V.

[0134] After obtaining the query weight matrix Q, the key weight matrix K, and the value weight matrix V, the matrix data of each column in the query weight matrix Q, the key weight matrix K, and the value weight matrix V can be extracted to obtain data such as matrix data Q1, matrix data Q2, matrix data K1, matrix data K2, matrix data V1, and matrix data V2. Then, each matrix data is interleaved and stored in the target array according to a preset fusion order to obtain the target weight matrix.

[0135] Specifically, interleaved storage can be a process of storing each matrix data based on the storage priority of the matrix data. It should be understood that interleaved storage can be expressed as a process of storing data sequentially according to an interleaving rule. In this embodiment, the interleaving rule can be expressed as a rule of arranging matrix data sequentially according to a preset arrangement order. Then, the process of storing each matrix data in the target array in sequence according to the storage priority can be a way of interleaved storage.

[0136] At the same time, after obtaining the static computational graph related to the self-attention mechanism, the quantization parameters corresponding to each weight matrix can be determined from the static computational graph to obtain the quantization parameters of the query weight matrix, the quantization parameters of the key weight matrix, and the quantization parameters of the value weight matrix. Then, according to the quantization parameters of each weight matrix, quantization operators are configured to store the configured quantization operators and the target weight matrix in the original model file to obtain a simplified model file.

[0137] In an embodiment of the present application, by extracting the query weight matrix, key weight matrix, and value weight matrix related to the self-attention mechanism from the original model file, based on the self-attention mechanism, setting the fusion order of the query weight matrix, key weight matrix, and value weight matrix, and combining the query weight matrix, key weight matrix, and value weight matrix according to the fusion order to obtain the target weight matrix. Since the target weight matrix is composed of the query weight matrix, key weight matrix, and value weight matrix, it is only necessary to input the input data into the quantized target weight matrix to obtain the self-attention score of the input data, without repeatedly accessing the input data and inputting the input data into the query weight matrix, key weight matrix, and value weight matrix respectively, reducing the situation of repeatedly reading and accessing the same input data, thereby reducing the number of memory accesses, reducing the memory access time consumption, and improving the efficiency of calculating the self-attention score. Also, by extracting the static computation graph related to the self-attention mechanism from the original model file, and determining the quantization parameters corresponding to each weight matrix from the static computation graph, obtaining the quantization parameters of the query weight matrix, the quantization parameters of the key weight matrix, and the quantization parameters of the value weight matrix, configuring the quantization operator according to the quantization parameters of each weight matrix, and quantizing the target weight matrix based on the configured quantization operator, it is possible to reduce the computing resources and bandwidth required for the target weight matrix to process the input data, and further improve the operation speed of processing the input data and the processing efficiency of the input data.

[0138] See Figure 6 , Figure 6 which shows a schematic flowchart of a process of applying a target weight matrix provided by an embodiment of the present application, and specifically may include the following steps:

[0139] Step 601, obtain the target weight matrix and obtain the input data;

[0140] Step 602, based on a pre-configured quantization operator, quantize the target weight matrix to obtain a quantized target weight matrix;

[0141] Step 603, input the input data into the quantized target weight matrix to obtain the self-attention score of the input data.

[0142] In an embodiment of the present application, by obtaining the target weight matrix, obtaining the input data, and based on a pre-configured quantization operator, quantizing the target weight matrix to obtain a quantized target weight matrix, it is then possible to input the input data into the quantized target weight matrix to obtain the self-attention score of the input data. And quantizing the target weight matrix based on the configured quantization operator can reduce the computing resources and bandwidth required for the target weight matrix to process the input data, and further improve the operation speed of processing the input data and the processing efficiency of the input data.

[0143] See Figure 7 , Figure 7 which shows a schematic structural diagram of a data processing device provided in an embodiment of the present application, and may specifically include the following modules:

[0144] An acquisition module 701, configured to acquire a target weight matrix; wherein, the target weight matrix is obtained by combining a plurality of weight matrices related to the self-attention mechanism; the target weight matrix includes each column of data in each weight matrix, and the data in each column of each weight matrix are arranged in sequence according to a preset interleaving rule;

[0145] An input module 702, configured to acquire input data;

[0146] A determination module 703, configured to input the input data into the target weight matrix to obtain the self-attention score of the input data.

[0147] In one implementation, the above-mentioned acquisition module 701 may also be used for:

[0148] Acquire a plurality of weight matrices related to the self-attention mechanism;

[0149] Based on the self-attention mechanism, determine an interleaving rule for arranging each weight matrix; wherein, the interleaving rule includes the arrangement order of each weight matrix;

[0150] Arrange and combine each column of data in each weight matrix according to the arrangement order in the interleaving rule to obtain the target weight matrix.

[0151] In one implementation, the above-mentioned acquisition module 701 may also be used for:

[0152] Sequentially store the elements in the i-th column of each weight matrix in the target array according to the arrangement order in the interleaving rule; wherein, i is greater than or equal to 1, and i + 1 is less than or equal to N, and N represents the number of columns of any weight matrix;

[0153] Sequentially store the elements in the (i + 1)-th column of each weight matrix in the target array according to the arrangement order, until all the data in all the weight matrices are stored in the target array to obtain the target weight matrix.

[0154] In one implementation, the weight matrix includes a query weight matrix, a key weight matrix, and a value weight matrix, and the arrangement order indicates that the storage priority of the matrix data of the query weight matrix in the target array is higher than the storage priority of the matrix data of the key weight matrix in the target array, and the storage priority of the matrix data of the key weight matrix in the target array is higher than the storage priority of the matrix data of the value weight matrix in the target array.

[0155] In one implementation, the apparatus may further include the following modules:

[0156] A quantization module, configured to input the target weight matrix into a quantization algorithm to obtain a quantized target weight matrix;

[0157] Process the input data based on the quantized target weight matrix; wherein, the quantization algorithm includes quantization parameters corresponding to each weight matrix.

[0158] In one implementation, the above quantization module may also be used for:

[0159] Traverse the target weight matrix and sequentially extract the matrix data of the target weight matrix;

[0160] Input each matrix data into the quantization algorithm respectively to obtain a quantized target weight matrix; wherein, the quantization algorithm is used to quantize each matrix data through quantization parameters, and the quantization parameters are determined based on the matrix type of each matrix data.

[0161] In one implementation, the above quantization module may also be used for:

[0162] Before inputting each matrix data into the quantization algorithm respectively to obtain a quantized target weight matrix, determine the storage position of each matrix data in the target weight matrix according to the fusion order;

[0163] Determine the traversal order of traversing the target weight matrix, and determine the currently traversed and extracted matrix data and its corresponding matrix type according to the traversal order and the storage position;

[0164] Determine the quantization parameters corresponding to the matrix type, so as to quantize each matrix data based on the quantization parameters.

[0165] In the embodiments of the present application, by obtaining a target weight matrix, which is obtained by combining multiple weight matrices related to the self-attention mechanism, the target weight matrix includes each column data in each weight matrix, and each column data in each weight matrix is arranged in sequence according to a preset interleaving rule, and by obtaining input data and inputting the input data into the target weight matrix to obtain the self-attention score of the input data, it is only necessary to input the input data into the target weight matrix to obtain the self-attention score of the input data, without accessing the input data multiple times and inputting the input data into the query weight matrix, the key weight matrix, and the value weight matrix respectively, reducing the situation of multiple read accesses to the same input data, reducing the memory access time consumption, and improving the efficiency of calculating the self-attention score.

[0166] It should be noted that for the information interaction, execution process, etc. between the above devices, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0167] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.

[0168] See Figure 8 , Figure 8 FIG. shows a structural block diagram of a terminal device provided in an embodiment of the present application. As Figure 8 shown, this embodiment provides a terminal device 81, and the terminal device 81 includes: at least one processor 811, a memory 812, and a computer program 8121 stored in the memory 812 and executable on at least one processor 811. When the processor 811 executes the computer program 8121, the steps in any of the above method embodiments are implemented.

[0169] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments can be implemented.

[0170] The embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above method embodiments when executed.

[0171] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium.

[0172] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for data processing, characterized in that: The method comprises: Obtain a target weight matrix; wherein the target weight matrix is ​​obtained by combining multiple weight matrices related to the self-attention mechanism; the target weight matrix includes each column of data in each of the weight matrices, and the data in each column of each of the weight matrices are arranged in sequence according to a preset interleaving rule; Get input data; The input data is input into the target weight matrix to obtain a self-attention score of the input data.

2. The data processing method according to claim 1, characterized in that: The obtaining of the target weight matrix comprises: Obtaining a plurality of weight matrices associated with the self-attention mechanism; Based on the self-attention mechanism, determining an interleaving rule for arranging each of the weight matrices; wherein the interleaving rule includes an arrangement order of each of the weight matrices; Each column of data in each weight matrix is ​​arranged and combined according to the arrangement order in the interleaving rule to obtain the target weight matrix.

3. The data processing method according to claim 2, characterized in that: The step of arranging and combining each column of data in each weight matrix according to the arrangement order in the interleaving rule to obtain the target weight matrix includes: According to the arrangement order in the interleaving rule, the elements of the i-th column in each of the weight matrices are stored in the target array in sequence; wherein i is greater than or equal to 1, and i+1 is less than or equal to N, and N represents the number of columns of any of the weight matrices; The elements of the i+1th column in each of the weight matrices are sequentially stored in the target array according to the arrangement order in the interleaving rule until the data in all the weight matrices are stored in the target array, thereby obtaining the target weight matrix.

4. The data processing method according to claim 3, characterized in that: The weight matrix includes a query weight matrix, a key weight matrix, and a value weight matrix. The arrangement order indicates that the storage priority of the matrix data of the query weight matrix in the target array is higher than the storage priority of the matrix data of the key weight matrix in the target array, and the storage priority of the matrix data of the key weight matrix in the target array is higher than the storage priority of the matrix data of the value weight matrix in the target array.

5. The data processing method according to any one of claims 2 to 4, characterized in that: The method further comprises: Inputting the target weight matrix into a quantization algorithm to obtain a quantized target weight matrix; The input data is processed based on the quantized target weight matrix; wherein the quantization algorithm includes quantization parameters corresponding to each of the weight matrices.

6. The data processing method according to claim 5, characterized in that: The step of inputting the target weight matrix into a quantization algorithm to obtain a quantized target weight matrix includes: Traversing the target weight matrix and extracting matrix data of the target weight matrix in sequence; Input each of the matrix data into the quantization algorithm respectively to obtain a quantized target weight matrix; wherein the quantization algorithm is used to quantize each of the matrix data through the quantization parameter, and the quantization parameter is determined based on the matrix type of each of the matrix data.

7. The data processing method according to claim 6, characterized in that: Before inputting each of the matrix data into the quantization algorithm to obtain the quantized target weight matrix, the method further includes: Determine, according to the arrangement order, a storage location for each of the matrix data in the target weight matrix; Determine a traversal order for traversing the target weight matrix, and determine the matrix data currently traversed and extracted and its corresponding matrix type according to the traversal order and the storage location; A quantization parameter corresponding to the matrix type is determined to quantize each of the matrix data based on the quantization parameter.

8. A data processing device, characterized in that: The device comprises: An acquisition module, used to acquire a target weight matrix; wherein the target weight matrix is ​​obtained by combining multiple weight matrices related to the self-attention mechanism; the target weight matrix includes each column of data in each of the weight matrices, and the data in each column of each of the weight matrices are arranged in sequence according to a preset interleaving rule; Input module, used to obtain input data; A determination module is used to input the input data into the target weight matrix to obtain a self-attention score of the input data.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.