A multi-head attention quantization method based on task mode perception

By employing pattern recognition technology and different precision quantization methods for the attention heads in multi-head attention mechanisms, the problems of low processing efficiency and poor compression effect of multi-head attention mechanism tasks are solved, thereby accelerating and compressing model tasks.

CN120525004BActive Publication Date: 2025-11-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511028139.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-18
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Multi-head attention mechanisms have low task processing efficiency and poor model compression performance. Existing mass-production methods ignore the differences in importance of attention heads, resulting in additional computation time and reduced model processing performance.

Method used

By using pattern recognition technology, based on the distribution pattern of attention heads in the attention weight matrix, different precision quantization processing methods are adopted for different types of model tasks to identify global, positional, extended, and uniform patterns, and to perform first-precision or second-precision calculations respectively.

Benefits of technology

It reduces the total computation time, improves the processing efficiency of model tasks, and enhances the processing effect of model tasks, achieving quantization acceleration and compression of models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525004B_ABST
    Figure CN120525004B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of deep learning, and discloses a multi-head attention weight quantification method based on task mode perception. The method comprises the following steps: for each type of model task, obtaining an attention weight matrix of each attention head in a multi-head attention mechanism when the multi-head attention mechanism is used to process input data of the model task; performing mode identification on each attention head of the model task according to the large-value distribution of all attention weights in the attention weight matrix; wherein the large-value distribution of the attention weights is used to indicate the attention distribution of all tokens in the input data; and performing quantification processing when the multi-head attention mechanism is used to process the input data of the model task according to the mode identification result of each attention head of the model task, so that the task processing efficiency and processing effect of the multi-head attention mechanism are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a multi-head attention quantization method, apparatus, device, and medium based on task pattern awareness. Background Technology

[0002] The Transformer architecture for deep neural networks has demonstrated exceptional performance in fields such as natural language processing and computer vision. The attention mechanism, as the core of this performance enhancement, enables the model to integrate global information and dynamically focus on content at specific locations. The use of multiple attention heads in the attention mechanism allows each attention head to learn and find different relationships between tokens, thereby further enhancing the model's capabilities. The computation process is as follows:

[0003] X is the feature vector x of each token in the input sequence. i Composition, with dimensions (N,d) model ), where N represents the number of tokens in the input sequence, and d model This represents the dimension of the eigenvector. It is determined by X and the weight matrix W. Q W K and W V Perform matrix multiplication, with the weight matrix having dimension (d). model ,d k This generates a query matrix (Q), a key matrix (K), and a value matrix (V), each with dimensions (N, d). k ), d k This represents the dimension of the query (q), key (k), and value (v) vector for each token. Then, Q, K, and V are divided into multiple groups, often called "headers," and the Q, K, and V are arranged along d... k Dimensional segmentation into Q i K i and V i .

[0004] Suppose there are n heads, and Q in each head i K i and V i The dimension will be (N,d) k / n), denoted as (N,D). Within each head, first through Q i and The original attention scores are calculated using matrix multiplication between the matrix D values. These scores are then scaled by the square root of D and passed through a softmax function to obtain the attention weight matrix A. i These attention weights are used to weight V. i Through A i and V iThe matrix multiplication generates the output of each head. The outputs of all heads are concatenated and passed through a final linear layer transformation to obtain the final output of the multi-head attention mechanism. Here, matrix A... i The values ​​in the matrix are considered as a measure of the association between different tokens. If the value at position (x, y) is larger, it means that the y-th token has a greater influence on the x-th token. Here, x and y are integers in the range [1, N]. Matrix A in different attention heads i The different distributions of the data capture various relationships between them, thus making the model more powerful.

[0005] Deep neural networks based on multi-head attention mechanisms have performed well in various fields, but at the cost of complexity dependence relative to the square of the number of input tokens. This is widely recognized as the main efficiency bottleneck for accelerating inference in such models, thus affecting the processing efficiency of multi-head attention-based model tasks. Quantization is one solution to accelerate the attention module and thus accelerate model task implementation.

[0006] Currently, the quantization method for multi-head attention modules is as follows: First, the Q and K matrices are set to a low-precision data format, and a coarse attention weight matrix A is calculated. i Then, elements with large values ​​in the matrix are considered important. Next, the q and k vectors of the tokens corresponding to the positions of the large values ​​are found, their data precision is improved, these important elements are recalculated, and the corresponding low-precision elements are replaced with the high-precision elements to form the final A. i This method requires adding an additional processing stage to determine A during the reasoning process for each input. i The distribution of these features increases the overall computation time for the model task. Furthermore, the attention module uses multiple heads to capture various dependencies, each contributing differently to the final output. This method uses the same quantization scheme for all attention heads—a mixture of high and low precision—ignoring the differences in importance between attention heads, reducing model compression, and thus lowering the overall performance of the model task. Summary of the Invention

[0007] The purpose of this invention is to provide a multi-head attention quantization method, apparatus, device, and medium based on task pattern awareness, which can solve the problems of low task processing efficiency and poor model compression effect for multi-head attention mechanisms.

[0008] To address the aforementioned technical problems, embodiments of the present invention provide a multi-head attention quantization method based on task pattern awareness, comprising the following steps:

[0009] For each type of model task, obtain the attention weight matrix of each attention head in the multi-head attention mechanism when processing the input data of the model task using the multi-head attention mechanism;

[0010] Based on the distribution of the largest values ​​of all attention weights in the attention weight matrix, pattern recognition is performed on each attention head of the model task; where the distribution of the largest values ​​is used to indicate the attention distribution of all tokens in the input data, and the largest value in the distribution of the largest values ​​is the value that is located at the beginning of the attention weight matrix after all values ​​in the attention weight matrix are sorted in descending order.

[0011] If the attention of all tokens in the input data is focused on at least one token, the attention head pattern is determined to be a global attention pattern; if the attention of each token in the input data is focused on the previous or next token, the attention head pattern is determined to be a positional attention pattern; if the attention of each token in the input data is focused on the token itself, on the same token, or on a token related to the token, the attention head pattern is determined to be an extended attention pattern; if the attention of all tokens in the input data is evenly distributed across all tokens in the input data, the attention head pattern is determined to be a uniform pattern.

[0012] Based on the type of the target model task, determine the pattern of each attention head for the target model task, and when using a multi-head attention mechanism to process the target input data of the target model task, perform the following quantization processing on the target input data according to the pattern of each attention head for the target model task:

[0013] If the attention head is in global attention mode or uniform mode, then the first precision calculation is performed on all target input data; if the attention head is in positional attention mode or extended attention mode, then the second precision calculation is performed on the target input data located at the position of the largest value in the attention weight matrix, and the first precision calculation is performed on the remaining positions; wherein, the first precision is less than the second precision.

[0014] Optionally, the step of performing pattern recognition on each attention head of the model task based on the maximum value distribution of all attention weights in the attention weight matrix includes:

[0015] Obtain the first percentage of the average attention weight of each column in the attention weight matrix relative to the sum of the average attention weights of all columns in the entire attention weight matrix;

[0016] The average attention weight of each diagonal line diagonally downwards to the right at 45° from positions (2,1), (1,1), and (1,2) in the attention weight matrix is ​​obtained sequentially. The average attention weight of each diagonal line diagonally downwards to the right at 45° from positions (1,N-2), (1,N-3), ..., (1,3) is also obtained, along with the diagonal line pairs formed by the diagonal lines diagonally downwards to the right at positions (N-2,1), (N-3,1), ..., (3,1).

[0017] Get the sum of all average attention weights, and get the percentage of each average attention weight to this sum of average attention weights; where the percentage of the average attention weight corresponding to each diagonal line diagonally downward to the right of (2,1) and (1,2) is the second percentage, and the rest are the third percentage;

[0018] Sort all first percentages, all second percentages, and all third percentages in descending order, and select the target percentages that are greater than a preset threshold and are among the top three.

[0019] If the target percentage is the first percentage, the attention head mode is determined to be the global attention mode; if the target percentage is the second percentage, the attention head mode is determined to be the positional attention mode; if the second target percentage is the third percentage, the attention head mode is determined to be the extended attention mode; if the target percentage includes both the first and second percentages, the attention head mode is determined to be the positional attention mode; if the target percentage includes both the first and third percentages, the attention head mode is determined to be the extended attention mode; if the target percentage includes both the second and third percentages, the attention head mode is determined to be the extended attention mode; if the target percentage includes the first, second, and third percentages, the attention head mode is determined to be the extended attention mode.

[0020] If none of the first, second, and third percentages exceeds a preset threshold, then the attention head mode is determined to be uniform mode.

[0021] Optionally, based on the pattern of each attention head in the target model task, the target input data is quantized as follows:

[0022] If the attention head is in global attention mode or uniform mode, then it will use the input data and weight matrix W Q W K and W V The generated matrix Q i K i and V iThe data is quantized to the first precision for computation, and the quantized Q is processed using an attention mechanism. i K i and V i After obtaining the attention weight matrix A i Then, A i After being quantized into a first-precision data format, it is then compared with V. i Perform matrix multiplication calculations.

[0023] Optionally, based on the pattern of each attention head in the target model task, the target input data is quantized as follows:

[0024] If the attention head pattern is the positional attention pattern, and A i The largest values ​​of all attention weights are distributed along a 45° downward-right slope starting from position (1,2), or along a 45° downward-right slope starting from position (2,1); in calculating A i When calculating, the vectors q and k corresponding to the values ​​on the diagonal lines are quantized to a second-precision data format for calculation, while the q and k corresponding to the values ​​at other positions are quantized to a first-precision data format for calculation; after obtaining A... i Then, the values ​​on the corresponding diagonal lines are quantized to a second-precision data format, while the remaining positions are quantized to a first-precision data format for calculation; in A i With V i During matrix multiplication, A will be involved. i The values ​​at the slash positions correspond to the vector v being calculated, which is also quantized to the second-precision data format for calculation, while the rest are quantized to the first-precision data format for calculation.

[0025] Optionally, based on the pattern recognition results of each attention head in the model task, the following quantization processing is performed when using a multi-head attention mechanism to process the input data of the model task:

[0026] If the attention head pattern is the extended attention pattern, then determine A. i Is the oblique line of the large value distribution of all attention weights a pair of oblique lines symmetrical about the diagonal or the diagonal itself?

[0027] If it is a diagonal, then in calculating A i At that time, the attention weight vectors q and k corresponding to the diagonal are quantized into a second-precision data format for calculation, while the remaining positions are quantized into a first-precision data format for calculation; and in A i With V i During matrix multiplication, the attention weights on the diagonal and the vector v used for calculation are quantized into a second-precision data format for calculation, while the other positions and the vector v are quantized into a first-precision data format for calculation.

[0028] If it is a pair of diagonal lines that are symmetrical about the diagonal, then determine whether this is the first time that attention has been drawn to this situation;

[0029] If so, then Q i and K i The data is quantized to the first precision for calculation. The average value of each pair of diagonal lines symmetrical about the diagonal is calculated, and the pair of diagonal lines with the largest average value is used as the position for calculation in the second precision data format. In calculating A... i At that time, the vectors q and k corresponding to the attention weights of the diagonal pair with the largest average value are quantized into a second-precision data format for calculation, while the remaining positions are quantized into a first-precision data format for calculation; in A i With V i During matrix multiplication, the attention weights on the diagonal pair with the largest average value and the vector v used for calculation are quantized into a second-precision data format for calculation, while the other positions and vector v are quantized into a first-precision data format for calculation.

[0030] If not, then calculate A i At that time, the vectors q and k corresponding to the attention weights of the diagonal pair with the largest average value are quantized into a second-precision data format for calculation, while the remaining positions are quantized into a first-precision data format for calculation; in A i With V i During matrix multiplication, the attention weights on the diagonal pairs with the largest average value and the vector v used for calculation are quantized into a second-precision data format for calculation, while the remaining positions and vector v are quantized into a first-precision data format for calculation.

[0031] Optionally, the pattern recognition results may also include a global attention pattern mixed with a location attention pattern and a global attention pattern mixed with an extended attention pattern;

[0032] Based on the pattern of each attention head in the target model task, the target input data is quantized as follows:

[0033] If the attention head mode is a mixture of global attention mode and positional attention mode, then the quantization method corresponding to the positional attention mode shall be used;

[0034] If the attention head mode is a mixture of global attention mode and extended attention mode, then the quantization method corresponding to the extended attention mode is used.

[0035] Optionally, the step of performing pattern recognition on each attention head of the model task based on the maximum value distribution of all attention weights in the attention weight matrix includes:

[0036] For each attention head, obtain the pattern recognition results corresponding to multiple input data for the model task, and take the pattern recognition result with the highest frequency as the target pattern recognition result of the attention head.

[0037] Embodiments of the present invention also provide a multi-head attention quantization device based on task pattern awareness, comprising:

[0038] The weight acquisition module is used to obtain the attention weight matrix of each attention head in the multi-head attention mechanism when processing the input data of the model task using the multi-head attention mechanism for each type of model task.

[0039] The pattern recognition module is used to perform pattern recognition on each attention head of the model task based on the distribution of the largest values ​​of all attention weights in the attention weight matrix. The distribution of the largest values ​​is used to indicate the attention distribution of all tokens in the input data. The largest value in the distribution of the largest values ​​is the value that is located at the beginning of the attention weight matrix after all values ​​in the attention weight matrix are sorted in descending order.

[0040] If the attention of all tokens in the input data is focused on at least one token, the attention head pattern is determined to be a global attention pattern; if the attention of each token in the input data is focused on the previous or next token, the attention head pattern is determined to be a positional attention pattern; if the attention of each token in the input data is focused on the token itself, on the same token, or on a token related to the token, the attention head pattern is determined to be an extended attention pattern; if the attention of all tokens in the input data is evenly distributed across all tokens in the input data, the attention head pattern is determined to be a uniform pattern.

[0041] The task quantization module is used to determine the pattern of each attention head of the target model task based on the type of the target model task, and to perform the following quantization processing on the target input data when using a multi-head attention mechanism to process the target input data:

[0042] If the attention head is in global attention mode or uniform mode, then the first precision calculation is performed on all target input data; if the attention head is in positional attention mode or extended attention mode, then the second precision calculation is performed on the target input data located at the position of the largest value in the attention weight matrix, and the first precision calculation is performed on the remaining positions; wherein, the first precision is less than the second precision.

[0043] Embodiments of the present invention also provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described multi-head attention quantization method based on task pattern awareness.

[0044] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multi-head attention quantization method based on task pattern awareness.

[0045] The multi-head attention quantization method based on task pattern awareness provided by this invention has at least the following beneficial effects:

[0046] Because when facing a specific model task, the attention weight matrix A i The positions of important elements in the model do not change dynamically with the input data, i.e., they follow a fixed pattern. Therefore, in this invention, when using a multi-head attention mechanism to process the input data of different types of model tasks, pattern recognition is performed on each attention head through the attention weight matrix of each attention head in the multi-head attention mechanism. This yields the pattern recognition results for each attention head corresponding to different types of model tasks, and these pattern recognition results reflect the aforementioned attention weight matrix A. i The important element positions in the attention head follow a fixed pattern. Then, based on the pattern recognition results of the attention head, the input data processed by the attention head is processed with corresponding precision.

[0047] Therefore, for the same type of model task, even if the input data are different, the attention heads at the same layer and position follow the same pattern during the processing of each input data. Only one pattern recognition of the attention heads is needed. After determining the patterns of different attention heads, for different input data of the model task, it is not necessary to divide the first precision calculation into a separate calculation stage to predict the position of important elements. Instead, the corresponding precision (second precision or first precision) of each attention head is calculated directly according to the pattern of each attention head, thereby reducing the total computation time, realizing the quantitative acceleration of model tasks oriented towards multi-head attention mechanisms, and improving the processing efficiency of model tasks.

[0048] Furthermore, by dividing the attention heads into different modes and providing different precision schemes for each mode, it is not necessary to perform high first-precision mixed calculations on all attention heads. Some attention heads can be calculated with only first precision, thus achieving a larger model compression ratio and improving the processing performance of the model task. Attached Figure Description

[0049] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.

[0050] Figure 1 This is a flowchart of a multi-head attention quantization method based on task pattern awareness provided according to an embodiment of the present invention. Figure 1 ;

[0051] Figure 2 This is an attention weight matrix A for different attention head modes provided according to an embodiment of the present invention. i Visual diagram;

[0052] Figure 3 This is an example of an embodiment of the present invention, showing the attention weight matrix A of each attention head of the BERTbase model on an MRPC dataset. i A visualized heatmap;

[0053] Figure 4 This is a flowchart of a multi-head attention quantization method based on task pattern awareness provided according to an embodiment of the present invention. Figure 2 . Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the embodiments of the present invention to facilitate a better understanding of the invention. However, the technical solutions claimed in the present invention can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with and referenced by each other without contradiction.

[0055] One embodiment of the present invention relates to a task-pattern-aware multi-head attention quantization method. The implementation details of the task-pattern-aware multi-head attention quantization method of this embodiment are described in detail below. The following content is only for the convenience of understanding and is not necessary for implementing this solution.

[0056] The specific process of the multi-head attention quantization method based on task pattern awareness in this embodiment can be described as follows: Figure 1 As shown, it includes:

[0057] Step 101: For each type of model task, obtain the attention weight matrix of each attention head in the multi-head attention mechanism when processing the input data of the model task using the multi-head attention mechanism.

[0058] Specifically, this embodiment assumes that the BERTbase model, used for semantic similarity evaluation between sentence pairs, particularly determining whether two sentences are synonyms, is chosen as a model task. This task uses the MRPC dataset as its input data. The attention weight matrix A for each attention head in the multi-head attention mechanism is obtained when processing each sample data in the MRPC dataset using the multi-head attention mechanism. i .

[0059] Step 102: Based on the distribution of the largest values ​​of all attention weights in the attention weight matrix, perform pattern recognition on each attention head of the model task; wherein, the distribution of the largest values ​​is used to indicate the attention distribution of all tokens in the input data, and the largest value in the distribution of the largest values ​​is the value that is located at the beginning of the attention weight matrix after all values ​​in the attention weight matrix are sorted in descending order.

[0060] Specifically, when dealing with the same model task, even if the inputs are different, the A in the attention head at the same position will still be the same. i The distribution of A will also follow a similar pattern, namely A i The locations of medium and large values ​​are similar, and A i The distribution pattern followed by these patterns can be considered the pattern of attention heads. For example... Figure 2 (a) to Figure 2 The different modes of A shown in (d) i The visualization diagrams are shown, where: (a) is the global attention mode, (b) is the location attention mode, (c) is the extended attention mode, and (d) is the uniform mode.

[0061] In one example, for each attention head, the pattern recognition results corresponding to multiple input data for the model task are obtained, and the pattern recognition result with the highest frequency is taken as the target pattern recognition result of the attention head.

[0062] The following is a detailed explanation of these attention head modes:

[0063] Global attention pattern: This refers to a pattern where all tokens focus their attention on one or a few tokens (i.e., at least one token). For A... iVisualizing this, in a graph, large values ​​in the matrix are distributed along one or several vertical lines. Specific tokens that typically garner global attention are [CLS], [SEP], and punctuation marks. Therefore, in this embodiment, when performing pattern recognition on each attention head for each model task based on the attention distribution of all tokens in the input data indicated by the distribution of large values ​​of all attention weights in the attention weight matrix, if the attention of all tokens in the input data is concentrated on one or several tokens, then the pattern of the attention head is determined to be a global attention pattern.

[0064] Positional attention pattern: This refers to the current token's attention being focused on the previous / next token. In A... i This is represented as A in the visualization. i The large values ​​in the attention weight matrix are distributed along a 45° downward-right slope starting from position (1,2) or along a 45° downward-right slope starting from position (2,1). The former is denoted as sub-mode b-1, and the latter as sub-mode b-2; the two sub-modes will not appear simultaneously. Therefore, in this embodiment, when performing pattern recognition on each attention head for each model task based on the attention distribution of all tokens in the input data indicated by the large value distribution of all attention weights in the attention weight matrix, if the attention of each token in the input data is on the previous or next token, then the mode of the attention head is determined to be a positional attention mode.

[0065] Extended attention pattern: This refers to focusing attention on the same / related tokens (including situations where attention is focused on the token itself). This pattern is found in A. i In the visualization, this is represented as follows: large values ​​in the matrix are distributed along a 45° downward-southward diagonal starting from positions (1,2) or (2,1), usually appearing in pairs (except when the token's attention is focused on itself), and this pair of diagonals is symmetrically distributed about the diagonal starting from position (1,1) at a 45° downward-southward diagonal (the matrix diagonal). The case where large values ​​are distributed along the diagonal starting from position (1,1) at a 45° downward-southward diagonal is denoted as sub-pattern c-1, and the case where large values ​​are distributed along the diagonal starting from positions (1,2), (2,1), or (1,1) at a 45° downward-southward diagonal is denoted as sub-pattern c-2. Therefore, in this embodiment, when performing pattern recognition on each attention head for each model task based on the attention distribution of all tokens in the input data indicated by the large value distribution of all attention weights in the attention weight matrix, if the attention of each token in the input data is on the token itself, on a token identical to the token, or on a token associated with the token, then the pattern of the attention head is determined to be an extended attention pattern.

[0066] Uniform pattern: This refers to the situation where attention is evenly distributed across the entire graph. Therefore, in this embodiment, when performing pattern recognition on each attention head for each model task based on the attention distribution of all tokens in the input data indicated by the maximum value distribution of all attention weights in the attention weight matrix, if the attention of all tokens in the input data is evenly distributed across each token in the input data, then the pattern of the attention head is determined to be uniform pattern.

[0067] In this specific implementation, based on the distribution of the maximum values ​​of all attention weights in the attention weight matrix of each attention head, the pattern recognition result of each attention head is obtained in the following way:

[0068] First, obtain the first percentage of the average attention weight of each column in the attention weight matrix relative to the sum of the average attention weights of all columns in the entire attention weight matrix, that is, A for each attention head. i The values ​​of each column are summed and divided by the number of tokens N, where N is an integer greater than 2, to obtain the average attention weight for each column. These averages are then summed, and the percentage of each column's average to the total is calculated. Columns with larger percentages indicate that their values ​​are larger compared to other columns.

[0069] Because positional attention and extended attention patterns are in A i The visualizations show the same shape, all being diagonal lines that slope downwards at a 45° angle from a certain position, and the two never appear simultaneously. The positional attention pattern always consists of a single diagonal line, while sub-pattern c-1 in the extended attention pattern is also a single diagonal line, while sub-pattern c-2 consists of paired diagonal lines. Therefore, when calculating the average attention weight of each diagonal line in sub-pattern c-2, a pair of diagonal lines symmetrical about the diagonal of the matrix must be taken at once for averaging. Specifically, first, obtain the diagonal lines symmetrical about the diagonal of the attention weight matrix from all diagonal lines sloping downwards at 45° from positions (1,N-2), (1,N-3), ..., (1,1), (2,1) in the attention weight matrix. Then, take the two symmetrical diagonal lines from all the diagonal lines symmetrical about the diagonal of the attention weight matrix as a pair of symmetrical diagonal lines, and obtain the average attention weight of each pair of symmetrical diagonal lines. For example, take two diagonal lines starting at positions (1, N-2) and (N-2, 1), add their values, and divide by the total number of elements on the two diagonal lines (6) to get the average. Then, obtain the third percentage of the average attention weight of each group of symmetrical diagonal lines relative to the sum of the average attention weights of all symmetrical diagonal line groups in the entire attention weight matrix.

[0070] Finally, obtain the second percentage of the average attention weight of each diagonal line except for all symmetrical diagonal line groups as a percentage of the sum of the average attention weights of all diagonal lines in the entire attention weight matrix. This is because for sub-pattern c-1 of the positional attention pattern and the extended attention pattern, when calculating the average attention weight of the diagonal lines starting from positions (1,2), (1,1), and (2,1), it is not necessary to calculate in pairs; only the average value of each individual diagonal line is needed, and then the percentage is calculated.

[0071] That is, for each attention head A i The average attention weight of each diagonal line diagonally downwards to the right at 45° from positions (2,1), (1,1), and (1,2) in the attention weight matrix is ​​obtained sequentially. The average attention weight of each diagonal line diagonally downwards to the right at 45° from positions (1,N-2), (1,N-3), ..., (1,3) is also obtained, along with the diagonal line symmetrical about the diagonal, forming a pair of diagonal lines symmetrical about the diagonal from positions (N-2,1), (N-3,1), ..., (3,1). The sum of all average attention weights is obtained, and the percentage of each average attention weight in the sum of all average attention weights is obtained. The percentage corresponding to the average attention weight of each diagonal line diagonally downwards to the right at 45° from positions (2,1) and (1,2) is the second percentage, and the rest are the third percentage.

[0072] Sort all first percentages, second percentages, and third percentages in descending order, and select the top three target percentages that are greater than a preset threshold. If the target percentage is the first percentage, the attention head mode is determined to be global attention mode; if the target percentage is the second percentage, the attention head mode is determined to be positional attention mode; if the second target percentage is the third percentage, the attention head mode is determined to be extended attention mode; if the target percentage includes both the first and second percentages, the attention head mode is determined to be positional attention mode; if the target percentage includes both the first and third percentages, the attention head mode is determined to be extended attention mode; if the target percentage includes both the second and third percentages, the attention head mode is determined to be extended attention mode; if the target percentage includes all three percentages, the attention head mode is determined to be extended attention mode.

[0073] If none of the first, second, and third percentages exceeds a preset threshold, then the attention focus mode is determined to be uniform.

[0074] In some embodiments, different types of attention patterns may appear mixed in the same attention head, such as Figure 2The mixture of global attention and positional attention patterns shown in (e) is manifested in A i In the visualization, the large values ​​are distributed along the vertical line and the diagonal line starting from position (2,1) and sloping downwards at a 45° angle to the right. Therefore, in this case, when performing attention head pattern recognition, if the large values ​​of all attention weights in the attention weight matrix are distributed along the vertical column of the attention weight matrix and the diagonal line starting from position (2,1) and sloping downwards at a 45° angle to the right, then the attention head pattern is determined to be a hybrid pattern of global attention and positional attention.

[0075] Furthermore, global attention may also occur in combination with extended attention patterns. However, it should be noted that positional attention and extended attention patterns typically do not appear simultaneously in a single attention head.

[0076] In one example, this embodiment uses sample data from the dataset under the model task for inference to obtain the pattern recognition result of each attention head for the model task. Specifically, each sample data is used as input data to perform the above steps, the pattern recognition results of attention heads at the same position during the inference process of these sample data are statistically analyzed, and the pattern with the highest frequency is selected as the final pattern recognition result of the attention head at that position.

[0077] At this point, the fixed pattern followed by different input data under the same model task (i.e., the pattern of different attention heads in a model task) has been determined.

[0078] In a specific example, using BERT base The MRPC dataset of the model is used as input data. Pattern recognition is performed on each attention head of the model task using the method described above, validating the attention head pattern recognition scheme of this embodiment. The pattern recognition results are shown in Table 1, where the horizontal axis represents different layers in the model, and the vertical axis represents different attention heads within the same layer. 'a' represents the global attention pattern, 'b1' represents the positional attention pattern starting from position (1,2) and moving diagonally downwards at 45°, 'b2' represents the positional attention pattern starting from position (2,1) and moving diagonally downwards at 45°, 'c1' represents the extended attention pattern starting from position (1,1) and moving diagonally downwards at 45°, 'c2' represents the extended attention pattern consisting of diagonal pairs symmetrical about the diagonal except for positions (1,2), (2,1), or (1,1), and 'd' represents the uniform pattern. The actual attention matrix A is shown in Table 1. i Visual effects such as Figure 3 (Each column represents a different attention head at the same level, and each row represents an attention head at a different level.) It can be seen that the pattern recognition results in this embodiment are consistent with the real situation.

[0079] Table 1

[0080]

[0081] Step 103: Based on the type of the target model task, determine the pattern of each attention head of the target model task, and when using a multi-head attention mechanism to process the target input data of the target model task, quantize the target input data according to the pattern of each attention head of the target model task.

[0082] Specifically, if the attention head is in global attention mode or uniform mode, then the first precision calculation is performed on all target input data; if the attention head is in positional attention mode or extended attention mode, then the second precision calculation is performed on target input data located at the position of the largest value in the attention weight matrix, and the first precision calculation is performed on the remaining positions. Here, the first precision is less than the second precision. The first precision refers to low precision (i.e., low bit width), and the second precision refers to high precision (i.e., high bit width). The low bit width and high bit width can be 4 bits and 8 bits, or 32 bits and 64 bits, etc. That is, this embodiment uses a mixed bit width approach to process the model task data to achieve task quantization, and the specific bit width is not limited.

[0083] In the specific implementation, if the attention head is in global attention mode or uniform mode, then the target input data and weight matrix W will be used. Q W K and W V The generated matrix Q i K i and V i The data is quantized to the first precision format, and the quantized Q is processed using an attention mechanism. i K i and V i After obtaining the attention weight matrix A i Then, A i It is also quantized into a first-precision data format, and then compared with V. i Perform matrix multiplication calculations.

[0084] If the attention head pattern is the positional attention pattern, then determine A. i The largest values ​​of all attention weights are distributed along a 45° downward-right slope starting from position (1,2) or starting from position (2,1). This determines whether the pattern recognition result is sub-pattern b-1 or sub-pattern b-2, thus determining A. i Specifically, on the diagonal line where the medium-to-large value begins, what is the exact position? (This is related to calculating A.) i When calculating, the vectors q and k corresponding to the positions with large values ​​on the diagonal are quantized to a second-precision data format for calculation, while the vectors q and k corresponding to the remaining positions are quantized to a first-precision data format for calculation; after obtaining A iThen, the q and k corresponding to the larger values ​​on the diagonal are quantized to the second-precision data format, while the remaining positions are quantized to the first-precision data format; in A i With V i During matrix multiplication, A will be involved. i The vector v calculated from the large values ​​q and k at the slant positions also uses the second-precision data format, while the rest use the first-precision data format.

[0085] If the attention head pattern is the extended attention pattern, then determine A. i Does the sloping line of the large value distribution of all attention weights in A exist? i The diagonal symmetrical line indicates whether sub-pattern c-2 exists in the attention head based on the pattern recognition result. If so, then Q... i K i and V i The data is quantized to the first precision format, and a coarse A value is obtained. i Then calculate about A i The average value of each pair of symmetrical diagonal lines is used, and the pair with the largest average value is selected as the position for the second-precision calculation. That is, if there is no sub-pattern c-2, only the calculations at the diagonal line position corresponding to c-1 are performed with second precision, and the rest with first precision. If there is a sub-pattern c-2, then for the attention head of sub-pattern c-2 calculated for the first time under the input data, the Q-value of this attention head needs to be... i K i Quantize to first precision and obtain a coarse A. i The result is then calculated, and the average value of each pair of diagonal lines contained in sub-pattern c-2 is calculated. The diagonal line pair with the largest average value is the position of the second precision calculation. For the attention head of sub-pattern c-2 that is not calculated for the first time, the judgment result of the first time is used to directly perform mixed precision calculation without first precision judgment.

[0086] If the attention head mode is a mixture of global attention mode and positional attention mode, then the quantization method corresponding to the positional attention mode is used; if the attention head mode is a mixture of global attention mode and extended attention mode, then the quantization method corresponding to the extended attention mode is used.

[0087] In this embodiment, since the attention weight matrix A is used for a specific model task... iThe location of large values ​​in the matrix does not change dynamically with the input samples, i.e., it follows a fixed pattern. Therefore, for different types of model tasks, before using a multi-head attention mechanism to process the input data of the model task, we first extract the attention weight matrix of each attention head in the training set data of the model task, perform pattern recognition on each attention head, and obtain the pattern recognition result of each attention head corresponding to different types of model tasks. This pattern recognition result reflects the aforementioned attention weight matrix A. i The model identifies fixed patterns governing the positions of important elements. Then, based on the pattern recognition results of the attention heads, the input data processed by the attention heads is processed with corresponding precision. Therefore, for the same type of model task, even if the input data differs, the attention heads at the same layer and position follow the same pattern during each input data processing. After determining the patterns of different attention heads, for different input data of the model task, it is not necessary to divide the first precision calculation into a separate calculation stage to predict the positions of important elements. Instead, the corresponding precision (second precision or first precision) is calculated directly for each attention head according to its pattern, thereby reducing the total computation time and achieving accelerated multi-head attention quantization based on task pattern awareness. Furthermore, by dividing the attention heads into different patterns and providing different precision schemes according to different patterns, it is not necessary to perform high first precision mixed calculations for all attention heads. Some attention heads can only perform first precision calculations, thus achieving a larger model compression ratio.

[0088] In one embodiment, the task-pattern-aware multi-head attention quantization method of the present invention can be implemented as follows: Figure 4 The steps shown are to be implemented as follows:

[0089] Step 1: Model Task-Oriented Pattern Recognition. Several task-adaptive patterns for attention heads are proposed. These are patterns that attention heads at the same position within the same layer will follow when the same model is facing the same task, even with different inputs. Based on the characteristics of these patterns, corresponding accuracy schemes are proposed, and different attention head patterns are recognized on the training dataset of the target task.

[0090] Step 2: Quantization and computation for each input. When deploying and executing the model task, for each specific input, the recognition results obtained in Step 1 are used to find the pattern of the current attention head, and the computation of the attention head is accelerated by quantization using the accuracy scheme corresponding to that pattern.

[0091] Specifically, step one further includes:

[0092] Step 1: Task-adaptive modes and accuracy schemes for attention heads. This includes various attention head modes and their corresponding accuracy schemes.

[0093] Step 2: Calculate the probability distribution of the global attention pattern of the current attention head, i.e., calculate the first percentage.

[0094] Step 3: Calculate the probability distribution of the current attention head's positional attention and the extended attention pattern, i.e., calculate the second percentage and the third percentage.

[0095] Step 4: Determine the current attention pattern.

[0096] Step 5: Determine the pattern of each attention head under the model task.

[0097] Step two further includes:

[0098] Step 1: Based on the position of the currently calculated attention head (i.e., which layer and which attention head), find its corresponding pattern recognition result.

[0099] Step 2: Determine the accuracy scheme based on the attention head pattern, and begin quantization and calculation.

[0100] The specific operation of each step in this embodiment can be found in the above embodiments, and will not be repeated here.

[0101] In one specific embodiment, a second precision of 16-bit integer and a first precision of 4-bit integer are used for quantization. Compared with quantization using only the second precision (16-bit integer), this demonstrates that the task-pattern-aware multi-head attention quantization method of the present invention can achieve a speedup of 14.9 times while ensuring the model's task accuracy.

[0102] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this invention. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the protection scope of this invention.

[0103] Another embodiment of the present invention relates to a task-mode-aware multi-head attention quantization device. The implementation details of this task-mode-aware multi-head attention quantization device are described below. The following details are provided for ease of understanding and are not essential for implementing this solution. The task-mode-aware multi-head attention quantization device of this embodiment includes:

[0104] The weight acquisition module is used to obtain the attention weight matrix of each attention head in the multi-head attention mechanism when processing the input data of the model task using the multi-head attention mechanism for each type of model task.

[0105] The pattern recognition module is used to perform pattern recognition on each attention head of the model task based on the distribution of the largest values ​​of all attention weights in the attention weight matrix. The distribution of the largest values ​​is used to indicate the attention distribution of all tokens in the input data. The largest value in the distribution of the largest values ​​is the value that is located at the beginning of the attention weight matrix after all values ​​in the attention weight matrix are sorted in descending order.

[0106] If the attention of all tokens in the input data is focused on at least one token, the attention head pattern is determined to be a global attention pattern; if the attention of each token in the input data is focused on the previous or next token, the attention head pattern is determined to be a positional attention pattern; if the attention of each token in the input data is focused on the token itself, on the same token, or on a token related to the token, the attention head pattern is determined to be an extended attention pattern; if the attention of all tokens in the input data is evenly distributed across all tokens in the input data, the attention head pattern is determined to be a uniform pattern.

[0107] The task quantization module is used to determine the pattern of each attention head of the target model task based on the type of the target model task, and to perform the following quantization processing on the target input data when using a multi-head attention mechanism to process the target input data:

[0108] If the attention head is in global attention mode or uniform mode, then the first precision calculation is performed on all target input data; if the attention head is in positional attention mode or extended attention mode, then the second precision calculation is performed on the target input data located at the position of the largest value in the attention weight matrix, and the first precision calculation is performed on the remaining positions; wherein, the first precision is less than the second precision.

[0109] It is not difficult to see that this embodiment is a device embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.

[0110] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.

[0111] Another embodiment of the present invention relates to a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the task-pattern-aware multi-head attention quantization method described in the above embodiments.

[0112] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0113] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0114] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0115] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0116] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of the present invention.

Claims

1. A multi-head attention quantization method based on task pattern awareness, characterized in that, include: For each type of model task, obtain the attention weight matrix of each attention head in the multi-head attention mechanism when processing the input data of the model task using the multi-head attention mechanism; Based on the distribution of the largest values ​​of all attention weights in the attention weight matrix, pattern recognition is performed on each attention head of the model task; where the distribution of the largest values ​​is used to indicate the attention distribution of all tokens in the input data, and the largest value in the distribution of the largest values ​​is the value that is located at the beginning of the attention weight matrix after all values ​​in the attention weight matrix are sorted in descending order. If the attention of all tokens in the input data is focused on at least one token, the attention head pattern is determined to be a global attention pattern; if the attention of each token in the input data is focused on the previous or next token, the attention head pattern is determined to be a positional attention pattern; if the attention of each token in the input data is focused on the token itself, on the same token, or on a token related to the token, the attention head pattern is determined to be an extended attention pattern; if the attention of all tokens in the input data is evenly distributed across all tokens in the input data, the attention head pattern is determined to be a uniform pattern. Based on the type of the target model task, determine the pattern of each attention head for the target model task, and when using a multi-head attention mechanism to process the target input data of the target model task, perform the following quantization processing on the target input data according to the pattern of each attention head for the target model task: If the attention head mode is global attention mode or uniform mode, then the first precision calculation is performed on all target input data; if the attention head mode is positional attention mode or extended attention mode, then the second precision calculation is performed on the target input data located at the position of the largest value in the attention weight matrix, and the first precision calculation is performed on the other positions; wherein, the target input data is the MRPC dataset used for the semantic similarity evaluation model task between sentence pairs, and the first precision is less than the second precision. The step of performing pattern recognition on each attention head of the model task based on the maximum value distribution of all attention weights in the attention weight matrix includes: For each attention head, obtain the pattern recognition results corresponding to multiple input data for the model task, and take the pattern recognition result with the highest frequency as the target pattern recognition result of the attention head.

2. The multi-head attention quantization method based on task pattern awareness according to claim 1, characterized in that, The step of performing pattern recognition on each attention head of the model task based on the maximum value distribution of all attention weights in the attention weight matrix includes: Obtain the first percentage of the average attention weight of each column in the attention weight matrix relative to the sum of the average attention weights of all columns in the entire attention weight matrix; The average attention weight of each diagonal line diagonally downwards to the right at 45° from positions (2,1), (1,1), and (1,2) in the attention weight matrix is ​​obtained sequentially. The average attention weight of each diagonal line diagonally downwards to the right at 45° from positions (1,N-2), (1,N-3), ..., (1,3) is also obtained, along with the diagonal line pairs formed by the diagonal lines diagonally downwards to the right at positions (N-2,1), (N-3,1), ..., (3,1). Get the sum of all average attention weights, and get the percentage of each average attention weight to this sum of average attention weights; where the percentage of the average attention weight corresponding to each diagonal line diagonally downward to the right of (2,1) and (1,2) is the second percentage, and the rest are the third percentage; Sort all first percentages, all second percentages, and all third percentages in descending order, and select the target percentages that are greater than a preset threshold and are among the top three. If the target percentage is the first percentage, the attention head mode is determined to be the global attention mode; if the target percentage is the second percentage, the attention head mode is determined to be the positional attention mode; if the second target percentage is the third percentage, the attention head mode is determined to be the extended attention mode; if the target percentage includes both the first and second percentages, the attention head mode is determined to be the positional attention mode; if the target percentage includes both the first and third percentages, the attention head mode is determined to be the extended attention mode; if the target percentage includes both the second and third percentages, the attention head mode is determined to be the extended attention mode; if the target percentage includes the first, second, and third percentages, the attention head mode is determined to be the extended attention mode. If none of the first, second, and third percentages exceeds a preset threshold, then the attention head mode is determined to be uniform mode.

3. The multi-head attention quantization method based on task pattern awareness according to claim 1, characterized in that, Based on the pattern of each attention head in the target model task, the target input data is quantized as follows: If the attention head is in global attention mode or uniform mode, then the target input data and weight matrix W will be used. Q W K and W V The generated matrix Q i K i and V i The data is quantized to the first precision for computation, and the quantized Q is processed using an attention mechanism. i K i and V i After obtaining the attention weight matrix A i Then, A i After being quantized into a first-precision data format, it is then compared with V. i Perform matrix multiplication calculations.

4. The multi-head attention quantization method based on task pattern awareness according to claim 3, characterized in that, Based on the pattern of each attention head in the target model task, the target input data is quantized as follows: If the attention head pattern is the positional attention pattern, and A i The largest values ​​of all attention weights are distributed along a 45° downward-right slope starting from position (1,2), or along a 45° downward-right slope starting from position (2,1); in calculating A i When calculating, the vectors q and k corresponding to the values ​​on the diagonal lines are quantized to a second-precision data format for calculation, while the q and k corresponding to the values ​​at other positions are quantized to a first-precision data format for calculation; after obtaining A... i Then, the values ​​on the corresponding diagonal lines are quantized to a second-precision data format, while the remaining positions are quantized to a first-precision data format for calculation; in A i With V i During matrix multiplication, A will be involved. i The values ​​at the slash positions correspond to the vector v being calculated, which is also quantized to the second-precision data format for calculation, while the rest are quantized to the first-precision data format for calculation.

5. The multi-head attention quantization method based on task pattern awareness according to claim 4, characterized in that, Based on the pattern of each attention head in the target model task, the target input data is quantized as follows: If the attention head pattern is the extended attention pattern, then determine A. i Is the oblique line of the large value distribution of all attention weights a pair of oblique lines symmetrical about the diagonal or the diagonal itself? If it is a diagonal, then in calculating A i At that time, the attention weight vectors q and k corresponding to the diagonal are quantized into a second-precision data format for calculation, while the remaining positions are quantized into a first-precision data format for calculation; and in A i With V i During matrix multiplication, the attention weights on the diagonal and the vector v used for calculation are quantized into a second-precision data format for calculation, while the other positions and the vector v are quantized into a first-precision data format for calculation. If it is a pair of diagonal lines that are symmetrical about the diagonal, then determine whether this is the first time that attention has been drawn to this situation; If so, then Q i and K i The data is quantized to the first precision for calculation. The average value of each pair of diagonal lines symmetrical about the diagonal is calculated, and the pair of diagonal lines with the largest average value is used as the position for calculation in the second precision data format. In calculating A... i At that time, the vectors q and k corresponding to the attention weights of the diagonal pair with the largest average value are quantized into a second-precision data format for calculation, while the remaining positions are quantized into a first-precision data format for calculation; in A i With V i During matrix multiplication, the attention weights on the diagonal pair with the largest average value and the vector v used for calculation are quantized into a second-precision data format for calculation, while the other positions and vector v are quantized into a first-precision data format for calculation. If not, then calculate A i At that time, the vectors q and k corresponding to the attention weights of the diagonal pair with the largest average value are quantized into a second-precision data format for calculation, while the remaining positions are quantized into a first-precision data format for calculation; in A i With V i During matrix multiplication, the attention weights on the diagonal pairs with the largest average value and the vector v used for calculation are quantized into a second-precision data format for calculation, while the remaining positions and vector v are quantized into a first-precision data format for calculation.

6. The multi-head attention quantization method based on task pattern awareness according to claim 5, characterized in that, The pattern recognition results also include a global attention pattern mixed with a location attention pattern and a global attention pattern mixed with an extended attention pattern; Based on the pattern of each attention head in the target model task, the target input data is quantized as follows: If the attention head mode is a mixture of global attention mode and positional attention mode, then the quantization method corresponding to the positional attention mode shall be used; If the attention head mode is a mixture of global attention mode and extended attention mode, then the quantization method corresponding to the extended attention mode is used.

7. A multi-head attention quantization device based on task pattern awareness, characterized in that, include: The weight acquisition module is used to obtain the attention weight matrix of each attention head in the multi-head attention mechanism when processing the input data of the model task using the multi-head attention mechanism for each type of model task. The pattern recognition module is used to perform pattern recognition on each attention head of the model task based on the distribution of the largest values ​​of all attention weights in the attention weight matrix. The distribution of the largest values ​​is used to indicate the attention distribution of all tokens in the input data. The largest value in the distribution of the largest values ​​is the value that is located at the beginning of the attention weight matrix after all values ​​in the attention weight matrix are sorted in descending order. If the attention of all tokens in the input data is focused on at least one token, the attention head pattern is determined to be a global attention pattern; if the attention of each token in the input data is focused on the previous or next token, the attention head pattern is determined to be a positional attention pattern; if the attention of each token in the input data is focused on the token itself, on the same token, or on a token related to the token, the attention head pattern is determined to be an extended attention pattern; if the attention of all tokens in the input data is evenly distributed across all tokens in the input data, the attention head pattern is determined to be a uniform pattern. The task quantization module is used to determine the pattern of each attention head of the target model task based on the type of the target model task, and to perform the following quantization processing on the target input data when using a multi-head attention mechanism to process the target input data: If the attention head mode is global attention mode or uniform mode, then the first precision calculation is performed on all target input data; if the attention head mode is positional attention mode or extended attention mode, then the second precision calculation is performed on the target input data located at the position of the largest value in the attention weight matrix, and the first precision calculation is performed on the other positions; wherein, the target input data is the MRPC dataset used for the semantic similarity evaluation model task between sentence pairs, and the first precision is less than the second precision. The pattern recognition module is further configured to, for each attention head, obtain the pattern recognition results corresponding to multiple input data under the model task, and take the pattern recognition result with the highest frequency as the target pattern recognition result of the attention head.

8. A computer device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the task-pattern-aware multi-head attention quantization method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-head attention quantization method based on task pattern awareness as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image restoration method and system based on multiple attention mechanisms

    CN119130863A

  • Text processing method and system based on deep learning model attention mechanism

    CN120123499A