Task execution method and device, electronic equipment, storage medium and program product

By converting fixed-point key features into floating-point types and using inverse quantization to scale parameters, combined with matrix multiplication tasks, the problem of high computing and storage resource overhead in deep learning models is solved, efficient computing and storage resource utilization is achieved, and the model's reasoning efficiency and accuracy are improved.

CN120704896APending Publication Date: 2025-09-26BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510899848.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Deep learning models consume a lot of computing and storage resources, which limits the length of input sequences and affects the efficiency of model training and inference.

Method used

By converting fixed-point key features into floating-point types and using inverse quantization scaling parameters to improve the accuracy of query features, combined with matrix multiplication tasks, the complex inverse quantization calculation process is simplified and hardware resource overhead is reduced.

Benefits of technology

While ensuring model performance, the consumption of computing and storage resources is reduced, computing efficiency and data accuracy are improved, and dependence on hardware resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704896A_ABST
    Figure CN120704896A_ABST
Patent Text Reader

Abstract

The invention provides a task execution method which comprises the steps that according to a first query feature and a first key feature, a computing unit is used for executing a first multiplication task of a target attention task queue, a first multiplication result feature is obtained, the first key feature is obtained by converting a second key feature of a fixed point type into a floating point type, and the second key feature is obtained by converting a second key feature of a floating point type; the numerical precision of the first key feature is greater than that of the second key feature, and the numerical precision of the first query feature is less than that of the second query feature; according to the first multiplication result feature and the first value feature, executing a plurality of target tasks of the target attention task queue by using a computing unit to obtain attention features; the first query feature is obtained through the second query feature and an inverse quantization scaling parameter for the second key feature, so that the hardware resource overhead of executing the target attention task queue by the computing unit is smaller than a preset hardware resource overhead value. The invention further provides a task execution device and equipment, a medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of large language models (LLMs), image processing, text processing, and audio processing. More specifically, the present disclosure provides a task execution method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] With the continuous development of artificial intelligence technology, the number of parameters in deep learning models is increasing. The computing and storage resource overhead required for training or inferencing deep learning models is also increasing. The computational time and memory complexity of deep learning models also limit the length of the input sequence. Summary of the Invention

[0003] The present disclosure provides a task execution method, apparatus, electronic device, storage medium, and program product.

[0004] According to one aspect of the present disclosure, a task execution method is provided, including: according to a first query feature and a first key feature, using a computing unit to execute a first multiplication operation task of a target attention task queue to obtain a first multiplication operation result feature, wherein the first key feature is obtained by converting a second key feature of a fixed-point type into a floating-point type, the numerical precision of the first key feature is greater than the numerical precision of the second key feature, and the numerical precision of the first query feature is less than the numerical precision of the second query feature; according to the first multiplication operation result feature and the first value feature, using a computing unit to execute multiple target tasks of the target attention task queue to obtain an attention feature; the first query feature is obtained by the second query feature and the inverse quantization scaling parameter for the second key feature, so that the hardware resource overhead of the computing unit executing the target attention task queue is less than a preset hardware resource overhead value.

[0005] According to another aspect of the present disclosure, a task execution device is provided, including: a first execution module, for executing a first multiplication operation task of a target attention task queue using a computing unit according to a first query feature and a first key feature, to obtain a first multiplication operation result feature, wherein the first key feature is obtained by converting a second key feature of a fixed-point type into a floating-point type, the numerical precision of the first key feature is greater than the numerical precision of the second key feature, and the numerical precision of the first query feature is less than the numerical precision of the second query feature; a second execution module, for executing multiple target tasks of the target attention task queue using a computing unit according to the first multiplication operation result feature and the first value feature, to obtain an attention feature; the first query feature is obtained by the second query feature and the inverse quantization scaling parameter for the second key feature, so that the hardware resource overhead of the computing unit executing the target attention task queue is less than a preset hardware resource overhead value.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to execute the prompt information determination method provided according to an embodiment of the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the prompt information determination method provided by the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the prompt information determination method provided by the present disclosure when executed by a processor.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure; wherein:

[0011] Figure 1 is a flowchart of a task execution method according to an embodiment of the present disclosure.

[0012] Figure 2 is a flowchart of a task execution method according to another embodiment of the present disclosure.

[0013] Figure 3 is a schematic diagram of a task execution method according to an embodiment of the present disclosure.

[0014] Figure 4 1 is a diagram comparing the architectures of different large models according to an embodiment of the present disclosure.

[0015] Figure 5 is a schematic diagram of a task execution method according to another embodiment of the present disclosure.

[0016] Figure 6 is a block diagram of a task execution apparatus according to an embodiment of the present disclosure; and

[0017] Figure 7 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0018] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0019] A large model refers to a deep learning model with large-scale model parameters. A large model usually contains hundreds of millions, billions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. Large models can include large language models (LLM), GPT (Generative Pre-trained Transformer), large visual models, multimodal large models, and so on. The large model involved in the embodiments of the present disclosure can be a general large model, or it can also be an expert large model obtained by fine-tuning (Fine Tune) based on needs. The embodiments of the present disclosure are not limited to this.

[0020] The large model involved in the embodiments of the present disclosure may include one or more Transformer networks. The Transformer network can process data based on the attention mechanism. The implementation process of the attention mechanism includes the generation of query features, key features, and value feature matrices, and the use of the softmax function to process them to obtain attention weights. The computational complexity of the standard attention mechanism is , where N is the length of the input sequence. This is because for each element in the input sequence, its similarity (or attention weight) with all other elements must be calculated. When N is large, this calculation becomes very time-consuming and consumes a lot of computing resources, resulting in inefficient model training and inference, and becoming a performance bottleneck in long-text inference.

[0021] In large models, CacheKV (cache key-value) technology is often introduced to improve model inference performance. Cache KV caches the reusable results of round i. When performing round i+1 calculations, the cached round i results are directly read and concatenated with the round i+1 results to obtain the final round i+1 results. Introducing Cache KV into large models improves model inference speed while also increasing the number of model parameters. Therefore, longer input sequences require more video memory (such as that provided by a graphics processing unit (GPU)) to store Cache KV, significantly increasing machine costs and restricting the processing of longer contexts.

[0022] Therefore, in order to reduce the hardware resource overhead of model training or inference, the present disclosure provides a task execution method, which will be described below.

[0023] Figure 1 is a flowchart of a task execution method according to an embodiment of the present disclosure.

[0024] like Figure 1 As shown, the method 100 may include operations S110 to S120.

[0025] In operation S110, a first multiplication operation task of the target attention task queue is executed by a computing unit according to the first query feature and the first key feature to obtain a first multiplication operation result feature.

[0026] The first key feature is obtained by converting the second key feature of the fixed-point type into the floating-point type. The numerical precision of the first key feature is greater than the numerical precision of the second key feature. The numerical precision of the first query feature is less than the numerical precision of the second query feature.

[0027] In operation S120 , a plurality of target tasks of a target attention task queue are executed by a computing unit according to the first multiplication result feature and the first value feature to obtain an attention feature.

[0028] The first query feature is obtained through the second query feature and the inverse quantization scaling parameter for the second key feature, so that the hardware resource overhead of the computing unit executing the target attention task queue is less than a preset hardware resource overhead value.

[0029] For example, the first query feature may be obtained by multiplying the input feature by the query weight matrix. The first key feature may be obtained by multiplying the input feature by the key weight matrix.

[0030] The computing unit may be an artificial intelligence computing unit, which may be various hardware computing units such as a general-purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), and a neural network processing unit (NPU).

[0031] The target attention task queue can be a task queue that processes data based on an attention mechanism. The attention mechanism can be a multi-head self-attention mechanism. For example, in a large model consisting of multiple cascaded Transformer networks, the target attention task queue can include multiple tasks to be executed by the attention encoder module (Transformerencoder) of the target Transformer network.

[0032] General Matrix to Matrix Multiplication (GEMM) is a fundamental operation involved in neural networks and multi-head self-attention mechanisms, and is also a crucial foundational operation for running large models. GEMM multiplies two matrices together to produce an output matrix. The input matrices can be either a weight matrix or a data matrix, and the output matrix is ​​obtained after performing the matrix multiplication.

[0033] The first multiplication task may instruct the computing unit to perform a general matrix multiplication task, for example, the operation instructed is: multiplying the first query feature and the first key transposed feature to obtain a first multiplication result feature. The first key transposed feature is obtained by transposing the first key feature.

[0034] For example, the feature score of the first multiplication result can be determined using the following formula:

[0035] (1)

[0036] It can be the first query feature. Features can be transposed for the first key.

[0037] To perform a general matrix multiplication on the first query feature and the first key feature, ensure that the first query feature and the first key feature have the same data type and precision. For example, both can be floating-point values ​​of any type, including 8-bit floating-point values ​​(FP8), 16-bit floating-point values ​​(FP16), and 32-bit floating-point values ​​(FP32). "FP" refers to float, and floating-point types can also include "BF" (brain float).

[0038] Fixed-point types can include 32-bit integers (int32), 16-bit integers (int16), 8-bit integers (int8), and 4-bit integers (int4).

[0039] For example, the first key feature can be obtained by converting the second key feature of type int4 to type FP8. FP8 has greater numerical precision than int4. For example, the first query feature of type FP8 can also be obtained based on the second query feature of 16-bit floating-point value (FP16 or BF16) and the dequantization scaling parameter used for the second key feature.

[0040] When training or inferring large models, high-precision data (such as FP32 floating-point) is quantized to lower precision (such as int4) for storage compression. This process is called quantization. During actual computations (such as matrix multiplication), the low-precision data (such as int4) must be restored to a usable precision (such as FP8). This is called dequantization. The dequantization scaling parameter is used during the dequantization process to compensate for errors with the original data.

[0041] Exemplarily, the Transformer network may include an attention encoding module and a feedforward (FFN) module. The feedforward module of the front Transformer network includes multiple fully connected layers. The feedforward module can execute tasks in the feedforward task queue. The quantization scaling parameters or dequantization scaling parameters for query features, key features or value features of the embodiment of the present disclosure can be obtained according to the maximum value, such as the maximum value in the execution result of the feedforward task queue before the target attention task queue or the maximum value that can be represented by the fixed-point type or floating-point type; the quantization scaling parameters or dequantization scaling parameters can also be obtained by static statistics, such as giving multiple sets of quantization scaling parameters or dequantization scaling parameters by experiment, and writing them into the computing unit after obtaining them according to the statistics of the quantization or dequantization effect, and directly calling them during inference; the quantization scaling parameters or dequantization scaling parameters can also be obtained through expert experience.

[0042] For example, the second key feature can be stored in the video memory (CacheK2). The fixed-point type of the second key feature can effectively save video memory, and is converted to a floating-point type during multiplication operations to participate in the calculation, while improving the calculation accuracy.

[0043] For example, the multiple target tasks in the target attention task queue may include one or more of a preset function processing task, a second multiplication operation task, a mask processing task, etc. It is understandable that based on the multi-head self-attention mechanism, the mask processing task may be optional.

[0044] The first value feature may be obtained by multiplying the input feature by the value weight matrix. For example, the first value feature may also be any floating point type including an 8-bit floating point value (FP8), a 16-bit floating point value (FP16), and a 32-bit floating point value (FP32).

[0045] According to an embodiment of the present disclosure, the dequantization process of the second key feature is decomposed, and the accuracy of the key feature is improved by converting the fixed-point type to the floating-point type. In the process of obtaining the first query feature, the dequantization scaling parameter for the second key feature is introduced to improve the computational efficiency and perform precision adaptation. This decomposition method can simplify the complex dequantization calculation process, further improve the computational efficiency, reduce the computational resource consumption required for dequantizing the second key feature, and effectively improve the accuracy of data processed based on the attention mechanism. At the same time, under the premise of ensuring model performance, the use of fixed-point second key features can reduce the dependence on and occupation of storage resources (such as video memory storage space, etc.).

[0046] Thus, while reducing the computing resources required for dequantization and the storage resources required for storing the second key features, the hardware resource overhead required for the computing resources to execute the target attention task queue can be less than a preset hardware resource overhead value. The preset hardware resource overhead value may include the computing resource overhead and storage resource overhead required to perform multi-head self-attention calculations based on the unquantized key features and the unquantized query features.

[0047] Figure 2 is a flowchart of a task execution method according to another embodiment of the present disclosure.

[0048] like Figure 2 As shown, the method 200 may include operations S210 to S240.

[0049] In operation S210, a first multiplication operation task of a target attention task queue is executed by a computing unit according to the first query feature and the first key feature to obtain a first multiplication operation result feature.

[0050] The first key feature is obtained by converting the second key feature of the fixed-point type into the floating-point type. The numerical precision of the first key feature is greater than the numerical precision of the second key feature. The first query feature is obtained by using the second query feature and the inverse quantization scaling parameter for the second key feature. The numerical precision of the first query feature is less than the numerical precision of the second query feature.

[0051] It can be understood that the above descriptions of the first query feature, the first key feature, the computing unit, the target attention task queue, the first multiplication task, the first multiplication result feature, the fixed-point type, the second key feature, the floating-point type, the second query feature, etc. are also applicable to this embodiment and will not be repeated here.

[0052] By executing operations S220 to S240, the calculation unit can be used to execute multiple target tasks of the target attention task queue based on the first multiplication operation result feature and the first value feature to obtain the attention feature; the multiple target tasks of the target attention task queue include a preset function processing task and a second multiplication operation task.

[0053] In operation S220, a preset function processing task is executed by using a computing unit to process the first multiplication result feature by using the preset function to obtain a feature to be processed.

[0054] In operation S230 , a second multiplication operation task is performed by using a computing unit according to the feature to be processed and the first value feature to obtain a second multiplication result feature.

[0055] In operation S240 , an attention feature is determined based on the feature to be processed and the second multiplication result feature.

[0056] The preset function processing task can instruct the computing unit to perform a corresponding computing operation using a preset function. The second multiplication task can be a matrix multiplication task, which can instruct the computing unit to perform a general matrix multiplication task based on the feature to be processed and the feature of the second multiplication result.

[0057] To implement a general matrix multiplication operation based on the feature to be processed and the feature of the second multiplication result, the feature to be processed and the feature of the second multiplication result can have the same data type and the same data precision. For example, both can be any floating-point type including 8-bit floating-point values ​​(FP8), 16-bit floating-point values ​​(FP16), and 32-bit floating-point values ​​(FP32).

[0058] In some embodiments, the attention feature can be calculated by the calculation unit according to the following formula:

[0059] (2)

[0060] Att can be an attention feature. Softmax() is a soft maximum function, which can be a preset function. It can be a first value feature. Can be the dimension of the first value feature.

[0061] According to the embodiments of the present disclosure, the target attention task queue can be executed by the computing unit, and the calculation of attention features can be completed efficiently by executing matrix multiplication tasks. In addition, the complex dequantization calculation process can be simplified by decomposition, further improving the operation efficiency, reducing the consumption of computing resources, and effectively improving the accuracy of data processed based on the attention mechanism. At the same time, under the premise of ensuring model performance, the use of fixed-point second key features can reduce the dependence on and occupation of hardware resources (such as computing power, video memory storage space, etc.).

[0062] The process of obtaining the first query feature is further described below, and the dequantization process for the second key feature is described in combination with the process of converting the fixed-point second key feature to the floating-point type to obtain the first key feature.

[0063] Figure 3 is a schematic diagram of a task execution method according to an embodiment of the present disclosure.

[0064] In some embodiments, the first query feature is obtained by multiplying the second query feature, the quantization scaling parameter for the second query feature, and the inverse quantization scaling parameter for the second key feature to obtain the first query feature.

[0065] For example, refer to Figure 3 , the second query feature Q312 of the 16-bit floating point value (BF16) can first be quantized into a floating point type (FP8), and then multiplied by the inverse quantization scaling parameter K323 for the second key feature K322 to obtain the first query feature Q311 of the FP8 type.

[0066] The first query feature can be calculated according to the following formula:

[0067] (3)

[0068] in, It can be the second query feature. The quantization scaling parameter for the second query feature may be used. The parameters may be scaled for the inverse quantization of the second key feature.

[0069] Combining Formula 1 and Formula 2, the derivation process of Formula 3 is described below.

[0070] The features to be processed can be determined based on the following formula:

[0071] (4)

[0072] in, The features to be processed. is the zero point parameter for the second bond feature. It can be a generalized query feature.

[0073] Formula 4 is simplified by using the Softmax function's property that "subtracting the maximum value does not affect the result":

[0074] (5)

[0075] in, It can be a generalized key feature. For example, Q can be based on get, Can be based on get.

[0076] In formula 5, the Softmax function's property of "subtracting the maximum value does not affect the result" or translation invariance is used to eliminate This constant term.

[0077] It can be understood that in Formula 4, each dequantization performs matrix subtraction and matrix multiplication operations on each eigenvalue of the key feature. Based on Formula 3 and Formula 5, for example, and The multiplication operation can be integrated into the operation of quantizing the second query feature, a 16-bit floating-point value (BF16), to a floating-point type (FP8), as shown in Formula 3. This can be combined with the fast conversion operation of converting the second key feature, which is of type int4, to FP8 to obtain the first key feature. This allows dequantization of the second key feature to require only shifts and AND operations, reducing the computational resources consumed by FMA (Fused-Multiply-Add) operations and data type conversion. This simplified dequantization process is therefore more suitable for accelerating operations through matrix multiplication tasks (such as those performed by GPU Tensor Cores).

[0078] According to an embodiment of the present disclosure, the "scaling" in the inverse quantization process is coordinated with the fast conversion operation between the second key feature and the first key feature, and the redundant matrix subtraction and floating-point operations related to the "zero point" are eliminated in the inverse quantization process, and the matrix multiplication tasks related to the "scaling" are retained and placed in the quantization process of the second query feature, thereby improving the computational efficiency and realizing fast inverse quantization.

[0079] The following further describes the quick conversion operation of "converting the second key feature of type int4 to type FP8 to obtain the first key feature".

[0080] In some embodiments, the first key feature is obtained by converting a second key feature of a fixed-point type into a floating-point type, including: placing multiple feature values ​​of the second key feature at preset positions of multiple first data respectively to obtain the first key feature, where the preset positions are multiple consecutive bits in the first data; wherein the numerical precision of the first data is consistent with the numerical precision of the first key feature.

[0081] For example, refer to Figure 3 , the second key feature K322 is of int4 type, and its multiple feature values ​​correspond to multiple numerical values ​​in int4. The first data may include byte data. The second key feature K322 of int4 type can be converted into uint4 type by adding a preset value (for example, 8) as a whole. Then, each element of the second key feature K322 of uint4 type is placed in the last 4 bits of a byte (i.e., the preset position, continuous bits) and regarded as the first key feature K321 of FP8 type. The numerical precision of the byte is 8 bits, which is consistent with the numerical precision of the first key feature K321 of FP8 type.

[0082] Taking FP8 (E4M3) as an example, by placing uint4 data in the last four bits of FP8 (E4M3) data, a mapping relationship represented by the following formula 6 can be obtained. Therefore, each element of the second key feature of the uint4 type can be placed in the last four bits of a byte and regarded as the first key feature of the FP8 type.

[0083] (6)

[0084] in, Belongs to FP8 (E4M3) data, It belongs to uint4 data. FP8 (E4M3) includes 1 sign bit, 4 exponent bits, and 3 mantissa bits.

[0085] The mapping relationship represented by Formula 6 is inferred as follows:

[0086] Binary to decimal conversion formula: (7)

[0087] The exponential offset value Bias of FP8 (E4M3) can be: (8)

[0088] The number of exponent bits that can be FP8 (E4M3) (4).

[0089] The decimal number converted from the FP8 (E4M3) exponent is recorded as , the decimal number converted from the last digit is recorded as

[0090] The calculation formula for non-normalized FP8 (E4M3) (not considering the sign bit) is:

[0091] (9)

[0092] Normalized FP8 (E4M3) calculation formula (not considering the sign bit):

[0093] (10)

[0094] After adding 8 to int4 type data, it is converted to uint4 type data, and the range can be changed to 0~15. The decimal number of uint4 is recorded as , the decimal number of FP8 is recorded as , put uint4 into the last four bits of FP8, then:

[0095] 0~7: FP8 (E4M3) sign bits are all 0, denormalized

[0096] (11)

[0097] 8:00-15:00:

[0098] (12)

[0099]

[0100]

[0101] (13)

[0102] Among them, 0~7 are non-standard numbers. When the last 4 bits of uint4 are placed on the last 4 bits of fp8, they do not occupy the exponent bit, so let a equal to 0. 8~15 are standard numbers. When the last 4 bits of uint4 are placed on the last 4 bits of fp8, they occupy 1 exponent bit and 3 mantissa bits, so let a equal to 1.

[0103] In summary:

[0104] It can be understood that based on the above reasoning and mapping relationship, not only can fast conversion from in4 to FP8 be achieved, but this method can also be used for conversion from uint8 to BF16 or other data types.

[0105] In some embodiments, when CacheK2 is based on 4-bit storage, [0,1,2,3,4,5,6,7] can be stored in the order of [0,4,1,5,2,6,3,7]. In this way, 8 4-bit numbers can be converted using a preset number of "AND operations" (such as 1 time) and loop control instructions (such as loop instructions, looping 3 times). The number of instructions required for each value conversion can be effectively reduced, for example, to 0.25 in the case of 3 loops.

[0106] According to the embodiments of the present disclosure, fixed-point data can be used to store the second key features to save video memory. Through fast conversion, floating-point data can be used to obtain the first key features, thereby improving both computational efficiency and accuracy. Furthermore, by separating the "zero offset" and "scaling" operations during the dequantization process, the efficiency of dequantization is improved, effectively reducing the computational resource overhead of the attention mechanism.

[0107] The process of obtaining the first value feature is further explained below.

[0108] In some embodiments, the first value feature is obtained by converting the second value feature of a fixed-point type into a floating-point type, and the numerical precision of the first value feature is greater than the numerical precision of the second value feature.

[0109] For example, refer to Figure 3 , the first value feature V331 can be obtained by converting the second value feature V332 of type int4 to type FP8. For example, the second value feature V332 can be stored in the video memory (CacheV2). The fixed-point type second value feature V332 can effectively save video memory, and is converted to floating-point type for multiplication calculations, while improving calculation accuracy.

[0110] The following further describes the fast conversion operation of "converting the second value feature of the int4 type to the FP8 type to obtain the first value feature".

[0111] In some embodiments, the first value feature is obtained by converting a second value feature of a fixed-point type into a floating-point type, including: placing multiple feature values ​​of the second value feature at preset positions of multiple second data respectively to obtain the first value feature, where the preset positions are multiple consecutive bits in the second data; wherein the numerical precision of the second data is consistent with the numerical precision of the first value feature.

[0112] For example, the second value feature is of type int4, and its multiple feature values ​​correspond to multiple numerical values ​​in int4. The second data may include byte data. The second value feature of type int4 can be converted into type uint4 by adding 8 as a whole. Then, each element of the second value feature of type uint4 is placed in the last 4 bits (i.e., the preset position, continuous bits) of a byte and regarded as the first value feature of type FP8. The numerical precision of the byte is regarded as 8 bits, which is consistent with the numerical precision of the first value feature of type FP8. The mapping relationship of this embodiment conforms to the representation content of formula 6 above and will not be repeated here.

[0113] In some embodiments, when CacheV2 is based on 4-bit storage, interleaved storage can also be used, so that 8 4-bit numbers can be converted using one "AND operation" and loop3 instruction, and the number of instructions required for each value conversion can be only 0.25.

[0114] According to an embodiment of the present disclosure, a fixed-point type can be used to store the second value feature to save video memory, and a floating-point type first value feature can be obtained through fast conversion, thereby improving both calculation efficiency and calculation accuracy.

[0115] The following describes the dequantization process for the second-value feature in conjunction with the fast conversion operation of "converting the second-value feature of type int4 to type FP8 to obtain the first-value feature".

[0116] In some embodiments, reference Figure 3 , based on the feature to be processed P30 and the second multiplication result feature score342, determining the attention feature Att361 includes: obtaining the first intermediate feature Inter351 based on the second multiplication result feature score342 and the inverse quantization scaling parameter V333 for the second value feature V332; obtaining the second intermediate feature Inter352 based on the row summation result P31 of the feature to be processed P30, the zero point parameter V334 for the second value feature V332 and the inverse quantization scaling parameter V333 for the second value feature V332; subtracting the second intermediate feature Inter352 from the first intermediate feature Inter351 to obtain the attention feature Att361.

[0117] The attention feature can be calculated according to the following formula:

[0118] (14)

[0119] in, is the characteristic of the second multiplication result, Can be based on The inverse quantization scaling parameter for the second value feature can be The zero point parameter for the second value feature can be . It can be the first intermediate feature. The results can be summed row by row. It can be a second intermediate feature.

[0120] Combining Formulas 1 to 5, the derivation process of Formula 14 is described below.

[0121] The original dequantization formula for value features is:

[0122] (15)

[0123] in, Can be a generalized value feature.

[0124] Formula 15 can be transformed into:

[0125] (16)

[0126] because To perform the Softmax result, can be transformed into:

[0127]

[0128] It can be understood that in Formula 8, each dequantization performs matrix subtraction and matrix multiplication operations on each eigenvalue of the value feature. Based on Formulas 7 and 9, the first intermediate feature and the second intermediate feature can be calculated by performing a matrix multiplication task, and then a one-time subtraction operation is performed. Therefore, the dequantization of the second value feature can be placed in the calculation process of the attention feature, and coordinated with the fast conversion operation from the second value feature to the first value feature. The simplified dequantization process is more suitable for accelerating the operation by performing matrix multiplication tasks (such as the matrix processing unit of the artificial intelligence chip), thereby improving computing efficiency and achieving fast dequantization.

[0129] In some embodiments, reference Figure 3 The preset function includes an exponential function, and the calculation unit is used to execute the preset function processing task to use the preset function to process the first multiplication result feature score341 to obtain the feature P30 to be processed, including: using an approximate calculation function to replace the exponential function, processing the first multiplication result feature score341, and obtaining the feature P30 to be processed. The approximate calculation function indicates the logical operation process of approximating the exponential function by calculating the limit.

[0130] For example, the preset function may be the softmax function. The formula of the softmax function is as follows:

[0131] (17)

[0132] in, Can be .

[0133] It can be seen that the Softmax operation involves the exp (exponential) calculation process. The approximate calculation function is as follows:

[0134] (18)

[0135] From the perspective of sequence limits, when n approaches infinity, Equation 18 holds. The exponential function can be approximated using logical operations for finding limits, where the nth power can be calculated using fast exponentiation.

[0136] According to the embodiments of the present disclosure, the exponential function can be approximated by using the logical operation of finding the limit, which can save computing resources and improve computing efficiency.

[0137] Figure 4 1 is a diagram comparing the architectures of different large models according to an embodiment of the present disclosure.

[0138] like Figure 4 As shown, the large model LLM1 uses BF16 type data in the Transformer network to participate in Attention calculations, where BF16 type is used to store Cache KV (i.e., cache key value). This 16-bit floating point calculation method causes a performance bottleneck in long text reasoning. The large model LLM2 uses FP8 type data in the Transformer network to participate in Attention calculations, where int4 type is used to store Cache KV. The use of int4 type can achieve 4-bit storage of Cache KV, effectively reducing the demand for video memory. At the same time, the use of FP8 type data in Attention calculations can use lower bit calculations, which can greatly improve computing efficiency and accuracy and save computing resources. Figure 5 The task execution method based on the large model LLM2 is further explained.

[0139] Figure 5 is a schematic diagram of a task execution method according to another embodiment of the present disclosure.

[0140] like Figure 5 As shown, first, the second query feature of the BF16 type is quantized, and an inverse quantization operation is performed for the key feature. For example, the second query feature, the quantization parameter _q for the second query feature, and the inverse quantization scaling parameter _k for the second key feature are multiplied to obtain the first query feature. The above formula 3 can be referred to.

[0141] Then, the second key cache uses 4 bits to store data, for example, the second key feature of the int4 type can be quickly converted to the first key feature of the FP8 type to obtain the first key cache.

[0142] Then, the FP8 GEMM 401 is used to perform a matrix multiplication task based on the first query feature and the first key feature in the first key cache, and obtain a first multiplication result feature of the FP8 type.

[0143] Then, the first multiplication result feature of FP32 type is obtained by using the first multiplication result feature and the inverse quantization scaling parameter _q for the second query feature, and divided by , continue to use the soft maximum function for processing.

[0144] Next, the exponential function in the softmax function is approximated by calculating limits to obtain a first intermediate result of FP32 type (not shown in the figure). This intermediate result of FP32 type is then converted to FP8 type for processing. This allows high-precision data to be used in calculations using the softmax function, and then converted to FP8 type to facilitate subsequent low-precision matrix multiplication tasks.

[0145] Then, the second value cache uses 4 bits to store data, for example, the second value feature of the int4 type can be quickly converted into the first value feature of the FP8 type to obtain the first value cache.

[0146] Then, the FP8 GEMM 402 is used to perform a matrix multiplication operation task based on the feature to be processed and the first value feature in the first value cache to obtain a second intermediate result of the FP8 type (not shown in the figure), which is then converted into a second multiplication result feature of the FP32 type to adapt to the result output by the soft maximum function, that is, the row-wise summation result of the feature to be processed, where the soft maximum function outputs FP32 type data.

[0147] Then, multiply the second multiplication result feature of the FP32 type by the inverse quantization scaling parameter _v for the second value feature to obtain the first intermediate feature, multiply the row summation result of the FP32 type, the zero point parameter _v for the second value feature, and the inverse quantization scaling parameter _v for the second value feature to obtain the second intermediate feature, and subtract the second intermediate feature from the first intermediate feature to obtain the attention feature, referring to the above formula 14. Figure 5 The float data refers to FP32 type data.

[0148] In the embodiment of the present disclosure, at least FP8 GEMM 401 and FP8 GEMM 402 are used to convert the matrix multiplication operation task of the first query feature and the first key feature, as well as the matrix multiplication operation task of the feature to be processed and the first value feature, into FP8 calculation, which greatly reduces the time consumption of Attention calculation and ensures that the model accuracy effect is not lost.

[0149] In the embodiments of the present disclosure, the quantization scaling parameters or inverse quantization scaling parameters for query features, key features or value features can be obtained by static statistics, thereby avoiding the degradation of inference performance. Figure 5 In the task execution method shown, the query features can initially be BF16 data. The 8-bit scaling parameter of Head Wise (attention head dimension) can be used to set the corresponding scaling parameter for each attention head. The Cache KV is stored as 4-bit data. The 4-bit scaling parameter of Channel Wise (channel dimension) can be used to set the scaling parameter for each channel, taking into account storage saving, high computational accuracy, and high computational efficiency.

[0150] The embodiments of this disclosure provide efficient, low-bit Attention quantization inference technology. Through sophisticated algorithm optimization and data processing, this technology significantly reduces the time required for model inference calculations, thereby greatly improving user experience and satisfaction. Furthermore, while ensuring model performance, it significantly reduces reliance on and usage of hardware resources (such as computing power and storage space), enabling rational allocation and efficient utilization of resources, and promoting the widespread and widespread adoption of AI technology in various application scenarios.

[0151] It is understandable that, referring to Figure 5 For the task execution method of any embodiment, any one of the dequantization process of the second key feature, the softmax-based approximate calculation, and the dequantization process of the second value feature of the embodiment of the present disclosure can be selected for execution alone, or two or more of them can be combined for execution.

[0152] For example, in some embodiments of the task execution method, based on the first query feature and the first key feature, the computing unit is used to execute the first multiplication operation task of the target attention task queue to obtain the first multiplication operation result feature; the computing unit is used to execute the preset function processing task to process the first multiplication operation result feature using the preset function to obtain the feature to be processed; based on the feature to be processed and the first value feature, the computing unit is used to execute the second multiplication operation task to obtain the second multiplication operation result feature; based on the second multiplication operation result feature and the inverse quantization scaling parameter for the second value feature, the first intermediate feature is obtained; based on the row summation result of the feature to be processed, the zero point parameter for the second value feature and the inverse quantization scaling parameter for the second value feature, the second intermediate feature is obtained; the first intermediate feature is subtracted from the second intermediate feature to obtain the attention feature.

[0153] For example, in some embodiments of the task execution method, based on the first query feature and the first key feature, the computing unit is used to execute the first multiplication operation task of the target attention task queue to obtain the first multiplication operation result feature; the computing unit is used to execute the preset function processing task, and the approximate calculation function is used to replace the exponential function to process the first multiplication operation result feature to obtain the feature to be processed, and the approximate calculation function indicates the logical operation process of approximating the exponential function by calculating the limit; based on the feature to be processed and the first value feature, the computing unit is used to execute the second multiplication operation task to obtain the second multiplication operation result feature; based on the feature to be processed and the second multiplication operation result feature, the attention feature is determined.

[0154] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0155] Figure 6 is a block diagram of a task execution apparatus according to an embodiment of the present disclosure.

[0156] like Figure 6 As shown, the task execution device 600 may include a first execution module 610 and a second execution module 620 .

[0157] The first execution module 610 can execute operation S210, which is used to use the computing unit to execute the first multiplication task of the target attention task queue based on the first query feature and the first key feature to obtain a first multiplication result feature, wherein the first key feature is obtained by converting the second key feature of the fixed-point type into the floating-point type, the numerical precision of the first key feature is greater than the numerical precision of the second key feature, and the numerical precision of the first query feature is less than the numerical precision of the second query feature.

[0158] The second execution module 620 can perform operation S220 to obtain an attention feature by using a computing unit to execute multiple target tasks in the target attention task queue according to the first multiplication result feature and the first value feature.

[0159] The first query feature is obtained through the second query feature and the inverse quantization scaling parameter for the second key feature, so that the hardware resource overhead of the computing unit executing the target attention task queue is less than a preset hardware resource overhead value.

[0160] In some embodiments, the first query feature is obtained by the second query feature and the inverse quantization scaling parameter for the second key feature, including: multiplying the second query feature, the quantization scaling parameter for the second query feature and the inverse quantization scaling parameter for the second key feature to obtain the first query feature.

[0161] In some embodiments, the first key feature is obtained by converting a second key feature of a fixed-point type into a floating-point type, including: placing multiple feature values ​​of the second key feature at preset positions of multiple first data respectively to obtain the first key feature, where the preset positions are multiple consecutive bits in the first data; wherein the numerical precision of the first data is consistent with the numerical precision of the first key feature.

[0162] In some embodiments, the multiple target tasks include a preset function processing task and a second multiplication operation task, and the second execution module 620 also includes: a preset function unit, used to use the calculation unit to execute the preset function processing task, so as to use the preset function to process the first multiplication operation result feature to obtain the feature to be processed; a first execution unit, used to use the calculation unit to execute the second multiplication operation task according to the feature to be processed and the first value feature, to obtain the second multiplication operation result feature; an attention calculation unit, used to determine the attention feature based on the feature to be processed and the second multiplication operation result feature.

[0163] In some embodiments, the preset function includes an exponential function, and the preset function unit is used to use an approximate calculation function to replace the exponential function, process the first multiplication operation result feature, and obtain the feature to be processed. The approximate calculation function indicates the logical operation process of approximating the exponential function by calculating the limit.

[0164] In some embodiments, the first value feature is obtained by converting the second value feature of a fixed-point type into a floating-point type, and the numerical precision of the first value feature is greater than the numerical precision of the second value feature.

[0165] In some embodiments, the first value feature is obtained by converting a second value feature of a fixed-point type into a floating-point type, including: placing multiple feature values ​​of the second value feature at preset positions of multiple second data respectively to obtain the first value feature, where the preset positions are multiple consecutive bits in the second data; wherein the numerical precision of the second data is consistent with the numerical precision of the first value feature.

[0166] In some embodiments, the attention calculation unit is further used to: obtain a first intermediate feature based on the second multiplication result feature and the inverse quantization scaling parameter for the second value feature; obtain a second intermediate feature based on the row summation result of the feature to be processed, the zero point parameter for the second value feature and the inverse quantization scaling parameter for the second value feature; subtract the second intermediate feature from the first intermediate feature to obtain the attention feature.

[0167] For the parts not mentioned in the apparatus part, they can be understood with reference to the various embodiments of the above-mentioned method. That is, the apparatus part includes modules for executing the various steps of any one of the method embodiments described above. In addition, the implementation methods, technical problems solved, functions achieved, and technical effects achieved of each module / unit / subunit, etc. in the apparatus part embodiment are respectively the same or similar to the implementation methods, technical problems solved, functions achieved, and technical effects achieved of each corresponding step in the method part embodiment, and will not be repeated here.

[0168] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0169] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0170] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.

[0171] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the above method when executed by a processor.

[0172] According to an embodiment of the present disclosure, the large model includes a computer program, and the computer program implements the above method when executed by a processor.

[0173] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0174] like Figure 7 As shown, electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from storage unit 178 into a random access memory (RAM) 703. Various programs and data required for operation can also be stored in RAM 703. Computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0175] Multiple components in the electronic device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0176] The computing unit 701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the methods and processes described above. For example, in some embodiments, the methods described above may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the methods described above by any other suitable means (e.g., via firmware).

[0177] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0178] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0179] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0181] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0182] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0183] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of this disclosure can be achieved, and this document is not limited here.

[0184] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A task execution method, comprising: executing, using a computing unit, a first multiplication operation task of the target attention task queue according to the first query feature and the first key feature, to obtain a first multiplication operation result feature, wherein the first key feature is obtained by converting a second key feature of a fixed-point type into a floating-point type, a numerical precision of the first key feature is greater than a numerical precision of the second key feature, and a numerical precision of the first query feature is less than a numerical precision of the second query feature; executing, by the computing unit, a plurality of target tasks in the target attention task queue according to the first multiplication result feature and the first value feature, to obtain an attention feature; The first query feature is obtained through the second query feature and the inverse quantization scaling parameter for the second key feature, so that the hardware resource overhead of the computing unit executing the target attention task queue is less than a preset hardware resource overhead value.

2. The method according to claim 1, wherein The first query feature is obtained by using the second query feature and a dequantized scaling parameter for the second key feature, including: The second query feature, the quantization scaling parameter for the second query feature, and the inverse quantization scaling parameter for the second key feature are multiplied to obtain the first query feature.

3. The method according to claim 1, wherein The first key feature is obtained by converting the second key feature of the fixed-point type into a floating-point type, including: Placing the plurality of characteristic values ​​of the second key characteristic at a plurality of preset positions of the first data respectively to obtain the first key characteristic, wherein the preset positions are a plurality of consecutive bits in the first data; The numerical precision of the first data is consistent with the numerical precision of the first key feature.

4. The method according to claim 1, wherein The multiple target tasks include a preset function processing task and a second multiplication operation task, The step of obtaining the attention feature by executing the plurality of target tasks of the target attention task queue using the computing unit according to the first multiplication result feature and the first value feature includes: Utilizing the computing unit to execute the preset function processing task, so as to process the first multiplication result feature using the preset function to obtain a feature to be processed; performing the second multiplication operation task using a computing unit according to the feature to be processed and the first value feature to obtain a second multiplication operation result feature; The attention feature is determined based on the feature to be processed and the second multiplication result feature.

5. The method according to claim 4, wherein The preset function includes an exponential function, and the using the calculation unit to perform the preset function processing task to process the first multiplication result feature using the preset function to obtain the feature to be processed includes: The exponential function is replaced by an approximate calculation function to process the first multiplication result feature to obtain a feature to be processed, wherein the approximate calculation function indicates a logical operation process of approximately calculating the exponential function by calculating the limit.

6. The method according to claim 4, wherein: The first value feature is obtained by converting a second value feature of a fixed-point type into a floating-point type, and a numerical precision of the first value feature is greater than a numerical precision of the second value feature.

7. The method according to claim 6, wherein: The first value feature is obtained by converting the second value feature of the fixed-point type into a floating-point type, including: Placing the plurality of characteristic values ​​of the second value characteristic at preset positions of a plurality of second data respectively to obtain the first value characteristic, wherein the preset positions are a plurality of consecutive bits in the second data; The numerical precision of the second data is consistent with the numerical precision of the first value feature.

8. The method according to claim 6, wherein: The determining the attention feature based on the feature to be processed and the second multiplication result feature includes: Obtaining a first intermediate feature based on the second multiplication result feature and a dequantization scaling parameter for the second value feature; Obtaining a second intermediate feature based on a row-wise summation result of the feature to be processed, a zero point parameter for the second-valued feature, and a dequantization scaling parameter for the second-valued feature; Subtract the second intermediate feature from the first intermediate feature to obtain the attention feature.

9. A task execution device, comprising: a first execution module, configured to execute, using a computing unit, a first multiplication operation task of the target attention task queue based on the first query feature and the first key feature, to obtain a first multiplication operation result feature, wherein the first key feature is obtained by converting a second key feature of a fixed-point type into a floating-point type, a numerical precision of the first key feature is greater than a numerical precision of the second key feature, and a numerical precision of the first query feature is less than a numerical precision of the second query feature; a second execution module, configured to execute, by using the calculation unit, a plurality of target tasks in the target attention task queue according to the first multiplication result feature and the first value feature, to obtain an attention feature; The first query feature is obtained through the second query feature and the inverse quantization scaling parameter for the second key feature, so that the hardware resource overhead of the computing unit executing the target attention task queue is less than a preset hardware resource overhead value.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.