Method for performing matrix multiplication operation by processor comprising plurality of computing units

By dividing the quantized parameter matrix into multiple parameter blocks and assigning it to multiple calculation units for solution quantization and matrix multiplication operations, the problem of waste of computing resources caused by multiple solutions in the big model inference process is solved, and the efficiency and processor speed of matrix multiplication operations are improved.

CN120234516AActive Publication Date: 2025-07-01SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510725346.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

In the inference process of large models, the quantized model parameters need to be solved and quantized multiple times, resulting in waste of computing resources and reduced inference speed. How to improve the computational efficiency of matrix multiplication operations has become an urgent problem.

Method used

The quantized parameter matrix is ​​divided into multiple parameter blocks and is assigned to multiple calculation units in sequence without repeated repetition, so that each calculation unit can dequantize the allocated parameter block and perform matrix multiplication with the input matrix to avoid multiple solutions of the quantized parameters globally.

Benefits of technology

The calculation efficiency of matrix multiplication operation is improved, the calculation speed of the processor is accelerated, the computing resources are saved, and the inference efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234516A_ABST
    Figure CN120234516A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a method for executing matrix multiplication by a processor comprising a plurality of computing units. In the method, a quantization parameter matrix is divided into a plurality of parameter blocks, and the parameter blocks are sequentially and respectively distributed to a plurality of calculation units in a non-repeated manner, so that at most one calculation unit is distributed with less than a first number of continuous parameter blocks in the plurality of parameter blocks, each computing unit is allocated a first number of consecutive parameter blocks in the plurality of parameter blocks; each calculation unit is used for carrying out dequantization on the distributed parameter blocks; each of the plurality of calculation units obtains a portion of the input matrix that should be multiplied by the allocated parameter block, and performs matrix multiplication on the dequantized allocated parameter block and the portion of the input matrix to complete matrix multiplication of the input matrix and the model parameter. Therefore, only one-time solution quantization needs to be carried out on the quantized parameter matrix globally, calculation resources are saved, and the reasoning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for a processor including multiple computing units to perform matrix multiplication operations. Background Art

[0002] With the development of artificial intelligence, large models such as large language models (LLMs) based on the Transformer architecture have been widely used in many fields. For example, large models can be used to process query tasks, generate natural language texts, generate images or videos, etc. Generally, a large amount of general matrix multiply (GEMM) operations may be included in the inference process of large models, and may involve parameter amounts of billions or more (such as model weights), which will occupy a large amount of storage space and pose huge challenges to the actual deployment and application of the models.

[0003] Currently, in order to reduce the storage space occupied by model parameters, model parameter quantization algorithms can be used to reduce the bit width of model parameters. For example, the Adaptive Weighted Quantization (AWQ) algorithm can quantize the 16-bit parameters of a pre-trained model to 4 bits. For LLMs, the number of parameters is reduced by approximately three-quarters, thus greatly reducing the storage, memory access, and other overheads brought by model parameters.

[0004] Although using model parameter quantization algorithms to quantize and compress model parameters can save storage space, during the inference process of the model, before performing the matrix multiplication operation between the input data and the model parameters, it is necessary to first dequantize the quantized model parameters (for example, restore from 4 bits to 16 bits), resulting in a long time-consuming matrix multiplication operation and reducing the inference efficiency of the model. Therefore, how to improve the computational efficiency of matrix multiplication operations for quantized model parameters to improve the inference efficiency is an urgent problem to be solved. Summary of the Invention

[0005] The inventors of the present disclosure have noticed that during the inference process of large models, since the scales of the input data (such as a 32×4096 matrix) and model parameters (such as a 4096×4096 matrix) that need to perform matrix multiplication are relatively large, the input data and model parameters are usually divided into multiple blocks and allocated to multiple computing units (such as Streaming Multiprocessors (SMs) in a Graphics Processing Unit (GPU)), so that each computing unit performs matrix multiplication on the allocated input data block and the corresponding model parameter block to obtain an output data block as the operation result, and the output data blocks of all multiple computing units can form the complete output data. In the related art, the computing tasks of the computing units are usually mapped to the output data, that is, for an output data block, the input data block and the corresponding model parameter block that can operate to obtain the output data block are allocated to one computing unit. Since the operations for different output data blocks often need to use the same model parameter block, when the model parameters are quantized, the computing units allocated with the same model parameter block all need to dequantize the same model parameter block before performing matrix multiplication with the corresponding input data block. This results in the need to dequantize the model parameters globally multiple times during the matrix multiplication operation of the input data and model parameters, wasting computing resources, reducing the utilization rate and computing speed of the processor, and reducing the inference speed of the model.

[0006] In view of this, the present disclosure proposes a method for a processor including multiple computing units to perform matrix multiplication operations, which can avoid dequantizing the quantized model parameters multiple times during the inference process when the model parameters of the large model are quantized, so as to improve the computing efficiency of the matrix multiplication operation, thereby accelerating the computing speed of the processor, saving computing resources, and improving the inference efficiency.

[0007] According to an aspect of the present disclosure, there is provided a method for a processor including multiple computing units to perform matrix multiplication operations, the method comprising:

[0008] Dividing the quantized parameter matrix into multiple parameter blocks such that each parameter block has a preset first length and first width;

[0009] Sequentially and non-repeatedly allocating the multiple parameter blocks to each of the multiple computing units, such that except for at most one computing unit being allocated less than the first number of consecutive parameter blocks among the multiple parameter blocks, each computing unit is allocated the first number of consecutive parameter blocks among the multiple parameter blocks;

[0010] Each of the multiple computing units dequantizes the allocated parameter blocks;

[0011] Each of the multiple computing units obtains a part of the input matrix that should be multiplied by the allocated parameter block, and performs a matrix multiplication operation on the dequantized allocated parameter block and the part of the input matrix to complete the matrix multiplication operation between the input matrix and the model parameters of the large model.

[0012] Wherein, the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, the first quantity is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the multiple computing units, and

[0013] The input matrix is a two-dimensional matrix that is the input data in the decoding stage of the inference of the large model.

[0014] In a possible implementation, the allocated parameter blocks of each of the multiple computing units are distributed across the entire length of the quantization parameter matrix, and the part of the input matrix obtained by each computing unit is the entire input matrix.

[0015] In a possible implementation, the performing the matrix multiplication operation on the dequantized allocated parameter block and the part of the input matrix includes:

[0016] For each parameter block in the dequantized allocated parameter block of each of the multiple computing units, using a preset matrix multiplication operation instruction, sequentially and without repetition, performs a matrix multiplication operation on a sub-parameter block with a second length and a second width in the parameter block and a sub-input block with a third length and a third width in the input matrix that should be multiplied by the sub-parameter block to complete the matrix multiplication operation between the allocated parameter block and the input matrix.

[0017] In a possible implementation, the method further includes:

[0018] When the length of the input matrix indicates the batch size of the input data and is less than the length and width of the multiplicand matrix specified by the matrix multiplication operation instruction, and the width of the multiplier matrix specified by the matrix multiplication operation instruction is less than the length and width of the multiplicand matrix and the length of the multiplier matrix specified by the matrix multiplication operation instruction,

[0019] When performing the matrix multiplication operation between the sub-parameter block and the sub-input block using the matrix multiplication operation instruction, taking the sub-parameter block as the multiplicand matrix and the sub-input block as the multiplier matrix, and

[0020] The second length, the second width, and the third width are respectively equal to the length, width, and length of the multiplicand matrix specified by the matrix multiplication operation instruction, and the second width is equal to the third length.

[0021] In a possible implementation, when the length of the input matrix is equal to or less than the width of the multiplier matrix specified by the matrix multiplication operation instruction, the third length is equal to the length of the input matrix.

[0022] In a possible implementation, the processor is a graphics processing unit, the computing unit is a streaming multiprocessor, and

[0023] Each computing unit runs a thread block, and in the thread block, the matrix multiplication operation of the sub-parameter block and the sub-input block is sequentially performed for each parameter block in the allocated parameter blocks after dequantization.

[0024] In a possible implementation, the matrix multiplication operation instruction is a matrix multiply-accumulate (MMA) instruction in the parallel thread execution (PTX) instruction set.

[0025] According to another aspect of the present disclosure, there is provided a processor including a plurality of computing units for performing matrix multiplication operations, the processor being configured to:

[0026] Divide the quantization parameter matrix into a plurality of parameter blocks such that each parameter block has a preset first length and a first width;

[0027] Allocate the plurality of parameter blocks to each of the plurality of computing units in sequence without repetition, such that except for at most one computing unit being allocated a consecutive number of parameter blocks less than the first number in the plurality of parameter blocks, each computing unit is allocated a consecutive number of the first number of parameter blocks in the plurality of parameter blocks, and

[0028] Each of the plurality of computing units is configured to:

[0029] Dequantize the allocated parameter blocks;

[0030] Obtain the part of the input matrix that should be multiplied by the allocated parameter blocks, and perform a matrix multiplication operation on the dequantized allocated parameter blocks and the part of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model,

[0031] wherein the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, the first number is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the plurality of computing units, and

[0032] The input matrix is a two-dimensional matrix that is input data in the decoding stage of the inference of the large model.

[0033] According to another aspect of the present disclosure, there is provided an apparatus for a processor including a plurality of computing units to perform matrix multiplication operations, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the following operations:

[0034] The processor executes the computer program to implement the following operations:

[0035] Divide the quantization parameter matrix into a plurality of parameter blocks such that each parameter block has a preset first length and first width;

[0036] Allocate the plurality of parameter blocks to each of the plurality of computing units included in the processor in sequence without repetition, such that except for at most one computing unit being allocated a consecutive number of less than the first number of parameter blocks among the plurality of parameter blocks, each computing unit is allocated a consecutive number of the first number of parameter blocks among the plurality of parameter blocks;

[0037] Cause each of the plurality of computing units to dequantize the allocated parameter blocks;

[0038] Cause each of the plurality of computing units to obtain the part of the input matrix that should be multiplied by the allocated parameter blocks, and perform matrix multiplication operations on the dequantized allocated parameter blocks and the part of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model,

[0039] wherein the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, the first number is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the plurality of computing units, and

[0040] The input matrix is a two-dimensional matrix that is input data in the decoding stage of the inference of the large model.

[0041] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0042] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0043] In a method for performing matrix multiplication operations by a processor including multiple computing units according to the present disclosure, a quantization parameter matrix is divided into multiple parameter blocks; the multiple parameter blocks are sequentially and non-repeatedly assigned to each of the multiple computing units such that each computing unit is assigned a consecutive first number of parameter blocks among the multiple parameter blocks, except that at most one computing unit is assigned a consecutive number of parameter blocks less than the first number among the multiple parameter blocks; each of the multiple computing units dequantizes the assigned parameter blocks; each of the multiple computing units obtains a part of the input matrix that should be multiplied by the assigned parameter blocks, and performs a matrix multiplication operation on the dequantized assigned parameter blocks and the part of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model. Since the quantization parameter blocks assigned to each computing unit are non-repeated, globally, the quantization parameter matrix only needs to be dequantized once, which can avoid dequantizing the quantized model parameters multiple times during the inference process when the model parameters of the large model are quantized, improve the computing efficiency of the matrix multiplication operation, speed up the computing speed of the processor, save computing resources, and improve the inference efficiency.

[0044] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure together with the specification.

[0046] Figure 1 A flowchart showing a method for performing matrix multiplication operations by a processor including multiple computing units according to an embodiment of the present disclosure;

[0047] Figure 2 A schematic diagram showing a process of assigning parameter blocks to a computing unit according to an embodiment of the present disclosure;

[0048] Figure 3 A schematic diagram showing a matrix multiplication operation according to an embodiment of the present disclosure;

[0049] Figure 4 A schematic diagram showing a matrix multiplication operation according to another embodiment of the present disclosure;

[0050] Figure 5 A block diagram showing a device for performing matrix multiplication operations by a processor including multiple computing units according to an embodiment of the present disclosure;

[0051] Figure 6 A block diagram showing a device for performing matrix multiplication operations by a processor including multiple computing units according to another embodiment of the present disclosure. Detailed Implementation Modes

[0052] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0053] As used herein, the terms "comprising," "including," "having," or variations thereof are open-ended and include one or more stated features, integers, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integers, elements, steps, components, functions, or groups thereof.

[0054] When an element is referred to as being "connected," "coupled," "responsive," or variations thereof with respect to another element, it can be directly connected, coupled, or responsive to the other element, or intervening elements may be present.

[0055] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0056] The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0057] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation modes. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0058] The inference process of a large model refers to the process of using a trained model to generate output data based on input data. The inference stage of large models such as Generative Pre-trained Transformer (GPT), Open Pre-trained Transformer (OPT), and Meta AI Large Language Model (LLaMA) includes the following two processes: prefill and decoding. Prefill is the computational process in the inference process of a large model before the model generates the first output token after receiving an input token sequence. Decoding is the computational process in which the model iteratively generates subsequent tokens until the end-of-generation flag is reached or the maximum sequence length is reached after the first token is generated in the prefill stage. In actual implementation, the inference stage of a large model may also include other processes, such as input processing before prefill, post-processing after decoding, etc., which are not listed one by one in this disclosure.

[0059] In the inference stage of a large model, matrix multiplication operations take a relatively long time, for example, up to 60% or more of the total inference time. To improve the computational efficiency of matrix multiplication operations, in some implementation manners, the large model can be deployed on a processor including multiple computing units (such as a graphics processing unit GPU). Since the scales of the input data (such as a 32×4096 matrix) and model parameters (such as a 4096×4096 matrix) that need to perform matrix multiplication are both large, the input data and model parameters are usually divided into multiple blocks and allocated to multiple computing units of the processor, so that each computing unit performs matrix multiplication operations on the allocated input data block and the corresponding model parameter block to obtain an output data block as the operation result, and the output data blocks of all multiple computing units can form complete output data. In the related art, the computing tasks of the computing units are usually mapped to the output data, that is, for an output data block, the input data block and the corresponding model parameter block that can operate to obtain the output data block are allocated to a computing unit. Since the operations for different output data blocks often need to use the same model parameter block, in the case where the model parameters are quantized, the computing units allocated with the same model parameter block all need to dequantize the same model parameter block before performing matrix multiplication operations with the corresponding input data blocks.

[0060] For example, as a processor capable of deploying large models, the Nvidia GPU has a powerful Compute Unified Device Architecture (CUDA) ecosystem and extremely high computing power. The CUDA Kernel is a function that executes in parallel on the Nvidia GPU and consists of multiple (e.g., 131,070) threads. A fixed number of (e.g., 512) threads can be organized into thread blocks. The Tensor Core is a dedicated computing unit in the NVIDIA GPU, focusing on accelerating Matrix Multiply-Accumulate (MMA) operations. It supports mixed-precision computing (such as FP16 input and FP32 accumulation). A single instruction can complete the multiplication and addition operations of a small matrix (such as 16×16), and its throughput is greater than that of traditional CUDA Kernels. When the model parameters are quantized, during matrix multiplication, the CUDA kernel of the Nvidia GPU processor caches the data in stages based on the multi-stage buffering technology, divides the parameter matrix into blocks based on the Loop Tiling technology, and dequantizes the divided parameter blocks based on the dequantization magic number optimization technology (e.g., restoring from 4 bits to 16 bits). Then, it calls the Tensor Core component of the Nvidia GPU to perform matrix multiplication on the dequantized parameter blocks and the input data. When performing matrix multiplication through the CUDA kernel, generally, the computing tasks of the thread blocks are mapped to the output data as described above, and the same model parameter blocks are repeatedly assigned to multiple thread blocks. This results in multiple thread blocks needing to dequantize the same model parameter block before executing MMA. From a global perspective, the model parameters are dequantized multiple times, causing a large amount of duplicate work, wasting computing resources, reducing GPU utilization and computing speed, and reducing the inference speed of the model.

[0061] More specifically, taking the Attention Layer of the LLaMA2-7B model as an example, during matrix multiplication, the size (shape) of the input matrix A of the attention layer is batch_size×4096, where batch_size is the number of samples input to the large language model each time. Generally, the value of batch size is much smaller than 4096. The size of the weight matrix B of the attention layer is 4096×4096. When performing matrix multiplication C = A×B between the input matrix A and the weight matrix B, the size of the resulting output matrix C is batch_size×4096.

[0062] Before performing the matrix multiplication operation of the input matrix A and the weight matrix B, in the related art, for example, according to the calculation logic of the CUDA kernel, the input matrix is usually divided into multiple input data blocks a of 16×4096, and the weight matrix is divided into multiple model parameter blocks b of 32×4096. By multiplying the input data block a and the model parameter block b, an output data block c of 16×32 is obtained, i.e., c = a×b. During actual operation, the calculation tasks of the thread blocks are mapped to the output data, that is, the calculation tasks of obtaining each output data block c are assigned to each thread block, so that each thread block obtains the output data block c by calculating c = a×b. If batch_size = 32, the size of the output matrix C is 32×4096, and the number of output data blocks c is (32 / 16)×(4096 / 32) = 256, that is, 256 thread blocks are required. According to the properties of matrix multiplication operations, different input data blocks a need to be multiplied by the same model parameter block b to obtain different output data blocks c. Therefore, in the traditional method of mapping the calculation tasks of thread blocks to output data, the same model parameter block b and different input data blocks a that need to be multiplied by this model parameter block are usually assigned to multiple different thread blocks.

[0063] In addition to the application scenarios of Nvidia GPU processors, in other processors including multiple computing units, there will usually be a situation where the entire processor repeatedly performs multiple dequantization operations on the same model parameters. For example, in AMD GPUs, the calculation tasks of obtaining each output data block c are usually assigned to each computing unit, resulting in multiple different computing units repeatedly performing multiple dequantization operations on the same model parameter block b.

[0064] Thus, in a processor deployed with a large model, according to the traditional method, multiple different computing units will repeatedly perform multiple dequantization operations on the same model parameter block. From a global perspective, the entire processor repeatedly performs multiple dequantization operations on the same model parameter, wasting computing resources, reducing GPU utilization and computing speed, and reducing the inference speed of the model.

[0065] Based on the above technical problems, the present disclosure proposes a method for a processor including multiple computing units to perform matrix multiplication operations. In the case where the model parameters of a large model are quantized, during inference, the quantized parameter matrix is divided into multiple parameter blocks and sequentially and non-repeatedly assigned to multiple computing units, which can avoid performing multiple dequantizations on the quantized model parameters during the inference process, reduce the computational overhead caused by repeated dequantization in the traditional method, improve the computational efficiency of matrix multiplication operations, accelerate the computing speed of the processor, save computing resources, and improve the inference efficiency. Next, the method provided by the present disclosure for a processor including multiple computing units to perform matrix multiplication operations will be introduced in detail.

[0066] Figure 1 A flowchart showing a method for a processor including a plurality of computing units to perform matrix multiplication operations. This method is for a processor having N computing units. N is an integer greater than 1.

[0067] In one example, the processor is an Nvidia GPU. Correspondingly, the computing units include Streaming Multiprocessors (SMs) in the Nvidia GPU. Optionally, the number of SMs is the same as the number of thread blocks of the CUDA Kernel function executed in the Nvidia GPU. At this time, one thread block is allocated to each SM. The thread block is used to dequantize the parameter block and perform matrix multiplication operations using each dequantized parameter block.

[0068] In another example, the processor can also be an AMD GPU platform. Correspondingly, the computing units include Compute Units (CUs) in the AMD GPU, and one Work-Group is allocated to each CU. The Work-Group is used to dequantize the parameter block and perform matrix multiplication operations using each dequantized parameter block.

[0069] In practical applications, the processor can also be of other types. The present disclosure does not limit the implementation manner of the processor, as long as it can apply the method for a processor including a plurality of computing units to perform matrix multiplication operations according to the present disclosure.

[0070] As Figure 1 shown, the method includes:

[0071] Step 101: Divide the quantized parameter matrix into a plurality of parameter blocks such that each parameter block has a preset first length and first width.

[0072] The quantized parameter matrix is a two-dimensional matrix obtained by quantizing model parameters. Exemplarily, the quantized parameter matrix is obtained by quantizing the model parameters in a large model based on a model parameter quantization algorithm. The model parameters can be, for example, the weight matrix of the Attention Layer of the large model, but are not limited thereto. It can also be other parameter matrices that need to perform matrix multiplication operations with the input data (described later) in the decoding stage of the large model. Optionally, the model parameter quantization algorithm can be the AWQ algorithm, uniform quantization, etc. The present disclosure does not limit the type of the model parameter quantization algorithm.

[0073] The length (e.g., the number of rows) of each parameter block, i.e., the first length, is less than the length (e.g., the number of rows) of the quantization parameter matrix, and the width (e.g., the number of columns) of the quantization parameter matrix, i.e., the first width, is less than the width (e.g., the number of columns) of the quantization parameter matrix. The first length and the first width may be the same or different. Taking the case where the first length and the first width are the same as an example, the first length and the first width may be 16, 32, 64, etc. Accordingly, the size (shape) of each parameter block may be 16×16, 32×32, 64×64, etc. In actual implementation, the values of the first length and the first width may also be other values, and the first length and the first width may also be different. This embodiment does not limit the implementation manners of the first length and the first width. In addition, the first length and the first width may be fixed values, or may be values dynamically adjusted according to, for example, the decoding progress, the processor load, etc.

[0074] Step 102: Sequentially and non-repeatedly allocate multiple parameter blocks to each of multiple computing units, so that except that at most one computing unit is allocated less than the first number of consecutive parameter blocks among the multiple parameter blocks, each computing unit is allocated the first number of consecutive parameter blocks among the multiple parameter blocks. Wherein, the first number may be calculated based on the length and width of the quantization parameter matrix, the first length and the first width of each parameter block, and the number of multiple computing units.

[0075] Exemplarily, the determination method of the first number is represented by the following formula:

[0076] ;

[0077] Wherein, K represents the length of the quantization parameter matrix; N represents the width of the quantization parameter matrix; k represents the length of the parameter block, i.e., the first length, and n represents the width of the parameter block, i.e., the first width; x represents the number of computing units, represents rounding up.

[0078] Taking the column direction (vertical) of the quantization parameter matrix as the length direction and the row direction (horizontal) as the width direction as an example, allocating multiple parameter blocks to each of multiple computing units in sequence without repetition, including: starting from the parameter block in the first row and the first column, continuously selecting a first number of parameter blocks along the column direction to obtain the first computing task; continuing to start from the next parameter block after the last parameter block selected in the previous time, continuously selecting a first number of parameter blocks along the column direction to obtain the second computing task, and so on, cycling in this way until starting from the next parameter block after the last parameter block selected in the (x - 1)-th time, continuously selecting all the remaining parameter blocks (the number of this parameter block is less than or equal to the first number) along the column direction to obtain the x-th computing task. During the process of selecting a first number of parameter blocks along the column direction, if all the parameter blocks in a column are selected and the number of selected parameter blocks does not reach the first number, then jump to the next column and continue to select parameter blocks along the column direction until the first number is reached. The first to x-th computing tasks obtained by partitioning are respectively allocated to different computing units so that each computing unit uses the allocated parameter blocks to perform subsequent matrix multiplication calculations.

[0079] For example, referring to Figure 2 , if both the length and width of the quantization parameter matrix are 4096, that is, the size of the quantization parameter matrix is 4096×4096. At this time, K = 4096, N = 4096, and assume that both the first length and the first width are 32, that is, the size of the parameter block is 32×32. At this time, k = 32, n = 32. At this time, after partitioning according to the first length and the first width, 128 parameter blocks in the row direction and 128 parameter blocks in the column direction can be obtained (that is, each rectangular block called a "small block" in Figure 2 ). If the number of computing units x = 80, then according to the above determination method of the first number, the first number is 128×128 / 80 = 205. That is, each computing unit is responsible for calculating the matrix multiplication of 205 consecutive parameter blocks.

[0080] Assume that the numbers of 80 computing units are integers from 0 to 79, then:

[0081] For the computing unit numbered 0, the computing task of this computing unit is: 128 parameter blocks in the first column + 77 parameter blocks in the second column, a total of 205 parameter blocks ( Figure 2 the green rectangular blocks in

[0082] For the computing unit numbered 1, the computing task of this computing unit is: 51 remaining parameter blocks in the second column + 128 parameter blocks in the third column + 26 parameter blocks in the fourth column, a total of 205 parameter blocks ( Figure 2 the red rectangular blocks in

[0083] ......

[0084] Repeat this process;

[0085] For the computing unit numbered 79, the computing task of this computing unit is: all the remaining unselected parameter blocks, that is, 128×128 - 79×205 = 189 parameter blocks.

[0086] Step 103: Each computing unit among multiple computing units dequantizes the allocated parameter blocks.

[0087] The computing unit uses a dequantization algorithm adapted to the model parameter quantization algorithm to dequantize the allocated parameter blocks. For example, if the model parameter quantization algorithm is the AWQ algorithm, the dequantization algorithm is the AWQ inverse quantization algorithm; if the model parameter quantization algorithm is the linear quantization algorithm, the dequantization algorithm is the linear inverse quantization algorithm. The input matrix is never quantized, so dequantization is not required.

[0088] Step 104: Each computing unit among multiple computing units obtains the part of the input matrix that should be multiplied by the allocated parameter blocks, and performs matrix multiplication on the dequantized allocated parameter blocks and the part of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model.

[0089] Exemplarily, the matrix multiplication operation of the input matrix and the model parameters of the large model is used to perform inference on the inference request input to the large model. For example, it is used to perform inference on a prompt to obtain an output token sequence as the inference result.

[0090] Among them, the input matrix is a two-dimensional matrix that is the input data in the decoding stage of the inference of the large model, and its size can be M×K. Among them, M = batch_size, and batch_size represents the number of input vectors in the same batch (i.e., the batch size) input to, for example, the attention layer in the decoding stage of the large model, which can be 64, 32, 16, 8, etc. The value of M is not limited in this embodiment. Exemplarily, the decoding stage of the inference of the large model is used to generate an output token sequence token by token, and each input vector in the decoding stage can be the vector corresponding to the previous token that has been inferred for the token sequence of each prompt input to the large model. Correspondingly, the input matrix can be understood as a set of M vectors corresponding to the previous token that has been inferred for the token sequences of M prompts input to the large model, and the length of each vector is K. The value of K is the same as the length K of the quantization parameter matrix with size K×N.

[0091] In one example, the assigned parameter blocks of each of the multiple computing units are distributed across the entire length of the quantization parameter matrix (which can also be understood as including a complete column of parameter blocks), and the part of the input matrix obtained by each computing unit is the entire input matrix.

[0092] For example: Figure 2 The parameter blocks assigned to the computing unit numbered 0 in are distributed across the entire length of the quantization parameter matrix, that is, including all 128 parameter blocks in the first column. These 128 parameter blocks need to be multiplied by the entire input matrix. Assuming the size of the input matrix is 64×4096, the computing unit numbered 0 needs to obtain the entire 64×4096 input matrix. In addition, the parameter blocks assigned to the computing unit numbered 0 ( Figure 2 the green rectangular blocks in ) span two columns of parameter blocks in the width direction, and the size in the width direction is 64. Therefore, in the computing unit numbered 0, a calculation result with a size of 64×64 can be obtained. In this 64×64 calculation result, the left half with a size of 64×32 is the result of multiplying the entire 64×4096 input matrix by the 128 parameter blocks in the first column (with a size of 4096×32), and can be directly output to the global memory (Global Memory) as part of the final result of the matrix multiplication operation of the input matrix and the parameter matrix. And in this 64×64 calculation result, the right half with a size of 64×32 is the result of multiplying the left part with a size of 64×2464 of the input matrix by the upper 77 parameter blocks in the second column (with a size of 2464×32), and needs to be added to the left part with a size of 64×32 in the calculation result of the computing unit numbered 1 below to obtain the multiplication result of the entire input matrix and the 128 parameter blocks in the second column. The calculation result of the right half with a size of 64×32 of the computing unit numbered 0 can be stored in a buffer, for example, waiting for the corresponding part in the calculation result of the computing unit numbered 1 to arrive and be accumulated with it.

[0093] Another example: Figure 2 The parameter blocks assigned to the computing unit numbered 1 in are also distributed across the entire length of the quantization parameter matrix, that is, including all 128 parameter blocks in the third column. These 128 parameter blocks also need to be multiplied by the entire input matrix. Assuming the size of the input matrix is also 64×4096, the computing unit numbered 1 needs to obtain the entire 64×4096 input matrix. In addition, the parameter blocks assigned to the computing unit numbered 1 ( Figure 2The red rectangular block (in the figure) spans three columns of parameter blocks in the width direction and has a size of 96 in the width direction. Therefore, in the computing unit numbered 1, a computing result with a size of 64×96 can be obtained. In this 64×96 computing result, the left part with a size of 64×32 is the result of multiplying the right part of the input matrix with a size of 64×1632 by the lower 51 parameter blocks (with a size of 1632×32) in the second column. For example, it can be stored in the above buffer and accumulated with the computing result of the right half with a size of 64×32 in the computing unit numbered 0. The middle part with a size of 64×32 in the computing result is the result of multiplying the entire input matrix with a size of 64×4096 by all 128 parameter blocks (with a size of 4096×32) in the third column, and can be directly output to the global memory as part of the final result of the matrix multiplication operation between the input matrix and the parameter matrix. The right part with a size of 64×32 in the computing result is the result of multiplying the left part of the input matrix with a size of 64×832 by the upper 26 parameter blocks (with a size of 832×32) in the fourth column, and can be similarly stored in the above buffer, waiting to be added to the left part with a size of 64×32 in the computing result of the computing unit numbered 2 to obtain the multiplication result of the entire input matrix with 128 parameter blocks in the fourth column.

[0094] And so on, each computing unit in multiple computing units of the processor obtains the part of the input matrix that should be multiplied by the assigned parameter blocks (where when the assigned parameter blocks are distributed across the entire length of the quantized parameter matrix, the entire input matrix is obtained), and performs a matrix multiplication operation on the dequantized assigned parameter blocks and the part of the input matrix, thereby completing the matrix multiplication operation between the input matrix and the model parameters of the large model.

[0095] In summary, in the method for performing matrix multiplication operations by a processor including multiple computing units provided in this embodiment, the quantization parameter matrix is divided into multiple parameter blocks; the multiple parameter blocks are sequentially and non-repetitively assigned to each of the multiple computing units, such that except for at most one computing unit being assigned a consecutive number of parameter blocks less than the first quantity among the multiple parameter blocks, each computing unit is assigned a consecutive number of parameter blocks equal to the first quantity among the multiple parameter blocks; each of the multiple computing units dequantizes the assigned parameter blocks; each of the multiple computing units obtains the part of the input matrix that should be multiplied by the assigned parameter blocks, and performs matrix multiplication operations on the dequantized assigned parameter blocks and the part of the input matrix, thereby completing the matrix multiplication operation of the input matrix and the model parameters of the large model. Since different computing units are assigned different parts (different parameter blocks) of the quantization parameter matrix, globally, dequantization only needs to be performed once on all parameter blocks, saving computing resources compared with the technical solution in the related art that needs to perform dequantization on the same parameter blocks multiple times, improving the computing efficiency of matrix multiplication operations, accelerating the computing speed of the processor, and improving the inference efficiency.

[0096] In one example, performing matrix multiplication operations on the dequantized assigned parameter blocks and the part of the input matrix includes: for each parameter block in the dequantized assigned parameter blocks of each of the multiple computing units, using a preset matrix multiplication operation instruction, sequentially and non-repetitively performing matrix multiplication operations on the sub-parameter block with the second length and the second width in the parameter block and the sub-input block with the third length and the third width in the input matrix that should be multiplied by the sub-parameter block, so as to complete the matrix multiplication operation of the assigned parameter block and the input matrix.

[0097] In one example, the processor is a graphics processing unit (GPU), the computing unit is a streaming multiprocessor (SM), and each computing unit runs a thread block. Accordingly, each of the multiple computing units sequentially and non-repetitively performing matrix multiplication operations on the sub-parameter block with the second length and the second width in the assigned parameter block and the sub-input block with the third length and the third width in the input matrix that should be multiplied by the sub-parameter block includes: in the thread block, sequentially performing matrix multiplication operations on the sub-parameter block and the sub-input block for each parameter block in the dequantized assigned parameter blocks.

[0098] Among them, the matrix multiplication operation instruction is a hardware-level instruction generated by the processor when matrix multiplication operation is required. For example, in the case where the processor including multiple computing units is an Nvidia GPU, the computing unit is a streaming multiprocessor SM, and each computing unit runs a thread block, the matrix multiplication operation instruction can be a Matrix Multiply-Accumulate (MMA) instruction in the Parallel Thread Execution (PTX) instruction set. The length of the sub-parameter block, i.e., the second length, is less than or equal to the length of the parameter block, i.e., the first length, and the width of the sub-parameter block, i.e., the second width, is less than or equal to the width of the parameter block, i.e., the first width. When the sub-parameter block is used as the multiplicand matrix (the matrix in front of the multiplication sign) and the sub-input block is used as the multiplier matrix (the matrix behind the multiplication sign), the second width is equal to the third length, and when the sub-input block is used as the multiplicand matrix and the sub-parameter block is used as the multiplier matrix, the third width is equal to the second length. The specific values of the second length, the second width, the third length, and the third width can be set according to the regulations of the matrix multiplication operation instruction, for example.

[0099] Exemplarily, still using the example of the computing unit numbered 0 given above. In this computing unit, assuming that both the second length and the second width are set to 16, the size of the sub-parameter block is 16×16; the third length is set to 8 and the third width is set to 16, then the size of the sub-input block is 8×16. In this case, for each parameter block after dequantization in this computing unit, for example, in each iteration step of the thread block, the matrix multiplication operation is sequentially and non-repeatedly performed on each 16×16-sized sub-parameter block in the parameter block and all 8×16 sub-input blocks in the part of the input matrix that should be multiplied with the sub-parameter block. For example, in Figure 3 , still assuming that the size of the input matrix is 64×4096, the size of the parameter matrix is 4096×4096, and the computing unit numbered 0 is allocated 128 parameter blocks in the first column and 77 parameter blocks in the second column (the size of each parameter block is 32×32) as in the example in Figure 2 , as shown by the green rectangular blocks in Figure 3 . In the computing unit numbered 0, in the first iteration step, for the first parameter block, the matrix multiplication operation is respectively performed on the four 16×16-sized sub-parameter blocks in the parameter block (shown as dark green in Figure 3 ) and all 8×16 sub-input blocks in the input matrix that should be multiplied with the corresponding sub-parameter blocks. Specifically, for each of the two sub-parameter blocks in the upper row in Figure 3 , all the sub-input blocks in the input matrix that should be multiplied with this sub-parameter block are the eight sub-input blocks in the left column shown as yellow in Figure 3 in the input matrix, and forFigure 3 For each of the two sub-parameter blocks in the lower row in Figure 3 the input matrix, the sub-input blocks in the input matrix that should be multiplied by this sub-parameter block are the 8 sub-input blocks in the rightmost column shown in yellow. For example, for the first time, the first sub-input block (with a size of 8×16) in the leftmost column of the 8 sub-input blocks can be multiplied by the upper-left sub-parameter block (with a size of 16×16) using a matrix multiplication operation instruction. For the second time, the second sub-input block (with a size of 8×16) in the leftmost column of the 8 sub-input blocks can be multiplied by the upper-left sub-parameter block (with a size of 16×16) using a matrix multiplication operation instruction, and so on, to complete the matrix multiplication operation of the first sub-parameter block with each of the 8 sub-input blocks in the leftmost column. Next, for example, the matrix multiplication operation of the first sub-input block in the rightmost column of the 8 sub-input blocks with the lower-left sub-parameter block, the matrix multiplication operation of the second sub-input block in the rightmost column of the 8 sub-input blocks with the lower-left sub-parameter block can be sequentially performed using a matrix multiplication operation instruction, and so on, to complete the matrix multiplication operation of the second sub-parameter block with each of the 8 sub-input blocks in the rightmost column. Next, similarly, the matrix multiplication operation of the upper-right sub-parameter block with each of the 8 sub-input blocks in the leftmost column is performed, and the matrix multiplication operation of the lower-right sub-parameter block with each of the 8 sub-input blocks in the rightmost column is performed. Thus, in each computing unit, for each parameter block in the dequantized and assigned parameter block, using a preset matrix multiplication operation instruction, the matrix multiplication operation of each sub-parameter block in this parameter block with all the sub-input blocks in the input matrix that should be multiplied by this sub-parameter block is sequentially and without repetition performed to complete the matrix multiplication operation of the assigned parameter block with the input matrix.

[0100] When performing the matrix multiplication operation of a sub-parameter block with a sub-input block using a preset matrix multiplication operation instruction, the matrix multiplication operation instruction usually specifies the sizes of the two matrices participating in the operation. That is, the matrix multiplication operation instruction can specify the values of m, k, and n to multiply an m×k matrix by a k×n matrix. For example, the MMA instruction in the PTX instruction set can specify m = 16, n = 8, and k = 16. In this case, referring to Figure 4 the matrix multiplication operation instruction can multiply a matrix A with a size of 16×16 by a matrix B with a size of 16×8 by calculating C = A×B to obtain a matrix C with a size of 16×8.

[0101] In the related art, when performing matrix multiplication on a model parameter block and an input data block using a matrix multiplication instruction, the input data block (or its sub-block) is usually corresponded to the above matrix A as the multiplicand matrix of the matrix multiplication instruction, and the model parameter block (or its sub-block) is corresponded to the above matrix B as the multiplier matrix of the matrix multiplication instruction. In this case, if the length of the input data indicates the batch size (for example, in the case where the size of the input matrix A in the attention layer above is batch_size×4096), and the length (batch size) of the input data is less than the length m and width k of the multiplicand matrix specified by the matrix multiplication instruction, then when the input data block is used as the multiplicand matrix A and the model parameter block is used as the multiplier matrix B according to the related art, the size of the input data block is smaller than the size of the multiplicand matrix A supported by the hardware, resulting in waste of hardware resources. For example, when m = 16, n = 8, k = 16 and the batch size is 1 to 15, the maximum length of the input data block is the batch size from 1 to 15, and the size of the input data block is correspondingly from 1×16 to 15×16 at most. If the input data block with a size of at most 1×16 to 15×16 is used as the multiplicand matrix A, then in actual operation, for example, the remaining part of less than 16×16, i.e., 15×16 to 1×16, will be filled with zeros to meet the size specified by the matrix multiplication instruction. In this way, in each operation, the operation of the data volume from 15×16 to 1×16 is invalid, wasting hardware resources. In addition, it can also be seen that the smaller the batch size, the smaller the size of the valid data, the larger the size of the invalid data, and the more hardware resources are wasted.

[0102] To solve such a technical problem, in some embodiments of the present disclosure, when the length of the input matrix indicates the batch size of the input data and is less than the length and width of the multiplicand matrix specified by the matrix multiplication operation instruction, and the width of the multiplier matrix specified by the matrix multiplication operation instruction is less than the length of the multiplicand matrix and the length of the multiplier matrix specified by the matrix multiplication operation instruction, when performing the matrix multiplication operation of the sub-parameter block and the sub-input block using the matrix multiplication operation instruction, the sub-parameter block is used as the multiplicand matrix, the sub-input block is used as the multiplier matrix, and the second length, the second width, and the third width are respectively equal to the length, width, and length of the multiplicand matrix specified by the matrix multiplication operation instruction, and the second width is equal to the third width. That is, when n < m and n < k are specified by the matrix multiplication operation instruction, the sub-parameter block is corresponded to the above matrix A with the size of m × k, the sub-input block is corresponded to the above matrix B with the size of k × n, the length m of the multiplicand matrix specified by the matrix multiplication operation instruction is determined as the length of the sub-parameter block, i.e., the second length, and the width k of the multiplicand matrix and the length of the multiplier matrix specified by the matrix multiplication operation instruction are determined as the width of the sub-parameter block, i.e., the second width, and the width of the sub-input block, i.e., the third width. For the length of the sub-input block, i.e., the third length, when the length of the input matrix, i.e., the batch size, is greater than the width n of the multiplier matrix specified by the matrix multiplication operation instruction, the input matrix can be divided into sub-input blocks with a maximum size of n × k, and at this time, the length of the sub-input block, i.e., the third length, is equal to or less than the width n of the multiplier matrix specified by the matrix multiplication operation instruction. When the length of the input matrix, i.e., the batch size, is equal to or less than the width n of the multiplier matrix specified by the matrix multiplication operation instruction, there is no need to divide the input matrix in the length direction and only need to divide it in the width direction to obtain sub-input blocks with the size of batch_size × k, and at this time, the length of the sub-input block, i.e., the third length, is equal to the length of the input matrix. By mapping the length of the input matrix indicating the batch size to the smallest one of the length, width, length, and width of the multiplicand matrix and the multiplier matrix specified by the matrix multiplication operation instruction, it is possible to reduce the waste of hardware resources for performing matrix multiplication operations in the case of a small batch size, and improve the hardware utilization rate and calculation efficiency.

[0103] Exemplarily, for the example where the size of the input matrix A of the attention layer in the above text is batch_size × 4096, the length of the input matrix indicates the batch size batch_size of the input data. In the case of m = 16, n = 8, k = 16, n < m and n < k. At this time, if batch_size < m and batch_size < k, that is, batch_size < 16, the sub-parameter block is used as the multiplicand matrix A, and the sub-input block is used as the multiplier matrix B, so that the length of the sub-parameter block, that is, the second length = m = 16, the width of the sub-parameter block, that is, the second width = k = 16, and the width of the sub-input block, that is, the third width = k = 16. When the length of the input matrix, that is, the batch size, is greater than the width n of the multiplier matrix specified by the matrix multiplication operation instruction (8 < batch_size < 16), during the matrix multiplication operation, the operation can be performed on the sub-input block with a maximum size of n × k = 8 × 16. For example, when batch_size = 14, the matrix multiplication operation instruction can be used to perform the matrix multiplication operation on the sub-parameter block with a size of 16 × 16 and the sub-input block with a size of 8 × 16 or 6 × 16 each time. When the length of the input matrix, that is, the batch size, is equal to or less than the width n of the multiplier matrix specified by the matrix multiplication operation instruction (batch_size ≤ 8), during the matrix multiplication operation, the operation can be performed on the sub-input block with a size of batch_size × k. For example, when batch_size = 4, the matrix multiplication operation instruction can be used to perform the matrix multiplication operation on the sub-parameter block with a size of 16 × 16 and the sub-input block with a size of 4 × 16 each time. It should be understood that inside the processor, sometimes there is a mechanism to use k as the internal dimension and automatically align the internal dimension during the matrix multiplication operation. Therefore, there is no need to manually transpose the sub-input block, and the alignment of the k dimension between the sub-parameter block and the sub-input block can be achieved during the actual operation.

[0104] For example, when performing matrix multiplication on the parameter block with a size of 32×32 in the above text, in the case of batch_size = 14, each parameter block needs to be multiplied by the 14×32 part of the input matrix. If the method in the related art is used, for a 32×32 parameter block, in each matrix multiplication operation, the 14×16 sub-input block needs to be used as matrix A and multiplied by the 16×8 sub-parameter block as matrix B respectively, and a total of 8 matrix multiplication operations are required. At this time, as described above, 2×16 data volume is invalid in each matrix multiplication operation, and a total of 16×16 invalid data volume in 8 operations, and the corresponding hardware resources (computing power) are wasted. If the method proposed in the present disclosure is adopted, for a 32×32 parameter block, in each matrix multiplication operation, the 16×16 sub-parameter block needs to be used as matrix A and multiplied by the 8×16 or 6×16 sub-input block as matrix B, and 8 matrix multiplication operations are also required. However, in these 8 matrix multiplication operations, 4 operations are to multiply the 16×16 sub-parameter block by the 8×16 sub-input block, and no invalid data is generated. Only in the 4 operations of multiplying the 16×16 sub-parameter block by the 6×16 sub-input block, 2×16 invalid data is generated each time, and a total of 8×16 invalid data volume in 8 operations. Compared with the method in the related art, the waste of hardware resources is reduced by half, and the hardware utilization rate and calculation efficiency are improved.

[0105] When batch_size < n, this advantage of the method of the present disclosure will be more obvious. For example, when performing matrix multiplication on the parameter block with a size of 32×32 in the above text, in the case of batch_size = 4, each parameter block needs to be multiplied by the 4×32 part of the input matrix. If the method in the related art is used, for a 32×32 parameter block, in each matrix multiplication operation, the 4×16 sub-input block needs to be used as matrix A and multiplied by the 16×8 sub-parameter block as matrix B respectively, and a total of 8 matrix multiplication operations are required. At this time, as described above, 12×16 data volume is invalid in each matrix multiplication operation, and a total of 96×16 invalid data volume in 8 operations, and the corresponding hardware resources (computing power) are wasted. If the method proposed in the present disclosure is adopted, for a 32×32 parameter block, in each matrix multiplication operation, the 16×16 sub-parameter block needs to be used as matrix A and multiplied by the 4×16 sub-input block as matrix B, and 4 matrix multiplication operations are required. In these 4 matrix multiplication operations, 4×16 invalid data is generated each time, and a total of 16×16 invalid data volume in 4 operations, which is only 1 / 6 of that when calculated by the method in the related art. Compared with the method in the related art, the waste of hardware resources is greatly reduced, and the hardware utilization rate and calculation efficiency are improved.

[0106] In summary, according to the embodiments of the present disclosure, when the length of the input matrix indicates the batch size of the input data and is less than the length and width of the multiplicand matrix specified by the matrix multiplication operation instruction, and the width of the multiplier matrix specified by the matrix multiplication operation instruction is less than the length of the multiplicand matrix and the width and length of the multiplier matrix specified by the matrix multiplication operation instruction, when performing the matrix multiplication operation of the sub-parameter block and the sub-input block using the matrix multiplication operation instruction, the sub-parameter block is used as the multiplicand matrix, the sub-input block is used as the multiplier matrix, and the second length, the second width, and the third width are respectively equal to the length, width, and length of the multiplicand matrix specified by the matrix multiplication operation instruction, and the second width is equal to the third width. Thus, when the batch size is less than the length and width of the multiplicand matrix specified by the matrix multiplication operation instruction, especially less than any one of the length and width of the multiplicand matrix and the length and width of the multiplier matrix specified by the matrix multiplication operation instruction, the waste of hardware resources can be reduced, the hardware utilization rate and the computing efficiency can be improved, and thus the inference efficiency can be improved.

[0107] Figure 5 FIG. shows a block diagram of an apparatus for performing matrix multiplication operations by a processor including a plurality of computing units according to an embodiment of the present disclosure. As Figure 5 shown, the apparatus includes: a block partitioning module 510, a block allocation module 520, a dequantization module 530, and a matrix operation module 540.

[0108] The block partitioning module 510 is configured to partition the quantized parameter matrix into a plurality of parameter blocks such that each parameter block has a preset first length and a first width;

[0109] The block allocation module 520 is configured to sequentially and non-repetitively allocate the plurality of parameter blocks to each of the plurality of computing units such that each computing unit is allocated a consecutive first number of the plurality of parameter blocks, except that at most one computing unit is allocated a consecutive number of parameter blocks less than the first number among the plurality of parameter blocks;

[0110] The dequantization module 530 is configured to perform dequantization on the parameter blocks allocated to each of the plurality of computing units;

[0111] The matrix operation module 540 is configured to obtain, for each of the plurality of computing units, a portion of the input matrix that should be multiplied by the allocated parameter block, and perform a matrix multiplication operation on the dequantized allocated parameter block and the portion of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model.

[0112] Among them, the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, and the first quantity is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the multiple computing units, and

[0113] The input matrix is a two-dimensional matrix that is input data in the decoding stage of the inference of the large model.

[0114] For relevant details, see the above method embodiments.

[0115] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.

[0116] The embodiments of the present disclosure further provide a processor including multiple computing units for performing matrix multiplication operations. The processor (specifically, by executing a computer program) is configured to:

[0117] Divide the quantization parameter matrix into multiple parameter blocks such that each parameter block has a preset first length and first width;

[0118] Allocate the multiple parameter blocks to each of the multiple computing units in sequence without repetition, such that except for at most one computing unit being allocated a continuous number of parameter blocks less than the first quantity among the multiple parameter blocks, each computing unit is allocated a continuous first quantity of parameter blocks among the multiple parameter blocks, and

[0119] Each of the multiple computing units is configured to:

[0120] Dequantize the allocated parameter blocks;

[0121] Obtain the part of the input matrix that should be multiplied by the allocated parameter blocks, and perform matrix multiplication operations on the dequantized allocated parameter blocks and the part of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model,

[0122] Among them, the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, the first quantity is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the multiple computing units, and

[0123] The input matrix is a two-dimensional matrix that is input data in the decoding stage of the inference of the large model.

[0124] For relevant details, see the above method embodiments.

[0125] An embodiment of the present disclosure also provides an apparatus for a processor including multiple computing units to perform matrix multiplication operations, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the above method.

[0126] An embodiment of the present disclosure also provides a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0127] An embodiment of the present disclosure also provides a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0128] Figure 6 is a block diagram of an apparatus 1900 for a processor including multiple computing units to perform matrix multiplication operations shown according to an exemplary embodiment. For example, the apparatus 1900 may be provided as a server or a terminal device. Referring to Figure 6 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0129] The apparatus 1900 may further include a power supply component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output interface 1958 (I / O interface). The apparatus 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0130] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the apparatus 1900 to complete the above method and implement the following operations:

[0131] Divide the quantization parameter matrix into multiple parameter blocks such that each parameter block has a preset first length and first width;

[0132] Assign the multiple parameter blocks to each of the multiple computing units included in the processor in sequence without repetition, such that each computing unit is assigned a consecutive number of the first quantity of the multiple parameter blocks, except that at most one computing unit is assigned a consecutive number of less than the first quantity of the multiple parameter blocks;

[0133] Cause each of the multiple computing units to dequantize the assigned parameter blocks;

[0134] Cause each of the multiple computing units to obtain the portion of the input matrix that should be multiplied by the assigned parameter blocks, and perform a matrix multiplication operation on the dequantized assigned parameter blocks and the portion of the input matrix, so as to complete the matrix multiplication operation of the input matrix and the model parameters of the large model,

[0135] wherein, the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, the first quantity is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the multiple computing units, and

[0136] the input matrix is a two-dimensional matrix that is input data in the decoding stage of the inference of the large model.

[0137] A computer-readable storage medium may be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0138] The computer programs (or computer-readable program instructions) described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0139] The computer programs (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0140] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0141] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0142] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0143] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.

[0144] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for a processor including a plurality of computing units to perform matrix multiplication operations, characterized in that, The method includes: Dividing the quantization parameter matrix into a plurality of parameter blocks such that each parameter block has a preset first length and a first width; Successively and non - repetitively allocating the plurality of parameter blocks to each of the plurality of computing units such that each computing unit is allocated a consecutive number of the first quantity of the plurality of parameter blocks, except that at most one computing unit is allocated a consecutive number of less than the first quantity of the plurality of parameter blocks; Each of the plurality of computing units de - quantizes the allocated parameter blocks; Each of the plurality of computing units obtains a portion of the input matrix that should be multiplied by the allocated parameter block, and performs a matrix multiplication operation on the de - quantized allocated parameter block and the portion of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model, wherein, the quantization parameter matrix is a two - dimensional matrix obtained by quantizing the model parameters, the first quantity is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the plurality of computing units, and the input matrix is a two - dimensional matrix that is the input data in the decoding stage of the inference of the large model.

2. The method according to claim 1, wherein The allocated parameter blocks of each of the plurality of computing units are distributed across the entire length of the quantization parameter matrix, and the portion of the input matrix obtained by each computing unit is the entire input matrix.

3. The method according to claim 2, wherein The performing the matrix multiplication operation on the de - quantized allocated parameter block and the portion of the input matrix includes: For each of the plurality of computing units, for each parameter block in the de - quantized allocated parameter block, using a preset matrix multiplication operation instruction, successively and non - repetitively performing a matrix multiplication operation on a sub - parameter block having a second length and a second width in the parameter block and a sub - input block having a third length and a third width in the input matrix that should be multiplied by the sub - parameter block to complete the matrix multiplication operation of the allocated parameter block and the input matrix.

4. The method according to claim 3, characterized in that, The method further includes: When the length of the input matrix indicates the batch size of the input data and is less than the length and width of the multiplicand matrix specified by the matrix multiplication operation instruction, and the width of the multiplier matrix specified by the matrix multiplication operation instruction is less than the length and width of the multiplicand matrix and the length of the multiplier matrix specified by the matrix multiplication operation instruction, When performing the matrix multiplication operation of the sub - parameter block and the sub - input block using the matrix multiplication operation instruction, taking the sub - parameter block as the multiplicand matrix and the sub - input block as the multiplier matrix, and the second length, second width, and third width are respectively equal to the length, width, and length of the multiplicand matrix specified by the matrix multiplication operation instruction, and the second width is equal to the third width.

5. The method according to claim 4, wherein when the length of the input matrix is equal to or less than the width of the multiplier matrix specified by the matrix multiplication operation instruction, the third length is equal to the length of the input matrix.

6. The method according to any one of claims 3 to 5, characterized in that The processor is a graphics processing unit, the computing unit is a streaming multiprocessor, and each computing unit runs a thread block, and in the thread block, matrix multiplication operations of the sub-parameter block and the sub-input block are sequentially performed for each parameter block in the allocated parameter blocks after dequantization.

7. The method according to claim 6, characterized in that The matrix multiplication operation instruction is a matrix multiply-accumulate (MMA) instruction in the parallel thread execution (PTX) instruction set.

8. A processor including a plurality of computing units for performing matrix multiplication operations, characterized in that, The processor is configured to: divide the quantization parameter matrix into a plurality of parameter blocks such that each parameter block has a preset first length and a first width; allocate the plurality of parameter blocks to each of the plurality of computing units in sequence without repetition, such that except for at most one computing unit being allocated a consecutive number of parameter blocks less than the first quantity among the plurality of parameter blocks, each computing unit is allocated a consecutive number of the first quantity of parameter blocks among the plurality of parameter blocks, and each of the plurality of computing units is configured to: dequantize the allocated parameter blocks; obtain a part of the input matrix that should be multiplied by the allocated parameter blocks, and perform matrix multiplication operations on the dequantized allocated parameter blocks and the part of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model, wherein the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, the first quantity is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the plurality of computing units, and the input matrix is a two-dimensional matrix that is input data in the decoding stage of the inference of the large model.

9. An apparatus for a processor including multiple computing units to perform matrix multiplication operations, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the following operations: divide the quantization parameter matrix into a plurality of parameter blocks such that each parameter block has a preset first length and a first width; allocate the plurality of parameter blocks to each of the plurality of computing units included in the processor in sequence without repetition, such that except for at most one computing unit being allocated a consecutive number of parameter blocks less than the first quantity among the plurality of parameter blocks, each computing unit is allocated a consecutive number of the first quantity of parameter blocks among the plurality of parameter blocks; cause each of the plurality of computing units to dequantize the allocated parameter blocks; cause each of the plurality of computing units to obtain a part of the input matrix that should be multiplied by the allocated parameter blocks, and perform matrix multiplication operations on the dequantized allocated parameter blocks and the part of the input matrix to complete the matrix multiplication operation of the input matrix and the model parameters of the large model, wherein the quantization parameter matrix is a two-dimensional matrix obtained by quantizing the model parameters, the first quantity is calculated based on the length and width of the quantization parameter matrix, the first length and first width of each parameter block, and the number of the plurality of computing units, and the input matrix is a two-dimensional matrix that is input data in the decoding stage of the inference of the large model.

10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Matrix multiplication matrix operation method and device

    CN110222308A

  • Model compression method, training method, text data processing method and device

    CN117195978A

  • Matrix vector multiplication calculation method and system and storage medium

    CN118445534A

  • Matrix processing method, processor, system on chip, electronic equipment and storage medium

    CN118656575A

  • Data processing device, chip, method, equipment and medium

    CN118861502A

Cited By

  • Matrix multiplication parameter obtaining method and device, model training method and device, equipment and medium

    CN121167308A