Code generation method and device, electronic equipment and storage medium
By partitioning and fusion of output tensors and automatically generating an MMA intermediate representation, the problem of difficulty in developing high-dimensional tensor matrix multiplication operations in the existing technology is solved, and efficient code generation with low error rates is achieved.
Patent Information
- Application Number
- CN202510765489.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the prior art, the intermediate representation interface of MMA requires 2-dimensional shapes of input tensors and output tensors, which leads to developers need to manually write loop codes for high-dimensional matrix multiplication operations, which is prone to errors and increases the difficulty of development.
By determining the target information, including the MMA operation granularity and tensor information supported by the chip, partitioning the output tensor, generating code segments, and fusing the code segments of each area, automatically generating an intermediate representation of the MMA, reducing the difficulty of development.
Without the need for developers to understand the tensor memory layout, they can automatically generate MMA intermediate representations suitable for tensors in different dimensions, reducing development difficulty and error rate.
Smart Images

Figure CN120276719A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of software development, and particularly to a code generation method, apparatus, electronic device, and storage medium. Background Art
[0002] In order to efficiently develop operators, an assembly generation program or a compiler usually encapsulates some commonly used instructions in an operator into an intermediate representation. For example, matrix multiply accumulate (MMA) is encapsulated into an intermediate representation (IR), and an interface is provided for developers to call.
[0003] In the prior art, the interface of the intermediate representation of MMA has some restrictions on input tensors and output tensors. For example, it is required that the shapes of both the input tensor and the output tensor are two-dimensional. In this way, if a developer needs to perform matrix multiplication operations on higher-dimensional tensors, they need to develop loop code for traversal by themselves, and the developed code is very prone to errors. Summary of the Invention
[0004] Embodiments of this application provide a code generation method, apparatus, electronic device, and storage medium, which are used to reduce the restrictions of the intermediate representation of MMA on input tensors and output tensors and reduce the development difficulty of the intermediate representation of MMA.
[0005] In a first aspect, an embodiment of this application provides a code generation method, including: In response to an instruction for the intermediate representation of matrix multiply accumulate (MMA) of a tensor, determining target information, where the target information includes the operation granularity information of MMA supported by a chip, and the tensor information of an input tensor and an output tensor, the input tensor includes a left tensor and a right tensor, and the tensor information of each tensor includes at least shape information; Partitioning the output tensor according to the tensor information of the input tensor and the output tensor, and the operation granularity information; Based on the left data block corresponding to each region in the left tensor and the right data block corresponding to each region in the right tensor, generating a code segment corresponding to the region, where the code segment is used to write the matrix multiplication operation result of the left data block and the right data block into the region, and each data block in the left data block and the right data block is determined according to the position information of the region and the tensor information of the tensor to which the data block belongs; Fusing the code segments corresponding to each region to obtain the intermediate representation of MMA for the tensor.
[0006] In some embodiments, the input tensor further includes a bias tensor, and further includes: After generating the code segment corresponding to the region, target code is added to the code segment. The target code is used to accumulate the bias data block corresponding to the region in the bias tensor to the matrix multiplication operation result of the left data block and the right data block. The bias data block is determined according to the position information of the region and the tensor information of the bias tensor.
[0007] In some embodiments, generating the code segment corresponding to the region based on the left data block corresponding to each region in the left tensor and the right data block corresponding to each region in the right tensor includes: According to the operation granularity information, the left data block and the right data block are respectively blocked. Based on each group of sub-blocks with the same blocking order in the left data block and the right data block, a set of MMA instructions is generated. The set of MMA instructions is used to perform matrix multiplication operations on the group of sub-blocks. When the group of sub-blocks is not the first group of sub-blocks corresponding to the region, the set of MMA instructions is further used to accumulate the matrix multiplication operation result of the group of sub-blocks to the matrix multiplication operation result of the previous group of sub-blocks corresponding to the region. The groups of MMA instructions are fused to obtain the code segment.
[0008] In some embodiments, generating a set of MMA instructions based on each group of sub-blocks with the same blocking order in the left data block and the right data block includes: According to the address representation information of each sub-block in the group of sub-blocks, an address calculation instruction for the sub-block is generated. Based on the sub-block addresses obtained from the address calculation instructions of each sub-block and the region address obtained from the address calculation instruction of the region, an MMA main instruction is generated. The address calculation instruction of the region is generated according to the address representation information of the region. According to the configuration information required for performing matrix multiplication operations on the group of sub-blocks in the chip, multiple configuration instructions for the MMA main instruction are generated.
[0009] In some embodiments, the tensor information of each tensor further includes step information and address information. When the group of sub-blocks is the first group of sub-blocks corresponding to the first region, the address representation information of each sub-block includes the position information of the sub-block in the corresponding tensor, the step information of the corresponding tensor, and the address information. The address representation information of the region includes the position information of the region in the output tensor, the step information of the output tensor, and the address information. When the group of sub-blocks is not the first group of sub-blocks corresponding to the first region, the address representation information of each sub-block is the offset address of the sub-block compared to the previous sub-block in the corresponding tensor, and the address representation information of the region is the offset address of the region compared to the previous region in the output tensor.
[0010] In some embodiments, when the group of sub - blocks is not the first group of sub - blocks corresponding to the first region, it further includes: When there is a duplication between the configuration information corresponding to the group of sub - blocks and the configuration information corresponding to the previous group of sub - blocks, skip generating the configuration instruction corresponding to the duplicated configuration information when generating the multiple configuration instructions.
[0011] In some embodiments, the input tensor further includes a bias tensor, and the target information further includes input synchronization information. It further includes: After each address calculation instruction in the first code segment, insert a first synchronization instruction, where the first synchronization instruction is used to indicate waiting for the data of at least one of the left tensor, the right tensor, and the bias tensor to be ready, and / or is used to indicate waiting for the memory corresponding to the output tensor to be unlocked and available.
[0012] In some embodiments, the input tensor further includes a bias tensor, and the target information further includes output synchronization information. It further includes: Insert a second synchronization instruction at a non - tail position of the last code segment, where the second synchronization instruction is used to indicate that the memory occupied by at least one of the left tensor, the right tensor, and the bias tensor is unlocked and available, and is also used to indicate that the data of the output tensor is ready; or, Configure synchronization code in the last code segment, where the synchronization code is used to indicate that the memory occupied by at least one of the left tensor, the right tensor, and the bias tensor is unlocked and available, and is also used to indicate that the data of the output tensor is ready.
[0013] In a second aspect, an embodiment of the present application provides a code generation device, including: A determination module, configured to determine target information in response to an instruction of an intermediate representation for performing matrix multiplication and accumulation (MMA) on tensors, where the target information includes operation granularity information of MMA supported by the chip, and tensor information of the input tensor and the output tensor, the input tensor includes a left tensor and a right tensor, and the tensor information of each tensor includes at least shape information; A block - dividing module, configured to partition the output tensor according to the tensor information of the input tensor and the output tensor, and the operation granularity information; A first generation module, configured to generate a code segment corresponding to a region based on the left data block corresponding to the region in the left tensor and the right data block corresponding to the region in the right tensor, where the code segment is used to write the matrix multiplication operation result of the left data block and the right data block into the region, and each data block in the left data block and the right data block is determined according to the position information of the region and the tensor information of the tensor to which the data block belongs; A second generation module, configured to fuse the code segments corresponding to each region to obtain an intermediate representation for performing MMA on tensors.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein: The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute any of the above code generation methods.
[0015] In a fourth aspect, an embodiment of the present application provides a storage medium, and when a computer program in the storage medium is executed by a processor of an electronic device, the electronic device is enabled to execute any of the above code generation methods.
[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, any of the above code generation methods is implemented.
[0017] In the embodiment of the present application, in response to an instruction for an intermediate representation of performing MMA on a tensor, target information is determined. The target information includes operation granularity information of MMA supported by a chip, and tensor information of an input tensor and an output tensor. The input tensor includes a left tensor and a right tensor. Then, according to the tensor information of the input tensor and the output tensor, and the operation granularity information, the output tensor is partitioned. Based on the left data block corresponding to each region in the left tensor and the right data block corresponding to each region in the right tensor, a code segment corresponding to this region is generated. The code segment is used to write the matrix multiplication operation result of the left data block and the right data block into this region. Furthermore, the code segments corresponding to each region are fused to obtain an intermediate representation for performing MMA on the tensor. In this way, by partitioning the output tensor and performing matrix multiplication operations on the data blocks corresponding to each region in each input tensor, regardless of the dimensions of each tensor, according to the target information required for the intermediate representation of performing MMA on the tensor, an intermediate representation of MMA can be automatically generated. Developers do not have to develop loop code for traversal by themselves, nor do they need to understand the memory layout of the tensor. Therefore, the development difficulty of the intermediate representation of MMA can also be reduced. Description of the Drawings
[0018] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings: Figure 1 It is an application scenario diagram of a code generation method provided by an embodiment of the present application; Figure 2Schematic diagram of an intermediate representation of MMA provided by an embodiment of this application; Figure 3 Flowchart of a code generation method provided by an embodiment of this application; Figure 4 Schematic diagram of performing matrix multiplication operation on a tensor provided by an embodiment of this application; Figure 5 Flowchart of another code generation method provided by an embodiment of this application; Figure 6 Schematic diagram of another matrix multiplication operation on a tensor provided by an embodiment of this application; Figure 7 Schematic diagram of the structure of a code generation device provided by an embodiment of this application; Figure 8 Schematic diagram of the hardware structure of an electronic device for implementing the code generation method provided by an embodiment of this application. Detailed implementation manners
[0019] To reduce the restrictions of the intermediate representation of MMA on input tensors and output tensors and reduce the development difficulty of the intermediate representation of MMA, embodiments of this application provide a code generation method, device, electronic device, and storage medium.
[0020] The following describes the preferred embodiments of this application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain this application and are not used to limit this application. And without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0021] To facilitate the understanding of this application, in the technical terms involved in this application: 1. Tensor, a basic concept in a deep learning model, generally designed for a graphics processing unit (GPU) and can run on the GPU to accelerate the calculation efficiency. In an AI model, a tensor can be regarded as a multi-dimensional array. Generally, a tensor will include multiple dimensions (or called logical axis dimensions). The tensor information of a tensor can include shape information (shape), stride information (stride), address information (address), etc. The shape information is used to describe the number of elements of the tensor in each dimension. The stride information is used to describe the element interval of the tensor in each dimension and can characterize the memory layout of the tensor (that is, characterize how the data in the tensor is stored in the memory). The address information such as the base address is used to describe the physical storage address of the tensor in the memory.
[0022] Suppose the tensor information of a certain tensor X is: shape X= (shapeN, shapeH, shapeW) = (2, 128, 448); stride X = (strideN, strideH, strideW) = (448 × 128 = 57344, 448, 1); address = address X 。
[0023] Then, the tensor X has three dimensions: N, H, W. The size in the N dimension is 2, the size in the H dimension is 128, and the size in the W dimension is 448. Also, the element spacing in the W dimension of the tensor X is 1, the element spacing in the H dimension is 448, and the element spacing in the N dimension is 57344. The starting address of the tensor X is address X 。
[0024] Generally, the stride information is arranged in descending order. For example, the shape and stride of an output tensor can be expressed as (for convenience of representation, it has been pre-sorted by stride in descending order, and it also supports tensors passed in by the user without sorting): (shape):(stride) = (sh1, sh2, sh3, sh4):(st1, st2, st3, st4) = (2, 2, 128, 64):(16384, 8192, 64, 1); In the above case, the traversal order of the output tensor is st4 => st3 => st2 => st1. Combining with the instruction granularity information of the MMA instruction supported by the chip, it traverses from the coordinate (0, 0, 0, 0) to (sh1, sh2, sh3, sh4).
[0025] 2. The operation granularity information of MMA refers to the data size that an MMA instruction can process, such as 64×32×32, 64×128×32, etc. Generally, the operation granularity of MMA supported by different chips is different.
[0026] Assume that D = A × B, where D is the output tensor, A is the left tensor, and B is the right tensor. If the operation granularity information of 64×128×32 is adopted, then 64×128 data in the left tensor A can be processed at one time, 128×32 data in the right tensor B can be processed, and 64×32 data in the output tensor D can be obtained.
[0027] 3. Intermediate Representation: During the operator development process, to improve development efficiency and reduce development costs, developers usually do not directly write low-level assembly code but use higher-level abstractions such as intermediate representation to construct their operators. Intermediate representation is an abstract code representation form between source code and target machine code, allowing code generation tools to analyze, optimize, and transform the instructions it contains without having to concern themselves with specific hardware details. Developers can achieve complex computational logic by combining different intermediate representations.
[0028] In the prior art, the interfaces of the intermediate representation of MMA have some restrictions on input tensors and output tensors. For example, it is required that the shapes of both the input tensor and the output tensor are two-dimensional. Thus, if developers need to perform matrix multiplication operations on higher-dimensional tensors, they need to develop loop code for traversal by themselves, and the developed code is prone to errors.
[0029] Therefore, the embodiments of this application provide a new scheme for the intermediate representation of MMA on tensors. The output tensor is partitioned, and matrix multiplication operations are performed on the corresponding data blocks in each input tensor for each region in the output tensor. Regardless of the dimensions of each tensor, according to the target information required for the intermediate representation of MMA on tensors, the intermediate representation of MMA can be automatically generated. Developers do not need to develop loop code for traversal by themselves, nor do they need to understand the memory layout of tensors. Therefore, the development difficulty of the intermediate representation of MMA can also be reduced.
[0030] To introduce the method of the embodiments of this application more clearly, the application scenarios of the embodiments of this application will be introduced first below.
[0031] See Figure 1 , Figure 1 which is an application scenario provided by the embodiments of this application, including an electronic device 130 and a chip 200. Among them, the electronic device 130 such as a desktop computer, a laptop computer, a server, etc., can install an assembly generation program or a compiler, and provide the interface for the intermediate representation of MMA on tensors in the embodiments of this application through the assembly generation program or the compiler. Developers pass in the target information required for the intermediate representation of MMA on tensors through the interface, such as the tensor information of each tensor, the operation granularity information of MMA supported by the chip, etc. The assembly generation program or the compiler can then automatically generate the intermediate representation of MMA on tensors according to the target information. Developers do not need to concern themselves with the dimensions of tensors, nor do they need to understand the memory layout of tensors, and the development difficulty of the intermediate representation of MMA is relatively low.
[0032] Subsequently, developers or other technical personnel can import the intermediate representation of MMA into the chip 200, parse and execute the intermediate representation in the chip 200, and then perform matrix multiplication operations on each tensor in the chip 200.
[0033] First, the principle of the intermediate representation of MMA in the embodiments of this application will be introduced below.
[0034] Assume that we want to calculate D = A × B + C, where A is the left tensor, B is the right tensor, C is the bias tensor (C is optional, i.e., it may or may not exist), A, B, and C are all input tensors, and D is the output tensor.
[0035] Refer to Figure 2 , which is a schematic diagram of the intermediate representation of MMA provided by the embodiments of this application. When implementing the intermediate representation of MMA, it is necessary to determine the target information, which at least includes the tensor information of each tensor and the operation granularity information of MMA supported by the chip, and may also include input synchronization information and output synchronization information. Among them, the input synchronization information can be used to configure waiting for at least one of the tensors A, B, and C to be data-ready, and can also be used to configure waiting for the memory unlock corresponding to D to be available; the output synchronization information is used to configure sending the information that D is data-ready, and is also used to configure sending the information that the memory unlock corresponding to at least one of the tensors A, B, and C is available.
[0036] For the input synchronization information, a first synchronization instruction such as at least one wait instruction can be generated. The first synchronization instruction is used to indicate waiting for at least one of the tensors A, B, and C to be data-ready, and can also be used to indicate waiting for the memory unlock corresponding to D to be available; Moreover, according to the tensor information of each tensor and the operation granularity information corresponding to the chip, D can be divided into regions 1, 2... n. According to the region information of region i (such as size, position, etc.) and the tensor information of each input tensor, the data blocks corresponding to region i in A, B, and C are determined. Based on the data blocks corresponding to region i in A, B, and C, code segment i is generated. The code segments 1, 2... n are concatenated in sequence and placed after the first synchronization instruction, where i ranges from 1 to n. Among them, code segment i may include address calculation instructions, MMA main instructions, and configuration instructions. There are generally multiple address calculation instructions, which are used to calculate the physical address of region i and the physical addresses of the data blocks corresponding to region i in A, B, and C; there are generally also multiple MMA main instructions, which are used to perform matrix multiplication and accumulation calculations on the data blocks corresponding to region i in A, B, and C; there are generally also multiple configuration instructions, which are used to update the data in the status register in the chip, and the data in the status register is used to describe various parameters required for executing the MMA main instructions, such as the data size processed each time, the memory layout, etc.
[0037] For the input synchronization information, a second synchronization instruction such as at least one arrive instruction can be generated. The second synchronization instruction is used to indicate waiting for D to be data-ready, and can also be used to indicate that the memory corresponding to at least one of the tensors A, B, and C is unlocked and available. The second synchronization instruction can be placed at the end.
[0038] It should be noted that when the number of elements in D is small, only one region may be divided. Moreover, when multiple regions are divided from D, the sizes of the regions can be the same or different. For the data blocks corresponding to region i in D in A, B, and C, the data blocks may be directly capable of matrix multiplication operations, or may need to be divided into blocks and then perform MMA block by block.
[0039] Next, the code generation method proposed in the embodiments of the present application will be described in conjunction with the flowchart.
[0040] See Figure 3 , Figure 3 which is a flowchart of a code generation method provided in the embodiments of the present application. This method is applied to Figure 1 electronic device 130, and this method includes the following steps.
[0041] In step 301, in response to an instruction for the intermediate representation of performing MMA on a tensor, target information is determined. The target information includes the operation granularity information of MMA supported by the chip, and the tensor information of the input tensor and the output tensor. The input tensor includes a left tensor and a right tensor.
[0042] In practical applications, developers can pass in the target information through an interface or input the target information through a configuration page.
[0043] Among them, the operation granularity information of MMA supported by the chip may be only one or there may be multiple. Assume that the operation granularity information of MMA supported by the chip is: 64×32×32, 64×128×32.
[0044] See Figure 4 , assume that the left tensor is A, the right tensor is B, and the output tensor is D. It is necessary to perform the intermediate representation of MMA on D = A×B, and the tensor information of the left tensor A is: Shape information: shape A = (shapeL, shapeM, shapeK) = (2, 128, 448), Stride information: stride A = (strideL, strideM, strideK) = (448×128 = 57344, 448, 1), Base address: addressA; The tensor information of the right tensor B is: Shape information: shape B = (shapeL, shapeK, shapeN) = (2, 448, 64), Stepping information: stride B = (strideL, strideK, strideN) = (512 × 64 = 32768, 1, 512), where strideN = 512 indicates that the data of the right tensor B is non - continuously stored on the N - axis. After storing 448 data of the K - axis, 64 gaps are left and then 448 data of the K - axis are stored continuously. Base address: addressB; The tensor information of the output tensor D is as follows: Shape information: shape D = (shapeL, shapeM, shapeN) = (2, 128, 64), Stepping information: stride D = (strideL, strideM, strideN) = (64 × 128 = 8192, 64, 1), Base address: addressD.
[0045] In some embodiments, cross - logical checks on the tensor information of each tensor can also be pre - performed, such as checking whether the matrix multiplication conditions are met, whether the actual tensor shape is consistent with the tensor information, etc., to prompt developers of possible errors.
[0046] In step 302, the output tensor is partitioned according to the tensor information of the input tensor and the output tensor, and the operation granularity information.
[0047] Specifically, according to the tensor information of the input tensor and the output tensor, and the operation granularity information, an operation granularity information can be selected, and then the output tensor is partitioned according to this operation granularity information.
[0048] Continuing with the above example, the operation granularity information supported by the chip for MMA is: 64 × 32 × 32, 64 × 128 × 32. Both of these operation granularity information use the 64 × 32 size to partition the output tensor D, and the shape and stride of the output tensor D also support this size. The left tensor A and the right tensor B both have 448 elements in the K dimension. To reduce the number of blocks, the larger - granularity 64 × 128 × 32 can be selected. Then, based on the selected 64 × 128 × 32, the output tensor D is partitioned using 64 × 32. The separated regions can be seen in Figure 4 , there are a total of 8 regions, and the gray area of the output tensor D represents the first region.
[0049] It should be noted that the number of partitions of the output tensor is positively correlated with the number of instructions included in the finally generated intermediate representation, while the number of instructions included in the intermediate representation is negatively correlated with the speed at which the chip parses and executes the intermediate representation. To improve the chip operation speed, it is necessary to reduce the number of partitions of the output tensor. Therefore, when there are multiple operation granularity information options available, the maximum operation granularity can be selected. In this way, each region can be maximized, thereby reducing the number of partitions.
[0050] In step 303, based on the left data block corresponding to each region in the output tensor in the left tensor and the right data block corresponding to each region in the right tensor, a code segment corresponding to this region is generated. The code segment is used to write the matrix multiplication result of the left data block and the right data block into this region. Among them, the left data block is determined according to the position information of this region and the tensor information of the left tensor, and the right data block is determined according to the position information of this region and the tensor information of the right tensor.
[0051] Take Figure 4 the first region of the output tensor D in [Example] as an example. According to the position information of the first region in the output tensor D and the tensor information of the left tensor A, it can be determined that the left data block corresponding to the first region in the left tensor A is the first 64×448 elements ( Figure 4 the gray region of the left tensor A in [Example]), and according to the position information of the first region in the output tensor D and the tensor information of the right tensor B, it can be determined that the right data block corresponding to the first region in the right tensor B is the first 448×32 elements ( Figure 4 the gray region of the right tensor B in [Example]).
[0052] In some embodiments, the left data block corresponding to each region in the left tensor and the right data block corresponding to each region in the right tensor just match the selected operation granularity information, and the matrix multiplication operation can be directly performed on the left data block and the right data block.
[0053] Take the left data block being 64×128 and the right data block being 128×32 as an example. When the operation granularity information is 64×128×32, the left data block and the right data block just match the selected operation granularity information, and a set of MMA instructions can be directly generated based on the left data block and the right data block. This set of MMA instructions is used to perform the matrix multiplication operation on the left data block and the right data block, and this set of MMA instructions is the code segment corresponding to this region.
[0054] In some embodiments, the left data block corresponding to each region in the left tensor and the right data block corresponding to each region in the right tensor do not match the selected operation granularity information, and need to be partitioned before MMA can be performed. In this case, the code segment corresponding to this region can be generated according to the following steps.
[0055] Step 1: Partition the left data block and the right data block corresponding to each region respectively according to the operation granularity information.
[0056] Take Figure 4 Figure 4 As an example, for the left data block and the right data block corresponding to the first region of the output tensor D. When the operation granularity information is 64×128×32, the left data block can be partitioned into 4 sub-blocks according to 64×128: sub-block 1, sub-block 2, sub-block 3, and sub-block 4, and the right data block can be partitioned into 4 sub-blocks according to 128×32: sub-block 1′, sub-block 2′, sub-block 3′, and sub-block 4′.
[0057] It should be noted that for the left data block, when partitioning according to 64×128 until the end, only 64×64 data remains. Therefore, the size of the last sub-block is 64×64. For the right data block, when partitioning according to 128×32 until the end, only 64×32 data remains. Therefore, the size of the last sub-block is 64×32.
[0058] Step 2: Based on each group of sub-blocks with the same partitioning order in the left data block and the right data block corresponding to each region, generate a group of MMA instructions for this group of sub-blocks. This group of MMA instructions is used to perform matrix multiplication operations on this group of sub-blocks. Among them, when this group of sub-blocks is not the first group of sub-blocks corresponding to this region, this group of MMA instructions is also used to accumulate the matrix multiplication operation results of this group of sub-blocks to the matrix multiplication operation results of the previous group of sub-blocks corresponding to this region.
[0059] Still take Figure 4 Figure 4 As an example, for the left data block and the right data block corresponding to the first region of the output tensor D, there are 4 groups of sub-blocks (shown by the dotted lines) in the left data block and the right data block: the first group of sub-blocks: sub-block 1 and sub-block 1′, the second group of sub-blocks: sub-block 2 and sub-block 2′, the third group of sub-blocks: sub-block 3 and sub-block 3′, the fourth group of sub-blocks: sub-block 4 and sub-block 4′.
[0060] Performing MMA on these 4 groups of sub-blocks in order can obtain the matrix multiplication operation results of the left data block and the right data block. Specifically, for the first group of sub-blocks, the first group of MMA instructions can be generated, and the first group of MMA instructions is used to perform matrix multiplication operations on the first group of sub-blocks. For the second group of sub-blocks, the second group of MMA instructions is generated, and the second group of MMA instructions is used to perform matrix multiplication operations on the second group of sub-blocks and accumulate the matrix multiplication operations to the matrix multiplication operation results of the first group of sub-blocks. For the third group of sub-blocks, the third group of MMA instructions is generated, and the third group of MMA instructions is used to perform matrix multiplication operations on the third group of sub-blocks and accumulate the matrix multiplication operations to the matrix multiplication operation results of the second group of sub-blocks. For the fourth group of sub-blocks, the fourth group of MMA instructions is generated, and the fourth group of MMA instructions is used to perform matrix multiplication operations on the fourth group of sub-blocks and accumulate the matrix multiplication operations to the matrix multiplication operation results of the third group of sub-blocks.
[0061] The process of generating a set of MMA instructions is introduced below.
[0062] 1. Generate address calculation instructions.
[0063] During actual calculation, the chip needs to know the address of each region and the addresses of each sub-block in each group of sub-blocks corresponding to the region. Therefore, according to the address characterization information of each region, the address calculation instructions for this region can be generated, and according to the address characterization information of each sub-block in each group of sub-blocks, the address calculation instructions for this sub-block can be generated.
[0064] First, the method of generating the address calculation instructions for each region is introduced below.
[0065] Method 1: The address characterization information of each region includes the position information of this region in the output tensor, the stride information of the output tensor, and address information such as the base address.
[0066] When the output tensor is stored in the global shared memory (GSM), the address calculation instruction for each region is: Address of the region = Position information of the region in the output tensor × Stride information of the output tensor + Base address of the output tensor.
[0067] When the output tensor is stored in the thread local register (TLR) memory, the calculation is relatively complex. Usually, by calling a library function function(), inputting the position information of each region in the output tensor, and then combining the stride information of the output tensor and the thread arrangement information, calculate the offset of this region relative to the base address of the output tensor, and then add this offset to the base address of the output tensor to obtain the address of this region.
[0068] That is, when the output tensor is stored in the thread local register memory, the address calculation instruction for each region is: Address of the region = function(Position information of the region in the output tensor, Stride information of the output tensor, Thread arrangement information) + Base address of the output tensor.
[0069] Method 2: The address characterization information of the first region includes the position information of this region in the output tensor, the stride information of the output tensor, and address information. The address characterization information of each region except the first region includes the offset address of this region compared to the previous region in the output tensor.
[0070] The address calculation instruction for the first region is the same as Method 1 and will not be elaborated here.
[0071] For each region except the first region, regardless of whether the output tensor is stored in the shared memory or the thread local register memory, the address calculation instruction is: Address of the region = Address of the previous region + Offset address.
[0072] Since the calculation method of the offset address is relatively simple, Method 2 can reduce the execution complexity of the address calculation instruction, thereby improving the computing speed of the chip.
[0073] In practical applications, the stride information of the output tensor can characterize the memory layout of the output tensor. Regardless of the memory layout of the output tensor, the address of the region can be calculated according to the stride information of the output tensor. In this way, the representation of the data storage layout of the output tensor is more flexible.
[0074] Next, the method of generating the address calculation instruction for each sub-block in each group of sub-blocks will be introduced.
[0075] Method 1: The address representation information of each sub-block in each group of sub-blocks includes the position information of the sub-block in the corresponding tensor, the stride information of the corresponding tensor, and the address information such as the base address.
[0076] For each sub-block in each group of sub-blocks, when the tensor to which the sub-block belongs is stored in the shared memory, the address calculation instruction of the sub-block is: the address of the sub-block = the position information of the sub-block in the corresponding tensor × the stride information of the corresponding tensor + the base address of the corresponding tensor.
[0077] Taking the calculation of the address of sub-block 1' in the right tensor B as an example, The address of sub-block 1' = the coordinates of sub-block 1' in the right tensor B × stride B T + addressB = (0, 0, 0) × (32768, 1, 512) T + addressB = 0 × 32768 + 0 × 1 + 0 × 512 + addressB.
[0078] For each sub-block in each group of sub-blocks, when the tensor to which the sub-block belongs is stored in the thread-local register memory, the calculation is relatively complex. Usually, by calling a library function, inputting the position information of the sub-block in the corresponding tensor, and then combining the stride information of the corresponding tensor and the thread arrangement information, the offset of the sub-block relative to the base address of the corresponding tensor is calculated, and then this offset is added to the base address of the corresponding tensor to obtain the address of the sub-block.
[0079] That is, when the tensor to which the sub-block belongs is stored in the thread-local register memory, the address calculation instruction of the sub-block is: the address of the sub-block = function (the position information of the sub-block in the corresponding tensor, the stride information of the corresponding tensor, the thread arrangement information) + the base address of the corresponding tensor.
[0080] Taking the calculation of the address of sub-block 1 in the left tensor A as an example, Address of Sub - block 1 = function(Coordinates of Sub - block 1 in left tensor A, stride A , thread arrangement information)+addressA.
[0081] Method 2: The address representation information of each sub - block in the first group of sub - blocks corresponding to the first region includes the position information of the sub - block in the corresponding tensor, the stride information of the corresponding tensor, and the address information. For each group of sub - blocks except the first group of sub - blocks corresponding to the first region (such as the non - first group of sub - blocks corresponding to the first region, each group of sub - blocks corresponding to non - first regions), the address representation information of each sub - block in this group of sub - blocks includes the offset address of the sub - block compared to the previous sub - block in the corresponding tensor.
[0082] The address calculation instruction for each sub - block in the first group of sub - blocks corresponding to the first region is the same as Method 1, which will not be elaborated here.
[0083] For each group of sub - blocks except the first group of sub - blocks corresponding to the first region, regardless of whether the tensor to which this group of sub - blocks belongs is stored in shared memory or thread - local register memory, the address calculation instruction for each sub - block in this group of sub - blocks is: Address of sub - block = Address of previous sub - block+Offset address.
[0084] Since the calculation method of the offset address is relatively simple, Method 2 can reduce the execution complexity of the address calculation instruction and improve the computing speed of the chip.
[0085] In practical applications, for any input tensor in the left tensor and the right tensor, the stride information of the input tensor can represent the memory layout of the input tensor. Regardless of the memory layout of the input tensor, the address of the corresponding sub - block can be calculated according to the stride information of the input tensor. In this way, the representation of the data storage layout of the input tensor is more flexible.
[0086] 2. Generate MMA main instructions.
[0087] For example, generate MMA main instructions based on the sub - block addresses obtained from the address calculation instructions of each sub - block in each group of sub - blocks and the region addresses obtained from the address calculation instructions of the corresponding regions.
[0088] 3. Generate multiple configuration instructions for MMA main instructions.
[0089] To ensure the normal operation of the chip, multiple configuration instructions corresponding to the MMA main instructions of each group of sub - blocks can also be generated according to the configuration information required for matrix multiplication operations on each group of sub - blocks in the chip. These multiple configuration instructions are used to update the data in the status register in the chip. The data in the status register is used to describe various parameters required for executing the MMA main instruction, such as the size of the sub - block (i.e., the data size processed each time), the offset address, the memory layout, etc.
[0090] Considering that information such as the size and offset address of each sub-block needs to be configured for each group of sub-blocks, when the chip calculates each group of sub-blocks, relevant data is copied into the status register according to these configurations, and then calculations are performed based on the data in the status register. If the data in the status register has not changed, there is no need to repeatedly copy the relevant data into the register, and the data in the status register can be continued to be used.
[0091] Therefore, for non-first groups of sub-blocks, if the configuration information corresponding to this group of sub-blocks duplicates the configuration information corresponding to the previous group of sub-blocks (corresponding to the same area or different areas as this group of sub-blocks), when generating multiple configuration instructions corresponding to this group of sub-blocks, the configuration instructions corresponding to the duplicate configuration information can also be skipped. Among them, the duplicate configuration information includes the size and offset address of the configured sub-blocks, etc.
[0092] In this way, by de-duplicating the configuration instructions corresponding to the current group of sub-blocks and the configuration instructions corresponding to the previous group of sub-blocks, the total number of instructions can be reduced, the hardware operations that the chip needs to execute can be reduced, and thus the operation speed of the chip can be improved.
[0093] Step 3: Fuse the MMA instructions corresponding to each group of sub-blocks in each area to obtain the code segment corresponding to this area.
[0094] For example, according to the block division order of each group of sub-blocks, the MMA instructions corresponding to each group of sub-blocks are spliced to obtain the code segment corresponding to this area.
[0095] In step 304, the code segments corresponding to each area are fused to obtain the intermediate representation for performing MMA on the tensor.
[0096] For example, according to the division order of each area, the code segments corresponding to each area are spliced to obtain the intermediate representation for performing MMA on the tensor, and this intermediate representation is used to write the matrix multiplication operation result of the left tensor and the right tensor into the output tensor.
[0097] In some embodiments, the input tensor may further include a bias tensor. After generating the code segment corresponding to each area, target code can also be added to this code segment. The target code is used to accumulate the bias data block corresponding to this area in the bias tensor to the matrix multiplication operation result of the left data block and the right data block corresponding to this area, where the bias data block is determined according to the position information of this area and the tensor information of the bias tensor.
[0098] In some embodiments, developers are also supported in configuring input synchronization information (generally information that the current intermediate representation needs to know), such as configuring waiting for at least one of the left tensor, right tensor, and bias tensor to have its data ready, and configuring waiting for the memory corresponding to the output tensor to be unlocked and available. At this time, the target information will also include the input synchronization information. Then, a first synchronization instruction, such as at least one wait instruction, can be inserted at the beginning of the first code segment, or, after each address calculation instruction in the first code segment, a first synchronization instruction can be inserted. The first synchronization instruction is used to indicate waiting for at least one of the left tensor, right tensor, and bias tensor to have its data ready, and / or is used to indicate waiting for the memory corresponding to the output tensor to be unlocked and available.
[0099] In practical applications, when the chip waits for at least one of the left tensor, right tensor, and bias tensor to have its data ready, it is actually executing a copy instruction (i.e., copying a tensor to the corresponding address). The copy instruction and each address calculation instruction can be parallel in the chip. However, when the first synchronization instruction is placed before each address calculation instruction in the first code segment (i.e., at the very beginning), since the first synchronization instruction itself blocks the execution of subsequent instructions, the chip needs to wait for the corresponding tensor to have its data ready before executing each address calculation instruction during actual instruction execution. In fact, there is no dependency between these two types of instructions. Therefore, here the first synchronization instruction can no longer be placed at the very beginning of the first code segment, but is appropriately postponed to after the address calculation instructions that have no dependency on it. In this way, the number of instructions blocked by the first synchronization instruction can be reduced, and the first synchronization instruction can be executed in parallel with each address calculation instruction in the chip, which can further improve the chip calculation speed.
[0100] In some embodiments, developers are also supported in configuring output synchronization information (generally information that the current intermediate representation needs to tell other intermediate representations), such as configuring waiting for the memory occupied by at least one of the left tensor, right tensor, and bias tensor to be unlocked and available, and configuring waiting for the data of the output tensor to be ready. At this time, the target information also includes the output synchronization information. Then, a second synchronization instruction, such as an arrive instruction, can be inserted at a non-tail position in the last code segment. The second synchronization instruction is used to send that the memory occupied by at least one of the left tensor, right tensor, and bias tensor is unlocked and available, and is also used to send that the data of the output tensor is ready. Alternatively, synchronization code can be configured in the last code segment. The synchronization code is used to send that the memory occupied by at least one of the left tensor, right tensor, and bias tensor is unlocked and available, and is also used to send that the data of the output tensor is ready.
[0101] That is to say, for the second synchronization instruction placed in the middle to represent the last one, it can also not be placed at the very end, but can be appropriately advanced to before the instructions that have no dependence on it, so as to send out the corresponding completion information in advance and enable the subsequent calculations in the chip to start earlier. Moreover, for some hardware architectures that support configuring the arrive instruction on the MMA main instruction or its configuration instruction, it can also be directly configured on the last MMA main instruction or its configuration instruction to reduce the number of generated instructions and further improve the chip calculation speed.
[0102] See Figure 5 , Figure 5 is a flowchart of another code generation method provided by an embodiment of the present application. This method is applied to Figure 1 electronic device 130, and this method includes the following steps.
[0103] In step 501, in response to an instruction for the intermediate representation of adding MMA to a tensor, determine the operation granularity information of the MMA supported by the chip, and the tensor information of the input tensor and the output tensor. The input tensor includes a left tensor, a right tensor, and a bias tensor.
[0104] For example, for the intermediate representation of performing MMA on D = A × B + C, where D is the input tensor, A is the left tensor, B is the right tensor, and C is the bias tensor.
[0105] In step 502, divide the output tensor according to the tensor information of the input tensor and the output tensor, and the operation granularity information.
[0106] In step 503, for region i in the output tensor, determine the corresponding left data block i in the left tensor, determine the corresponding right data block i in the right tensor, and determine the corresponding bias data block i in the bias tensor.
[0107] Initially, region i represents the first region in the output tensor, and the value of i is a preset value such as 0, 1, etc.
[0108] In step 504, according to the operation granularity information corresponding to region i, block the left data block i and the right data block i respectively, and use two sub-blocks with the same block order in the left data block i and the right data block i as a group of sub-blocks.
[0109] In step 505, generate the j-th group of MMA instructions for the j-th group of sub-blocks in the left data block i and the right data block i.
[0110] Initially, the j-th group of sub-blocks is the first group of sub-blocks in the left data block i and the right data block i, and j is a set value such as 0, 1. Moreover, the j-th group of MMA instructions includes an address calculation instruction, an MMA main instruction, and multiple configuration instructions.
[0111] In specific implementation, when the j-th group of sub-blocks is the first group of sub-blocks corresponding to region i, the j-th group of MMA instructions is used to perform matrix multiplication operations on the j-th group of sub-blocks in the left data block i and the right data block i. When the j-th group of sub-blocks is not the first group of sub-blocks corresponding to region i, the j-th group of MMA instructions is used to perform matrix multiplication operations on the j-th group of sub-blocks in the left data block i and the right data block i, and the result of the matrix multiplication operation is accumulated onto the result of the matrix multiplication operation of the (j - 1)-th group of sub-blocks in the left data block i and the right data block i.
[0112] In practical applications, in order to improve the computing speed of subsequent chips, the j-th group of MMA instructions can also be optimized.
[0113] For example, if there is also input synchronization information determined in response to an instruction for adding MMA to the intermediate representation of a tensor, when the j-th group of sub-blocks is the first group of sub-blocks corresponding to the first region in the output tensor, a first synchronization instruction such as at least one wait instruction can also be inserted after each address calculation instruction. This group of synchronization instructions is used to indicate waiting for the data of at least one of the left tensor, the right tensor, and the bias tensor to be ready, and / or is used to indicate waiting for the memory corresponding to the output tensor to be unlocked and available.
[0114] For another example, when the j-th group of sub-blocks is not the first group of sub-blocks corresponding to the first region, when the configuration information corresponding to the j-th group of sub-blocks is repeated with the configuration information corresponding to the previous group of sub-blocks (which may correspond to the same region as the j-th group of sub-blocks or may correspond to a different region from the j-th group of sub-blocks), the configuration instructions corresponding to the repeated configuration information are skipped when generating multiple configuration instructions in the j-th group of sub-blocks. That is, duplicate removal processing can be performed on the configuration instructions of the j-th group of sub-blocks and the previous group of sub-blocks.
[0115] In step 506, it is determined whether the j-th group of sub-blocks is the last group of sub-blocks corresponding to region i. If not, step 507 is entered; if so, step 508 is entered.
[0116] In step 507, j is updated to j + 1, and the process returns to step 505.
[0117] In step 508, the MMA instructions corresponding to the left data block i and the right data block i are fused to obtain a code segment corresponding to region i. The code segment corresponding to region i is used to write the matrix multiplication operation results of the left data block i and the right data block i into region i.
[0118] In specific implementation, if there is also input synchronization information determined in response to an instruction for adding MMA to the intermediate representation of a tensor, when region i is the last region, synchronization code can also be configured at the end of the code segment corresponding to region i. The synchronization code is used to indicate that the memory occupied by at least one of the left tensor, the right tensor, and the bias tensor has been unlocked and is available, and is also used to indicate that the data of the output tensor is ready.
[0119] In step 509, target code is added to the code segment corresponding to region i, and the target code is used to accumulate the bias data block i onto the matrix multiplication result of the left data block i and the right data block i.
[0120] In step 510, it is determined whether region i is the last region. If not, step 511 is entered; if so, step 512 is entered.
[0121] In step 511, i is updated to i + 1, and step 503 is returned.
[0122] In step 512, the code segments corresponding to each region are fused to obtain an intermediate representation for performing MMA on the tensor.
[0123] The solution of the embodiment of the present application will be introduced below in conjunction with specific examples.
[0124] See Figure 6 , assuming that MMA calculations are to be performed on the left tensor A, the right tensor B, and the bias tensor C to obtain the output tensor D, i.e., D = A × B + C. And assume that the tensor information of each tensor is as follows.
[0125] The tensor information of the left tensor A is: Shape information: shape A = (shapeL, shapeM, shapeK) = (2, 128, 448), Stride information: stride A = (strideL, strideM, strideK) = (448 × 128 = 57344, 448, 1), Base address: addressA; The tensor information of the right tensor B is: Shape information: shape B = (shapeL, shapeK, shapeN) = (2, 448, 64), Stride information: stride B = (strideL, strideK, strideN) = (512 × 64 = 32768, 1, 512), Base address: addressB; The tensor information of the bias tensor C is: Shape information: shape C = (shapeL, shapeM, shapeN) = (2, 128, 64), Stride information: stride C=(strideL, strideM, strideN)=(64×128 = 8192, 64, 1), Base address: addressC; The tensor information of the output tensor D is as follows: Shape information: shape D = (shapeL, shapeM, shapeN)=(2, 128, 64), Stride information: Stride D =(strideL, strideM, strideN)=(64×128 = 8192, 64, 1), Base address: addressD.
[0126] After obtaining the tensor information of each tensor passed in by the user, start tensor traversal and instruction generation, and the steps are as follows.
[0127] Step 1. Refer to the shape information and stride information of each tensor, as well as the operation granularity information supported by the chip, and traverse the output tensor D (the first traversal of the outer layer, that is, traverse to the first area of the output tensor D), starting from the coordinates (0, 0, 0) of the output tensor D.
[0128] Step 2. For any one of the left tensor A and the right tensor B, refer to the shape information and stride information of the tensor, and traverse the tensor B (the first traversal of the outer layer, the first traversal of the inner layer). Logically, the coordinates (0, 0, 0) in the output tensor D correspond to the coordinates (0, 0, X) in the left tensor A and the coordinates (0, X, 0) in the right tensor B, where X represents the number of elements in the K dimension to be accumulated. At this time, it is the first traversal of the inner layer. Therefore, the actual coordinates of the left tensor A used are (0, 0, 0), and the coordinates of the right tensor B are (0, 0, 0).
[0129] Step 3. Assume that according to the current hardware limitations, such as the operation granularity information of MMA supported by the chip, determine that the granularity of a single traversal of the output tensor D is M×N = 64×32. Combining the current hardware limitations, as well as the number of elements of the left tensor A and the right tensor B on the K axis, determine that the granularity of a single traversal of the left tensor A is M×K = 64×128, and the granularity of a single traversal of the right tensor B is K×N = 128×32. Therefore, the granularity of a single MMA instruction is M×N×K = 64×32×128.
[0130] Step 4. Generate a single MMA and its related instructions.
[0131] Step 4.1 First, calculate the addresses of the small blocks sliced from the left tensor A, the right tensor B, and the output tensor D.
[0132] For the GSM memory, taking the right tensor B as an example, the address calculation formula for the divided small blocks is as follows: Address of the sub-block in the right tensor B = Coordinates of the sub-block in the right tensor B × stride B T + Base address of the right tensor B.
[0133] Taking the first sub-block in the right tensor B as an example, Address of the first sub-block = (0, 0, 0) × (32768, 1, 512) T + addressB; = 0 × 32768 + 0 × 1 + 0 × 512 + addressB.
[0134] For the TLR memory, its calculation is relatively complex. By calling an assembly generation program or a library function function() in the compiler, inputting the coordinates of the currently traversed sub-block, and combining the stride information of the tensor to which the sub-block belongs and the thread arrangement information, calculate the offset of the sub-block relative to the base address of the corresponding tensor, and then add this offset and the base address to obtain the address of the small block.
[0135] Taking the left tensor A as an example, the address of the divided small blocks is: Address of the sub-block in the left tensor A = function(Coordinates of the sub-block in the left tensor A, stride A , Thread arrangement information) + Base address of the left tensor A.
[0136] Step 4.2 To reduce the complexity of address calculation, for sub-blocks that are not the first sub-block, the offset address relative to the previous sub-block can be used for address calculation.
[0137] Taking the right tensor B using the offset address as an example, after calculating the address and before generating the specific MMA main instruction and configuration instruction, attempt the next traversal. At this time, the traversal reaches the first time in the outer layer and the second time in the inner layer (i.e., reaching the sub-block 2' of the right tensor). Combining the coordinates (0, 0, 0) of the previous traversal of the right tensor B and the granularity K = 128, the coordinates of the right tensor B reached at this time are (0, 128, 0). Therefore, using the previous calculation formula, the address of the second sub-block in the right tensor B in the next MMA calculation = Coordinates of the sub-block in the right tensor B × stride B T = (0, 128, 0) × (32768, 1, 512) T = 0 × 32768 + 128 × 1 + 0 × 512 + addressB, compared with the current traversal, the memory offset is 128.
[0138] The left tensor A is similar, so it will not be elaborated here.
[0139] Step 4.3 Assume that the user configures input synchronization information, such as waiting for the data of the left tensor A to be ready. Then, a waitA instruction needs to be inserted into the intermediate representation. Since the current traversal corresponds to the first MMA instruction, the waitA instruction needs to be inserted into its instruction block. Also, because only the MMA instruction body depends on the matrix A being ready, and the previous address calculation does not depend on the matrix A being ready, the waitA instruction will be inserted after the instruction corresponding to the address calculation. Subsequently, when the chip executes, the processes of "waiting" and "address calculation" are actually processed in parallel, which can improve the calculation speed and optimize the chip performance.
[0140] Step 4.4 Combine the granularity information in Step 3, the address information calculated in Step 4.1, the bias information in Step 4.2, the input synchronization information in Step 4.3, and some other information (such as the data type of the tensor, etc.) to generate a set of instructions including address calculation instructions, synchronization instructions, MMA main instructions, and configuration instructions.
[0141] Step 5. After completing the first traversal of the outer layer and the first traversal of the inner layer (i.e., for the first region in the output tensor D, traversing sub-block 1 in the left tensor A and sub-block 1' in the right tensor B), considering the sizes of each tensor and the traversal granularity, it is found that for the first region in the output tensor D, only 128 elements of the K dimension have been traversed and the calculation is not yet complete. Therefore, continue the traversal. Similar to Step 3, in the current traversal (the first time in the outer layer and the second time in the inner layer, i.e., for the first region in the output tensor D, traversing sub-block 2 in the left tensor A and sub-block 2' in the right tensor B), the execution granularity of the MMA instruction is still M×N×K = 64×32×128. At this time, it is still the first traversal of the outer layer, the coordinates of the first region are still (0, 0, 0), and the inner layer is the second traversal. The coordinates of sub-block 2 in the left tensor A are (0, 0, 128), and the coordinates of sub-block 2' in the right tensor B are (0, 128, 0). Then, similar to the above Step 4, a set of instructions including address calculation instructions, MMA instructions, and configuration instructions is generated (since this traversal does not correspond to the first or last MMA instruction, the first synchronization instruction or the second synchronization instruction does not need to be inserted), and this set of instructions is concatenated behind the set of instructions generated in Step 4.
[0142] Step 5.1 Before generating the specific MMA main instruction and configuration instruction, duplicate elimination can be performed on the configuration instruction to be generated and the previously generated configuration instruction to reduce the number of instructions executed by the chip and improve the calculation performance.
[0143] Step 6. Continue to determine whether the traversal of the K dimension is completed. Since it is found that the traversal is still not completed, repeat Step 5 to continue the traversal until the traversal of the K dimension is completed.
[0144] Step 6.1 It should be noted that when traversing the outermost layer for the first time and the innermost layer for the fourth time (i.e., for the first region in the output tensor D, after traversing sub-block 4 in the left tensor A and sub-block 4' in the right tensor B), there are still 64 numbers left in the K dimension. Therefore, the computational granularity of the last MMA instruction is M×N×K = 64×32×64, and the address calculation method is the same as before, which will not be elaborated here.
[0145] Step 6.2 The user also provides the bias tensor C. Therefore, during the fourth and final traversal of the inner layer, the address of the first bias data block 1'' in the bias tensor C can also be configured on the relevant instructions to add the bias tensor C while performing the MMA calculation. The coordinates of the first bias data block in the bias tensor C are the same as those of the first region in the output tensor D, and the address calculation method is also the same, which will not be elaborated here.
[0146] Step 7. After completing the first traversal of the outermost layer, perform the second traversal of the outermost layer in ascending order according to the step sequence of the D matrix. At this time, the coordinates of the second region in the output tensor D are (0, 0, 32). In this traversal, except for not inserting the first synchronization instruction, the other address calculations, MMA, and its configuration instruction generation logics are the same as those in Steps 3 - 6, which will not be elaborated here.
[0147] Step 8. Next, similar to Step 7, continue the outermost layer traversal. The coordinates (l, m, n) of the regions in the output tensor D are as follows: (0, 64, 0) => (0, 64, 32) => (1, 0, 0) => (1, 0, 32) => (1, 64, 0) => (0, 64, 32); In the inner layer traversal of each outermost layer traversal, the coordinates of the sub-block in the left tensor A are (l, m, k), and the coordinates of the sub-block in the right tensor B are (l, k, n). The l, m, n in these two coordinates are the same as the coordinates (l, m, n) of the region in the output tensor D, and k in these two coordinates starts from 0 and keeps increasing until the accumulation is completed.
[0148] Step 9. When the user configures the output synchronization information, after traversing the outermost layer for the last time and the innermost layer for the last time (i.e., after all calculations are completed), the second synchronization instruction can be inserted at the non-tail position of the last group of MMA instructions, or directly add synchronization code in the last group of MMA, without inserting additional synchronization instructions, to reduce the total number of instructions.
[0149] Based on the same inventive concept, an embodiment of the present application further provides a code generation device. The principle of the code generation device for solving problems is similar to the above-mentioned code generation method. Therefore, for the implementation of the code generation device, reference can be made to the implementation of the code generation method, and the repeated parts will not be elaborated.
[0150] Figure 7 FIG. 4 is a schematic structural diagram of a code generation device provided by an embodiment of the present application, including: A determination module 701, configured to determine target information in response to an instruction for performing matrix multiplication accumulation (MMA) on an intermediate representation of a tensor. The target information includes operation granularity information supported by the chip for MMA, and tensor information of an input tensor and an output tensor. The input tensor includes a left tensor and a right tensor, and the tensor information of each tensor includes at least shape information; A partitioning module 702, configured to partition the output tensor according to the tensor information of the input tensor and the output tensor, and the operation granularity information; A first generation module 703, configured to generate a code segment corresponding to the region based on a left data block corresponding to the region in the left tensor and a right data block corresponding to the region in the right tensor. The code segment is used to write the matrix multiplication operation result of the left data block and the right data block into the region, where each data block in the left data block and the right data block is determined according to the position information of the region and the tensor information of the tensor to which the data block belongs; A second generation module 704, configured to fuse the code segments corresponding to each region to obtain an intermediate representation of performing MMA on the tensor.
[0151] In some embodiments, the input tensor further includes a bias tensor, and the first generation module 703 is further configured to: After generating the code segment corresponding to the region, add target code to the code segment. The target code is used to accumulate a bias data block corresponding to the region in the bias tensor to the matrix multiplication operation result of the left data block and the right data block. The bias data block is determined according to the position information of the region and the tensor information of the bias tensor.
[0152] In some embodiments, the first generation module 703 is specifically configured to: Partition the left data block and the right data block respectively according to the operation granularity information; Based on each group of sub-blocks with the same partitioning order in the left data block and the right data block, generate a group of MMA instructions for performing matrix multiplication operations on the group of sub-blocks, where When the group of sub - blocks is not the first group of sub - blocks corresponding to the region, the group of MMA instructions is further used to accumulate the matrix multiplication operation result of the group of sub - blocks to the matrix multiplication operation result of the previous group of sub - blocks corresponding to the region; Fuse each group of MMA instructions to obtain the code segment.
[0153] In some embodiments, the first generation module 703 is specifically configured to: Generate an address calculation instruction for the sub - block according to the address characterization information of each sub - block in the group of sub - blocks; Generate an MMA main instruction according to the sub - block address obtained from the address calculation instruction of each sub - block and the region address obtained from the address calculation instruction of the region, where the address calculation instruction of the region is generated according to the address characterization information of the region; Generate multiple configuration instructions for the MMA main instruction according to the configuration information required for matrix multiplication operation on the group of sub - blocks in the chip.
[0154] In some embodiments, the tensor information of each tensor further includes stride information and address information; When the group of sub - blocks is the first group of sub - blocks corresponding to the first region, the address characterization information of each sub - block includes the position information of the sub - block in the corresponding tensor, the stride information of the corresponding tensor, and the address information, and the address characterization information of the region includes the position information of the region in the output tensor, the stride information of the output tensor, and the address information; When the group of sub - blocks is not the first group of sub - blocks corresponding to the first region, the address characterization information of each sub - block is the offset address of the sub - block compared to the previous sub - block in the corresponding tensor, and the address characterization information of the region is the offset address of the region compared to the previous region in the output tensor.
[0155] In some embodiments, when the group of sub - blocks is not the first group of sub - blocks corresponding to the first region, the first generation module 703 is further configured to: When there is a duplicate in the configuration information corresponding to the group of sub - blocks and the configuration information corresponding to the previous group of sub - blocks, skip generating the configuration instruction corresponding to the duplicate configuration information when generating the multiple configuration instructions.
[0156] In some embodiments, the input tensor further includes a bias tensor, the target information further includes input synchronization information, and there is also an optimization module 705, which is used to: Insert a first synchronization instruction after each address calculation instruction in the first code segment, where the first synchronization instruction is used to indicate waiting for the data of at least one of the left tensor, the right tensor, and the bias tensor to be ready, and / or, is used to indicate waiting for the memory corresponding to the output tensor to be unlocked and available.
[0157] In some embodiments, the input tensor further includes a bias tensor, the target information further includes output synchronization information, and an optimization module 705 is further included, configured to: Insert a non-tail of the last code segment into a second synchronization instruction, where the second synchronization instruction is used to indicate that the memory occupied by at least one of the left tensor, the right tensor, and the bias tensor is unlocked and available, and is further used to indicate that the data of the output tensor is ready; or, Configure synchronization code in the last code segment, where the synchronization code is used to indicate that the memory occupied by at least one of the left tensor, the right tensor, and the bias tensor is unlocked and available, and is further used to indicate that the data of the output tensor is ready.
[0158] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, each functional module in the embodiments of the present application may be integrated in a processor, or may exist physically separately, or two or more modules may be integrated in one module. The coupling between each module can be implemented through some interfaces, and these interfaces are usually electrical communication interfaces, but it does not exclude the possibility of being mechanical interfaces or other forms of interfaces. Therefore, the modules described as separate components may or may not be physically separated, and may be located in one place, or may be distributed to different positions of the same or different devices. The above integrated modules can be implemented in the form of hardware or in the form of software function modules.
[0159] After introducing the code generation method and apparatus of the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.
[0160] Next, with reference to Figure 8 an electronic device 130 implemented according to this embodiment of the present application is described. Figure 8 The shown electronic device 130 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0161] As Figure 8 shown, the electronic device 130 is presented in the form of a general-purpose electronic device. The components of the electronic device 130 may include, but are not limited to: the above at least one processor 131, the above at least one memory 132, and a bus 133 connecting different system components (including the memory 132 and the processor 131).
[0162] The bus 133 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.
[0163] The memory 132 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 1321 and / or cache memory 1322, and may further include read-only memory (ROM) 1323.
[0164] The memory 132 may also include a program / utilities 1325 having a set (at least one) of program modules 1324. Such program modules 1324 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0165] The electronic device 130 may also communicate with one or more external devices 134 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 130, and / or may communicate with any device that enables the electronic device 130 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 135. Moreover, the electronic device 130 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 136. As shown in the figure, the network adapter 136 communicates with other modules for the electronic device 130 through a bus 133. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0166] In an exemplary embodiment, the electronic device of the present application may at least include at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the steps of any code generation method provided by the embodiments of the present application.
[0167] In an exemplary embodiment, a storage medium is also provided. When the computer program in the storage medium is executed by the processor of the electronic device, the electronic device can execute any of the above code generation methods. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.
[0168] In an exemplary embodiment, there is also provided a computer program product which, when executed by a processor, implements any of the exemplary methods provided by the present application.
[0169] It should be noted that although several modules or sub - modules of the device are mentioned in the above - detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above - described modules can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0170] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the illustrated operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0171] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk memory, CD - ROM, optical memory, etc.) that contain computer - usable program code.
[0172] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0173] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also includes these changes and modifications.
Claims
1. A code generation method, characterized in that, Including: In response to an instruction for matrix multiplication accumulation (MMA) of a tensor in intermediate representation, determining target information, where the target information includes operation granularity information of MMA supported by the chip, and tensor information of an input tensor and an output tensor, the input tensor includes a left tensor and a right tensor, and the tensor information of each tensor includes at least shape information; Partitioning the output tensor according to the tensor information of the input tensor and the output tensor, and the operation granularity information; Based on a left data block corresponding to each region in the left tensor and a right data block corresponding to each region in the right tensor, generating a code segment corresponding to the region, the code segment being used to write the matrix multiplication operation result of the left data block and the right data block into the region, where each data block in the left data block and the right data block is determined according to the position information of the region and the tensor information of the tensor to which the data block belongs; Fusing the code segments corresponding to each region to obtain an intermediate representation of MMA for the tensor.
2. The method according to claim 1, wherein The input tensor further includes a bias tensor, and further includes: After generating the code segment corresponding to the region, adding target code to the code segment, the target code being used to accumulate the bias data block corresponding to the region in the bias tensor to the matrix multiplication operation result of the left data block and the right data block, and the bias data block is determined according to the position information of the region and the tensor information of the bias tensor.
3. The method according to claim 1, wherein Based on a left data block corresponding to each region in the left tensor and a right data block corresponding to each region in the right tensor, generating a code segment corresponding to the region includes: Chunking the left data block and the right data block respectively according to the operation granularity information; Based on each group of sub-chunks with the same chunking order in the left data block and the right data block, generating a group of MMA instructions, the group of MMA instructions being used to perform matrix multiplication operations on the group of sub-chunks, where when the group of sub-chunks is not the first group of sub-chunks corresponding to the region, the group of MMA instructions is further used to accumulate the matrix multiplication operation result of the group of sub-chunks to the matrix multiplication operation result of the previous group of sub-chunks corresponding to the region; Fusing each group of MMA instructions to obtain the code segment.
4. The method according to claim 3, characterized in that, Based on each group of sub-chunks with the same chunking order in the left data block and the right data block, generating a group of MMA instructions includes: Generating an address calculation instruction for the sub-chunk according to the address representation information of each sub-chunk in the group of sub-chunks; Generating an MMA main instruction according to the sub-chunk addresses obtained from the address calculation instructions of each sub-chunk and the region address obtained from the address calculation instruction of the region, and the address calculation instruction of the region is generated according to the address representation information of the region; Generating multiple configuration instructions for the MMA main instruction according to the configuration information required for performing matrix multiplication operations on the group of sub-chunks in the chip.
5. The method according to claim 4, characterized in that The tensor information of each tensor further includes step information and address information; When the group of sub - blocks is the first group of sub - blocks corresponding to the first region, the address representation information of each sub - block includes the position information of the sub - block in the corresponding tensor, the stride information of the corresponding tensor, and the address information, and the address representation information of the region includes the position information of the region in the output tensor, the stride information of the output tensor, and the address information; When the group of sub - blocks is not the first group of sub - blocks corresponding to the first region, the address representation information of each sub - block is the offset address of the sub - block compared to the previous sub - block in the corresponding tensor, and the address representation information of the region is the offset address of the region compared to the previous region in the output tensor.
6. The method according to claim 4, wherein When the group of sub - blocks is not the first group of sub - blocks corresponding to the first region, it further includes: When there is a duplication between the configuration information corresponding to the group of sub - blocks and the configuration information corresponding to the previous group of sub - blocks, skip generating the configuration instruction corresponding to the duplicated configuration information when generating the multiple configuration instructions.
7. The method according to any one of claims 1 to 6, characterized in that, The input tensor further includes a bias tensor, the target information further includes input synchronization information, and it further includes: After each address calculation instruction in the first code segment, insert a first synchronization instruction, which is used to indicate waiting for the data of at least one of the left tensor, the right tensor, and the bias tensor to be ready, and / or, is used to indicate waiting for the memory corresponding to the output tensor to be unlocked and available.
8. The method according to any one of claims 1 to 6, characterized in that, The input tensor further includes a bias tensor, the target information further includes output synchronization information, and it further includes: Insert a second synchronization instruction at the non - tail of the last code segment, which is used to indicate that the memory occupied by at least one of the left tensor, the right tensor, and the bias tensor is unlocked and available, and is also used to indicate that the data of the output tensor is ready; or, Configure synchronization code in the last code segment, which is used to indicate that the memory occupied by at least one of the left tensor, the right tensor, and the bias tensor is unlocked and available, and is also used to indicate that the data of the output tensor is ready.
9. A code generation device, characterized in that, It includes: A determination module, which is used to respond to an instruction for the intermediate representation of matrix multiplication accumulation (MMA) of a tensor, and determine target information, where the target information includes the operation granularity information of MMA supported by the chip, and the tensor information of the input tensor and the output tensor. The input tensor includes a left tensor and a right tensor, and the tensor information of each tensor includes at least shape information; A block - dividing module, which is used to partition the output tensor according to the tensor information of the input tensor and the output tensor, and the operation granularity information; A first generation module, which is used to generate a code segment corresponding to the region based on the left data block corresponding to the region in the left tensor and the right data block corresponding to the region in the right tensor. The code segment is used to write the matrix multiplication operation result of the left data block and the right data block into the region, where each data block in the left data block and the right data block is determined according to the position information of the region and the tensor information of the tensor to which the data block belongs; A second generation module, which is used to fuse the code segments corresponding to each region to obtain the intermediate representation of MMA for the tensor.
10. An electronic device, characterized in that, It includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein: the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, enables the at least one processor to execute the method according to any one of claims 1-8.
11. A storage medium, characterized in that, When the computer program in the storage medium is executed by a processor of an electronic device, the electronic device is capable of executing the method according to any one of claims 1-8.
12. A computer program product, characterized in that, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Code generation method and device, equipment and storage medium
CN112947908A
Varying operand accuracy
CN117270816A
Data processing method, processor, electronic equipment and storage medium
CN118520210A
Tensor processing method and device, electronic equipment, storage medium and program product
CN118674072A
In-memory computing method and in-memory computing system for general matrix multiplication
CN118860964A
Cited By
Code generation method, electronic equipment and storage medium
CN121070322A