Model Program Compilation Method, Electronic Device, Program Product and Medium
By using block-level instructions during the compilation of neural network models, recording the unit type, operation data location and operation result location of model units, the problems of large size, high complexity and low execution efficiency caused by conventional instruction set compilation are solved, and more efficient compilation and execution are achieved.
Patent Information
- Application Number
- CN202510265416.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-07
AI Technical Summary
In the prior art, neural network models use conventional instruction sets during compilation, resulting in large size, high complexity and low execution efficiency of executable programs.
A model program compilation method is proposed. By matching the block-level instructions corresponding to the program block in the source program in the preset instruction set, the block-level instructions include the unit type, operation data position and operation result position of the model unit, thereby reducing the number of instructions required for compilation and improving the computing efficiency of the processor.
This method can reduce the size and complexity of the executable program, while improving the computing efficiency, and avoiding the problem of low execution efficiency caused by conventional instruction set compilation.
Smart Images

Figure CN119781776B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model program compilation method, an electronic device, a program product, and a medium. Background Art
[0002] As the functions of neural network models are continuously improved, their running complexity is also increasing. Therefore, how to reduce the running complexity of neural network models is an important part of improving the running performance of such models.
[0003] In related technologies, neural network models are generally compiled using conventional instruction sets during the compilation process. However, conventional instruction sets are not set in combination with the characteristics of neural network models, which easily leads to a relatively large volume, high complexity, and low execution efficiency of the compiled executable program. Summary of the Invention
[0004] This application provides a model program compilation method, an electronic device, a program product, and a medium, so as to at least solve the technical problems in related technologies that using conventional instruction sets to compile neural network models easily leads to a relatively large volume and low execution efficiency of the executable program.
[0005] To solve the above technical problems, this application can provide a model program compilation method, including:
[0006] During the process of compiling a neural network model from a source program into an executable program, matching block-level instructions corresponding to program blocks in the source program in a preset instruction set; where the program blocks correspond to model units pre-divided in the neural network model; the block-level instructions are machine instructions corresponding to the model units, including the unit type, operation data position, and operation result position corresponding to the model unit, the operation data position is the memory position corresponding to the operation data input to the model unit, and the operation result position is the memory position corresponding to the operation result output by the model unit;
[0007] Allocating an operation data position and an operation result position for the program block and filling them into the block-level instructions;
[0008] Adding the filled block-level instructions to the executable program.
[0009] Optionally, the block-level instructions include a unit type, an operation data address, an operation result address, an operation data scale, and a bias flag, the operation data address is the memory address corresponding to the operation data, the operation result address is the memory address corresponding to the operation result, the operation data scale is the size of the operation data corresponding to the operation data, and the bias flag is the usage information corresponding to the bias data in the operation data.
[0010] Optionally, after adding the filled block-level instructions to the executable program, it further includes:
[0011] When executing a block-level instruction, read the cell type in the block-level instruction to determine the combined arithmetic operation corresponding to the model cell according to the cell type; wherein, the combined arithmetic operation includes the arithmetic operation corresponding to the model cell and the operation order;
[0012] Assign values to the arithmetic data register according to the arithmetic data address and the bias flag, assign values to the arithmetic result register according to the arithmetic result address, and assign values to the arithmetic data scale register according to the arithmetic data scale;
[0013] Extract arithmetic data from the memory according to the arithmetic data register and the arithmetic data scale register, perform a combined arithmetic operation on the arithmetic data to obtain an arithmetic result, and write the arithmetic result into the memory according to the arithmetic result register and the arithmetic data scale register.
[0014] Optionally, the combined arithmetic operation has been set in the processor based on hardware code.
[0015] Optionally, the arithmetic data scale consists of multiple bits, and the bits correspond to the arithmetic data;
[0016] Extracting arithmetic data from the memory according to the arithmetic data register and the arithmetic data scale register includes:
[0017] Determine the arithmetic data size of each arithmetic data according to each bit in the arithmetic data scale register;
[0018] Extract arithmetic data from the memory according to the arithmetic data register corresponding to each arithmetic data and the arithmetic data size;
[0019] Writing the arithmetic result into the memory according to the arithmetic result register and the arithmetic data scale register includes:
[0020] Determine the arithmetic result size according to each bit in the arithmetic data scale register;
[0021] Write the arithmetic result into the memory according to the arithmetic result register and the arithmetic result size.
[0022] Optionally, assigning values to the arithmetic data register according to the arithmetic data address and the bias flag includes:
[0023] Judge whether the bias flag indicates the use of bias data;
[0024] If the bias flag indicates the use of bias data, assign values to the arithmetic data register corresponding to the bias data using the arithmetic data address of the bias data;
[0025] If the bias flag indicates not to use bias data, do not assign values to the arithmetic data register corresponding to the bias data.
[0026] Optionally, after matching the block-level instructions corresponding to the program blocks in the source program in the preset instruction set, it further includes:
[0027] If there is no block-level instruction corresponding to the program block in the preset instruction set, then for the code lines in the program block, determine the operation type corresponding to the code lines; the operation types include matrix operation types and scalar operation types;
[0028] If the operation type of the code line is a matrix operation type, allocate memory locations for the matrix operation data and matrix operation results corresponding to the code line, and fill the memory locations and the matrix operation information of the code line into the operator-level instruction; wherein, the operator-level instruction is the machine instruction corresponding to the matrix operation type, including matrix operation information, matrix operation data location, and matrix operation result location, the matrix operation data location is the memory location corresponding to the matrix operation data, and the matrix operation result location is the memory location corresponding to the matrix operation result;
[0029] If the operation type of the code line is a scalar operation type, allocate memory locations for the scalar operation data and scalar operation results corresponding to the code line, and fill the memory locations and the scalar operation information of the code line into the basic instruction; wherein, the basic instruction is the machine instruction corresponding to the scalar operation type, including scalar operation information, scalar operation data location, and scalar operation result location, the scalar operation data location is the memory location corresponding to the scalar operation data, and the scalar operation result location is the memory location corresponding to the scalar operation result;
[0030] Add the operator-level instruction or the basic instruction to the executable program.
[0031] Optionally, the operator-level instruction includes matrix operation information, matrix operation data address, matrix operation result address, and matrix operation data scale, the matrix operation data address is the memory address corresponding to the matrix operation data, the matrix operation result address is the memory address corresponding to the matrix operation result, and the matrix operation data scale is the matrix operation data size corresponding to the matrix operation data.
[0032] Optionally, after adding the filled block-level instruction to the executable program, it further includes:
[0033] When executing the operator-level instruction, read the matrix operation information in the operator-level instruction to determine the matrix operation to be executed;
[0034] Assign values to the operation data register according to the matrix operation data address, assign values to the operation result register according to the matrix operation result address, and assign values to the operation data scale register according to the matrix operation data scale;
[0035] Extract matrix operation data from memory according to the operation data register and the operation data size register, perform matrix operation operations using the matrix operation data to obtain a matrix operation result, and write the matrix operation result into memory according to the operation result register and the operation data size register.
[0036] Optionally, the basic instruction includes scalar operation operation information, scalar operation data address, scalar operation result address, and scalar operation data identifier. The scalar operation data address is the memory address corresponding to the scalar operation data, the scalar operation result address is the memory address corresponding to the scalar operation result, and the scalar operation data identifier is the number of scalar operation data participating in the scalar operation operation.
[0037] Optionally, after adding the completed block-level instruction to the executable program, it further includes:
[0038] When executing the basic instruction, read the scalar operation operation information to determine the scalar operation operation to be executed;
[0039] Assign values to the operation data register according to the scalar operation data address and the scalar operation data identifier, and assign values to the operation result register according to the scalar operation result address;
[0040] Extract scalar operation data from memory according to the operation data register, perform scalar operation operations using the scalar operation data to obtain a scalar operation result, and write the scalar operation result into memory according to the operation result register.
[0041] Optionally, assigning values to the operation data register according to the scalar operation data address and the scalar operation data identifier includes:
[0042] Judge whether the number of scalar operation data represented by the scalar operation data identifier is 1;
[0043] If so, assign values to the operation data register of the first scalar operation data according to the scalar operation data address of the first scalar operation data;
[0044] If not, assign values to the operation data register of the first scalar operation data and the operation data register of the second scalar operation data according to the scalar operation data addresses of the first scalar operation data and the second scalar operation data respectively.
[0045] Optionally, allocating a memory location for the matrix operation data corresponding to the code line includes:
[0046] Judge whether the matrix operation result of the previous code line is the matrix operation data of the current code line;
[0047] If so, use the memory location of the matrix operation result of the previous code line as the memory location of the matrix operation data of the current code line;
[0048] If not, allocate a memory location for the matrix operation data of the current code line;
[0049] Allocate a memory location for the scalar operation data corresponding to the code line, including:
[0050] Determine whether the scalar operation result of the previous code line is the scalar operation data of the current code line;
[0051] If so, use the memory location of the scalar operation result of the previous code line as the memory location of the scalar operation data of the current code line;
[0052] If not, allocate a memory location for the scalar operation data of the current code line;
[0053] Add the operator-level instruction or basic instruction to the executable program, including:
[0054] Form an instruction chain with the operator-level instructions or basic instructions corresponding to the code lines in the program block in the order of the code lines, and add the instruction chain to the executable program.
[0055] This application also provides an electronic device, including:
[0056] A memory for storing a computer program;
[0057] A processor for implementing the above model program compilation method when executing the computer program.
[0058] This application also provides a computer program product, including a computer program or instruction, where the computer program or instruction implements the above model program compilation method when executed by a processor.
[0059] This application also provides a non-volatile computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, the above model program compilation method is implemented.
[0060] With this application, since corresponding block-level instructions can be set for the model units pre-divided in the neural network model, and these block-level instructions are used to record the unit type of the model unit, the memory location of the operation data input to the model unit, and the memory location where the operation result output by the model unit is stored. Therefore, when compiling the neural network model from the source program into an executable program using these block-level instructions, one block-level instruction can represent a complex model unit, which can reduce the number of instructions required to compile the model unit. At the same time, it helps the processor to understand the corresponding operation operations, the storage locations of the operation data, and the storage locations of the operation results of the model unit at one time. In this way, both the volume and complexity of the executable program can be reduced, and at the same time, it helps to improve the operation efficiency, and can avoid the defects of large volume and low execution efficiency of the executable program easily caused by compiling the neural network model with conventional instruction sets. This application also provides an electronic device, a program product, and a medium, which have the above beneficial effects. Description of the Drawings
[0061] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0062] Figure 1 It is a flowchart of a model program compilation method provided by an embodiment of the present application;
[0063] Figure 2 It is a schematic diagram of a multi-layer perceptron block-level instruction provided by an embodiment of the present application;
[0064] Figure 3 It is a schematic diagram of an attention mechanism block-level instruction provided by an embodiment of the present application;
[0065] Figure 4 It is a schematic diagram of a binary variable operator-level instruction provided by an embodiment of the present invention;
[0066] Figure 5 It is a schematic diagram of a unary variable operator-level instruction provided by an embodiment of the present invention;
[0067] Figure 6 It is a schematic diagram of a basic instruction provided by an embodiment of the present invention;
[0068] Figure 7 It is a structural block diagram of a model program compilation device provided by an embodiment of the present application;
[0069] Figure 8 It is a structural block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0070] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0071] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0072] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0073] In the related art, a neural network model generally uses a conventional instruction set for compilation during the compilation process. For example, the Reduced Instruction Set Computer (RISC) is used to compile the neural network model. However, the machine instructions in the conventional instruction set contain a small number of fields and are suitable for controlling the execution of small tasks, while the neural network model contains a large number of arithmetic operations. Therefore, when using the conventional instruction set to compile the executable program of the neural network model, it is easy to cause the executable program to be large, complex, and have low execution efficiency. For example, when dealing with complex tasks in the neural network model, the above-mentioned executable program has an increased demand for cache and is prone to cache miss events; another example is that the execution of the above-mentioned executable program requires more hardware resources and more optimization designs, such as adding proprietary instructions, instruction pipelines, etc.
[0074] In view of this, in view of the technical problems in the related art that compiling a neural network model using a conventional instruction set is likely to result in a large volume of executable programs and low execution efficiency, the present application can provide a model program compilation method. Starting from the perspective of the instruction set, corresponding block-level instructions can be set for the model units in the neural network model, and the block-level instructions can be used to record the unit type of the model unit, the memory location where the operation data input to the model unit is located, and the memory location where the operation result output by the model unit is stored. In this way, it is possible to reduce the number of instructions required to compile the model unit, and at the same time, it helps the processor to understand the corresponding operation operations, the storage location of the operation data, and the storage location of the operation result of the model unit at one time. Furthermore, it is possible to reduce the volume and complexity of the executable program, and at the same time, it also helps to improve the processing efficiency.
[0075] It should be noted that the present application does not limit the hardware device and hardware architecture for executing this method, and can be set according to actual application requirements. For example, this method can be run on a personal computer, a server, etc.
[0076] The following introduces the model program compilation method provided by this method. For ease of understanding, please refer to Figure 1 , Figure 1 which is a flowchart of a model program compilation method provided by an embodiment of the present application. This method may include:
[0077] S101. During the process of compiling a neural network model from a source program into an executable program, match block-level instructions corresponding to program blocks in the source program in a preset instruction set; wherein, the program blocks correspond to model units pre-divided in the neural network model; the block-level instructions are machine instructions corresponding to the model units, including the unit type corresponding to the model unit, the operation data location, and the operation result location, the operation data location is the memory location corresponding to the operation data input to the model unit, and the operation result location is the memory location corresponding to the operation result output by the model unit.
[0078] The following first introduces the relevant terms in this step:
[0079] Neural Network Models and Model Units: A neural network model refers to a mathematical model set up based on artificial neurons. Common neural network models can include image processing models, language processing models, etc. Among them, an image processing model is a neural network model that performs image processing tasks such as image analysis and image generation; a language processing model is a neural network model that performs text processing tasks such as text analysis and text generation. A model unit is a pre-set and divided model structure, which is the basic component unit of a neural network model, and a neural network model can be constructed by connecting model units in series or in parallel. Common model units can include a multi-layer perceptron (MLP, Multilayer Perceptron), an attention mechanism module (Attention Mechanism, such as a self-attention mechanism module self-attention and a cross-attention mechanism module cross-attention), a Transformer module, etc. It should be noted that a neural network model is usually composed of multiple identical model units. For example, a language processing model is usually composed of multiple Transformer modules.
[0080] Source Program, Executable Program and Compilation: A source program refers to an uncompiled text file written according to certain programming language specifications. For example, a neural network model can be developed using the Python language. An executable program refers to the binary code generated after the source program is compiled by a compiler. Compilation refers to translating the human-oriented programming language in the source program into machine instructions oriented to machines in the executable program, and a instruction set is required for compilation.
[0081] Program Block: In this embodiment, the program block in the source program corresponds to the model unit and contains all the code of the model unit. As mentioned above, a neural network model is usually composed of multiple identical model units, so this program block will appear repeatedly in the source program of the neural network model.
[0082] Preset instruction set and block-level instructions: In this embodiment, to reduce the compilation complexity of the neural network model, a preset instruction set containing block-level instructions is specifically provided. Among them, the block-level instruction is the machine instruction corresponding to the model unit, including the unit type corresponding to the model unit, the operation data position, and the operation result position corresponding to the model unit. The unit type is used to prompt the processor which operation operations need to be performed for the corresponding model unit; the operation data position is the memory position corresponding to the operation data input to the model unit, used to prompt the processor to obtain the operation data from which memory position; the operation result position is the memory position corresponding to the operation result output by the model unit, used to prompt the processor to write the operation result to which memory position. It should be noted that the operation data includes not only the input data of the model unit, but also the data participating in the operation inside the model unit (such as bias data). It should be noted that this embodiment does not limit the form of the operation data. For example, it can be a matrix or a scalar. It is worth pointing out that a model unit may include several operation operations, and these operation operations need to be executed sequentially. And in this embodiment, based on the hardware code, the combined operation operations corresponding to the model unit (including the operation operations and operation order corresponding to the model unit) and the corresponding relationship between the combined operation operations and the unit type can be set in the processor, so that the processor can directly determine the combined operation operations to be executed according to the unit type in the block-level instruction.
[0083] Based on the above introduction of the terms, it can be seen that this application can set block-level instructions corresponding to the model units in the neural network model, and use these block-level instructions to record the unit type of the model unit, the memory position where the operation data input to the model unit is located, and the memory position where the operation result output by the model unit is stored. In this step, during the process of compiling the neural network model from the source program into an executable program, this application can match the block-level instructions corresponding to the program blocks in the source program in the preset instruction set to translate the program block into the block-level instruction. In this way, this application can not only use one block-level instruction to represent a complex model unit, instead of using several conventional instructions to represent a model unit, which can reduce the compilation complexity; at the same time, the processor can, by parsing one block-level instruction, understand at one time the operation operations corresponding to the model unit, the storage position of the operation data, and the storage position of the operation result. In this way, it can not only reduce the volume and complexity of the executable program, but also help to improve the operation efficiency.
[0084] It should be noted that the specific form of the block-level instruction in this embodiment is not limited and can be set according to actual application requirements. To record the above unit type, operation data location, and operation result location, and at the same time be closer to the common operations in the model unit, the block-level instruction can include the unit type, operation data address, operation result address, operation data scale, and bias flag. Among them, the operation data address is the memory address corresponding to the operation data, the operation result address is the memory address corresponding to the operation result, and the operation data scale is the size of the operation data corresponding to the operation data. The memory location corresponding to the operation data can be determined based on the operation data address and the operation data scale. In addition, since the operations in the model unit are usually matrix operations and scalar operations, and the scale of the matrix operation result can be determined according to the scale of the matrix participating in the matrix operation, the memory location corresponding to the operation data can also be determined based on the operation result address and the operation data scale. The bias flag is the usage information corresponding to the bias data in the operation data, and this usage information includes two types: using the bias data and not using the bias data.
[0085] Furthermore, the block-level instruction can also be divided into an instruction prefix and instruction numbers. The instruction prefix includes the above unit type, and the instruction numbers include the above operation data address, operation result address, operation data scale, and bias flag. For easy understanding, please refer to Figure 2 and Figure 3 . Figure 2 FIG. Figure 3 is a schematic diagram of a block-level instruction of a multi-layer perceptron provided by an embodiment of the present application. The block-level instruction includes an instruction prefix and 8 instruction numbers. Among them, the instruction prefix includes the unit type of the multi-layer perceptron; "Registers 0 to 6" are 7 instruction numbers that need to be assigned to registers 0 to 6, which are the matrix A address, matrix B1 address, bias C1 address, matrix B2 address, bias C2 address, result address, and data scale respectively; "Flag" is the bias flag.
[0086] It should also be noted that this embodiment does not limit how to match the block-level instruction corresponding to the program block in the source program in the preset instruction set, and relevant compilation techniques can be referred to. For example, the code content of the program block can be mapped to the block-level instruction, and this mapping relationship can be saved in the preset instruction set. In this way, when the program block is compiled, the corresponding block-level instruction can be matched according to the above mapping relationship and the code content of the program block.
[0087] S102. Allocate operation data positions and operation result positions for the program block and fill them into the block-level instructions.
[0088] In this step, after completing the block-level instruction matching, memory spaces can be allocated for the operation data and operation results in the program block, that is, allocate operation data positions and operation result positions, and fill the allocation results into the block-level instructions to complete the compilation of the program block.
[0089] It can be seen that since memory allocation can be performed during the compilation process in this embodiment, and memory allocation can be performed in units of a program block (i.e., a model unit), memory spaces can be reasonably allocated according to the model unit to reasonably use memory, and at the same time, the occurrence probability of memory miss events can be reduced.
[0090] S103. Add the filled block-level instructions to the executable program.
[0091] In this step, the filled block-level instructions can be added to the executable program to complete the compilation of the neural network model. It should be noted that since one block-level instruction can represent one model unit in this embodiment, and a neural network model usually consists of multiple model units with the same structure, the number of block-level instructions in the executable program is small, which can effectively reduce the volume and complexity of the executable program.
[0092] Based on the above embodiments, since this application can set corresponding block-level instructions for the model units pre-divided in the neural network model, and use the block-level instructions to record the unit type of the model unit, the memory position where the operation data input to the model unit is located, and the memory position where the operation result output by the model unit is stored, when compiling the neural network model from the source program using the block-level instructions, one block-level instruction can represent a complex model unit, which can reduce the number of instructions required to compile the model unit; at the same time, it helps the processor to understand the corresponding arithmetic operations, the storage positions of operation data, and the storage positions of operation results of the model unit at one time. In this way, both the volume and complexity of the executable program can be reduced, and at the same time, it helps to improve the operation efficiency and avoid the defects of large volume and low execution efficiency of the executable program easily caused by compiling the neural network model with conventional instruction sets.
[0093] Based on the above embodiments, the specific process of the processor executing the block-level instructions is introduced below. In one possible case, after adding the filled block-level instructions to the executable program, the method may further include:
[0094] S201. When executing the block-level instruction, read the unit type in the block-level instruction to determine the combined arithmetic operation corresponding to the model unit according to the unit type; where the combined arithmetic operation includes the arithmetic operation corresponding to the model unit and the operation order.
[0095] In this step, when the processor executes a block-level instruction, it can first read the unit type in the block-level instruction to determine the combined operation corresponding to the model unit according to the unit type. Among them, the combined operation includes all the operation operations and operation sequences corresponding to the model unit.
[0096] It should be noted that the specific combined operation of this embodiment is not limited and can be set according to the specific model unit.
[0097] S202. Assign values to the arithmetic data register according to the arithmetic data address and the bias flag, assign values to the arithmetic result register according to the arithmetic result address, and assign values to the arithmetic data scale register according to the arithmetic data scale.
[0098] In this step, the processor needs to assign values to the corresponding registers according to the arithmetic data address, arithmetic result address, arithmetic data scale, and bias flag in the block-level instruction, as follows:
[0099] First, it is necessary to assign values to the arithmetic data register according to the arithmetic data address and the bias flag. Among them, the arithmetic data can be divided into bias data and non-bias data according to the function. For the bias data, before assigning values, it is necessary to first determine whether the bias flag indicates the use of bias data. If the bias flag indicates the use of bias data, the register corresponding to the bias data can be assigned values; if the bias flag indicates the non-use of bias data, there is no need to assign values to the register corresponding to the bias data. For non-bias data, the register can be directly assigned values using the corresponding arithmetic data address.
[0100] Based on this, assigning values to the arithmetic data register according to the arithmetic data address and the bias flag can include:
[0101] Step 11: Determine whether the bias flag indicates the use of bias data;
[0102] Step 12: If the bias flag indicates the use of bias data, assign values to the arithmetic data register corresponding to the bias data using the arithmetic data address of the bias data;
[0103] Step 13: If the bias flag indicates the non-use of bias data, do not assign values to the arithmetic data register corresponding to the bias data.
[0104] Secondly, it is necessary to assign values to the arithmetic result register according to the arithmetic result address.
[0105] Finally, it is necessary to assign values to the operation data scale register according to the operation data scale. It is worth noting that in order to use one operation data scale to record the scales of each operation data, the operation data scale can be composed of multiple bits, and the bits correspond to the operation data and are used to record the scale of the corresponding operation data. For example, in the multi-layer perceptron block-level instruction, the operation data scale can be divided into left, middle, and right bits. The left bit stores the number of rows of matrix A, the middle bit stores the number of columns of matrix A (equal to the number of rows of matrix B1), and the right bit stores the number of columns of matrix B1. In this way, the data scales of bias C1, matrix B2, bias C2, and the result matrix can also be obtained from the data of this bit, thus realizing that all matrix data length information can be obtained from one register.
[0106] It should be noted that the above operation data register, operation result register, and operation scale register are all general-purpose registers.
[0107] S203. Extract operation data from the memory according to the operation data register and the operation data scale register, perform a combined operation on the operation data to obtain an operation result, and write the operation result into the memory according to the operation result register and the operation data scale register.
[0108] In this step, after the register assignment is completed, the processor can extract operation data from the memory according to the operation data register and the operation data scale register, perform a combined operation on the operation data to obtain an operation result, and write the operation result into the memory according to the operation result register and the operation data scale register, thereby efficiently completing the operation of the model unit.
[0109] It should be noted that since the operation data scale can be composed of multiple bits, and the bits correspond to the operation data and are used to record the scale of the corresponding operation data, when extracting operation data and writing the operation result, it is necessary to determine the sizes of each operation data and the operation result according to different bits in the operation data scale.
[0110] Based on this, extracting operation data from the memory according to the operation data register and the operation data scale register can include:
[0111] Step 21: Determine the operation data sizes of each operation data according to each bit in the operation data scale register;
[0112] Step 22: Extract operation data from the memory according to the operation data register corresponding to each operation data and the operation data size.
[0113] Writing the operation result into the memory according to the operation result register and the operation data scale register can include:
[0114] Step 31: Determine the size of the operation result according to each bit in the operation data size register;
[0115] Step 32: Write the operation result into the memory according to the operation result register and the operation result size.
[0116] Next, taking two types of block-level instructions as examples, the process of the processor executing block-level instructions will be introduced.
[0117] 1. For a multi-layer perceptron, it usually contains two matrix multiplications and activation function calculations, which can be specifically expressed as:
[0118] ;
[0119] Among them, D represents the operation result, A represents matrix A, B1 represents matrix B1, C1 represents bias C1, B2 represents matrix B2, C2 represents bias C2, and gelu() represents the activation function. The block-level instruction form of the multi-layer perceptron can refer to Figure 2 . Among them, the register 0, register 1, register 2, register 3, register 4, and register 5 fields respectively indicate the addresses of matrix A, B1, bias C1, matrix B2, bias C2, and the result matrix D; the identification field indicates whether bias data is included in the two matrix calculations; the data size field can use multiple bits, divided into left, middle, and right bits. The left bit stores the number of rows of matrix A, the middle bit stores the number of columns of matrix A (equal to the number of rows of matrix B1), and the right bit stores the number of columns of matrix B1. In this way, the data sizes of bias C1, matrix B2, bias C2, and the result matrix can also be obtained through the data of this bit, thus realizing that all matrix data length information can be obtained from one register. In the hardware calculation, when the instruction parser parses this block-level instruction, it will sequentially calculate the multiplication of matrix A and B1, the addition of the intermediate result matrix and bias C1, the activation function calculation of the result, the multiplication calculation of the intermediate result matrix and matrix B2, and the addition of the intermediate result matrix and bias C2.
[0120] In a practical scenario, the calculation scale of the multi-layer perceptron is the product of matrix A with a scale of token_length×896 and matrix B1 with a scale of 896×4864, then the matrix result is subjected to an activation function calculation, and then the activation function result is multiplied by a matrix with a scale of 4864×896 to obtain a result with a matrix scale of token_length×896. Among them, token_length represents the input length of the multi-layer perceptron. Then when generating this block-level instruction, the following is executed:
[0121] 1) The instruction prefix contains bits indicating that this instruction is a multi-layer perceptron block-level instruction;
[0122] 2) Give the starting addresses of matrix A and B1 and assign them to register 0 and register 1;
[0123] 3) At this time, there are no biases C1 and C2. Assign the flag bit as 0, indicating that the bias C does not participate in the calculation during the calculation;
[0124] 4) This process performs activation function calculation;
[0125] 5) Give the starting address of matrix B2 and assign it to register 3;
[0126] 6) Give the storage address of the result and assign it to register 5 for storing the result of the matrix product;
[0127] 7) Set the data scale field to store the value of token_length in the high 16-bit bits, store the value of 896 in the middle 16-bit bits, and store the value of 4864 in the low 32-bit bits, and assign the final value to register 6.
[0128] Among them, the flag bit is 2 bits, respectively indicating whether the product results of matrix A and B1, and the product results of the activation function and B2 need to be added with the bias matrices C1 and C2 to obtain the final result matrix; the advantage of such a block-level instruction is that the field values can be configured once to give assignments for a series of related instructions, saving the trouble of assigning values one by one. During hardware operation, the instruction parsing process can also be continuously executed according to the block-level instruction to avoid parsing instructions one by one, thereby increasing the actual execution time.
[0129] 2. For the attention mechanism calculation, the specific calculation formula is:
[0130] ;
[0131] Among them, D represents the operation result, Q, K, and V respectively represent matrix Q (query matrix), matrix K (key matrix), and matrix V (value matrix), T represents the transpose, represents the dimension of matrix K. softmax() represents the softmax function (normalization function). The block-level instruction form of the attention mechanism can refer to Figure 3 . Similar to the block-level instruction of the multi-layer perceptron, in actual use, configure the values of registers 0, 1, 3, and 4 corresponding to Q, K, V, and the result matrix address in sequence, and assign a value to register 2 , assign a value representing the data scale to register 5, and the instruction prefix contains a bit indicating that this instruction is for the attention mechanism. Therefore, this series of instructions are encapsulated into a block-level instruction to achieve the purpose of continuous instruction generation, thereby reducing the instruction configuration work and the instruction parsing work during hardware operation. During actual hardware operation, when this instruction is parsed, the hardware will calculate the transpose of matrix K, the product of matrix Q and the transpose of K, multiply each element of the intermediate result matrix by a scalar data, perform a softmax calculation on each row of data, and then multiply it by matrix V in sequence according to the information such as the address and data scale in the instruction.
[0132] Based on the above embodiments, considering that the model units of the neural network model are replaced relatively quickly, there may be a situation where the block-level instructions cannot fully cover all model units in the neural network model. For this reason, considering that most operations in the neural network model are matrix operations and scalar operations, operator-level instructions and basic instructions can also be set in the preset instruction set. The operator-level instructions can cover matrix operations, and the basic instructions can cover scalar operations. In this way, when the block-level instructions cannot fully cover the model units, the operator-level instructions and basic instructions can be used instead. Based on this, after matching the block-level instructions corresponding to the program block in the source program in the preset instruction set, this method may further include:
[0133] S301. If there is no block-level instruction corresponding to the program block in the preset instruction set, determine the type of operation corresponding to the code line in the program block; the type of operation includes matrix operation type and scalar operation type.
[0134] In this step, if there is no block-level instruction matching the program block in the preset instruction set, each code line in the program block needs to be compiled separately. As described above, this embodiment can also provide operator-level instructions corresponding to matrix operations and basic instructions corresponding to scalar operations. Therefore, before compiling the code line, it is first necessary to determine the type of operation corresponding to this code line.
[0135] It should be noted that this embodiment does not limit how to determine the type of operation corresponding to the code line. For example, it can be determined according to the content of the code line (such as parameter type, operation type).
[0136] S302. If the type of operation of the code line is matrix operation type, allocate memory locations for the matrix operation data and matrix operation result corresponding to the code line, and fill the memory locations and the matrix operation information of the code line into the operator-level instruction; where the operator-level instruction is the machine instruction corresponding to the matrix operation type, including matrix operation information, matrix operation data location, and matrix operation result location. The matrix operation data location is the memory location corresponding to the matrix operation data, and the matrix operation result location is the memory location corresponding to the matrix operation result.
[0137] In this step, when it is determined that the operation type of the code line is a matrix operation type, memory locations can be allocated for the matrix operation data and matrix operation result corresponding to the code line, and the memory locations and the matrix operation information of the code line are filled into the operator-level instruction, thereby completing the compilation of the code line. The operator-level instruction is introduced as follows:
[0138] The operator-level instruction is the machine instruction corresponding to the matrix operation type. This instruction can include matrix operation information, matrix operation data location, and matrix operation result location. The matrix operation information is the specific information of the matrix operation, which is used to prompt the processor which matrix operation needs to be executed. Common matrix operations include: normalization calculation, softmax, activation function calculation, matrix-vector multiplication, matrix-matrix addition, matrix-vector addition, matrix transpose, matrix left and right bisect and swap positions and take the inverse operation on the right part, vector summation, vector exp (perform exponential function operation on the vector), vector activation function, etc. The matrix operation data location is the memory location corresponding to the matrix operation data, which is used to prompt the processor to obtain the matrix operation data from which memory location. The matrix operation result location is the memory location corresponding to the matrix operation result, which is used to prompt the processor to write the matrix operation result to which memory location. It should be noted that in this embodiment, the matrix operation can be set in the processor based on the hardware code, so that the processor can directly determine the matrix operation to be executed according to the matrix operation information in the operator-level instruction.
[0139] It should be noted that this embodiment does not limit the specific form of the operator-level instruction, which can be set according to actual application requirements. For example, the operator-level instruction can include matrix operation information, matrix operation data address, matrix operation result address, and matrix operation data scale. The matrix operation data address is the memory address corresponding to the matrix operation data, the matrix operation result address is the memory address corresponding to the matrix operation result, and the matrix operation data scale is the matrix operation data size corresponding to the matrix operation data.
[0140] Furthermore, since some matrix operations occur between matrices, such as matrix-vector multiplication, matrix-matrix addition, matrix-vector addition, vector summation; some matrix operations are only performed on a single matrix, such as normalization calculation, softmax, activation function calculation, matrix transpose, matrix left and right bisect and swap positions and take the inverse operation on the right part, vector exp, vector activation function, so this embodiment can provide two forms of operator-level instructions. For ease of understanding, please refer to Figure 4 and Figure 5 . Figure 4A schematic diagram of a binary variable operator-level instruction provided by an embodiment of the present invention. This operator-level instruction includes an instruction prefix and 5 instruction numbers. Among them, "Registers 0-3" are 4 instruction numbers that need to be assigned to Registers 0-3, which are the addresses of matrix / vector A, the addresses of matrix / vector V, the data scale, and the result address respectively; "Identifier" is matrix calculation operation information. Figure 5 A schematic diagram of a unary variable operator-level instruction provided by an embodiment of the present invention. This operator-level instruction includes an instruction prefix and 4 instruction numbers. Among them, "Registers 0-2" are 3 instruction numbers that need to be assigned to Registers 0-2, which are the addresses of matrix / vector A, the data scale, and the result address respectively; "Identifier" is matrix calculation operation information. It can be seen that the main difference between the binary variable operator-level instruction and the unary variable operator-level instruction is the number of matrix operation data.
[0141] Furthermore, since the use of operator-level instructions increases the number of instructions, in order to minimize the negative impact on the model running performance, the coherence between instructions can be considered during memory allocation. Specifically, when allocating a memory location for the matrix operation data of the current code line, it can be determined whether the matrix operation result of the previous code line is the matrix operation data of the current code line, that is, it is determined whether the current code line will use the matrix operation result of the previous code line. If the matrix operation result of the previous code line is the matrix operation data of the current code line, the memory location of the matrix operation result of the previous code line can be used as the memory location of the matrix operation data of the current code line to avoid additional memory location allocation. If the matrix operation result of the previous code line is not the matrix operation data of the current code line, a memory location can be reallocated for the matrix operation data of the current code line. In this way, not only can the memory location be reasonably allocated, but also the occurrence probability of memory miss events can be reduced.
[0142] Based on this, allocating a memory location for the matrix operation data corresponding to a code line may include:
[0143] Step 41: Determine whether the matrix operation result of the previous code line is the matrix operation data of the current code line;
[0144] Step 42: If so, use the memory location of the matrix operation result of the previous code line as the memory location of the matrix operation data of the current code line;
[0145] Step 43: If not, allocate a memory location for the matrix operation data of the current code line.
[0146] S303. If the operation type of the code line is a scalar operation type, allocate memory locations for the scalar operation data and scalar operation result corresponding to the code line, and fill the memory locations and the scalar operation information of the code line into the basic instruction. The basic instruction is the machine instruction corresponding to the scalar operation type, including scalar operation information, scalar operation data location, and scalar operation result location. The scalar operation data location is the memory location corresponding to the scalar operation data, and the scalar operation result location is the memory location corresponding to the scalar operation result.
[0147] In this step, when it is determined that the operation type of the code line is a scalar operation type, memory locations can be allocated for the scalar operation data and scalar operation result corresponding to the code line, and the memory locations and the scalar operation information of the code line can be filled into the basic instruction, thereby completing the compilation of the code line. The following introduces the basic instruction:
[0148] The basic instruction is the machine instruction corresponding to the scalar operation type. This instruction can include scalar operation information, scalar operation data location, and scalar operation result location. The scalar operation information is the specific information of the scalar operation, which is used to prompt the processor which scalar operation needs to be executed. Common scalar operations include: reciprocal, scalar addition, multiplication, square root, exp (performing exponential function operation), activation function, etc. The scalar operation data location is the memory location corresponding to the scalar operation data, which is used to prompt the processor to obtain the scalar operation data from which memory location. The scalar operation result location is the memory location corresponding to the scalar operation result, which is used to prompt the processor to write the scalar operation result to which memory location. It should be noted that in this embodiment, the scalar operation can be set in the processor based on the hardware code, so that the processor can directly determine the scalar operation to be executed according to the scalar operation information in the basic instruction.
[0149] It should be noted that the specific form of the basic instruction is not limited in this embodiment and can be set according to actual application requirements. For example, the basic instruction includes scalar operation information, scalar operation data address, scalar operation result address, and scalar operation data identifier. The scalar operation data address is the memory address corresponding to the scalar operation data, the scalar operation result address is the memory address corresponding to the scalar operation result, and the scalar operation data identifier is the number of scalar operation data participating in the scalar operation. For easy understanding, please refer to Figure 6 . Figure 6A schematic diagram of a basic instruction provided by an embodiment of the present invention. The basic instruction includes an instruction prefix and 5 instruction numbers. Among them, "Registers 0-2" are 3 instruction numbers that need to be assigned to Registers 0-2, which are the addresses of data a, the address of data b, and the result address respectively; "Identifier 1" is scalar calculation operation information; "Identifier 2" is a scalar operation data identifier used to determine the number of scalar operation data. When the number of scalar operation data is 1, only Register 0 needs to be copied. When the number of scalar operation data is 2, Registers 0 and 1 need to be copied to distinguish the cases of unary variables and binary variables.
[0150] In addition, it is also worth noting that this embodiment can also use basic instructions to achieve the effect of operator-level instructions, only by converting matrix operations into scalar operations for each element in the matrix. For example, the operator-level instruction of multiplying vector a by vector b can be assembled through the following basic instructions: 1) Configure Figure 6 The basic instruction shown, where Register 0 and Register 1 respectively indicate the addresses of the first elements of vector a and vector b, Register 2 indicates the address of the result, Identifier 1 indicates performing multiplication calculation, and Identifier 2 is assigned 1 to indicate that data b participates in the calculation; 2) Repeat the calculation in step 1 for the specific number of times equal to the length of the vector; 3) Then use basic instructions to perform scalar addition calculation; the specific number of repetitions is the length of the vector minus 1. According to a series of basic instructions for scalar multiplication and scalar addition, the calculation effect of multiplying vector by vector of the operator-level instruction can be achieved.
[0151] Furthermore, since the use of basic instructions increases the number of instructions, in order to minimize the negative impact on the model running performance, the coherence between instructions can be considered during memory allocation. Specifically, when allocating a memory location for the scalar operation data of the current code line, it can be determined whether the scalar operation result of the previous code line is the scalar operation data of the current code line, that is, to determine whether the current code line will use the scalar operation result of the previous code line. If the scalar operation result of the previous code line is the scalar operation data of the current code line, the memory location of the scalar operation result of the previous code line can be used as the memory location of the scalar operation data of the current code line to avoid additional memory location allocation. If the scalar operation result of the previous code line is not the scalar operation data of the current code line, a new memory location can be allocated for the scalar operation data of the current code line. In this way, both the memory location can be reasonably allocated and the occurrence probability of memory miss events can be reduced.
[0152] Based on this, allocating a memory location for the scalar operation data corresponding to the code line includes:
[0153] Step 51: Determine whether the scalar operation result of the previous code line is the scalar operation data of the current code line;
[0154] Step 52: If so, use the memory location of the scalar operation result of the previous code line as the memory location of the scalar operation data of the current code line;
[0155] Step 53: If not, allocate a memory location for the scalar operation data of the current code line.
[0156] S304. Add the operator-level instruction or the basic instruction to the executable program.
[0157] In this step, the operator-level instruction and the basic instruction can be added to the executable program, so as to use the operator-level instruction and the basic instruction to replace the block-level instruction to compile the model unit, thereby improving the compilation flexibility.
[0158] Furthermore, it is worth noting that since the program block corresponds to the model unit, that is, the code lines in the program block have a coherent execution flow, in order to ensure that the corresponding operator-level instruction / basic instruction also has a smooth execution order, the operator-level instruction or the basic instruction corresponding to the code lines in the program block can be formed into an instruction chain according to the order of the code lines, and the instruction chain is added to the executable program. In this way, since the instructions in the instruction chain are interrelated through the memory location, the execution efficiency of the instruction chain can be improved.
[0159] Based on this, adding the operator-level instruction or the basic instruction to the executable program can include:
[0160] Step 61: Form the operator-level instruction or the basic instruction corresponding to the code lines in the program block into an instruction chain according to the order of the code lines, and add the instruction chain to the executable program.
[0161] Next, the situation of using the operator-level instruction to replace the block-level instruction to compile the model unit will be introduced. In one case, a multi-layer perceptron can be compiled by the following chain assembly method to achieve the same calculation effect:
[0162] 1) Configure the binary variable operator-level instruction shown in Figure 4 for matrices A and B1. The field registers 0, 1, and 3 are respectively assigned the starting addresses of matrices A, B1, and the result matrix D1. The data scale field is assigned the row and column dimensions required in the matrix calculation, and the identification field is configured with a value indicating the calculation type of matrix multiplication;
[0163] 2) If there is a bias C1, configure the operator-level instruction of the binary variable shown in Figure 4 The field registers 0, 1, and 3 are respectively assigned the starting addresses of matrix D1, the bias, and the result matrix D2. The data scale field is assigned the row and column dimensions required in the calculation, and the identification field is configured with a value indicating the calculation type of matrix addition / matrix-vector; if there is no bias C1, this step is omitted;
[0164] 3) Configuration Figure 5 The operator-level instruction of a unary variable shown performs an activation function calculation on matrix D2. The field register 0 and register 2 are respectively assigned the starting addresses of matrix D2 and the result matrix D3, the data scale field is assigned the row and column dimensions required in the calculation, and the identification field configures a numerical value to indicate the calculation type of the activation function;
[0165] 4) Configure for matrix D3 and B2 Figure 4 The operator-level instruction of a binary variable shown, the field registers 0, 1, and 3 are respectively assigned the starting addresses of matrix D3, B2, and the result matrix D4, the data scale field is assigned the row and column dimensions required in the matrix calculation, and the identification field configures a numerical value to indicate the calculation type of matrix multiplying matrix;
[0166] 5) If there is a bias C2, configure Figure 4 The operator-level instruction of a binary variable shown, the field registers 0, 1, and 3 are respectively assigned the starting addresses of matrix D4, bias C2, and the result matrix D5, the data scale field is assigned the row and column dimensions required in the calculation, and the identification field configures a numerical value to indicate the calculation type of matrix adding matrix / vector; if there is no bias C2, this step is omitted.
[0167] Based on the above embodiments, the specific process of the processor executing the operator-level instruction is introduced below. In a possible case, after adding the filled block-level instruction to the executable program, the method may further include:
[0168] S401. When executing the operator-level instruction, read the matrix operation operation information in the operator-level instruction to determine the matrix operation operation to be executed.
[0169] In this step, when the processor executes the operator-level instruction, it can first read the matrix operation operation information in the operator-level instruction to determine the matrix operation operation to be executed.
[0170] S402. Assign values to the operation data registers according to the matrix operation data address, assign values to the operation result registers according to the matrix operation result address, and assign values to the operation data scale registers according to the matrix operation data scale.
[0171] In this step, the processor needs to assign values to the corresponding registers according to the matrix operation data address, matrix operation result address, and matrix operation data scale in the operator-level instruction. For the operator-level instruction of a unary variable, only a single operation data register needs to be assigned; for the operator-level instruction of a binary variable, two operation data registers need to be assigned.
[0172] S403. Extract matrix operation data from memory according to the operation data register and the operation data size register, perform matrix operation operations using the matrix operation data to obtain a matrix operation result, and write the matrix operation result into memory according to the operation result register and the operation data size register.
[0173] In this step, after completing the register assignment, the processor can extract matrix operation data from memory according to the operation data register and the operation data size register, perform matrix operation operations using the matrix operation data to obtain a matrix operation result, and write the matrix operation result into memory according to the operation result register and the operation data size register, thereby efficiently completing the matrix operation operations.
[0174] Based on the above embodiments, the following introduces the specific process of the processor executing basic instructions. In a possible case, after adding the filled block-level instructions to the executable program, the method may further include:
[0175] S501. When executing a basic instruction, read the scalar operation operation information to determine the scalar operation operation to be executed.
[0176] In this step, when the processor executes a basic instruction, it can first read the scalar operation operation information in the basic instruction to determine the scalar operation operation to be executed.
[0177] S502. Assign values to the operation data register according to the scalar operation data address and the scalar operation data identifier, and assign values to the operation result register according to the scalar operation result address.
[0178] In this step, the processor needs to assign values to the corresponding registers according to the scalar operation data address, the scalar operation result address, and the scalar operation data identifier in the basic instruction. It is necessary to pay special attention to the value of the scalar operation data identifier. If the scalar operation data identifier indicates that the number of scalar operation data is 1, only the register of one scalar operation data needs to be assigned; if the scalar operation data identifier indicates that the number of scalar operation data is 2, only the registers of two scalar operation data need to be assigned.
[0179] Based on this, assigning values to the operation data register according to the scalar operation data address and the scalar operation data identifier may include:
[0180] Step 71: Determine whether the number of scalar operation data indicated by the scalar operation data identifier is 1;
[0181] Step 72: If so, assign values to the operation data register of the first scalar operation data according to the scalar operation data address of the first scalar operation data;
[0182] Step 73: If not, assign values to the operation data register of the first scalar operation data and the operation data register of the second scalar operation data according to the scalar operation data addresses of the first scalar operation data and the second scalar operation data, respectively.
[0183] S503. Extract scalar operation data from the memory according to the operation data register, perform a scalar operation on the scalar operation data to obtain a scalar operation result, and write the scalar operation result into the memory according to the operation result register.
[0184] In this step, after the register assignment is completed, the processor can extract scalar operation data from the memory according to the operation data register and the operation data size register, perform a scalar operation on the scalar operation data to obtain a scalar operation result, and write the scalar operation result into the memory according to the operation result register and the operation data size register, thereby efficiently completing the scalar operation.
[0185] Through the above improvement of the multi-level instruction design, for a typical model inference process, block-level instructions can be directly used to generate the instructions required for a typical program block. This modular instruction generation method can not only reduce the complexity of decoding multiple instructions in the later stage, but also improve the hit rate of the instruction cache and speed up the calculation process. For a changing large model without ready-made block-level instructions, operator-level instructions can be used to chain and assemble the process of generating the corresponding changed block-level instructions, which enables flexible generation of custom high-level task instructions through chaining of operator-level instructions. Similarly, for a transformed operator, if there are no ready-made operator-level instructions, the corresponding operator-level instruction requirements can be generated by chaining basic instructions. Therefore, the multi-level instruction design achieves a good balance between the flexibility of instructions and high-performance computing, supporting both the instruction generation of existing typical program blocks and the flexible generation of instructions after the model changes. By providing modular and ready-made chain assembly, this instruction design simplifies the complexity of instruction generation. Developers can focus more on high-level task logic without spending a lot of time and effort dealing with the technical details of underlying instructions.
[0186] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0187] Next, the model program compilation device, electronic device, computer program product, and non-volatile computer-readable storage medium provided by the embodiments of the present application will be introduced.
[0188] Please refer to Figure 7 , Figure 7The following is a structural block diagram of a model program compilation device provided by an embodiment of the present application. The device may include:
[0189] A matching module 701, configured to match block-level instructions corresponding to program blocks in a source program in a preset instruction set during the process of compiling a neural network model from a source program into an executable program; wherein, the program blocks correspond to model units pre-divided in the neural network model; the block-level instructions are machine instructions corresponding to the model units, including the unit type, operation data position, and operation result position corresponding to the model unit. The operation data position is the memory position corresponding to the operation data input to the model unit, and the operation result position is the memory position corresponding to the operation result output by the model unit;
[0190] A memory allocation module 702, configured to allocate an operation data position and an operation result position for the program block and fill them into the block-level instructions;
[0191] An instruction adding module 703, configured to add the filled block-level instructions to the executable program.
[0192] Optionally, the device may further include:
[0193] A unit type parsing module, configured to read the unit type in the block-level instruction when executing the block-level instruction, so as to determine the combined operation corresponding to the model unit according to the unit type; wherein, the combined operation includes the operation and operation order corresponding to the model unit;
[0194] A register assignment module, configured to assign values to operation data registers according to operation data addresses and bias identifiers, assign values to operation result registers according to operation result addresses, and assign values to operation data scale registers according to operation data scales;
[0195] An operation module, configured to extract operation data from memory according to operation data registers and operation data scale registers, perform combined operation operations on the operation data to obtain an operation result, and write the operation result into memory according to operation result registers and operation data scale registers.
[0196] Optionally, the combined operation has been set in the processor based on hardware code.
[0197] Optionally, the operation data scale consists of multiple bits, and the bits correspond to the operation data;
[0198] The operation module may include:
[0199] An operation data size parsing sub-module, configured to determine the operation data size of each operation data according to each bit in the operation data scale register;
[0200] An operation data extraction sub-module, configured to extract operation data from a memory according to an operation data register corresponding to each operation data and an operation data size;
[0201] An operation result size parsing sub-module, configured to determine an operation result size according to each bit in an operation data scale register;
[0202] An operation result writing sub-module, configured to write an operation result into a memory according to an operation result register and an operation result size.
[0203] Optionally, a register assignment module may include:
[0204] A bias assignment sub-module, configured to determine whether a bias identifier indicates using bias data; if the bias identifier indicates using bias data, assign a corresponding operation data register of the bias data by using an operation data address of the bias data; if the bias identifier indicates not using bias data, do not assign a corresponding operation data register of the bias data.
[0205] Optionally, the apparatus may further include:
[0206] An operation type determination module, configured to, if there is no block-level instruction corresponding to a program block in a preset instruction set, determine an operation operation type corresponding to a code line in the program block; the operation operation type includes a matrix operation type and a scalar operation type;
[0207] An operator-level instruction translation module, configured to, if the operation operation type of a code line is a matrix operation type, allocate memory locations for matrix operation data and a matrix operation result corresponding to the code line, and fill the memory locations and matrix operation operation information of the code line into an operator-level instruction; wherein, the operator-level instruction is a machine instruction corresponding to the matrix operation type, including matrix operation operation information, a matrix operation data location, and a matrix operation result location, the matrix operation data location is a memory location corresponding to the matrix operation data, and the matrix operation result location is a memory location corresponding to the matrix operation result;
[0208] A basic instruction translation module, configured to, if the operation operation type of a code line is a scalar operation type, allocate memory locations for scalar operation data and a scalar operation result corresponding to the code line, and fill the memory locations and scalar operation operation information of the code line into a basic instruction; wherein, the basic instruction is a machine instruction corresponding to the scalar operation type, including scalar operation operation information, a scalar operation data location, and a scalar operation result location, the scalar operation data location is a memory location corresponding to the scalar operation data, and the scalar operation result location is a memory location corresponding to the scalar operation result;
[0209] An instruction adding module 703 is further configured to add the operator-level instruction or the basic instruction to an executable program.
[0210] Optionally, the device may further include:
[0211] An operator parsing module, configured to read matrix operation operation information in an operator-level instruction to determine a matrix operation operation to be executed when executing the operator-level instruction;
[0212] A register assignment module, further configured to assign values to operation data registers according to matrix operation data addresses, assign values to operation result registers according to matrix operation result addresses, and assign values to operation data scale registers according to matrix operation data scales;
[0213] An operation module, further configured to extract matrix operation data from a memory according to operation data registers and operation data scale registers, execute a matrix operation operation using the matrix operation data to obtain a matrix operation result, and write the matrix operation result into the memory according to operation result registers and operation data scale registers.
[0214] Optionally, the device may further include:
[0215] A scalar operation operation parsing module, configured to read scalar operation operation information to determine a scalar operation operation to be executed when executing a basic instruction;
[0216] A register assignment module, further configured to assign values to operation data registers according to scalar operation data addresses and scalar operation data identifiers, and assign values to operation result registers according to scalar operation result addresses;
[0217] An operation module, further configured to extract scalar operation data from a memory according to operation data registers, execute a scalar operation operation using the scalar operation data to obtain a scalar operation result, and write the scalar operation result into the memory according to operation result registers.
[0218] Optionally, the register assignment module may include:
[0219] A scalar operation data assignment sub-module, configured to determine whether the number of scalar operation data represented by a scalar operation data identifier is 1; if so, assign a value to an operation data register of a first scalar operation data according to the scalar operation data address of the first scalar operation data; if not, assign values to operation data registers of the first scalar operation data and a second scalar operation data according to the scalar operation data addresses of the first scalar operation data and the second scalar operation data respectively.
[0220] Optionally, the operator-level instruction translation module may include:
[0221] The first memory allocation sub-module is used to determine whether the matrix operation result of the previous code line is the matrix operation data of the current code line; if so, use the memory location of the matrix operation result of the previous code line as the memory location of the matrix operation data of the current code line; if not, allocate a memory location for the matrix operation data of the current code line.
[0222] The basic instruction translation module may include:
[0223] The first memory allocation sub-module is used to determine whether the scalar operation result of the previous code line is the scalar operation data of the current code line; if so, use the memory location of the scalar operation result of the previous code line as the memory location of the scalar operation data of the current code line; if not, allocate a memory location for the scalar operation data of the current code line.
[0224] The instruction addition module 703 may include:
[0225] The instruction chain addition sub-module is used to form an instruction chain by arranging the operator-level instructions or basic instructions corresponding to the code lines in the program block in the order of the code lines, and add the instruction chain to the executable program.
[0226] For the description of the features in the corresponding embodiments of the model program compilation device, reference can be made to the relevant description in the corresponding embodiments of the model program compilation method, which will not be elaborated here.
[0227] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned model program compilation method embodiments.
[0228] Please refer to Figure 8 , Figure 8 FIG. is a structural block diagram of an electronic device provided by an embodiment of the present invention. An embodiment of the present invention provides an electronic device 10, including a processor 11 and a memory 12; wherein, the memory 12 is used to store a computer program; the processor 11 is used to execute the model program compilation method provided in the foregoing embodiment when executing the computer program.
[0229] For the specific process of the above model program compilation method, reference can be made to the corresponding content provided in the foregoing embodiments, and details will not be repeated here.
[0230] Moreover, as a carrier for resource storage, the memory 12 can be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the storage method can be temporary storage or permanent storage.
[0231] In addition, the electronic device 10 further includes a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. Among them, the power supply 13 is used to provide operating voltage for each hardware device on the electronic device 10. The communication interface 14 can create a data transmission channel between the electronic device 10 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of the present invention, and no specific limitation is imposed here. The input / output interface 15 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0232] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any of the above-described model program compilation method embodiments when running.
[0233] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0234] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described model program compilation method embodiments.
[0235] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described model program compilation method embodiments.
[0236] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0237] The above has introduced in detail a model program compilation method, an electronic device, a program product, and a medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A model program compiling method, characterized in that: include: In the process of compiling the neural network model from a source program to an executable program, a block-level instruction corresponding to a program block in the source program is matched in a preset instruction set; wherein the program block corresponds to a model unit pre-divided in the neural network model; the block-level instruction is a machine instruction corresponding to the model unit, including a unit type, a calculation data position and a calculation result position corresponding to the model unit, wherein the calculation data position is a memory position corresponding to the calculation data input to the model unit, and the calculation result position is a memory position corresponding to the calculation result output by the model unit; Allocating the operation data location and the operation result location to the program block and filling them into the block-level instructions; Block-level instructions that complete the padding are added to the executable program.
2. The model program compiling method according to claim 1, characterized in that: The block-level instruction includes the unit type, operation data address, operation result address, operation data scale and offset identifier, the operation data address is the memory address corresponding to the operation data, the operation result address is the memory address corresponding to the operation result, the operation data scale is the operation data size corresponding to the operation data, and the offset identifier is the usage information corresponding to the offset data in the operation data.
3. The model program compiling method according to claim 2, characterized in that: After adding the filled block-level instructions to the executable program, the method further includes: When executing the block-level instruction, the unit type in the block-level instruction is read to determine the combined operation corresponding to the model unit according to the unit type; wherein the combined operation includes the operation and operation sequence corresponding to the model unit; Assigning a value to an operation data register according to the operation data address and the offset identifier, assigning a value to an operation result register according to the operation result address, and assigning a value to an operation data scale register according to the operation data scale; The operation data is extracted from the memory according to the operation data register and the operation data scale register, the combined operation operation is performed using the operation data to obtain the operation result, and the operation result is written into the memory according to the operation result register and the operation data scale register.
4. The model program compiling method according to claim 3, characterized in that: The operation data scale is composed of a plurality of bits, and the bits correspond to the operation data; Extracting the operation data from the memory according to the operation data register and the operation data size register comprises: Determine the operation data size of each operation data according to each bit in the operation data size register; Extracting the operation data from the memory according to the operation data register corresponding to each operation data and the operation data size; Writing the operation result into the memory according to the operation result register and the operation data size register comprises: Determine the size of the operation result according to each bit in the operation data size register; The operation result is written into a memory according to the operation result register and the operation result size.
5. The model program compiling method according to claim 3, characterized in that: Assigning a value to an operation data register according to the operation data address and the offset identifier, comprising: Determining whether the bias identifier indicates using the bias data; If the bias identifier indicates that the bias data is used, then the operation data register corresponding to the bias data is assigned a value using the operation data address of the bias data; If the offset identifier indicates that the offset data is not used, no value is assigned to the operation data register corresponding to the offset data.
6. The model program compiling method according to any one of claims 1 to 5, characterized in that: After matching the block-level instructions corresponding to the program blocks in the source program in the preset instruction set, the method further includes: If there is no block-level instruction corresponding to the program block in the preset instruction set, determining the type of operation corresponding to the code line in the program block; the operation type includes a matrix operation type and a scalar operation type; If the operation type of the code line is the matrix operation type, a memory location is allocated for the matrix operation data and the matrix operation result corresponding to the code line, and the memory location and the matrix operation operation information of the code line are filled into the operator-level instruction; wherein the operator-level instruction is a machine instruction corresponding to the matrix operation type, comprising the matrix operation operation information, the matrix operation data location and the matrix operation result location, the matrix operation data location is the memory location corresponding to the matrix operation data, and the matrix operation result location is the memory location corresponding to the matrix operation result; If the operation type of the code line is the scalar operation type, a memory location is allocated for the scalar operation data and the scalar operation result corresponding to the code line, and the memory location and the scalar operation information of the code line are filled into the basic instruction; wherein the basic instruction is a machine instruction corresponding to the scalar operation type, comprising the scalar operation information, the scalar operation data location and the scalar operation result location, the scalar operation data location is the memory location corresponding to the scalar operation data, and the scalar operation result location is the memory location corresponding to the scalar operation result; The operator-level instruction or the basic instruction is added to the executable program.
7. The model program compiling method according to claim 6, characterized in that: The operator-level instruction includes the matrix operation operation information, the matrix operation data address, the matrix operation result address and the matrix operation data scale. The matrix operation data address is the memory address corresponding to the matrix operation data, the matrix operation result address is the memory address corresponding to the matrix operation result, and the matrix operation data scale is the matrix operation data size corresponding to the matrix operation data.
8. The model program compiling method according to claim 7, characterized in that: After adding the filled block-level instructions to the executable program, the method further includes: When executing the operator-level instruction, reading the matrix operation information in the operator-level instruction to determine the matrix operation to be performed; Assigning a value to an operation data register according to the matrix operation data address, assigning a value to an operation result register according to the matrix operation result address, and assigning a value to an operation data scale register according to the matrix operation data scale; The matrix operation data is extracted from the memory according to the operation data register and the operation data scale register, the matrix operation operation is performed using the matrix operation data to obtain the matrix operation result, and the matrix operation result is written into the memory according to the operation result register and the operation data scale register.
9. The model program compiling method according to claim 6, characterized in that: The basic instruction includes the scalar operation information, the scalar operation data address, the scalar operation result address and the scalar operation data identifier. The scalar operation data address is the memory address corresponding to the scalar operation data, the scalar operation result address is the memory address corresponding to the scalar operation result, and the scalar operation data identifier is the number of scalar operation data involved in the scalar operation operation.
10. The model program compiling method according to claim 9, characterized in that: After adding the filled block-level instructions to the executable program, the method further includes: When executing the basic instruction, reading the scalar operation information to determine the scalar operation to be performed; Assigning a value to an operation data register according to the scalar operation data address and the scalar operation data identifier, and assigning a value to an operation result register according to the scalar operation result address; The scalar operation data is extracted from the memory according to the operation data register, the scalar operation operation is performed using the scalar operation data to obtain the scalar operation result, and the scalar operation result is written into the memory according to the operation result register.
11. The model program compiling method according to claim 10, characterized in that: Assigning a value to an operation data register according to the scalar operation data address and the scalar operation data identifier includes: Determine whether the number of scalar operation data represented by the scalar operation data identifier is 1; If yes, assigning a value to the operation data register of the first scalar operation data according to the scalar operation data address of the first scalar operation data; If not, assign values to the operation data register of the first scalar operation data and the operation data register of the second scalar operation data respectively according to the scalar operation data addresses of the first scalar operation data and the second scalar operation data.
12. The model program compiling method according to claim 6, characterized in that: Allocating memory locations for matrix operation data corresponding to the code line includes: Determine whether the matrix operation result of the previous code line is the matrix operation data of the current code line; If yes, the memory location of the matrix operation result of the previous code line is used as the memory location of the matrix operation data of the current code line; If not, allocate memory locations for the matrix operation data of the current code line; Allocating a memory location for the scalar operation data corresponding to the code line includes: Determine whether the scalar operation result of the previous code line is the scalar operation data of the current code line; If yes, the memory location of the scalar operation result of the previous code line is used as the memory location of the scalar operation data of the current code line; If not, allocate memory locations for the scalar operation data for the current line of code; Adding the operator-level instruction or the basic instruction to the executable program includes: Operator-level instructions or basic instructions corresponding to the code lines in the program block are combined into an instruction chain in the order of the code lines, and the instruction chain is added to the executable program.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the model program compiling method as claimed in any one of claims 1 to 12 when executing the computer program.
14. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the model program compiling method according to any one of claims 1 to 12 is implemented.
15. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the model program compiling method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Neural network model compiling method and system, equipment and storage medium
CN113918163A
Reducing computation in neural networks using selfmodifying code
CN113939801A