Operation method and device based on artificial intelligence model, equipment, medium and product
By storing large-scale matrices into multiple small-block matrices and reading them in sequence for calculation, the problem of inefficient computing in large-scale artificial intelligence models is solved, and more efficient computing and hardware resource savings are achieved.
Patent Information
- Application Number
- CN202510805418.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Traditional matrix computing methods are inefficient in large-scale artificial intelligence models and have high demand for hardware resources, resulting in inefficient computing.
By storing large-scale matrices into multiple small-block matrices and reading these small-block matrices in turn for calculations during the calculation process, the address jump calculation and data acquisition delay are reduced, and the computing efficiency is improved.
Improve computing efficiency, save hardware resources, optimize the data reading process, and shorten the data acquisition delay.
Smart Images

Figure CN120337997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, equipment, medium and product for running an artificial intelligence model. Background Art
[0002] In the field of artificial intelligence, especially in the process of inference and training of large models, matrix operations play a core role. With the continuous increase in the scale of the model, traditional matrix operation methods face many challenges. Large-scale matrix operations need to process a large amount of data, have extremely high requirements for hardware resources, and result in low computing efficiency.
[0003] Therefore, how to improve the operation efficiency of artificial intelligence models is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] The present invention provides a method, device, electronic device, storage medium and computer program product for running an artificial intelligence model, which improves the operation efficiency of the artificial intelligence model and saves hardware resources.
[0005] The present invention provides a method for running an artificial intelligence model, including: Obtaining source data of a target operation; wherein, the source data includes a first matrix and a second matrix; Judging whether the target operation is a block calculation; if so, determining a first block size of the first matrix and a second block size of the second matrix; Storing the first matrix in blocks as a plurality of first block matrices according to the first block size, and storing the second matrix in blocks as a plurality of second block matrices according to the second block size; During the execution of the target operation, sequentially reading the plurality of first block matrices and the plurality of second block matrices to implement the target operation to obtain a target operation result.
[0006] The present invention further provides a device for running an artificial intelligence model, including: An obtaining module, configured to obtain source data of a target operation; wherein, the source data includes a first matrix and a second matrix; A determining module, configured to judge whether the target operation is a block calculation; if so, determining a first block size of the first matrix and a second block size of the second matrix; A block storage module, configured to store the first matrix in blocks as a plurality of first block matrices according to the first block size, and store the second matrix in blocks as a plurality of second block matrices according to the second block size; An operation module, configured to sequentially read the plurality of first block matrices and the plurality of second block matrices during the execution of the target operation to implement the target operation to obtain a target operation result.
[0007] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any one of the above-mentioned operation methods based on an artificial intelligence model when executing the computer program.
[0008] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any one of the above-mentioned operation methods based on an artificial intelligence model when executed by a processor.
[0009] The present invention also provides a computer program product including a computer program, and the computer program implements the steps of any one of the above-mentioned operation methods based on an artificial intelligence model when executed by a processor.
[0010] The beneficial effects of the present invention are as follows: In the operation method based on an artificial intelligence model provided by the present invention, before performing a target operation, the first matrix and the second matrix are stored in a block manner, and a large-scale matrix is split into multiple small block matrices. This block storage strategy makes the data reading more efficient during the calculation process, reduces the address jump calculation, shortens the data acquisition delay, thereby improving the operation efficiency and saving hardware resources. The present invention also discloses an operation device based on an artificial intelligence model, an electronic device, a computer-readable storage medium, and a computer program product, which can also achieve the above technical effects.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0013] Figure 1 FIG. is a flowchart of an operation method based on an artificial intelligence model shown according to an exemplary embodiment.
[0014] Figure 2 FIG. is a schematic diagram of the block storage of a weight matrix shown according to an exemplary embodiment.
[0015] Figure 3 FIG. is a schematic diagram of a direct memory read instruction shown according to an exemplary embodiment.
[0016] Figure 4 FIG. is a schematic diagram of an operation instruction shown according to an exemplary embodiment.
[0017] Figure 5 Schematic diagram of a configuration register instruction shown according to an exemplary embodiment.
[0018] Figure 6 Schematic diagram of a loop start instruction shown according to an exemplary embodiment.
[0019] Figure 7 Schematic diagram of a loop end instruction shown according to an exemplary embodiment.
[0020] Figure 8 Schematic diagram of a general register calculation instruction shown according to an exemplary embodiment.
[0021] Figure 9 Schematic diagram of a general register conditional judgment instruction shown according to an exemplary embodiment.
[0022] Figure 10 Flowchart of another operation method based on an artificial intelligence model shown according to an exemplary embodiment.
[0023] Figure 11 Schematic diagram of the first matrix division method shown according to an exemplary embodiment.
[0024] Figure 12 Schematic diagram of the second matrix division method shown according to an exemplary embodiment.
[0025] Figure 13 Structural diagram of an operation device based on an artificial intelligence model shown according to an exemplary embodiment.
[0026] Figure 14 Structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0028] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0029] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] An embodiment of the present invention provides a method for running an artificial intelligence model. The method will be described in detail in combination with the execution flow of the method for running an artificial intelligence model.
[0031] See Figure 1 , a flowchart of a method for running an artificial intelligence model shown according to an exemplary embodiment, as Figure 1 shown, includes: S101: Obtain the source data of the target operation; wherein, the source data includes a first matrix and a second matrix.
[0032] In specific implementation, before performing matrix operations, it is first necessary to obtain the source data required for the target operation. These source data usually include two matrices, namely the first matrix and the second matrix. These two matrices are the basic inputs for matrix operations. For example, in the inference process of an artificial intelligence model, the first matrix may represent the input feature data, and the second matrix may represent the weight parameters of the model.
[0033] S102: Determine whether the target operation is a block calculation; if so, proceed to S103.
[0034] In specific implementation, after obtaining the source data, it is necessary to determine whether the target operation requires block calculation. Block calculation is an optimization strategy used to process large-scale matrix operations. By decomposing the matrix into multiple small block matrices, the calculation efficiency can be improved. The basis for determining whether block calculation is required is usually the size of the matrix and the limitations of hardware resources. If the matrix scale is large, exceeding the capacity of the hardware cache, or the hardware resources are limited, then block calculation is particularly important. By block calculation, the latency of data transmission can be reduced, the utilization rate of the cache can be improved, and thus the overall operation performance can be enhanced. If the target operation does not require block calculation, the overall operation can be directly performed; if block calculation is required, then proceed to the next step to determine the block size.
[0035] S103: Determine the first block size of the first matrix and the second block size of the second matrix.
[0036] In a specific implementation, after determining that the target operation requires block calculation, the block sizes of the first matrix and the second matrix need to be determined next. The choice of block size has an important impact on the operation efficiency and the utilization of hardware resources. Determining the block size requires comprehensive consideration of the size of the matrix, the capacity of the hardware cache, and the complexity of the calculation logic. For example, if the hardware cache is small, a smaller block matrix should be selected as the block size to ensure that each block matrix can be fully stored in the cache, thereby reducing the number of accesses to the main memory. At the same time, the block size should be as close as possible to the size of the cache to make full use of the cache resources. In addition, the complexity of the calculation logic needs to be considered when determining the block size to ensure that the matrices after block division can be efficiently operated on.
[0037] S104: Store the first matrix in block form as multiple first block matrices according to the first block size, and store the second matrix in block form as multiple second block matrices according to the second block size.
[0038] In a specific implementation, after determining the block size, the first matrix is stored in block form as multiple first block matrices according to the determined first block size, and the second matrix is stored in block form as multiple second block matrices according to the second block size. The purpose of block storage is to decompose a large-scale matrix into multiple small block matrices so that data can be read and processed more efficiently during the calculation process. Block storage can reduce the latency of data transmission, improve the utilization rate of the cache, and thus enhance the overall operation performance. For example, in matrix multiplication operations, through block storage, the multiplication operation of a large-scale matrix can be decomposed into the multiplication operations of multiple small block matrices, and the multiplication operations of each small block matrix can be efficiently completed in the cache, thereby reducing the number of accesses to the main memory and increasing the operation speed.
[0039] S105: During the execution of the target operation, read multiple first block matrices and multiple second block matrices in sequence to perform the target operation to obtain the target operation result.
[0040] In a specific implementation, during the execution of the target operation, multiple first block matrices and multiple second block matrices are sequentially read, and operations are performed according to a predetermined calculation logic, and finally a target operation result is obtained. The purpose of sequentially reading the block matrices is to ensure that each block matrix can be efficiently operated in the cache, thereby reducing the latency of data transmission. By means of block storage and on-demand reading, cache resources can be better utilized and the operation efficiency can be improved. For example, in matrix multiplication operations, the first block matrix and the second block matrix after being divided into blocks are sequentially read, the multiplication operation of the block matrices is performed, and then the results are accumulated to obtain the final result. This on-demand reading method not only improves the data reading efficiency but also reduces the waste of hardware resources, making the entire operation process more efficient.
[0041] For example, when calculating the feature matrix A × weight matrix B, the weight matrix B can be stored in advance in blocks according to the scale of the block matrix required in the actual calculation process. In this way, during the calculation, when reading data, there is no need to skip the jump data length of each row after reading part of the data of each row and then read the data of the next row. The advance block storage of the weight matrix can enable continuous reading of data during actual calculation, thereby reducing the address jump calculation during reading and shortening the time delay caused by obtaining data. As Figure 2 shown, assuming that the weight data is continuously stored in memory, then during block calculation, when reading the second row of data B1, N - n data need to be skipped and then the next row is read, which will increase the time delay of reading. However, if the data is stored according to the calculation block method, as Figure 2 shown on the right, the block data B1 is continuously stored in memory, and there is no need to calculate the jump length, thus reducing the time delay of obtaining data. In actual calculation, the required block data can be stored in advance for the calculation and acquisition in the next step. After the calculation of this step is completed, it can be decided whether to store in blocks according to the actual calculation of the next step. If the next step is still block calculation, then the calculation result of this step can be stored according to the actual block scale; if the next step does not require block calculation but requires calculation of the entire large matrix, such as softmax, layrenorm (layer normalization), activation function calculation, then the calculation result of this step is stored as the original matrix scale, as Figure 2 shown on the left.
[0042] The operation method based on the artificial intelligence model provided by the embodiment of the present invention divides and stores the first matrix and the second matrix before executing the target operation, and splits the large-scale matrix into multiple small block matrices. This block storage strategy makes the data reading more efficient during the calculation process, reduces the address jump calculation, shortens the time delay of obtaining data, and thus improves the operation efficiency and saves hardware resources.
[0043] Based on the above embodiments, it further includes: constructing an instruction architecture to utilize instructions in the instruction architecture to perform matrix multiplication and / or non-linear function operations during the operation of the artificial intelligence model; wherein, the instruction architecture includes direct memory access instructions and control instructions, and the control instructions include any one or several combinations of arithmetic instructions, configuration register instructions, loop instructions, general register calculation instructions, and general register conditional judgment instructions.
[0044] The artificial intelligence model in this embodiment can be a Transformer model (a deep learning architecture based on the self-attention mechanism). During the inference process of the artificial intelligence model, the main computational blocks involved are multi-layer perceptrons, layer normalization, softmax, and activation function calculations. At a finer granularity, the operator-level calculations include matrix multiplication by matrix, matrix multiplication by vector, matrix transpose, matrix left and right bisection and position swapping with right part negation operation, vector summation, vector exp (exponential operation), and vector activation function calculation. The most basic operations at the bottom layer include scalar reciprocal calculation, scalar addition, multiplication, square root calculation, exp calculation, and activation function calculation. This embodiment summarizes and analyzes the above calculation types, induces common characteristics, and constructs an instruction architecture to achieve the goal of efficient execution with a simple design of instructions.
[0045] For the calculation types related to matrices and vectors, they can be summarized as data acquisition, calculation, and writing results. Analyzing at the instruction level, the main steps include: reading data, completion of fetching data, execution of calculations (involving loops and conditional judgments), completion of calculations, writing back results, and completion of writing data, etc. Design instructions for these steps to construct the instruction architecture. The instruction architecture can include arithmetic instructions, configuration register instructions, loop instructions, general register calculation instructions, general register conditional judgment instructions, etc. Using the instructions therein, matrix multiplication operations and non-linear function operations during the operation of the artificial intelligence model can be performed, for example, function operations such as layernorm, softmax, gelu (Gaussian Error Linear Unit).
[0046] As a feasible implementation, the direct memory access instructions include direct memory read instructions and direct memory write instructions; the direct memory read instructions include any one or any combination of the operator field, address field, data length field, line length field, skip length field, data type field, and maximum burst transfer length field. The operator field is used to indicate the direct memory access module for reading data, the address field is used to indicate the starting address of the data to be read, the data length field is used to indicate the data length of the data to be read, the line length field is used to indicate the length of each line of data when reading data in chunks, the skip length field is used to indicate the length of the data to be skipped at the end of each line when reading data in chunks, the data type field is used to indicate the data type of the data to be read, and the maximum burst transfer length field is used to indicate the maximum number of specified data scale units transferred in a single burst operation during the data reading process; the direct memory write instructions include any one or any combination of the operator field, address field, data length field, line length field, skip length field, direct memory access module management field, and maximum burst transfer length field. The operator field is used to indicate the direct memory access module for writing back data, the address field is used to indicate the starting address of the data to be written back, the data length field is used to indicate the data length of the data to be written back, the line length field is used to indicate the length of each line of data when writing back data in chunks, the skip length field is used to indicate the length of the data to be skipped at the end of each line when writing back data in chunks, the direct memory access module management field is used to indicate the status of the direct memory access module for writing back data, and the maximum burst transfer length field is used to indicate the maximum number of specified data scale units transferred in a single burst operation during the data writing back process.
[0047] The direct memory read instructions are used to implement data reading and write the data into the cache. For example, the source data of the target operation is read using the direct memory read instructions. The direct memory write instructions are used to implement result write-back and write the result back to the cache or DDR (Double Data Rate SDRAM). For example, the target operation result is written into the memory using the direct memory write instructions.
[0048] The direct memory read instruction contains multiple fields that together define the specific parameters and behavior of the data read operation. The operator field is used to indicate the direct memory access module to be used when reading data. It clearly specifies, through a specific number or identifier, which DMA (Direct Memory Access) module will perform the data read operation. The address field indicates the starting address of the data to be read, which is the specific location of the data in memory. The data length field defines the length of the data to be read. It specifies how many bytes or data units need to be continuously read starting from the starting address. The line length field is used when reading data in chunks. It indicates the length of each line of data, which is particularly important for processing matrices or two-dimensional data structures. The skip length field is used to indicate the length of the data to be skipped at the end of each line when reading data in chunks, which helps to process non-contiguous stored data. The data type field indicates the data type of the data to be read, which is crucial for ensuring the correct interpretation and processing of the data. The maximum burst transfer length field is used to indicate the maximum number of specified data size units that can be transferred in a single burst operation during the data read process, which helps to optimize the efficiency of data transfer. For example, as Figure 3 shown, the direct memory read instruction includes a prefix field, an operator field, an address field, a data length field, a line length field, a skip length field, a data type field, and a maximum burst transfer length field.
[0049] The direct memory write instruction also contains similar fields, but there are some specific differences. The operator field is used to indicate the direct memory access module to be used when writing back data. The address field indicates the starting address of the data to be written back. The data length field defines the length of the data to be written back. The line length field and the skip length field are used when writing back data in chunks. They respectively indicate the length of each line of data and the length of the data to be skipped at the end of each line. The direct memory access module management field is used to indicate the status of the direct memory access module to be used when writing back data, which may include the current status, availability, or other status information related to module management of the module. The maximum burst transfer length field is used to indicate the maximum number of specified data size units that can be transferred in a single burst operation during the data write-back process, which helps to optimize the efficiency of data write-back.
[0050] Through the combination of these fields, the direct memory access instruction can precisely control the data read and write-back operations, ensuring the efficiency and accuracy of data transfer. The design of these fields takes into account the complexity and diversity of data storage, enabling the DMA module to adapt to various different application scenarios and data structures.
[0051] As a feasible implementation, the arithmetic instruction includes any one or a combination of several of an operator field, an arithmetic type field, and an identifier field. The operator field is used to indicate an arithmetic operation, the arithmetic type field is used to indicate the arithmetic type, and the identifier field is used to indicate the arithmetic attribute.
[0052] The arithmetic instruction is used to initiate matrix multiplication, vector, and scalar operations, such as executing a target operation using the arithmetic instruction. The arithmetic instruction may include an operator field, an arithmetic type field, an identifier field, etc. Among them, the operator field is used to indicate an arithmetic operation. It clearly indicates that the current instruction is an arithmetic instruction through a specific encoding or identifier, enabling the hardware to recognize and execute the corresponding arithmetic operation. The arithmetic type field is used to indicate the specific type of arithmetic, such as addition, subtraction, multiplication, or division, etc. It provides detailed information for the hardware on what arithmetic operation to perform. The identifier field is used to indicate the arithmetic attribute, which may include additional information such as the priority of the arithmetic, whether special processing is required, etc., enabling the arithmetic to be optimized or adjusted according to specific attributes. For example, as Figure 4 shown, the arithmetic instruction includes a prefix field, an operator field, an arithmetic type field, an identifier field, and a reserved field.
[0053] As a feasible implementation, the configuration register instruction includes any one or a combination of several of an operator field, a register address field, and a register data field. The operator field is used to indicate a configuration register operation, the register address field is used to indicate the address of the register to be configured, and the register data field is used to indicate the data to be written into the register.
[0054] The configuration register instruction is used to configure data for a specified register address, which is widely used in computing. For example, by configuring the register value, it is used to indicate whether the data reading process is completed, whether the computing process is executed, whether the data writing process is completed, etc. The configuration register instruction may include an operator field, a register address field, a register data field, etc. Among them, the operator field is used to indicate a configuration register operation. It clearly identifies the current instruction as a configuration register instruction, enabling the hardware to recognize and execute the corresponding configuration operation. The register address field is used to indicate the address of the register to be configured. It provides the location information of the specific register to be configured, ensuring that the data can be written into the correct register. The register data field is used to indicate the data to be written into the register. It provides the specific data value to be stored in the register, thus completing the configuration of the register. For example, as Figure 5 shown, the configuration register instruction includes a prefix field, an operator field, a register address field, a register data field, and a reserved field.
[0055] As a feasible implementation, the loop instruction includes a loop start instruction and a loop end instruction; the loop start instruction includes an operator field, and the operator field is used to indicate the start of the loop; the loop end instruction includes any one or any combination of an operator field, a register address field, and a loop count field. The operator field is used to indicate the end of the loop, the register address field is used to indicate the address of the loop control register, and the loop count field is used to indicate the number of iterations of the loop.
[0056] Among them, the loop start instruction includes an operator field, and this operator field is used to indicate the start of the loop. It notifies the start of the loop operation to the hardware through specific encoding or identifiers, so that the hardware can initialize the loop control logic and prepare to execute the instructions in the loop body. For example, as Figure 6 shown, the loop start instruction includes a prefix field, an operator field, and a reserved field. The loop end instruction includes any one or any combination of an operator field, a register address field, and a loop count field. The operator field is used to indicate the end of the loop, and it clearly identifies the end of the loop operation, enabling the hardware to terminate the loop control logic. The register address field is used to indicate the address of the loop control register, and it provides the register location for storing loop control information (such as a loop counter). The loop count field is used to indicate the number of iterations of the loop, and it provides the specific number of times the loop needs to be executed, so that the hardware can control the execution of the loop according to this number. For example, as Figure 7 shown, the loop end instruction includes a prefix field, an operator field, a register address field, a loop count field, and a reserved field.
[0057] As a feasible implementation, the general register calculation instruction includes any one or any combination of an operator field, a general register operator field, a general register address field, and an immediate value field. The operator field is used to indicate a general register calculation operation, the general register operator field is used to indicate the type of calculation operation, the general register address field is used to indicate the address of the general register participating in the calculation, and the immediate value field is used to indicate the immediate value used in the calculation.
[0058] In a specific implementation, in order to ensure the consistency of generated instructions in loop calculations, it is necessary to configure general-purpose registers for the changing values in the loop and design general-purpose register calculation instructions to obtain the specific values required in the loop. The general-purpose register calculation instructions can include an operator field, a general-purpose register operator field, a general-purpose register address field, an immediate number field, etc. Among them, the operator field is used to indicate the general-purpose register calculation operation, which clearly identifies the current instruction as a general-purpose register calculation instruction, enabling the hardware to recognize and execute the corresponding calculation operation. The general-purpose register operator field is used to indicate the type of calculation operation, such as addition, subtraction, shift, etc., which provides the hardware with detailed information on what calculation operation to perform. The general-purpose register address field is used to indicate operating on the register values of the source addresses (saddr0, saddr1) and writing data to the destination register address (daddr). The immediate number field is used to indicate the immediate value used in the calculation, which provides the specific value directly used in the calculation, enabling the calculation to operate using the data or immediate number in the register as needed. For example, as Figure 8 shown, the general-purpose register calculation instruction includes a prefix field, an operator field, a general-purpose register operator field, a general-purpose register address field, and an immediate number field.
[0059] As a feasible implementation, the general-purpose register conditional judgment instruction, the general-purpose register conditional judgment instruction includes a combination of any one or several of an operator field, a general-purpose register operator field, a conditional operator, a general-purpose register address field, and an immediate number field. The operator field is used to indicate the general-purpose register conditional judgment operation, the general-purpose register operator field is used to indicate the type of conditional judgment operation, the conditional operator is used to indicate the type of conditional judgment, the general-purpose register address field is used to indicate the address of the general-purpose register participating in the conditional judgment, and the immediate number field is used to indicate the immediate value used in the conditional judgment.
[0060] In a specific implementation, the conditional logic in the calculation process can be realized by a general register conditional judgment instruction to execute different tasks when different conditions are met. The general register conditional judgment instruction may include an operator field, a general register operator field, a conditional symbol, a general register address field, an immediate number field, etc. Among them, the operator field is used to indicate the general register conditional judgment operation, which clearly identifies the current instruction as a conditional judgment instruction, enabling the hardware to recognize and execute the corresponding conditional judgment operation. The general register operator field is used to indicate the operation type of the conditional judgment, such as if, elseif, else, end, ifi, elseifi, etc., which provides detailed information for the hardware on what conditional judgment operation to execute. The conditional symbol is used to indicate the conditional judgment type, such as equal to, greater than, less than, etc., which defines the specific conditional judgment logic. The general register address field is used to indicate the address of the general register participating in the conditional judgment, which provides the specific position information of the register participating in the conditional judgment. The immediate number field is used to indicate the immediate value used in the conditional judgment, which provides the specific value directly used in the conditional judgment process, so that the conditional judgment can operate according to the data or immediate number in the register as needed. For example, as Figure 9 shown, the general register conditional judgment instruction includes a prefix field, an operator field, a general register operator field, a conditional symbol, a general register address field, and an immediate number field.
[0061] An embodiment of the present invention discloses a method for running an artificial intelligence model. Compared with the previous embodiment, this embodiment further describes and optimizes the technical solution. Specifically: See Figure 10 , a flowchart of another method for running an artificial intelligence model shown according to an exemplary embodiment, as Figure 10 shown, includes: S201: Obtain the source data of the target operation; wherein, the source data includes a first matrix and a second matrix.
[0062] S202: Determine whether the target operation is a block calculation; if so, enter S203.
[0063] S203: Determine multiple candidate block methods; wherein, each candidate block method includes a first candidate block scale of the first matrix and a second candidate block scale of the second matrix; wherein, both the first candidate block scale and the second candidate block scale are smaller than the cache size.
[0064] In a specific implementation, in the process of determining the first block size of the first matrix and the second block size of the second matrix, it is first necessary to determine multiple candidate block division methods. Each candidate block division method includes the first candidate block size of the first matrix and the second candidate block size of the second matrix. The setting of these candidate block sizes is based on the limitation of the cache size to ensure that the size of each block is smaller than the capacity of the cache. Such a design is to optimize the data reading efficiency because when the block size fits the cache, the number of accesses to the main memory can be reduced, thereby increasing the data processing speed.
[0065] S204: Select a target block division method from multiple candidate block division methods according to the hardware accelerator resources used by the artificial intelligence model, the number of reads of the block matrix, and the calculation amount of block calculation, and determine the first block size of the first matrix and the second block size of the second matrix according to the target block division method.
[0066] In a specific implementation, after determining multiple candidate block division methods, the next step is to select the most suitable target block division method from these candidate block division methods according to the hardware accelerator resources used by the artificial intelligence model, the number of reads of the block matrix, and the calculation amount of block calculation. The hardware accelerator can be designed based on a systolic array. The consideration of hardware accelerator resources involves the performance and limitations of the hardware, such as the size of the cache, the number of processing units, etc. The number of reads of the block matrix is a key factor because fewer reads usually mean higher efficiency. And the calculation amount of block calculation involves the number of operations that need to be performed for each block during the calculation process. A smaller calculation amount usually means faster processing speed. After comprehensively considering these factors, the selected target block division method will determine the first block size of the first matrix and the second block size of the second matrix. This process is a process of weighing and optimization, aiming to find a block division strategy that achieves the best balance among hardware resource utilization, data reading efficiency, and calculation efficiency. The block sizes determined in this way can ensure that the hardware resources are fully utilized during the execution of matrix operations, and both the data reading and calculation processes are as efficient as possible, thereby improving the performance of the entire system when processing the artificial intelligence model.
[0067] As a feasible implementation, selecting a target block division method from multiple candidate block division methods according to the hardware accelerator resources used by the artificial intelligence model, the number of reads of the block matrix, and the calculation amount of block calculation includes: when the hardware accelerator used by the artificial intelligence model supports parallel data reading and calculation, select the candidate block division method with the smallest calculation amount of block calculation as the target block division method; when the hardware accelerator used by the artificial intelligence model does not support parallel data reading and calculation, select a target block division method from multiple candidate block division methods according to the number of reads of the block matrix and the calculation amount of block calculation.
[0068] In specific implementation, during the process of selecting the target chunking method, different strategies are adopted according to the different characteristics of the hardware accelerator used by the artificial intelligence model. When the hardware accelerator supports parallel reading and computing of data, that is, the reading and computing operations of data can be carried out simultaneously, this parallel processing ability can effectively reduce the total processing time. In this case, the candidate chunking method with the smallest computational amount of chunking calculation is selected as the target chunking method, that is, the candidate chunking method with the smallest total computational amount of multiplication calculation and accumulation calculation is selected as the target chunking method. This is because when data reading and computing can be carried out in parallel, the computational amount becomes the main factor affecting the processing speed, and the computing time can cover the reading time. A smaller computational amount means that the computing task can be completed faster, thus improving the efficiency of the entire system.
[0069] On the contrary, when the hardware accelerator does not support parallel reading and computing of data, that is, the data reading and computing operations cannot be carried out simultaneously, it is necessary to comprehensively consider the reading times of the block matrix and the computational amount of chunking calculation to select the target chunking method. In this case, the reading times become particularly important because each reading operation will take up a certain amount of time, and too many reading times will significantly increase the total processing time. Therefore, it is necessary to find a balance between reducing the reading times and controlling the computational amount. By conducting a detailed analysis and comparison of different candidate chunking methods, to determine which method can minimize the reading times of the block matrix while ensuring a reasonable computational amount. In this way, the chunking strategy can be optimized under limited hardware resources to achieve the best performance.
[0070] As a feasible implementation method, the target chunking method is selected from multiple candidate chunking methods according to the reading times of the block matrix and the computational amount of chunking calculation, including: selecting the first candidate chunking method and the second candidate chunking method from multiple candidate chunking methods; judging whether the ratio between the computational amount of chunking calculation corresponding to the first candidate chunking method and the computational amount of chunking calculation corresponding to the second candidate chunking method is less than or equal to a preset value; if so, deleting the candidate chunking method with the largest reading times of the block matrix among the first candidate chunking method and the second candidate chunking method; if not, deleting the candidate chunking method with the largest computational amount among the first candidate chunking method and the second candidate chunking method; re-executing the step of selecting the first candidate chunking method and the second candidate chunking method from multiple candidate chunking methods until there is only one candidate chunking method left, and taking the remaining one candidate chunking method as the target chunking method.
[0071] In a specific implementation, during the process of selecting the target chunking method, first, two specific candidate chunking methods are selected from multiple candidate chunking methods, namely the first candidate chunking method and the second candidate chunking method. This selection process is based on a preliminary screening of different chunking methods, aiming to further compare and evaluate the performance metrics of these chunking methods. Next, it is necessary to determine whether the ratio of the computational amount of the chunk calculation corresponding to the first candidate chunking method to the computational amount of the chunk calculation corresponding to the second candidate chunking method is less than or equal to a preset value, that is, to determine whether the ratio of the total computational amount of the multiplication calculation amount and the accumulation calculation amount corresponding to the first candidate chunking method to the total computational amount of the multiplication calculation amount and the accumulation calculation amount corresponding to the second candidate chunking method is less than or equal to the preset value. This preset value is a key threshold used to measure whether the relative difference in computational amount between the two chunking methods is within an acceptable range. For example, 10 can be selected. If the ratio is less than or equal to the preset value, it means that the two chunking methods are relatively close in computational amount. At this time, the number of reads becomes the main decision-making factor. Therefore, the candidate chunking method with the largest number of reads of the block matrix in the first candidate chunking method and the second candidate chunking method will be deleted. This is because, in the case of similar computational amounts, the chunking method with fewer reads usually has higher efficiency. If the ratio is greater than the preset value, it means that there is a significant difference in computational amount between the two chunking methods. At this time, the computational amount becomes the main decision-making factor. Therefore, the candidate chunking method with the largest computational amount in the first candidate chunking method and the second candidate chunking method will be deleted, that is, the candidate chunking method with the largest total computational amount of the multiplication calculation amount and the accumulation calculation amount in the first candidate chunking method and the second candidate chunking method will be deleted. This is because, in the case of similar numbers of reads, the chunking method with a smaller computational amount usually can complete the calculation task faster. After deleting a candidate chunking method, the steps of selecting the first candidate chunking method and the second candidate chunking method are re-executed. This process will be repeated continuously, comparing and deleting one candidate chunking method each time, until only one candidate chunking method remains. Finally, the remaining one candidate chunking method is used as the target chunking method. This step-by-step screening and comparison method can ensure that, considering both the computational amount and the number of reads, the optimal chunking method is selected, so as to achieve efficient matrix operations under the condition of limited hardware resources.
[0072] For example, for calculating A×B, the size of matrix A is 32×896, the size of matrix B is 896×4864, and the cache size is 256×256. The first candidate chunking method: Matrix A is divided into block matrices with a size of 32×128 , matrix B is divided into block matrices with a size of 128×256 , so that the result matrix C is composed of block matrices , and , as Figure 11As shown. In the above-mentioned block matrix multiplication calculation, each time the block result matrix is obtained matrix A needs to be read once, and the block matrix needs to be read once. Therefore, for the overall calculation result matrix A, matrix A needs to be read 19 times, and matrix B needs to be read once. The total multiplication calculation amount is 32×128×256×7×19 = 139460608. For a block matrix with a size of 32×256, it needs to be accumulated 6×19 times. Therefore, the total accumulation calculation amount is 32×256×6×19 = 933888. The second candidate block division method: Matrix A is divided into block matrices with a size of 32×8 , and matrix B is divided into block matrices with a size of 8×4864 , so that the result matrix C is obtained from , as shown in Figure 12 . In the above-mentioned block matrix multiplication calculation, for the overall calculation result matrix C, matrix A needs to be read once, and matrix B needs to be read once. The total multiplication calculation amount is 32×8×4864×112 = 139460608. Each time the intermediate result matrix with a size of 32×4864 is obtained, it needs to be accumulated 111 times. Therefore, the total accumulation calculation amount is 32×4864×111 = 17276928. If the hardware accelerator used by the artificial intelligence model supports parallel data reading and calculation, then the first candidate block division method with less calculation amount is selected. If the hardware accelerator used by the artificial intelligence model does not support parallel data reading and calculation, the ratio between the calculation amount of the second candidate block division method (17276928 + 139460608) and the calculation amount of the first candidate block division method (933888 + 139460608) is approximately 1.1164, and this value is less than the preset value of 10. Then the second candidate block division method with fewer reading times is selected.
[0073] S205: Store the first matrix in multiple first block matrices according to the first block size, and store the second matrix in multiple second block matrices according to the second block size.
[0074] S206: During the execution of the target operation, read multiple first block matrices and multiple second block matrices in sequence to implement the target operation to obtain the target operation result.
[0075] It can be seen that in this embodiment, the block division methods of the first matrix and the second matrix are flexibly determined according to the hardware accelerator resources used by the artificial intelligence model, the reading times of the block matrices, and the calculation amount of the block calculation. On the premise of meeting the hardware resources provided by the hardware accelerator, the data reading times and the calculation amount are effectively reduced, and the operation efficiency is improved.
[0076] The following introduces an operation device based on an artificial intelligence model provided by an embodiment of the present invention. The operation device based on an artificial intelligence model described below can be referred to in mutual reference with the operation method based on an artificial intelligence model described above.
[0077] See Figure 13 , a structural diagram of an operation device based on an artificial intelligence model shown according to an exemplary embodiment, as Figure 13 shown, includes: An acquisition module 100, configured to acquire source data of a target operation; wherein, the source data includes a first matrix and a second matrix; A determination module 200, configured to determine whether the target operation is a block calculation; if so, determine a first block size of the first matrix and a second block size of the second matrix; A block storage module 300, configured to block and store the first matrix as a plurality of first block matrices according to the first block size, and block and store the second matrix as a plurality of second block matrices according to the second block size; An operation module 400, configured to sequentially read a plurality of first block matrices and a plurality of second block matrices during the execution of the target operation to implement the target operation to obtain a target operation result.
[0078] For the operation device based on an artificial intelligence model provided by an embodiment of the present invention, before executing the target operation, the first matrix and the second matrix are block-stored, and the large-scale matrix is split into a plurality of small block matrices. This block storage strategy makes the data reading more efficient during the calculation process, reduces the address jump calculation, shortens the data acquisition delay, thereby improving the operation efficiency and saving hardware resources.
[0079] On the basis of the above embodiment, as a preferred embodiment, it further includes: A construction module, configured to construct an instruction architecture to use instructions in the instruction architecture to execute matrix multiplication and / or non-linear function operations during the operation of the artificial intelligence model; wherein, the instruction architecture includes direct memory access instructions and control instructions, and the control instructions include any one or a combination of several of operation instructions, configuration register instructions, loop instructions, general register calculation instructions, and general register conditional judgment instructions.
[0080] On the basis of the above embodiment, as a preferred embodiment, the direct memory access instructions include direct memory read instructions and direct memory write instructions; The direct memory read instruction includes any one or a combination of several of an operator field, an address field, a data length field, a line length field, a skip length field, a data type field, and a maximum burst transfer length field. The operator field is used to indicate the direct memory access module for reading data. The address field is used to indicate the starting address of the data to be read. The data length field is used to indicate the data length of the data to be read. The line length field is used to indicate the length of each line of data when reading data in chunks. The skip length field is used to indicate the length of data to be skipped at the end of each line when reading data in chunks. The data type field is used to indicate the data type of the data to be read. The maximum burst transfer length field is used to indicate the maximum number of specified data scale units transferred in a single burst operation during the data reading process; The direct memory write instruction includes any one or a combination of several of an operator field, an address field, a data length field, a line length field, a skip length field, a direct memory access module management field, and a maximum burst transfer length field. The operator field is used to indicate the direct memory access module for writing back data. The address field is used to indicate the starting address of the data to be written back. The data length field is used to indicate the data length of the data to be written back. The line length field is used to indicate the length of each line of data when writing back data in chunks. The skip length field is used to indicate the length of data to be skipped at the end of each line when writing back data in chunks. The direct memory access module management field is used to indicate the status of the direct memory access module for writing back data. The maximum burst transfer length field is used to indicate the maximum number of specified data scale units transferred in a single burst operation during the data writing back process.
[0081] Based on the above embodiments, as a preferred implementation manner, the obtaining module 100 is specifically configured to: read the source data of the target operation by using the direct memory read instruction; Correspondingly, the apparatus further includes: A writing module, configured to write the target operation result into the memory by using the direct memory write instruction.
[0082] Based on the above embodiments, as a preferred implementation manner, the operation instruction includes any one or a combination of several of an operator field, an operation type field, and a flag field. The operator field is used to indicate an operation to be performed. The operation type field is used to indicate the operation type. The flag field is used to indicate the operation attribute.
[0083] Based on the above embodiments, as a preferred implementation manner, the operation module 400 is specifically configured to: execute the target operation by using the operation instruction.
[0084] Based on the above embodiments, as a preferred implementation, the configuration register instruction includes any one or any combination of an operator field, a register address field, and a register data field. The operator field is used to indicate a configuration register operation, the register address field is used to indicate the address of the register to be configured, and the register data field is used to indicate the data to be written into the register.
[0085] Based on the above embodiments, as a preferred implementation, the loop instruction includes a loop start instruction and a loop end instruction; The loop start instruction includes an operator field, and the operator field is used to indicate the start of the loop; The loop end instruction includes any one or any combination of an operator field, a register address field, and a loop count field. The operator field is used to indicate the end of the loop, the register address field is used to indicate the address of the loop control register, and the loop count field is used to indicate the number of iterations of the loop.
[0086] Based on the above embodiments, as a preferred implementation, the general register calculation instruction includes any one or any combination of an operator field, a general register operator field, a general register address field, and an immediate value field. The operator field is used to indicate a general register calculation operation, the general register operator field is used to indicate the type of calculation operation, the general register address field is used to indicate the address of the general register participating in the calculation, and the immediate value field is used to indicate the immediate value used in the calculation.
[0087] Based on the above embodiments, as a preferred implementation, the general register conditional judgment instruction includes any one or any combination of an operator field, a general register operator field, a conditional operator, a general register address field, and an immediate value field. The operator field is used to indicate a general register conditional judgment operation, the general register operator field is used to indicate the type of conditional judgment operation, the conditional operator is used to indicate the type of conditional judgment, the general register address field is used to indicate the address of the general register participating in the conditional judgment, and the immediate value field is used to indicate the immediate value used in the conditional judgment.
[0088] Based on the above embodiments, as a preferred implementation, the determination module 200 includes: A determination unit, configured to determine multiple candidate partitioning methods; wherein, each candidate partitioning method includes a first candidate partitioning scale of the first matrix and a second candidate partitioning scale of the second matrix; wherein, both the first candidate partitioning scale and the second candidate partitioning scale are smaller than the cache size; A selection unit is configured to select a target chunking method from multiple candidate chunking methods according to the hardware accelerator resources used by the artificial intelligence model, the number of reads of the block matrix, and the amount of calculation of the chunked calculation, and determine the first chunking scale of the first matrix and the second chunking scale of the second matrix according to the target chunking method.
[0089] Based on the above embodiments, as a preferred embodiment, the selection unit is specifically configured to: when the hardware accelerator used by the artificial intelligence model supports parallel reading of data and calculation, select the candidate chunking method with the smallest amount of calculation of the chunked calculation as the target chunking method.
[0090] Based on the above embodiments, as a preferred embodiment, the selection unit is specifically configured to: when the hardware accelerator used by the artificial intelligence model supports parallel reading of data and calculation, select the candidate chunking method with the smallest total amount of calculation of multiplication calculation amount and accumulation calculation amount as the target chunking method.
[0091] Based on the above embodiments, as a preferred embodiment, the selection unit is specifically configured to: when the hardware accelerator used by the artificial intelligence model does not support parallel reading of data and calculation, select a target chunking method from multiple candidate chunking methods according to the number of reads of the block matrix and the amount of calculation of the chunked calculation.
[0092] Based on the above embodiments, as a preferred embodiment, the selection unit is specifically configured to: when the hardware accelerator used by the artificial intelligence model does not support parallel reading of data and calculation, select a first candidate chunking method and a second candidate chunking method from multiple candidate chunking methods; determine whether the ratio between the amount of calculation of the chunked calculation corresponding to the first candidate chunking method and the amount of calculation of the chunked calculation corresponding to the second candidate chunking method is less than or equal to a preset value; if so, delete the candidate chunking method with the largest number of reads of the block matrix among the first candidate chunking method and the second candidate chunking method; if not, delete the candidate chunking method with the largest amount of calculation among the first candidate chunking method and the second candidate chunking method; re-execute the step of selecting a first candidate chunking method and a second candidate chunking method from multiple candidate chunking methods until there is one remaining candidate chunking method, and use the remaining one candidate chunking method as the target chunking method.
[0093] Based on the above embodiments, as a preferred embodiment, the selection unit is specifically configured to: when the hardware accelerator used by the artificial intelligence model does not support parallel data reading and computing, select a first candidate partitioning method and a second candidate partitioning method from various candidate partitioning methods; determine whether the ratio between the total computation amount of the multiplication computation amount and the accumulation computation amount corresponding to the first candidate partitioning method and the total computation amount of the multiplication computation amount and the accumulation computation amount corresponding to the second candidate partitioning method is less than or equal to a preset value; if so, delete the candidate partitioning method with the largest number of block matrix reads among the first candidate partitioning method and the second candidate partitioning method; if not, delete the candidate partitioning method with the largest total computation amount of the multiplication computation amount and the accumulation computation amount among the first candidate partitioning method and the second candidate partitioning method; re-execute the step of selecting the first candidate partitioning method and the second candidate partitioning method from various candidate partitioning methods until there is one remaining candidate partitioning method, and use the remaining one candidate partitioning method as the target partitioning method.
[0094] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0095] Embodiments of the present invention further provide an electronic device, Figure 14 which is a structural diagram of an electronic device shown according to an exemplary embodiment, as Figure 14 shown, the electronic device includes: A communication interface 1, capable of interacting with other devices such as network devices for information. A processor 2, connected to the communication interface 1 to achieve information interaction with other devices, and when used to run a computer program, executes the operation method based on the artificial intelligence model provided by the above one or more technical solutions. The computer program is stored on a memory 3.
[0096] Of course, in actual applications, the various components in the electronic device are coupled together through a bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. The bus system 4 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 14 all the various buses are labeled as the bus system 4.
[0097] The memory 3 in the embodiments of the present invention is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device.
[0098] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 3 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0099] The method disclosed in the embodiments of the present invention above can be applied to the processor 2 or implemented by the processor 2. The processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 2 or the instructions in the form of software. The above-mentioned processor 2 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 3. The processor 2 reads the program in the memory 3 and combines its hardware to complete the steps of the foregoing method.
[0100] When the processor 2 executes the program, it implements the corresponding processes in each method of the embodiments of the present invention. For the sake of brevity, it will not be elaborated here.
[0101] The embodiments of the present invention also provide a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any of the above-mentioned embodiments of the operation method based on the artificial intelligence model when running.
[0102] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs, etc., various media that can store computer programs.
[0103] The embodiments of the present invention also provide a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above-mentioned embodiments of the operation method based on the artificial intelligence model.
[0104] The embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above-mentioned embodiments of the operation method based on the artificial intelligence model.
[0105] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0106] The above has introduced in detail a running system, method, device, equipment, medium, and product based on an artificial intelligence model provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A running method based on an artificial intelligence model, characterized in that Including: Obtain the source data of the target operation; wherein, the source data includes a first matrix and a second matrix; Determine whether the target operation is a block calculation; if so, determine the first block size of the first matrix and the second block size of the second matrix; Block-store the first matrix into a plurality of first block matrices according to the first block size, and block-store the second matrix into a plurality of second block matrices according to the second block size; During the execution of the target operation, sequentially read the plurality of first block matrices and the plurality of second block matrices to implement the target operation to obtain a target operation result.
2. The operation method based on the artificial intelligence model according to claim 1, wherein Also including: Construct an instruction architecture to use the instructions in the instruction architecture to execute matrix multiplication and / or non-linear function operations during the operation of the artificial intelligence model; wherein, the instruction architecture includes a direct memory access instruction and a control instruction, and the control instruction includes any one or a combination of several of an operation instruction, a configuration register instruction, a loop instruction, a general register calculation instruction, and a general register condition judgment instruction.
3. The operating method based on the artificial intelligence model according to claim 2, characterized in that, The direct memory access instruction includes a direct memory read instruction and a direct memory write instruction; The direct memory read instruction includes any one or a combination of several of an operator field, an address field, a data length field, a row length field, a skip length field, a data type field, and a maximum burst transfer length field. The operator field is used to indicate the direct memory access module for reading data, the address field is used to indicate the starting address of the data to be read, the data length field is used to indicate the data length of the data to be read, the row length field is used to indicate the length of each row of data when reading data in blocks, the skip length field is used to indicate the length of the data to be skipped at the end of each row when reading data in blocks, the data type field is used to indicate the data type of the data to be read, and the maximum burst transfer length field is used to indicate the maximum number of specified data scale units transferred in a single burst operation during the data reading process. The direct memory write instruction includes any one or a combination of several of an operator field, an address field, a data length field, a row length field, a skip length field, a direct memory access module management field, and a maximum burst transfer length field. The operator field is used to indicate the direct memory access module for writing back data, the address field is used to indicate the starting address of the data to be written back, the data length field is used to indicate the data length of the data to be written back, the row length field is used to indicate the length of each row of data when writing back data in blocks, the skip length field is used to indicate the length of the data to be skipped at the end of each row when writing back data in blocks, the direct memory access module management field is used to indicate the status of the direct memory access module for writing back data, and the maximum burst transfer length field is used to indicate the maximum number of specified data scale units transferred in a single burst operation during the data writing back process.
4. The operating method based on the artificial intelligence model according to claim 3, wherein Obtaining the source data of the target operation includes: Read the source data of the target operation by using the direct memory read instruction; After obtaining the target operation result, it also includes: Write the target operation result into the memory by using the direct memory write instruction.
5. The operating method based on the artificial intelligence model according to claim 2, characterized in that, The arithmetic instruction includes any one or any combination of an operator field, an arithmetic type field, and an identifier field. The operator field is used to indicate an arithmetic operation, the arithmetic type field is used to indicate the arithmetic type, and the identifier field is used to indicate the arithmetic attribute.
6. The operating method based on the artificial intelligence model according to claim 2, wherein Executing the target arithmetic operation includes: Executing the target arithmetic operation by using the arithmetic instruction.
7. The operation method based on the artificial intelligence model according to claim 2, wherein, The configuration register instruction includes any one or any combination of an operator field, a register address field, and a register data field. The operator field is used to indicate a configuration register operation, the register address field is used to indicate the address of the register to be configured, and the register data field is used to indicate the data to be written into the register.
8. The operating method based on the artificial intelligence model according to claim 2, wherein The loop instruction includes a loop start instruction and a loop end instruction; The loop start instruction includes an operator field, and the operator field is used to indicate the start of the loop; The loop end instruction includes any one or any combination of an operator field, a register address field, and a loop count field. The operator field is used to indicate the end of the loop, the register address field is used to indicate the address of the loop control register, and the loop count field is used to indicate the number of iterations of the loop.
9. The operating method based on the artificial intelligence model according to claim 2, wherein The general register calculation instruction includes any one or any combination of an operator field, a general register operator field, a general register address field, and an immediate value field. The operator field is used to indicate a general register calculation operation, the general register operator field is used to indicate the type of the calculation operation, the general register address field is used to indicate the address of the general register participating in the calculation, and the immediate value field is used to indicate the immediate value used in the calculation.
10. The operation method based on the artificial intelligence model according to claim 2, wherein The general register conditional judgment instruction includes any one or any combination of an operator field, a general register operator field, a conditional operator, a general register address field, and an immediate value field. The operator field is used to indicate a general register conditional judgment operation, the general register operator field is used to indicate the type of the conditional judgment operation, the conditional operator is used to indicate the type of the conditional judgment, the general register address field is used to indicate the address of the general register participating in the conditional judgment, and the immediate value field is used to indicate the immediate value used in the conditional judgment.
11. The operation method based on the artificial intelligence model according to claim 1, wherein Determining the first block size of the first matrix and the second block size of the second matrix includes: Determining a plurality of candidate block division methods; wherein, each of the candidate block division methods includes a first candidate block size of the first matrix and a second candidate block size of the second matrix; wherein, both the first candidate block size and the second candidate block size are smaller than the cache size; Selecting a target block division method from the plurality of candidate block division methods according to the hardware accelerator resources used by the artificial intelligence model, the number of reads of the block matrix, and the calculation amount of the block calculation, and determining the first block size of the first matrix and the second block size of the second matrix according to the target block division method.
12. The operation method based on an artificial intelligence model according to claim 11, wherein Select a target block division method from multiple candidate block division methods according to the hardware accelerator resources used by the artificial intelligence model, the number of times of reading the block matrix, and the amount of calculation of the block calculation, including: When the hardware accelerator used by the artificial intelligence model supports parallel reading of data and calculation, select the candidate block division method with the smallest amount of calculation of the block calculation as the target block division method.
13. The operation method based on the artificial intelligence model according to claim 12, wherein Selecting the candidate block division method with the smallest amount of calculation of the block calculation as the target block division method includes: Select the candidate block division method with the smallest total amount of calculation of the multiplication calculation amount and the accumulation calculation amount as the target block division method.
14. The operation method based on the artificial intelligence model according to claim 11, wherein Select a target block division method from multiple candidate block division methods according to the hardware accelerator resources used by the artificial intelligence model, the number of times of reading the block matrix, and the amount of calculation of the block calculation, including: When the hardware accelerator used by the artificial intelligence model does not support parallel reading of data and calculation, select a target block division method from multiple candidate block division methods according to the number of times of reading the block matrix and the amount of calculation of the block calculation.
15. The operation method based on the artificial intelligence model according to claim 14, wherein Selecting a target block division method from multiple candidate block division methods according to the number of times of reading the block matrix and the amount of calculation of the block calculation includes: Select a first candidate block division method and a second candidate block division method from multiple candidate block division methods; Judge whether the ratio between the amount of calculation of the block calculation corresponding to the first candidate block division method and the amount of calculation of the block calculation corresponding to the second candidate block division method is less than or equal to a preset value; If so, delete the candidate block division method with the largest number of times of reading the block matrix among the first candidate block division method and the second candidate block division method; if not, delete the candidate block division method with the largest amount of calculation among the first candidate block division method and the second candidate block division method; Re-execute the step of selecting a first candidate block division method and a second candidate block division method from multiple candidate block division methods until there is only one remaining candidate block division method, and use the remaining one candidate block division method as the target block division method.
16. The operation method based on the artificial intelligence model according to claim 15, characterized in that, Judging whether the ratio between the amount of calculation of the block calculation corresponding to the first candidate block division method and the amount of calculation of the block calculation corresponding to the second candidate block division method is less than or equal to a preset value includes: Judge whether the ratio between the total amount of calculation of the multiplication calculation amount and the accumulation calculation amount corresponding to the first candidate block division method and the total amount of calculation of the multiplication calculation amount and the accumulation calculation amount corresponding to the second candidate block division method is less than or equal to a preset value; Correspondingly, the deleting the candidate block division method with the largest amount of calculation among the first candidate block division method and the second candidate block division method includes: Delete the candidate block division method with the largest total amount of calculation of the multiplication calculation amount and the accumulation calculation amount among the first candidate block division method and the second candidate block division method.
17. An operating device based on an artificial intelligence model, characterized in that, Includes: An acquisition module, configured to acquire source data of a target operation; wherein, the source data includes a first matrix and a second matrix; A determination module, configured to determine whether the target operation is a block calculation; if so, determine a first block size of the first matrix and a second block size of the second matrix; A block storage module, configured to block and store the first matrix into a plurality of first block matrices according to the first block scale, and block and store the second matrix into a plurality of second block matrices according to the second block scale; An operation module, configured to sequentially read the plurality of first block matrices and the plurality of second block matrices during the execution of the target operation, so as to implement the target operation to obtain a target operation result.
18. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the operation method based on the artificial intelligence model according to any one of claims 1 to 16 when executing the computer program.
19. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed, the steps of the operation method based on the artificial intelligence model according to any one of claims 1 to 16 are implemented.
20. A computer program product, characterized in that, Comprising a computer program, and when the computer program is executed, the steps of the operation method based on the artificial intelligence model according to any one of claims 1 to 16 are implemented.
Citation Information
Patent Citations
Data processing method and device and electronic equipment
CN112069460A
Data processing method and device for matrix multiplication
CN112182496A
Matrix multiplication execution method and device, electronic equipment and storage medium
CN117785114A
Matrix operation method, device and equipment based on multi-core hardware and medium
CN117892050A
Matrix multiplication performance optimization method, system, device, medium and program
CN117951436A