Instruction generation device, method, equipment, storage medium and computer program product

By designing an instruction generation device that generates control instructions in parallel, the problem of low execution efficiency of neural network operations in traditional technology is solved, and more efficient neural network processor operations are achieved.

CN117131911BActive Publication Date: 2025-05-09GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310949140.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-28
Publication Date
2025-05-09
Estimated Expiration
2043-07-28

AI Technical Summary

Technical Problem

The traditional instruction generation device controls the neural network processor to perform neural network computing, and has low execution efficiency.

Method used

An instruction generation device is designed, including a number-fetch instruction transmission module, a matrix instruction transmission module and a number-store instruction transmission module. By generating control instructions in parallel, the neural network processor is controlled to perform the operation of number-fetch, calculation and number-store operations respectively.

Benefits of technology

It effectively improves the efficiency of neural network processors in processing neural network operations, and avoids the inefficiency problem caused by sending control instructions one by one in traditional technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117131911B_ABST
    Figure CN117131911B_ABST
Patent Text Reader

Abstract

The present application relates to an instruction generation device, method, equipment, storage medium and computer program product. The instruction generation device includes: an access instruction transmission module, which is used to obtain an access decoding signal corresponding to a target neural network, generate an access instruction according to the access decoding signal, and control the neural network processor to obtain an input feature map and a convolution kernel according to the access instruction; a matrix instruction transmission module, which is used to obtain a matrix decoding signal, generate a matrix calculation instruction according to the matrix decoding signal, and control the neural network processor to perform convolution calculation operations on the input feature map and the convolution kernel according to the matrix calculation instruction; a storage instruction transmission module, which is used to obtain a storage decoding signal, generate a storage instruction according to the storage decoding signal, and control the neural network processor to store the output feature map corresponding to the target neural network to the target address. The instruction generation device can improve the efficiency of controlling the neural network processor to process neural network operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of control technology, and in particular to an instruction generation device, method, equipment, storage medium and computer program product. Background Art

[0002] As the problems that neural networks can solve become more and more complex, the scale of neural networks is gradually increasing, and the amount of computation required by neural network processors to handle neural network operations is extremely large.

[0003] In the conventional technology, the neural network processor performs a set of operations including fetching, calculating and storing data in sequence according to the control instructions issued by the corresponding instruction generation device. The instruction generation device in the conventional technology sends control instructions to the neural network processor one by one according to the instruction sequence, and the control processor processes the neural network operation.

[0004] However, the above instruction generation device may result in a relatively low execution efficiency of the neural network operation. Summary of the invention

[0005] Based on this, it is necessary to provide an instruction generation device, method, equipment, storage medium and computer program product that can improve the efficiency of neural network operations.

[0006] In a first aspect, the present application provides an instruction generation device, the device comprising: a fetch instruction emission module, a matrix instruction emission module and a store instruction emission module;

[0007] The data acquisition instruction transmitting module is used to obtain the data acquisition decoding signal corresponding to the target neural network and generate the data acquisition instruction according to the data acquisition decoding signal; the data acquisition instruction is used to control the neural network processor to obtain the input feature map and the convolution kernel according to the data acquisition instruction;

[0008] A matrix instruction transmitting module is used to obtain a matrix decoding signal corresponding to the target neural network and generate a matrix calculation instruction according to the matrix decoding signal; the matrix calculation instruction is used to control the neural network processor to perform a convolution calculation operation on the input feature map and the convolution kernel according to the matrix calculation instruction;

[0009] The storage instruction transmitting module is used to obtain the storage decoding signal corresponding to the target neural network and generate the storage instruction according to the storage decoding signal; the storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address.

[0010] In one embodiment, the data fetch decoding signal includes a first hardware loop signal and a first address generation signal; the data fetch instruction transmitting module includes a first hardware loop control unit and a first address generation unit;

[0011] A first hardware loop control unit, used for acquiring a first hardware loop signal, executing a first hardware loop process according to the first hardware loop signal, and obtaining a first sequence group and a first mask sequence;

[0012] The first address generating unit is used to obtain a first address generating signal, a first sequence group and a first mask sequence, and execute a first address generating process according to the first address generating signal, the first sequence group and the first mask sequence to obtain a data fetch instruction.

[0013] In one of the embodiments, the first hardware loop signal includes a first accumulated signal, a second accumulated signal, a third accumulated signal, and a fourth accumulated signal, and the first hardware loop control unit includes a first accumulator, a second accumulator, a third accumulator, a fourth accumulator, a first finite state machine, and a first mask generator;

[0014] A first accumulator, used for acquiring a first accumulation signal, performing a first accumulation operation according to the first accumulation signal, and outputting a first accumulation result corresponding to each first accumulation operation to the first finite state machine and the first mask generator;

[0015] A second accumulator, used to obtain a second accumulation signal and a first accumulation result, perform a second accumulation operation according to the second accumulation signal and the first accumulation result, and output the second accumulation result corresponding to each second accumulation operation to the first finite state machine;

[0016] A third accumulator is used to obtain a third accumulation signal and a second accumulation result, and perform a third accumulation operation according to the third accumulation signal and the second accumulation result, and output the third accumulation result corresponding to each third accumulation operation to the first finite state machine;

[0017] a fourth accumulator, configured to obtain a fourth accumulation signal and a third accumulation result, perform a fourth accumulation operation according to the fourth accumulation signal and the third accumulation result, and output a fourth accumulation result corresponding to each fourth accumulation operation to the first finite state machine;

[0018] A first finite state machine is used to obtain a first sequence group according to each first accumulation result, each second accumulation result, each third accumulation result and each fourth accumulation result;

[0019] The first mask generator is used to obtain a first mask sequence according to each first accumulation result.

[0020] In one of the embodiments, the first address generation signal includes an input feature map address, a convolution kernel address, a first address step, an upsampling enable signal, and a fill signal, and the first address generation unit includes a first address generation register and a second address generation register;

[0021] The first address generation register is used to obtain the input feature map address and the convolution kernel address, and generate a first base address according to the input feature map address and the convolution kernel address;

[0022] The second address generation register is used to obtain the first address step, the upsampling enable signal, the fill signal, the first base address, the first sequence group and the first mask sequence, and obtain the data fetch instruction according to the first address step, the upsampling enable signal, the fill signal, the first base address, the first sequence group and the first mask sequence.

[0023] In one embodiment, the matrix decoding signal includes a second hardware loop signal, a second address generation signal and a matrix configuration signal; the matrix instruction transmission module includes a second hardware loop control unit, a second address generation unit and a matrix configuration unit;

[0024] A matrix configuration unit, used for obtaining a matrix configuration signal and generating a matrix configuration result according to the matrix configuration signal;

[0025] A second hardware loop control unit, used for acquiring a second hardware loop signal, executing a second hardware loop process according to the second hardware loop signal, and obtaining a second sequence group and a second mask sequence;

[0026] The second address generating unit is used to obtain the matrix configuration result, the second sequence group, the second mask sequence and the second address generating signal, and execute the second address generating process according to the matrix configuration result, the second sequence group, the second mask sequence and the second address generating signal to obtain the matrix calculation instruction.

[0027] In one embodiment, the store decoding signal includes a third hardware loop signal and a third address generation signal; the store instruction emission module includes a third hardware loop control unit and a third address generation unit;

[0028] A third hardware loop control unit, used for acquiring a third hardware loop signal, executing a third hardware loop process according to the third hardware loop signal, and obtaining a third sequence group and a third mask sequence;

[0029] The third address generating unit is used to obtain a third sequence group and a third address generating signal, and execute a third address generating process according to the third address generating signal, the third sequence group and the third mask sequence to obtain a store instruction.

[0030] In a second aspect, the present application provides an instruction generation method, the method comprising:

[0031] Obtaining a data acquisition decoding signal, and generating a data acquisition instruction according to the data acquisition decoding signal, wherein the data acquisition instruction is used to control the neural network processor to acquire an input feature map and a convolution kernel according to the data acquisition instruction;

[0032] Obtaining a matrix decoding signal, and generating a matrix calculation instruction according to the matrix decoding signal, wherein the matrix calculation instruction is used to control the neural network processor to perform a convolution calculation operation on the input feature map and the convolution kernel according to the matrix calculation instruction;

[0033] A storage decoding signal is obtained, and a storage instruction is generated according to the storage decoding signal; the storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address.

[0034] In a third aspect, the present application further provides a computer device, wherein the computer device comprises the instruction generating device as described in the first aspect above.

[0035] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in the second aspect are implemented.

[0036] In a fifth aspect, the present application further provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the method described in the second aspect are implemented.

[0037] The above-mentioned instruction generation device, method, equipment, storage medium and computer program product obtain a data acquisition decoding signal corresponding to the target neural network, generate a data acquisition instruction according to the data acquisition decoding signal, the data acquisition instruction is used to control the neural network processor to obtain the input feature map and the convolution kernel according to the data acquisition instruction; obtain a matrix decoding signal, generate a matrix calculation instruction according to the matrix decoding signal, the matrix calculation instruction is used to control the neural network processor to perform convolution calculation operations on the input feature map and the convolution kernel according to the matrix calculation instruction; obtain a storage decoding signal, and generate a storage instruction according to the storage decoding signal; the storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address; in this way, the data acquisition instruction transmission module, the matrix instruction transmission module, the vector instruction transmission module and the storage instruction transmission module in the present application generate control instructions in parallel, respectively control the neural network processor to perform data acquisition, calculation and storage operations, and can control the neural network processor to perform multiple groups of operations including data acquisition, calculation and storage at the same time, avoiding the problem of low efficiency caused by the control device in the traditional technology sending control instructions to the neural network processor one by one and controlling the processor to perform operations one by one, and effectively improving the efficiency of controlling the neural network processor to process neural network operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 It is an application environment diagram of an instruction generation device according to an embodiment;

[0040] Figure 2 for Figure 1 A module structure of the instruction generating device;

[0041] Figure 3 for Figure 2 A module structure of a fetch instruction emission module;

[0042] Figure 4 for Figure 3 A unit structure of a first hardware loop control unit in;

[0043] Figure 5 for Figure 3 A unit structure of a first address generating unit in;

[0044] Figure 6 for Figure 2 A module structure of a matrix instruction transmission module;

[0045] Figure 7 for Figure 2 A module structure of a storage instruction emission module.

[0046] Description of reference numerals:

[0047] 100-data fetch instruction emission module, 102-first hardware loop control unit, 1022-first accumulator, 1024-second accumulator, 1026-third accumulator, 1028-fourth accumulator, 1032-first finite state machine, 1034-first mask generator, 104-first address generation unit, 1042-first address generation register, 1044-second address generation register, 200-matrix instruction emission module, 202-matrix configuration unit, 204-second hardware loop control unit, 206-second address generation unit, 300-storage instruction emission module, 302-third hardware loop control unit, 304-third address generation unit. DETAILED DESCRIPTION

[0048] In order to facilitate understanding of the present application, the present application will be described more fully below with reference to the relevant drawings. Embodiments of the present application are provided in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0050] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first resistor may be referred to as a second resistor, and similarly, a second resistor may be referred to as a first resistor. Both the first resistor and the second resistor are resistors, but they are not the same resistor.

[0051] It can be understood that the “connection” in the following embodiments should be understood as “electrical connection”, “communication connection”, etc. if the connected circuits, modules, units, etc. have electrical signals or data transmission between each other.

[0052] When used herein, the singular forms "a", "an", and "said / the" may also include plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "include / comprise" or "have" etc. specify the presence of stated features, wholes, steps, operations, components, parts or combinations thereof, but do not exclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, components, parts or combinations thereof.

[0053] The instruction generation device provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the instructions generated by the instruction generation device are sent to the storage module of the neural network processor, which is used to control the storage module to extract the input feature map and convolution kernel required for matrix calculation and send them to the matrix calculation module of the neural network processor; the instructions generated by the instruction generation device are also sent to the matrix calculation module of the neural network processor, which is used to control the matrix calculation module to perform convolution calculation on the received input feature map and convolution kernel, and send the output feature map generated by the calculation back to the storage module; and the instructions generated by the instruction generation device are also sent to the storage module of the neural network processor, which is used to control the storage module to store the received output feature map in a preset position.

[0054] In one embodiment, Figure 2 As shown, the instruction generation device includes a fetch instruction emission module 100 , a matrix instruction emission module 200 and a store instruction emission module 300 .

[0055] The data acquisition instruction transmitting module 100 is used to obtain the data acquisition decoding signal corresponding to the target neural network and generate the data acquisition instruction according to the data acquisition decoding signal; the data acquisition instruction is used to control the neural network processor to obtain the input feature map and the convolution kernel according to the data acquisition instruction.

[0056] Among them, the data acquisition decoding signal corresponding to the target neural network refers to the computer-executable data acquisition code signal formed by compiling the external data acquisition program.

[0057] Exemplarily, the external data acquisition procedure can be expressed as:

[0058] for k in[0,K):

[0059] for p in[0,P):

[0060] for y in[0,KY):,

[0061] Among them, K represents the width of the input feature map, k represents the kth element of the width of the input feature map for the current fetch operation; P represents the height of the input feature map, p represents the pth element of the height of the input feature map for the current fetch operation; KY represents the weight matrix corresponding to the convolution kernel, and y represents the pth element in the convolution kernel for the current fetch operation.

[0062] Compile the above external data acquisition program, and the resulting data acquisition decoding signal can be expressed as:

[0063] load.exec Ifmap

[0064] load.exec Weight,

[0065] Among them, the load instruction is used to load data from the external storage into the computer memory for use by the data fetch instruction emission module 100; the exec instruction is used to execute the load instruction with the execution prefix; Ifmap represents the input feature map; Weight represents the weight matrix corresponding to the convolution kernel.

[0066] The matrix instruction transmitting module 200 is used to obtain the matrix decoding signal corresponding to the target neural network and generate the matrix calculation instruction according to the matrix decoding signal; the matrix calculation instruction is used to control the neural network processor to perform the convolution calculation operation on the input feature map and the convolution kernel according to the matrix calculation instruction.

[0067] Among them, the matrix decoding signal corresponding to the target neural network refers to the computer-executable matrix calculation code signal formed by compiling the external matrix calculation program.

[0068] Exemplarily, the external matrix calculation procedure can be expressed as:

[0069]

[0070] Among them, O represents the output feature map obtained by executing the current matrix calculation operation, I represents the input feature map of the current matrix calculation operation, W represents the convolution kernel of the current matrix calculation operation, and j represents the processing variable of the neural network processor. When the neural network processor contains 8×8×8 INT8 multipliers, it means that the neural network processor can process a matrix multiplication operation of one 8×8 matrix by another 8×8 matrix in one cycle. At this time, the value range of j is [0,8].

[0071] Compiling the above external matrix calculation program, the resulting matrix decoding signal can be expressed as:

[0072] mat.exec,

[0073] Among them, the mat instruction can be used to perform matrix operations, such as matrix addition, multiplication, transposition, inversion, eigenvalue calculation, etc. The specific syntax and supported operations depend on the programming language or mathematical library used, and this application does not limit this.

[0074] The storage instruction transmitting module 300 is used to obtain the storage decoding signal corresponding to the target neural network and generate the storage instruction according to the storage decoding signal; the storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address.

[0075] The storage decoding signal refers to a computer-executable storage code signal formed by compiling an external program.

[0076] Exemplarily, the external program may be the output object O[k][p] specified in the above matrix calculation program. The above external program is compiled, and the resulting storage decoding signal may be expressed as:

[0077] store.exec Ofmap,

[0078] Among them, Ofmap represents the output feature map.

[0079] The instruction generation device provided in the above embodiment obtains a data acquisition decoding signal corresponding to the target neural network, generates a data acquisition instruction according to the data acquisition decoding signal, and the data acquisition instruction is used to control the neural network processor to obtain the input feature map and the convolution kernel according to the data acquisition instruction; obtains a matrix decoding signal, generates a matrix calculation instruction according to the matrix decoding signal, and the matrix calculation instruction is used to control the neural network processor to perform convolution calculation operations on the input feature map and the convolution kernel according to the matrix calculation instruction; obtains a storage decoding signal, and generates a storage instruction according to the storage decoding signal; the storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address; In this way, the data acquisition instruction transmitting module 100, the matrix instruction transmitting module 200, and the storage instruction transmitting module 300 in this embodiment generate control instructions in parallel, respectively control the neural network processor to perform data acquisition, calculation and storage operations, and can control the neural network processor to perform multiple groups of operations including data acquisition, calculation and storage at the same time, avoiding the problem of low efficiency caused by the control device in the traditional technology sending control instructions to the neural network processor one by one and controlling the processor to perform operations one by one, thereby effectively improving the efficiency of the neural network processor in processing neural network operations.

[0080] In one embodiment, based on Figure 2 The embodiment shown, as Figure 3 As shown, the fetch instruction transmitting module 100 includes a first hardware loop control unit 102 and a first address generating unit 104, and the fetch decoding signal includes a first hardware loop signal and a first address generating signal.

[0081] The first hardware loop signal refers to a signal in the data acquisition and decoding signal used to configure parameters for the first hardware loop control unit 102 to execute the first hardware loop process, including a start value, an end value and a step size.

[0082] Exemplarily, the first hardware loop signal may be expressed as:

[0083] load.loop[1]0,8,8

[0084] The loop[1] instruction indicates setting parameters for the first hardware loop process of level 1. The first hardware loop signal indicates that the starting value of the first hardware loop process of level 1 is 0, the step length is 8, and the ending value is 8.

[0085] The first address generation signal refers to a signal in the data acquisition decoding signal that is used to configure parameters for the first address generation unit 104 to perform the first address generation process.

[0086] Exemplarily, the first address generation signal may be expressed as:

[0087] load.opstep[1]8,x,x

[0088] load.exec,

[0089] Among them, the opstep instruction is used to set the address step size in the first address generation process. The above-mentioned first address generation signal load.opstep[1] indicates that the address step size corresponding to the level 1 first address generation process is 8; load.exec indicates that the level 1 first address generation process is executed according to the parameter setting of the load.opstep[1] signal.

[0090] The first hardware loop control unit 102 is used to obtain a first hardware loop signal, and execute a first hardware loop process according to the first hardware loop signal to obtain a first sequence group and a first mask sequence.

[0091] The first hardware loop process may have multiple stages; the first sequence group is obtained according to the loop results of the multiple stages of the first hardware loop process; and the elements in the first mask sequence are calculated according to the loop result of the 0th stage loop.

[0092] Exemplarily, the first hardware loop process may have three stages, and the first hardware loop signal may be expressed as:

[0093] load.loop[2]0,P,1

[0094] load.loop[1]0,KY,1

[0095] load.loop[0]0,8,8

[0096] It can be obtained that the first sequence of the first sequence group is [0,0,0], which corresponds to the starting values ​​of the first hardware loop process of level 2, level 1, and level 0 respectively, and the first value of the first mask sequence corresponds to the starting value of the first hardware loop process of level 0; if after one first hardware loop process, the second sequence of the first sequence group is [0,1,0], which corresponds to the loop results of the current level 2, level 1, and level 0 loop processes respectively, and the second value of the first mask sequence corresponds to the loop result of the current level 0 loop process.

[0097] The first address generating unit 104 is used to obtain a first address generating signal, a first sequence group and a first mask sequence, and execute a first address generating process according to the first address generating signal, the first sequence group and the first mask sequence to obtain a data fetch instruction.

[0098] Exemplarily, when the neural network processor includes 8×8×8 INT8 multipliers, it means that the neural network processor can process a matrix multiplication operation of an 8×8 matrix multiplied by another 8×8 matrix in one cycle. In this way, the storage module of the neural network processor can accept the 8 addresses generated by the first address generation unit 104 as fetch instructions, and control the storage module to obtain the input feature map and convolution kernel according to the address corresponding to the fetch instruction. The first address generation unit 104 sets the address step length in the first address generation process according to the first address generation signal, and obtains 8 addresses as fetch instructions according to the first address generation formula according to the address step length, the first sequence group and the first mask sequence.

[0099] In this embodiment, the first hardware loop control unit 102 in the data acquisition instruction transmitting module 100 is used to obtain the first hardware loop signal in the data acquisition decoding signal, and execute the first hardware loop process according to the first hardware loop signal to obtain the first sequence group and the first mask sequence; the first address generating unit 104 is used to obtain the data acquisition instruction according to the first address generating signal, the first sequence group and the first mask sequence in the data acquisition decoding signal. In this way, the data acquisition instruction transmitting module 100 independently generates the data acquisition instruction according to the data acquisition decoding signal, which is used to independently control the storage unit of the neural network processor to execute the operation of acquiring the input feature map and the convolution kernel, so that the neural network processor can enter the current data acquisition process after completing the previous data acquisition process, thereby improving the efficiency of controlling the neural network processor to process neural network operations.

[0100] In one embodiment, Figure 4 As shown, the first hardware loop control unit 102 includes a first accumulator 1022, a second accumulator 1024, a third accumulator 1026, a fourth accumulator 1028, a first finite state machine 1032 and a first mask generator 1034; the first hardware loop signal includes a first accumulated signal, a second accumulated signal, a third accumulated signal and a fourth accumulated signal.

[0101] The first accumulator 1022 is used to obtain a first accumulation signal, perform a first accumulation operation according to the first accumulation signal, and output a first accumulation result corresponding to each first accumulation operation to the first finite state machine 1032 and the first mask generator 1034 .

[0102] The first accumulation signal refers to a signal in the first hardware loop signal used to configure parameters of the first accumulator 1022 , including a start value, an end value, and a step size.

[0103] Exemplarily, the first accumulation operation performed by the first accumulator 1022 may be set to correspond to the first hardware loop process of level 0 of the first hardware loop control unit 102, and the first accumulation signal may be expressed as:

[0104] load.loop[0]0,8,8,

[0105] The first accumulation operation performed by the first accumulator 1022 has a starting value of 0, an ending value of 8, and a step length of 8. The starting value 0 is output as the first accumulation result of the 0th first address generation process to the first finite state machine 1032 and the first mask generator 1034. In the first first address generation process, the first accumulation operation is performed, and the ending value of the first accumulation operation is reached. The first accumulation result returns to the starting value of 0, and the first accumulation result is output to the first finite state machine 1032 and the first mask generator 1034.

[0106] The second accumulator 1024 is used to obtain the second accumulation signal and the first accumulation result, and perform a second accumulation operation according to the second accumulation signal and the first accumulation result, and output the second accumulation result corresponding to each second accumulation operation to the first finite state machine 1032 .

[0107] The second accumulation signal refers to a signal in the first hardware loop signal used to configure parameters of the second accumulator 1024, including a start value, an end value, and a step size.

[0108] Exemplarily, the second accumulation operation performed by the second accumulator 1024 may be set to correspond to the first hardware loop process of level 1 of the first hardware loop control unit 102, and the second accumulation signal may be expressed as:

[0109] load.loop[1]0,3,1,

[0110] The second accumulation operation performed by the second accumulator 1024 has a starting value of 0, an ending value of 3, and a step length of 1, and outputs the second accumulation result of the 0th first address generation process with the starting value 0 to the first finite state machine 1032. If the first accumulation result reaches the ending value of the first accumulation operation, the second accumulator 1024 can perform the second accumulation operation, and after one second accumulation operation, the second accumulation result can be 1, and the ending value of the second accumulation operation has not been reached, and the second accumulation result is output to the first finite state machine 1032.

[0111] The third accumulator is used to obtain a third accumulation signal and a second accumulation result, perform a third accumulation operation according to the third accumulation signal and the second accumulation result, and output the third accumulation result corresponding to each third accumulation operation to the first finite state machine.

[0112] The third accumulation signal refers to a signal in the first hardware loop signal used to configure parameters of the third accumulator 1026, including a start value, an end value, and a step size.

[0113] Exemplarily, the third accumulation operation performed by the third accumulator 1026 may be set to correspond to the 2-level first hardware loop process of the first hardware loop control unit 102, and the third accumulation signal may be expressed as:

[0114] load.loop[2]0,3,1,

[0115] The third accumulation operation performed by the third accumulator 1026 has a starting value of 0, an ending value of 3, a step length of 1, and outputs the starting value 0 as the third accumulation result of the 0th first address generation process to the first finite state machine 1032 .

[0116] If the second accumulation result does not reach the termination value of the second accumulation operation, the third accumulation operation is not started, and the starting value of the third accumulation operation is output to the first finite state machine 1032 as the third accumulation result of the current level 2 first hardware loop process.

[0117] If the second accumulation result reaches the termination value of the second accumulation operation, the third accumulator 1026 performs the third accumulation operation. After one third accumulation operation, the third accumulation result can be 1, which does not reach the termination value of the third accumulation operation. The third accumulation result is output to the first finite state machine 1032.

[0118] The fourth accumulator is used to obtain a fourth accumulation signal and a third accumulation result, and perform a fourth accumulation operation according to the fourth accumulation signal and the third accumulation result, and output the fourth accumulation result corresponding to each fourth accumulation operation to the first finite state machine.

[0119] The fourth accumulation signal refers to a signal in the first hardware loop signal used to configure parameters of the fourth accumulator 1028 , including a start value, an end value, and a step size.

[0120] Exemplarily, the fourth accumulation operation performed by the fourth accumulator 1028 may be set to correspond to the 3-level first hardware loop process of the first hardware loop control unit 102, and the third accumulation signal may be expressed as:

[0121] load.loop[3]0,8,1,

[0122] The fourth accumulation operation performed by the fourth accumulator 1028 has a starting value of 0, an ending value of 8, a step length of 1, and outputs the starting value 0 as the fourth accumulation result of the 0th first address generation process to the first finite state machine 1032 .

[0123] If the third accumulation result does not reach the termination value of the third accumulation operation, the fourth accumulation operation is not started, and the starting value of the fourth accumulation operation is output to the first finite state machine 1032 as the fourth accumulation result of the current 3-level first hardware loop process.

[0124] If the third accumulation result reaches the termination value of the third accumulation operation, the fourth accumulator 1028 performs the fourth accumulation operation. After one fourth accumulation operation, the fourth accumulation result can be 1, which does not reach the termination value of the fourth accumulation operation. The fourth accumulation result is output to the first finite state machine 1032.

[0125] The first finite state machine 1032 is used to obtain a first sequence group according to the first accumulation result, the second accumulation result, the third accumulation result and the fourth accumulation result.

[0126] The first sequence group refers to a set of sequences obtained according to the first accumulation result, the second accumulation result, the third accumulation result and the fourth accumulation result.

[0127] Exemplarily, in the 0th first address generation process, the first accumulation result, the second accumulation result, the third accumulation result and the fourth accumulation result are all 0, and the first value of the first sequence group can be obtained as [0,0,0,0], which corresponds to the starting values ​​of the first hardware loop processes of level 3, level 2, level 1 and level 0 respectively.

[0128] During the first first address generation process, the first accumulator 1022 performs the first accumulation operation and reaches the termination value of the first accumulation operation. The first accumulator 1022 returns to the starting value 0. The second accumulator 1024 performs the second accumulation operation and obtains a second accumulation result of 1. The third accumulator 1026 and the fourth accumulator 1028 do not perform the accumulation operation, and the second value of the first sequence group can be obtained as [0, 0, 1, 0], which correspond to the fourth accumulation result, the third accumulation result, the second accumulation result and the first accumulation result of the first hardware loop process of level 3, level 2, level 1 and level 0 respectively.

[0129] The first mask generator 1034 is used to obtain a first mask sequence according to each first accumulation result.

[0130] The values ​​in the first mask sequence are obtained by judging the first accumulation results obtained by the first accumulator 1022 in each first address generation process.

[0131] In this embodiment, the first hardware loop control unit includes four accumulators, which can implement a four-level first hardware loop process. The start and stop of the accumulator corresponding to the first hardware loop process at the current level is controlled according to the accumulation result of the accumulator corresponding to the first hardware loop process at the previous level, and the first sequence group is obtained according to the accumulation results of each level. In this way, only four accumulation signals are needed to quickly generate a first sequence group including multiple sequences, which can be used to generate multiple data fetch instructions, further reducing the scale of data fetch decoding signals and improving the efficiency of generating data fetch instructions, thereby improving the efficiency of controlling the neural network processor to process neural network operations.

[0132] In one embodiment, Figure 5 As shown, the first address generation signal includes an input feature map address, a convolution kernel address, a first address step, an upsampling enable signal, and a fill signal, and the first address generation unit 104 includes a first address generation register 1042 and a second address generation register 1044.

[0133] Among them, the first address step size includes the step size of the input feature map address and the step size of the convolution kernel address; the upsampling enable signal and the padding signal are determined according to the actual situation of the neural network processor. If the neural network processor uses an upsampling operation when processing the target neural network, the upsampling enable signal is 1, otherwise it is 0; if the neural network processor performs a padding operation when processing the target neural network, the padding signal is 1, otherwise it is 0.

[0134] The first address generation register 1042 is used to obtain the input feature map address and the convolution kernel address, and generate a first base address based on the input feature map address and the convolution kernel address.

[0135] Among them, the first base address is recalculated at the input feature map address and the convolution kernel address.

[0136] Exemplarily, the first address generation register 1042 obtains the input feature map address KP and the convolution kernel address KY, and sets the first base address to 0 on this basis.

[0137] Exemplarily, the first address generation register 1042 may be a SPM memory (ScratchPad Memory).

[0138] The second address generation register 1044 is used to obtain the first address step, the upsampling enable signal, the fill signal, the first base address, the first sequence group and the first mask sequence, and obtain the data access instruction according to the first address step, the upsampling enable signal, the fill signal, the first base address, the first sequence group and the first mask sequence.

[0139] The fetch instruction is obtained by the second address generation register 1044 performing the following processing on the first address step, the up-sampling enable signal, the fill signal, the first base address, the first sequence group and the first mask sequence:

[0140] The second address generation register 1044 obtains the first loop address loop_addr according to the first address step opstep[i], the first base address base_addr, the first sequence group {loop_index[i]}, and the up-sampling enable signal upsample_en[i]. The specific calculation can be expressed as:

[0141]

[0142] Among them, i represents the level of the first hardware loop process, and its value range is [0,3]. oft[i] represents the parity of i. When i is an odd number, oft[i]=1, and when i is an even number, oft[i]=0. The first loop address loop_addr can be expressed in binary form.

[0143] The second address generation register 1044 obtains a fetch instruction number stride_id[n] generated in the first hardware loop process according to the first mask sequence pad_mask[i] and the first loop address loop_addr. The specific calculation process can be expressed as:

[0144] stride_id[n]=pad_mask[n]? loop_addr[2:0]+(n-pad)×opstep[0]>>upsample_en[n]:0,

[0145] The value range of n is determined according to the actual situation of the neural network processor; the data access instruction number stride_id[n] can be expressed in binary form.

[0146] For a first hardware loop process, the second address generation register 1044 determines the first address jump variable stride_step[n] according to the fetch instruction number stride_id[n]:

[0147] If the fetch instruction number stride_id[n] is greater than or equal to 0, and the third and fourth bits of the fetch instruction number stride_id[n] in binary form are both 0, then the first address jump variable stride_step[n] is determined to be 0;

[0148] If the access instruction number stride_id[n] is greater than or equal to 0, and the third and fourth bits of the binary access instruction number stride_id[n] are not all 0, the first address jump variable stride_step[n] is determined according to the fourth bit of the binary access instruction number stride_id[n] and the first address step length;

[0149] If the data fetch instruction number stride_id[n] is less than 0, the first address jump variable stride_step[n] is determined according to the inverse of the first address step length.

[0150] For a first hardware loop process, the second address generation register 1044 generates a first sequence group {loop_index[i]} nThe intermediate address middle_addr[n] is determined by the fetch instruction number stride_id[n]. The 3rd to 12th bits of the intermediate address middle_addr[n] are determined by the first loop address loop_addr, and the 0th to 2nd bits are determined by the fetch instruction number stride_id[n]. The intermediate address middle_addr[n] and the first address jump variable stride_step[n] are summed to obtain the fetch instruction addr[n], which can be expressed as:

[0151] middle_addr[n]=loop_addr[12:3], stride_id[n][2:0]

[0152] addr[n]=middle_addr[n]+stride_step[n].

[0153] In a possible implementation, the second address generation register 1044 may be a VRF memory (Vector Register File, register file memory).

[0154] In this embodiment, the first address generation register 1042 in the first address generation unit 104 is used to obtain the input feature map address and the convolution kernel address, and generate a first base address according to the input feature map address and the convolution kernel address; the second address generation register 1044 is used to obtain the first address step, the upsampling enable signal, the padding signal, the first base address, the first sequence group and the first mask sequence, and obtain the data fetch instruction according to the first address step, the upsampling enable signal, the padding signal, the first base address, the first sequence group and the first mask sequence. In this way, the first address generation unit 104 takes into account the two neural network processing methods of upsampling and padding in the process of generating the data fetch instruction, so that the data fetch instruction in this embodiment has a wider control range for the neural network processor, thereby expanding the application scope of the instruction generation device in this embodiment.

[0155] In one embodiment, Figure 6 As shown, the matrix instruction transmitting module 200 includes a matrix configuration unit 202, a second hardware loop control unit 204 and a second address generating unit 206; the matrix decoding signal includes a second hardware loop signal, a second address generating signal and a matrix configuration signal.

[0156] The matrix configuration unit 206 is used to obtain a matrix configuration signal and generate a matrix configuration result according to the matrix configuration signal.

[0157] Exemplarily, the matrix configuration signal can be expressed as:

[0158] mat.config.bs 64,x,x

[0159] The config instruction can be used to configure the number system and data format in the matrix execution module. In the matrix configuration signal, bs represents binary, and 64 represents that the maximum step length of the data format is 64.

[0160] The second hardware loop control unit 204 is configured to obtain a second hardware loop signal, and execute a second hardware loop process according to the second hardware loop signal to obtain a second sequence group and a second mask sequence.

[0161] Among them, the structure of the second hardware loop control unit 204 is the same as that of the first hardware loop control unit 102. The second hardware loop process can have multiple levels. Each level of the second hardware loop process is controlled by an accumulator. The start and stop of the accumulator corresponding to the second hardware loop process at the current level is controlled according to the accumulation result of the accumulator corresponding to the second hardware loop process at the previous level, and the second sequence group is obtained according to the accumulation results of each level; the second mask sequence is calculated based on the loop result of the 0th level loop of multiple second hardware loop processes.

[0162] In a possible implementation, the second hardware loop process may have 4 stages. Exemplarily, the second hardware loop signal may be expressed as:

[0163] mat.loop[3]0,8,1

[0164] mat.loop[2]0,3,1

[0165] mat.loop[1]0,3,1

[0166] mat.loop[0]0,8,8,

[0167] The mat.loop[0] instruction indicates the parameter setting for the second hardware loop process of level 0, and the second hardware loop signal indicates that the starting value of the second hardware loop process of level 0 is 0, the step length is 8, and the ending value is 8.

[0168] It can be obtained that the first sequence of the second sequence group is [0,0,0,0], which corresponds to the starting values ​​of the first hardware loop process of level 3, level 2, level 1, and level 0 respectively, and the first value of the second mask sequence corresponds to the starting value of the second hardware loop process of level 0; if a second hardware loop process is carried out, the second sequence of the second sequence group is [0,0,1,0], which corresponds to the loop results of the current level 3, level 2, level 1, and level 0 loop processes respectively, and the second value of the second mask sequence corresponds to the loop result of the current level 0 loop process.

[0169] The second address generating unit 206 is used to obtain the matrix configuration result, the second sequence group, the second mask sequence and the second address generating signal, and execute the second address generating process according to the matrix configuration result, the second sequence group, the second mask sequence and the second address generating signal to obtain the matrix calculation instruction.

[0170] The structure of the second address generating unit 206 is the same as that of the first address generating unit 104, including a third address generating register and a fourth address generating register.

[0171] Among them, the second address generation signal includes the input feature map address, the convolution kernel address, the second address step, the upsampling enable signal and the filling signal.

[0172] The third address generation register is used to obtain the input feature map address and the convolution kernel address, and generate a second base address based on the input feature map address and the convolution kernel address.

[0173] The fourth address generation register is used to obtain the matrix configuration result, the second address step opstep2[j], the upsampling enable signal upsample_en[j], the padding signal pad, the second base address base_addr2, the second sequence group {loop_index2[j]} and the second mask sequence pad_mask2[j], and obtain the matrix calculation instruction according to the above data.

[0174] The specific calculation can be expressed as:

[0175]

[0176] Wherein, j represents the level of the second hardware loop process, and its value range is [0,3], oft2[j] represents the parity of j, and the second loop address loop_addr2 is expressed in binary form according to the matrix configuration result.

[0177] stride_id2[t]=pad_mask2[t]? loop_addr2[2:0]+(t-pad)×opstep2[0]>>upsample_en[t]:0,

[0178] The value range of t is determined according to the actual situation of the neural network processor; the matrix calculation instruction number stride_id2[t] is expressed in binary form according to the matrix configuration result.

[0179] stride_step2[t]=(stride_id2[t][5]==0)? (stride)id2[t][4:3]==0)? :0:big_step< <stride_id2[t][4]:-big_step,

[0180] Here, big_step refers to the maximum step size in the matrix configuration signal.

[0181] middle_addr2[t]=loop_addr2[12:3], stride_id2[t][2:0]

[0182] addr2[t]=middle_addr2[t]+stride_step2[t].

[0183] Example 1: When the neural network processor contains 8×8×8 INT8 multipliers, it means that the neural network processor can process a matrix multiplication operation of one 8×8 matrix by another 8×8 matrix in one cycle. In this way, the storage module of the neural network processor can accept 8 addresses as matrix calculation instructions and supply them to the matrix calculation module for calculation, that is, the value range of t is [0,7].

[0184] The second address generation signal can be expressed as:

[0185] mat.opstep[3]48,x,x

[0186] mat.opstep[2]16,x,x

[0187] mat.opstep[1]8,x,x

[0188] mat.opstep[0]1,x,x

[0189] mat.exec 0,x,x,

[0190] Among them, the third address generation register sets the second base address base_addr2 to 0 based on the input feature map address and the convolution kernel address, the second address step opstep2[j]=[48,16,8,1], and big_step is 64.

[0191] When one of the sequences in the second sequence group is [1,1,2,0]:

[0192] loop_addr2=0+1×48+1×16+2×8=80; stride_id2[t]=0+(t-0)×1=t;

[0193] stride_step2[t]=0; addr2[t]=80+t+0=80+t;

[0194] The matrix calculation instructions at this time are: [87,86,85,84,83,82,81,80].

[0195] Example 2: When the neural network processor processes the target neural network in a padding manner and sets the padding signal to 1, and other settings are the same as in Example 1, when one of the sequences in the second sequence group is [1,1,2,0]:

[0196] loop_addr2=0+1×48+1×16+2×8=80; stride_id2[t]=0+(t-1)×1=t-1;

[0197] stride_step2: stride_step2[0]=-64; stride_step2[t]=0

[0198] addr2: addr2[0]=80-64=16; addr2[t]=80+(t-1)+0=79+t;

[0199] The matrix calculation instructions at this time are: [86,85,84,83,82,81,80,16], where the instruction corresponding to the first number will be discarded because the first number is filled.

[0200] Example 3, when the neural network processor processes the target neural network in an upsampling manner and sets the upsampling enable signal of the 1st and 0th level loop processes to 1, and other settings are the same as Example 1, when one of the sequences in the second sequence group is [1,1,2,0]: oft2 of the 1st level loop process is 1;

[0201] loop_addr2=0+1×48+1×16+(2+1)×8>>1=0+48+16+8=72; stride_id2[t]=0+t×1>>1;

[0202] stride_step2[t]=0;addr2[t]=72+t×1>>1+0

[0203] The matrix calculation instructions at this time are: [75,75,74,74,73,73,72,72].

[0204] Example 4, when the matrix instruction emission module needs to emit instructions to control the neural network processor to process the target neural network, the convolution kernel sliding step is 2, and the 0-level loop of the second address generation signal can be expressed as: mat.opstep[0]2,x,x; other settings are the same as Example 1. When one of the sequences in the second sequence group is [1,1,2,0]:

[0205] loop_addr2=0+1×48+1×16+1×16=80; stride_id2[t]=0+t×2, t=0,1,2,3,4,5,6,7

[0206] stride_step2: When t=[0,3], stride_step2[t]=0; when t=[4,7], stride_step2[t]=big_step=64;

[0207] addr2: when t=[0,3], addr2[t]=80+t×2; when t=[4,7], addr2[t]=80+(t×2-8)+64=136+t×2;

[0208] The matrix calculation instructions at this time are: [150,148,146,144,86,84,82,80].

[0209] In the above embodiment, the matrix configuration unit included in the matrix instruction transmitting module is used to obtain a matrix configuration signal, and generate a matrix configuration result according to the matrix configuration signal, which is used to control the number system and shape of each data in the second hardware loop control unit and the second address generation unit. In this way, by setting different matrix configuration signals, the matrix calculation instructions emitted by the matrix instruction transmitting module can adapt to the neural network processors in different scenarios, thereby expanding the application scenarios of the instruction generation device.

[0210] In one embodiment, Figure 7 As shown, the store decoding signal includes a third hardware loop signal and a third address generation signal; the store instruction transmitting module 300 includes a third hardware loop control unit 302 and a third address generation unit 304;

[0211] The third hardware loop control unit 302 is used to obtain a third hardware loop signal, and execute a third hardware loop process according to the third hardware loop signal to obtain a third sequence group and a third mask sequence.

[0212] The structure of the third hardware loop control unit 302 is the same as that of the first hardware loop control unit 102 .

[0213] The third address generating unit 304 is used to obtain a third sequence group and a third address generating signal, and execute a third address generating process according to the third address generating signal, the third sequence group and the third mask sequence to obtain a store instruction.

[0214] The structure of the third address generating unit 304 is the same as that of the first address generating unit 104 .

[0215] In the above embodiment, the storage instruction transmitting module 300 independently generates a storage instruction according to the storage decoding signal, which is used to independently control the storage unit of the neural network processor to perform the operation of storing the output feature map, so that the neural network processor can enter the current storage process after completing the previous storage process, thereby improving the efficiency of controlling the neural network processor to process neural network operations.

[0216] In one embodiment, a method for generating an instruction is provided, the method comprising:

[0217] Obtain a data access decoding signal, generate a data access instruction according to the data access decoding signal, the data access instruction is used to control the neural network processor to obtain the input feature map and the convolution kernel according to the data access instruction; obtain a matrix decoding signal, generate a matrix calculation instruction according to the matrix decoding signal, the matrix calculation instruction is used to control the neural network processor to perform convolution calculation operations on the input feature map and the convolution kernel according to the matrix calculation instruction; obtain a data storage decoding signal, generate a data storage instruction according to the data storage decoding signal; the data storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address.

[0218] In one embodiment, the data fetch decoding signal includes a first hardware loop signal and a first address generation signal; the data fetch instruction transmitting module includes a first hardware loop control unit and a first address generation unit; the first hardware loop process is executed according to the first hardware loop signal to obtain a first sequence group and a first mask sequence; the first address generation process is executed according to the first address generation signal, the first sequence group and the first mask sequence to obtain a data fetch instruction.

[0219] In one embodiment, the first hardware loop signal includes a first accumulation signal, a second accumulation signal, a third accumulation signal and a fourth accumulation signal. A first accumulation operation is performed according to the first accumulation signal, and a first accumulation result corresponding to each first accumulation operation is output; a second accumulation operation is performed according to the second accumulation signal and the first accumulation result, and a second accumulation result corresponding to each second accumulation operation is output; a third accumulation operation is performed according to the third accumulation signal and the second accumulation result, and a third accumulation result corresponding to each third accumulation operation is output; a fourth accumulation operation is performed according to the fourth accumulation signal and the third accumulation result, and a fourth accumulation result corresponding to each fourth accumulation operation is output; a first sequence group is obtained according to each first accumulation result, each second accumulation result, each third accumulation result and each fourth accumulation result; a first mask sequence is obtained according to each first accumulation result.

[0220] In one embodiment, the first address generation signal includes an input feature map address, a convolution kernel address, a first address step, an upsampling enable signal, and a fill signal. A first base address is generated according to the input feature map address and the convolution kernel address; and a data fetch instruction is obtained according to the first address step, the upsampling enable signal, the fill signal, the first base address, the first sequence group, and the first mask sequence.

[0221] In one embodiment, the matrix decoding signal includes a second hardware loop signal, a second address generation signal and a matrix configuration signal; a matrix configuration result is generated according to the matrix configuration signal; a second hardware loop process is executed according to the second hardware loop signal to obtain a second sequence group and a second mask sequence; a second address generation process is executed according to the matrix configuration result, the second sequence group, the second mask sequence and the second address generation signal to obtain a matrix calculation instruction.

[0222] In one embodiment, the store decoding signal includes a third hardware loop signal and a third address generation signal; a third hardware loop process is executed according to the third hardware loop signal to obtain a third sequence group and a third mask sequence; a third address generation process is executed according to the third address generation signal, the third sequence group and the third mask sequence to obtain a store instruction.

[0223] In one embodiment, a computer device is provided, including an instruction generating device as described in the above-mentioned device embodiments.

[0224] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0225] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0226] In the description of this specification, the description with reference to the terms "some embodiments", "other embodiments", "ideal embodiments", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example.

[0227] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0228] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. An instruction generating device, characterized in that: include: A first hardware loop control unit and a first address generation unit, a matrix instruction emission module and a store instruction emission module; The first hardware loop control unit is used to obtain a first hardware loop signal, and execute a first hardware loop process according to the first hardware loop signal to obtain a first sequence group and a first mask sequence; The first address generation unit is used to obtain a first address generation signal, the first sequence group and the first mask sequence, and perform a first address generation process according to the first address generation signal, the first sequence group and the first mask sequence to obtain a data acquisition instruction, wherein the data acquisition instruction is used to control the neural network processor to acquire an input feature map and a convolution kernel according to the data acquisition instruction; The matrix instruction transmitting module is used to obtain a matrix decoding signal corresponding to the target neural network, and generate a matrix calculation instruction according to the matrix decoding signal; the matrix calculation instruction is used to control the neural network processor to perform a convolution calculation operation on the input feature map and the convolution kernel according to the matrix calculation instruction; The storage instruction transmitting module is used to obtain the storage decoding signal corresponding to the target neural network, and generate a storage instruction according to the storage decoding signal; the storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address; Wherein, the first hardware loop signal includes a first accumulation signal, a second accumulation signal, a third accumulation signal and a fourth accumulation signal, and the first hardware loop control unit includes a first accumulator, a second accumulator, a third accumulator, a fourth accumulator, a first finite state machine and a first mask generator; The first accumulator is used to obtain the first accumulation signal, perform a first accumulation operation according to the first accumulation signal, and output a first accumulation result corresponding to each first accumulation operation to the first finite state machine and the first mask generator; The second accumulator is used to obtain the second accumulation signal and the first accumulation result, perform a second accumulation operation according to the second accumulation signal and the first accumulation result, and output the second accumulation result corresponding to each second accumulation operation to the first finite state machine; The third accumulator is used to obtain the third accumulation signal and the second accumulation result, perform a third accumulation operation according to the third accumulation signal and the second accumulation result, and output the third accumulation result corresponding to each third accumulation operation to the first finite state machine; The fourth accumulator is used to obtain the fourth accumulation signal and the third accumulation result, perform a fourth accumulation operation according to the fourth accumulation signal and the third accumulation result, and output the fourth accumulation result corresponding to each fourth accumulation operation to the first finite state machine; The first finite state machine is used to obtain the first sequence group according to each of the first accumulation results, each of the second accumulation results, each of the third accumulation results, and each of the fourth accumulation results; The first mask generator is used to obtain a first mask sequence according to each of the first accumulation results.

2. The device according to claim 1, characterized in that The first address generation signal includes an input feature map address, a convolution kernel address, a first address step, an upsampling enable signal, and a padding signal, and the first address generation unit includes a first address generation register and a second address generation register; The first address generation register is used to obtain the input feature map address and the convolution kernel address, and generate a first base address according to the input feature map address and the convolution kernel address; The second address generation register is used to obtain the first address step, the upsampling enable signal, the fill signal, the first base address, the first sequence group and the first mask sequence, and obtain the data fetch instruction according to the first address step, the upsampling enable signal, the fill signal, the first base address, the first sequence group and the first mask sequence.

3. The device according to claim 1, characterized in that The matrix decoding signal includes a second hardware loop signal, a second address generation signal and a matrix configuration signal; the matrix instruction transmission module includes a second hardware loop control unit, a second address generation unit and a matrix configuration unit; The matrix configuration unit is used to obtain the matrix configuration signal and generate a matrix configuration result according to the matrix configuration signal; The second hardware loop control unit is used to obtain the second hardware loop signal, and execute a second hardware loop process according to the second hardware loop signal to obtain a second sequence group and a second mask sequence; The second address generation unit is used to obtain the matrix configuration result, the second sequence group, the second mask sequence and the second address generation signal, and perform a second address generation process according to the matrix configuration result, the second sequence group, the second mask sequence and the second address generation signal to obtain the matrix calculation instruction.

4. The device according to claim 1, characterized in that The store decoding signal includes a third hardware loop signal and a third address generation signal; the store instruction transmitting module includes a third hardware loop control unit and a third address generation unit; The third hardware loop control unit is used to obtain the third hardware loop signal, and execute a third hardware loop process according to the third hardware loop signal to obtain a third sequence group and a third mask sequence; The third address generation unit is used to obtain the third sequence group and the third address generation signal, and execute a third address generation process according to the third address generation signal, the third sequence group and the third mask sequence to obtain the store instruction.

5. A method for generating an instruction, characterized in that: Applied to the instruction generation device according to any one of claims 1 to 4, the method comprising: Acquire a first hardware loop signal, and execute a first hardware loop process according to the first hardware loop signal to obtain a first sequence group and a first mask sequence; Obtaining a first address generation signal, the first sequence group, and the first mask sequence, performing a first address generation process according to the first address generation signal, the first sequence group, and the first mask sequence, and obtaining a data acquisition instruction, wherein the data acquisition instruction is used to control the neural network processor to acquire an input feature map and a convolution kernel according to the data acquisition instruction; Acquire a matrix decoding signal, and generate a matrix calculation instruction according to the matrix decoding signal, wherein the matrix calculation instruction is used to control the neural network processor to perform a convolution calculation operation on the input feature map and the convolution kernel according to the matrix calculation instruction; Obtain a storage decoding signal, and generate a storage instruction according to the storage decoding signal; the storage instruction is used to control the neural network processor to store the output feature map corresponding to the target neural network to the target address; The first hardware loop signal includes a first accumulated signal, a second accumulated signal, a third accumulated signal, and a fourth accumulated signal, and the process of obtaining the first sequence group and the first mask sequence includes: Acquire the first accumulation signal, perform a first accumulation operation according to the first accumulation signal, and output a first accumulation result corresponding to each first accumulation operation; Acquire the second accumulation signal and the first accumulation result, perform a second accumulation operation according to the second accumulation signal and the first accumulation result, and output a second accumulation result corresponding to each second accumulation operation; Acquire the third accumulation signal and the second accumulation result, perform a third accumulation operation according to the third accumulation signal and the second accumulation result, and output a third accumulation result corresponding to each third accumulation operation; Acquire the fourth accumulation signal and the third accumulation result, perform a fourth accumulation operation according to the fourth accumulation signal and the third accumulation result, and output a fourth accumulation result corresponding to each fourth accumulation operation; Obtaining the first sequence group according to each of the first accumulation results, each of the second accumulation results, each of the third accumulation results, and each of the fourth accumulation results; A first mask sequence is obtained according to each of the first accumulation results.

6. A computer device, characterized in that: The invention comprises an instruction generating device according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method described in claim 5 are implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in claim 5 are implemented.

Citation Information

Patent Citations

  • Neural network accelerator compiling method and device

    CN113554161A