Matrix operation instruction controller of data stream processor
Through dual-state machine optimization timing control, parallel control of matrix calculation and result output is realized, solving the problem of inefficient convolution instructions in existing hardware accelerators, and significantly improving the efficiency of convolution operation.
Patent Information
- Application Number
- CN202510435161.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing convolutional neural network computing hardware accelerator, a single instruction can only implement the matrix multiplication and accumulation operation of three-fold loops, resulting in the convolutional instruction that needs to be split into multiple instructions for execution, which is inefficient.
Using a matrix operation instruction controller based on a dual state machine, the matrix calculation and result output are controlled in parallel by optimizing the timing control structure, realizing independent control of calculation and output, eliminating waiting bubbles, and improving convolution operation efficiency.
In the 4-layer convolution cycle, the computing efficiency is improved by more than 50%, which significantly improves the computing performance of the hardware accelerator.
Smart Images

Figure CN120353499A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of instruction controllers, and in particular to a matrix operation instruction controller for a data stream processor. Background Art
[0002] A data stream processor (DSP) is a hardware or software system specifically designed for efficiently processing continuously flowing data (i.e., data streams). In modern computing environments, data stream processors are commonly used in fields such as real-time data processing, streaming media analysis, sensor data processing, Internet of Things (IoT) applications, and financial market analysis. Their key feature is the ability to process continuously and rapidly generated data streams and perform analysis, processing, and response in real time. The core task of a data stream processor is to extract, process, store, and analyze useful information from continuously generated data streams; different from traditional batch processing methods, data stream processing emphasizes real-time and continuity, meaning it can process changing data sources and respond quickly.
[0003] A convolutional neural network (CNN) is a widely used network architecture in deep learning, and one of its core operations is convolutional operation. Although convolution itself is a local weighted summation, through the transformation of unfolding and matrix multiplication, the convolutional operation can be converted into matrix multiplication, so that the operation of the entire network can be accelerated through efficient matrix operations. And a data stream processor can accelerate these operations through parallelization and real-time processing, so it has significant advantages when high-efficiency processing of large-scale data is required.
[0004] There are several ways to implement the hardware acceleration of convolutional neural network operations, such as ASIC, GPU, and FPGA. Currently, most of the instructions in the hardware accelerators available on the market can only implement the matrix multiply-accumulate operation of 3 nested loops, and most complete convolutional instructions need to be split into multiple instructions to be implemented in the hardware accelerator. Summary of the Invention
[0005] Aiming at the existing hardware accelerators for convolutional neural network operations, the present invention proposes a matrix operation instruction controller for a data stream processor, which is the core module of an AI acceleration core (Neural Network Accelerator, NNA). Based on the normal control loop, the timing control structure is optimized, the parallel control of matrix calculation and result output is realized, and the efficiency of convolutional operation is improved.
[0006] The present invention protects a matrix operation instruction execution control method based on a dual state machine, and the dual state machine includes state machine Ⅰ and state machine Ⅱ.
[0007] State machine Ⅰ contains 5 states: Top_loop, Acc_ini, Uop_loop, idle, and finish, which are used to control the MAC data stream calculation. Among them,
[0008] The Top_loop state is the top-level loop state, which includes two-level loop control of the top-level loop top_loop and the outer loop out_loop for the convolution operation. The length of the top_loop loop is top_loop_cnt, and the loop step is 1. The length of the out_loop loop is outer_loop_cnt, and the loop step is 1.
[0009] The Uop_loop state is the inner loop state, which includes two-level loop control of the micro-operation loop uop_loop and the inner loop inner_loop for the convolution operation. The length of the uop_loop loop is uop_len, and the loop step is 1. The length of the inner_loop loop is inner_loop_cnt, and the loop step is 1. This state controls the address generator to perform address calculation or output, triggers the MAC output handshake signal MAC_out_en after the micro-operation loop ends, and returns to the Top_loop state.
[0010] The ACC_ini state realizes the initialization control of the ACC register, and the ACC initialization needs to be performed before each complete cycle of the Uop_loop.
[0011] The idle state is the idle state. When the state machine enters this state, it does not perform any operations and is in a waiting state.
[0012] The finish state is the end state of a complete cycle control of the state machine. When the execution of a convolution operation or a matrix multiplication instruction ends, the state machine enters the finish state.
[0013] State machine Ⅱ contains 2 states: Cal_wait and ACC_out. Among them,
[0014] The Cal_wait state is the calculation waiting state, which is used to wait for the PE to complete the multiply-accumulate operation before the result data is output. The Cal_wait state is started by the end flag of the Uop_loop. When waiting for the calculation result in the inner loop, the Uop_loop state machine has entered the carry control of the new round of MAC. The independent control of the MAC carry and output enables the output of the previous round and the calculation of the new round to be executed in parallel, and the intermediate calculation cycle waiting time can be hidden and executed in parallel within the whole loop.
[0015] The ACC_out status is the output status of ACC. After receiving the MAC output handshake signal MAC_out_en, it outputs the calculation result according to the output cycle requirement through the output enable and output address, and stores the output data in the storage unit.
[0016] When the incoming of the source operands required for the inner loop calculation is completed, it directly enters the next top-level / outer loop. The uop_loop end flag starts the waiting for the calculation result of the upper-level inner loop. After receiving the MAC output handshake signal mac_out_en, it controls the start of the multi-cycle result output, and the output process is parallel to the incoming and calculation of the new round of top-level / outer loop.
[0017] The present invention also protects a matrix operation instruction controller of a data flow processor, which includes the above-mentioned dual state machines and controls the execution of matrix operation instructions based on them. Specifically, the matrix operation instruction controller includes:
[0018] The Matrix instruction execution state machine control module controls the states of state machine Ⅰ and state machine Ⅱ.
[0019] The Matrix instruction read / write address generator receives the loop index output by the Matrix instruction execution state machine control module, and generates two source input addresses per cycle for the inner loop, 1 ACC initialization address at the start of the uop loop, and multi-cycle output addresses after the uop ends according to the address calculation formula.
[0020] The Matrix instruction cache module caches Matrix and parameter configuration instructions. This instruction cache module is a synchronous instruction fifo with a size of 16×138, a bit width of 138 bits, and a depth of 16.
[0021] The Uop instruction cache module caches Uop instructions. It is a synchronous fifo with a size of 1k×64, a bit width of 64 bits, and a depth of 1K. The Uop instruction is used to initialize the start address for reading / writing from the storage unit. The start addresses for different Uop loops are different. At the start of each Uop loop, it reads the Uop instruction from the Uop instruction cache module and obtains the start address for reading and writing.
[0022] The Matrix instruction decoder decodes the instructions read from the Matrix instruction cache module, extracts the calculation control information of the Matrix multiply-accumulate execution unit, and the output information includes: Uop loop end flag, instruction type, calculation mode, output data bit width, output mode, calculation data type, ACC initialization data bit width.
[0023] The Uop instruction decoder decodes the instructions read from the Uop instruction cache module, and extracts the read / write start addresses from the storage unit, that is, the initialization address information of the address generator, including: ACC initialization address offset, weightbuffer start address, inputbuffer start address, outputbuffer start address;
[0024] Furthermore, the matrix operation instruction controller further includes a Matrix instruction dependency module, which decodes the dependency relationship information between Matrix instructions and other instructions and sends it to the data dependency resolution module to resolve the dependencies of different instructions; when there is no dependency between the Matrix instruction and other instructions, it is immediately executed; when there is a dependency between the Matrix instruction and other instructions, the Matrix instruction dependency module sends the dependency information to the data dependency resolution module and waits to be executed after receiving the inform signal.
[0025] The present invention controls the execution of matrix operation instructions based on a dual state machine, realizes independent control of input, calculation, and output states, and squeezes out the bubbles in the waiting cycles in the operation sequence, thereby improving the efficiency of multiple loop convolution operations. Brief Description of the Drawings
[0026] Figure 1 It is the state transition diagram of State Machine I;
[0027] Figure 2 It is the state transition diagram of State Machine II;
[0028] Figure 3 It is the control timing waveform of the existing matrix operation instructions;
[0029] Figure 4 It is the control timing waveform of the matrix operation instructions provided by the present invention;
[0030] Figure 5 It is the structural diagram of the matrix operation instruction controller proposed by the present invention;
[0031] Figure 6 It is the structural diagram of the address generator. Detailed Embodiment
[0032] The present invention will be further described in detail below with reference to the drawings and specific embodiments. The embodiments of the present invention are given for the purpose of illustration and description, and are not exhaustive or limited to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are selected and described to better illustrate the principles and practical applications of the present invention, and enable those of ordinary skill in the art to understand the present invention and design various embodiments with various modifications suitable for specific purposes.
[0033] Example 1
[0034] A method for controlling the execution of matrix operation instructions based on a dual state machine, where the dual state machine includes State Machine I and State Machine II.
[0035] State Machine I contains 5 states: Top_loop, Acc_ini, Uop_loop, idle, and finish, which are used to control the MAC data stream calculation. Its state transition diagram is as Figure 1 shown.
[0036] The Top_loop state is the top-level loop state, which includes the two-layer loop control of the top-level loop top_loop and the outer loop out_loop of the convolution operation. The length of the top_loop loop is top_loop_cnt, and the loop step is 1. The length of the out_loop loop is outer_loop_cnt, and the loop step is 1. Refer to Figure 1 , the Top_loop state switches to the idle state, the finish state, and the ACC_ini state under different conditions.
[0037] The Uop_loop state is the inner loop state, which includes the two-layer loop control of the micro-operation loop uop_loop and the inner loop inner_loop of the convolution operation. The length of the uop_loop loop is uop_len, and the loop step is 1. The length of the inner_loop loop is inner_loop_cnt, and the loop step is 1. This state controls the address generator to perform address calculation or output, triggers the MAC output handshake signal MAC_out_en after the end of the micro-operation loop, and returns to the Top_loop state. Refer to Figure 1 .
[0038] The ACC_ini state realizes the initialization control of the ACC register. The ACC initialization needs to be performed before each complete cycle of the Uop_loop. The ACC initialization is divided into all-zero, deep learning mode, and scientific computing mode.
[0039] The idle state is the idle state. When the state machine enters this state, it does not perform any operations and is in a waiting state.
[0040] The finish state is the end state of a complete cycle control of the state machine. When the execution of a convolution operation or a matrix multiplication instruction ends, the state machine enters the finish state.
[0041] State Machine II contains 2 states: Cal_wait and ACC_out. Its state transition diagram is as Figure 2 shown.
[0042] The Cal_wait state is the calculation waiting state, which is used to wait for the PE to complete the multiply-accumulate operation before the result data is output. The Cal_wait state is started by the end flag of Uop_loop. After the Uop_loop loop ends, the ACC output handshake signal ACC_out_en becomes valid, and at the same time, state machine II enters the Cal_wait state until it receives the MAC output handshake signal MAC_out_en and starts to output the calculation result. When waiting for the calculation result in the inner loop, the Uop_loop state machine has entered the carry control of a new round of MAC. The independent control of the MAC carry and output enables the output of the previous round and the calculation of the new round to be executed in parallel, and the waiting time of the intermediate calculation cycle can be hidden and executed in parallel within the entire loop.
[0043] The ACC_out state is the output state of ACC. When it receives the MAC output handshake signal MAC_out_en, it outputs the calculation result according to the output cycle requirements through the output enable and output address, and stores the output data in the storage unit L1_buffer. When the output is 8 bits, it takes 1 cycle; when the output is 16 bits, it takes 2 cycles; when the output is 32 bits, it takes 4 cycles.
[0044] When the carry of the source operand required for the inner loop calculation ends, it directly enters the next top / outer loop. The end flag of uop_loop starts to wait for the calculation result of the previous inner loop. When it receives the MAC output handshake signal mac_out_en, it controls the start of the multi-cycle result output, and the output process is parallel with the carry and calculation of the new round of top / outer loop.
[0045] The combined implementation of state machine I and state machine II controls the MAC convolution operation. It contains four nested loops, namely the inner loop (inner_loop), uop loop (uop_loop), outer loop (outer_loop), and top loop (top_loop). The four loops are nested in sequence. At the start of the uop loop, the ACC data in the MAC array is initialized, and after the uop loop ends, the quantized truncated output of the calculation result ACC in the MAC array is performed.
[0046] It is assumed that the MAC calculation array controlled by the present invention is a two-dimensional array, which contains R×C PEs inside and is used to complete two-dimensional convolution or multiply-accumulate operation. The main features are as follows:
[0047] 1. The MAC array includes R×C PEs, arranged in R rows and C columns;
[0048] 2. Each PE contains N 8-bit multipliers and 16-bit multipliers inside;
[0049] 3. The operand types supported are int8, uint8, int16, uint16, bfloat16, and fp16;
[0050] 4. The MAC array receives the operands of the input buffer / output buffer and the weight buffer for dot product multiplication, accumulates them with the ACC stored inside the PE, and outputs the ACC data to the output buffer;
[0051] 5. The result supports saturation and rounding processing and is output according to the quantization requirements.
[0052] When processing single-batch two-dimensional convolution, the input feature map is stored in the input buffer, the convolution kernel weights are stored in the weight buffer, the convolution calculation result is stored in the output buffer, and the input buffer, weight buffer, and output buffer are combined into one storage unit, L1_buffer.
[0053] Normally, a matrix multiply-accumulate calculation is sequentially loop-controlled through steps such as data input, multiply-accumulate operation, and data output. Its control timing waveform is as Figure 3 , where clk represents the clock signal, customized_r_en represents the read enable, customized_w_en represents the write enable, input_buffer_index represents the read address of the input data for PE calculation, and output_buffer_index represents the write address of the MAC output.
[0054] To improve the efficiency of convolution operations, the present invention makes improvements on this basis. It adopts a control method that isolates calculation from output, cuts the loop body into two parts for parallel processing. Compared with sequential loop control, it can squeeze out the timing waiting bubbles between the operation in each calculation cycle and the data truncation of the result output. Its control timing waveform is as Figure 4 .
[0055] From Figure 2 it can be seen that the output waiting time (T_wait) is hidden between the ACC initialization time T_ini and the calculation time T_cal. Since the calculation method of the controller connecting the MAC array is relatively complex and there are many output data types, under the condition that the calculation data type is floating-point 16 bits and the output bit width is 32-bit quantized relu output, the calculation and output together take more than 20 operation cycles.
[0056] Operating cycle calculation formula: (load_acc_num + inner_loop_cnt * uop_len + 1) * outer_loop_cnt *
[0057] top_loop_cnt + pe_cal_num + store_acc_num + ACC_out_num, where load_acc_num represents the ACC initialization cycle, pe_cal_num represents the cycles required for PE calculation, store_acc_num represents the cycles required for quantization truncation, and ACC_out_num represents the ACC output cycle.
[0058] Through hardware simulation experiments, when calculating the data type as floating-point 16-bit and the output bit width as 32-bit quantization relu output in a 4-layer convolution loop, and when the product of the inner loop and the uop loop length is less than 23, compared with the sequential control method, the present invention can improve the hardware operation efficiency by more than 50%.
[0059] In the present invention, the number of times of the T_wait loop is related to the product of the outer loop and the top loop. When the lengths of the outer loop and the top loop are relatively large, the beneficial effect is obvious.
[0060] Embodiment 2
[0061] A matrix operation instruction controller for a data flow processor, including the dual state machine described in Embodiment 1, and performing matrix operation instruction execution control based on it.
[0062] This matrix operation instruction controller, as Figure 5 shown, includes:
[0063] 1. Matrix instruction execution state machine control module (Matrix_pipeline_state), which controls the states of state machine Ⅰ and state machine Ⅱ.
[0064] 2. Matrix instruction read / write address generator (Matrix_addr_generator), which receives the loop index output by the Matrix instruction execution state machine control module, and realizes the generation of two source input addresses per cycle for the inner loop, 1 ACC initialization address at the start of the uop loop, and multi-cycle output addresses after the end of the uop according to the address calculation formula.
[0065] The structural diagram of the Matrix instruction read / write address generator is as Figure 6 shown, and the address generator will be described in detail below.
[0066] (1) ACC initialization address (implemented in three cases)
[0067] ① Initialize to 0
[0068] In this case, there is no need to read data from the L1_buffer, and the ACC is initialized to all 0.
[0069] ② Deep learning mode
[0070] Formula 1:
[0071] Acc_addr
[0072] = uop_buffer[uop_offset].Uop_bias_start_addr + top_loop_index * Top_bias_stride * WEIGHT_W + outer_loop_index * Outer_bias_stride * WEIGHT_W;
[0073] Where, uop_offset is the Uop instruction offset, Uop_bias_start_addr is the ACC initialization address offset, top_loop_index is the top - level loop index, Top_bias_stride is the top - level loop increment of the initialized ACC in the weight buffer, outer_loop_index is the outer - layer loop index, Outer_bias_stride is the outer - layer loop increment of the initialized ACC in the weight buffer, and WEIGHT_W is the weight.
[0074] ③ Scientific computing mode
[0075] Formula 2:
[0076] Acc_addr
[0077] = uop_buffer[uop_offset].Uop_bias_startaddr + top_loop_index * Top_output_stride * OUTPUT_W + outer_loop_index * outer_output_stride * OUTPUT_W;
[0078] Where, Top_output_stride is the top - loop increment of the initialized ACC in the input data buffer, outer_output_stride is the outer - layer loop increment of the initialized ACC in the input data buffer, OUTPUT_W is the weight, and the rest is the same as Formula 1.
[0079] (2) Source operand address for MAC calculation
[0080] Formula 3:
[0081] Input_buffer_index = uop_buffer[uop_offset + uop_loop_index (initialized to 0, incremented by 1 each uop loop)].uop_input_offset
[0082] + top_loop_index (initialized to 0, incremented by 1 each top loop) * input_stride_top
[0083] + outer_loop_index (initialized to 0, incremented by 1 each outer loop) * input_stride_outer
[0084] + inner_loop_index (initialized to 0, incremented by 1 each inner loop) * input_stride_inner;
[0085] Wherein, uop_loop_index is the micro-operation loop index, uop_input_offset is the starting offset address of the input buffer, input_stride_top is the position increment of the input data in the input buffer in the top loop, input_stride_outer is the position increment of the input data in the input buffer in the outer loop, input_stride_inner is the position increment of the input data in the input buffer in the inner loop, inner_loop_index is the inner loop index, and the rest is the same as Formula 1.
[0086] Formula 4:
[0087] Weight_buffer_index = uop_buffer[uop_offset + uop_index (incremented by 1 each uop loop)].uop_weight_offset + top_loop_index * weight_stride_top + outer_loop_index * weight_stride_out er + inner_loop_index * weight_stride_inner;
[0088] Wherein, uop_weight_offset is the starting offset address in the weight buffer, weight_stride_top is the input address offset, weight_stride_outer is the position increment of the weight data in the weight buffer in the outer loop, weight_stride_inner is the position increment of the weight data in the weight buffer in the inner loop, and the rest is the same as Formula 1.
[0089] (3) ACC output address
[0090] Formula 5:
[0091] output_buffer_index = uop_buffer[uop_offset].uop_output_offset + top_loop_index * output_stride_top + outer_loop_index * output_stride_outer;
[0092] Among them, uop_output_offset is the starting offset address of the result buffer, output_stride_top is the position increment of the result data in the result buffer in the top-level loop; output_stride_outer is the position increment of the result data in the result buffer in the outer loop, and the rest is the same as Formula 1.
[0093] ACC is divided into single-cycle or multi-cycle output according to different data types, and is enabled after each Uop loop ends. The calculation method of the starting address is as shown in Formula 3, and the address increases by 256 in each subsequent cycle.
[0094] The source operands, ACC initialization, and output formulas in the formula are all composed of multi-step multiply-accumulate operations, and are implemented using multi-level adders during hardware implementation. According to the above formula, there are only 3 cases for the change of the loop address index each time: increment by 1, reset to zero, or remain unchanged. Therefore, the multiplication of the index and the step size can be achieved by adding the step size to the result of the previous multiplication. Note that when the state transitions, it is necessary to judge whether the address resets to zero or remains the same as the previous result under different conditions. In this way, the multiply-accumulate operations in the multi-layer loops in the hardware circuit are transformed into conditional addition operations, reducing the hardware implementation logic.
[0095] 3. Matrix instruction cache module (Matrix_fifo), which caches Matrix and parameter configuration instructions; this instruction cache module is a synchronous instruction fifo with a size of 16×138, a bit width of 138 bits, and a depth of 16.
[0096] 4. Uop instruction cache module (Uop_fifo), which caches Uop instructions, is a synchronous fifo with a size of 1k×64, a bit width of 64 bits, and a depth of 1K; Uop instructions are used to initialize the starting address for reading / writing from the storage unit. The starting addresses of different Uop loops are different, and can be continuous or discontinuous; at the start of each Uop loop, the Uop instruction cache module is read to obtain the Uop instruction and the starting address of the read / write.
[0097] 5. Matrix Instruction Decoder (Matrix_decoder) decodes the instructions read from the Matrix instruction cache module, extracts the calculation control information of the Matrix multiply-accumulate execution unit, and the output information includes: Uop loop end flag, instruction type, calculation mode (deep learning / scientific computing), output data width (int8 / int16 / int32 / fp16 / fp32 / bf16), output mode, calculation data type (int8 / uint8 / int16 / uint16 / bf16 / fp16), ACC initialization data width, etc.
[0098] 6. Uop Instruction Decoder (Uop_decoder) decodes the instructions read from the Uop instruction cache module, and extracts the read / write start addresses from the storage unit, that is, the initialization address information of the address generator: uop_bias_start_addr (ACC initialization address offset), uop_wgt_start_addr (start address of weight buffer), uop_inp_start_addr (start address of input buffer), uop_out_start_addr (start address of output buffer).
[0099] 7. Matrix Instruction Dependency Module (Matrix_depend) decodes the dependency relationship information between Matrix instructions and other instructions and sends it to the data dependency resolution module to resolve the dependency relationships of different instructions. When there is no dependency between Matrix instructions and other instructions, they are executed immediately; when there is a dependency between Matrix instructions and other instructions, the Matrix instruction dependency module sends the dependency information to the data dependency resolution module and waits to be executed after receiving the inform signal.
[0100] The following introduces the instruction control process in combination with the matrix operation instruction controller disclosed in this embodiment.
[0101] Step 1. First, perform Uop memory initialization, and the Uop initialization is completed through a dedicated load instruction. The internal bus of the processor writes the initialization data from the L1_bufer to the Uopmemory through the write enable, write data, and write address of the Uop, and the written data width corresponds to the Uop instruction machine code 64bit.
[0102] Step 2. Initialize the Matrix memory. The Matrix memory is a first-in-first-out 16×138bit fifo, and the initialization is to input the Matrix instructions and related configuration instructions into the Matrix fifo.
[0103] Step 3: Before executing the Matrix convolution instruction, the controller first fetches the Matrix instruction from the Matrix fifo, parses and obtains the necessary control information through the Matrix_decoder, such as parameters Weight_stride_top (input address offset), Input_stride_top (input address offset), Out_stride_top (output address offset), top_loop_cnt (top layer loop length), outer_loop_cnt (outer layer loop length), uop_len (Uop loop length), inner_loop_cnt (inner layer loop length), uop_offset (Uop instruction offset), ping_pang (ping-pong information), etc.
[0104] Step 4: Send the instruction dependency information to the depend arbitration module through Matrix_depend and wait to receive the inform information. When the controller receives the inform information, it starts to work. First, it sends a bus occupancy request customized_require to the L1_buffer and waits to receive the customized_rlock information sent from the L1_buffer, indicating that the current L1_buffer is ready for read and write operations.
[0105] Step 5: The module Matrix_pipeline_state control state machine starts to work, and realizes the counting control of a total of 4 layers of loops, namely inner_loop, uop_loop, outer_loop, and top_loop, through 5 states: Top_loop, ACC_ini, Uop_loop, finish, and idle. Realize the output cycle counting control through the two states Cal_wait and ACC_out.
[0106] This module finally generates counting enable control and the indexes of each layer of loop, Top_loop_index, out_loop_index, uop_loop_index, inner_loop_index, including other information in the address generator calculation formula. In addition, it also includes the index calculation read enable for the Uop memory. A read operation needs to be performed on the Uop memory during each uop calculation cycle, ACC initialization, and output to obtain the offset information for address calculation, and finally output the calculation index to the address generator.
[0107] Step 6: The address generator obtains the Uop read address by adding uop_offset (Uop instruction offset) and uop_loop_index (Uop loop index). According to the Uop memory read enable generated by the state machine, it reads the uop_fifo to obtain the uop instruction used in the current Uop loop. This instruction passes through the Uop instruction decoder to obtain the initialization address information of the address generator: uop_bias_start_addr (ACC initialization address offset), uop_wgt_start_addr (starting address of the weight buffer), uop_inp_start_addr (starting address of the input buffer), and uop_out_start_addr (starting address of the output buffer).
[0108] The information such as the indexes and starting addresses extracted in the above steps are finally summarized in the Matrix_addr_generator, and the address generator generates read and write addresses. The address generation process follows the calculation method of the address formula (see Address Calculation Formulas 1-5 in the detailed description of the address generator), and through the index information of four-layer loops, the address offset of the uop instruction, and the accumulation of various step information, the MAC initialization address ACC_addr, the read addresses Input_buffer_index and Weight_buffer_index for the PE to calculate the input data, and the write address output_buffer_index for the MAC output are obtained, etc.
[0109] The address information generated by the above address generator, together with the read and write enables customized_r_en and customized_w_en and the read and write selections customized_rdevice and customized_wdevice generated by the state machine, are sent to the L1_buffer module.
[0110] Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art and related fields based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
Claims
1. A matrix operation instruction execution control method based on a dual state machine, characterized in that The dual state machine includes state machine Ⅰ and state machine Ⅱ; State machine Ⅰ contains 5 states: Top_loop, Acc_ini, Uop_loop, idle, and finish, which are used to control the MAC data stream calculation. Among them, The Top_loop state is the top-level loop state, which includes the control of two nested loops: the top-level loop and the outer loop of the convolution operation; The Uop_loop state is the inner loop state, which includes the control of two nested loops: the micro-operation loop and the inner loop of the convolution operation. This state controls the address generator to perform address calculation or output, triggers the MAC output handshake signal MAC_out_en after the micro-operation loop ends, and returns to the Top_loop state; The ACC_ini state realizes the initialization control of the ACC register, and the ACC initialization needs to be performed before each complete cycle of the Uop_loop; The idle state is the idle state. When the state machine enters this state, it does nothing and is in a waiting state; The finish state is the end state of a complete cycle control of the state machine. When a convolution operation or matrix multiplication instruction execution ends, the state machine enters the finish state; State machine Ⅱ contains 2 states: Cal_wait and ACC_out. Among them, The Cal_wait state is the calculation waiting state, which is used to wait for the PE to complete the multiply-accumulate operation before the result data is output. The Cal_wait state is started by the end flag of the Uop_loop. When waiting for the calculation result in the inner loop, the Uop_loop state machine has entered the incoming number control of a new round of MAC. The independent control of the MAC incoming number and output enables the output of the previous round and the calculation of the new round to be executed in parallel, and the waiting time of the intermediate calculation cycle can be hidden and executed in parallel in the whole loop; The ACC_out state is the output state of the ACC. When receiving the MAC output handshake signal MAC_out_en, it outputs the calculation result according to the output cycle requirements through the output enable and output address, and stores the output data in the storage unit; After the incoming number of the source operands required for the inner loop calculation ends, it directly enters the next top-level / outer loop. The end flag of the uop_loop starts to wait for the calculation result of the upper-level inner loop. When receiving the MAC output handshake signal mac_out_en, it controls the start of the multi-cycle result output, and the output process is parallel with the incoming number and calculation of the new round of top-level / outer loop.
2. A matrix operation instruction controller for a data flow processor, characterized in that It includes the dual state machine described in claim 1, and controls the execution of matrix operation instructions based on it.
3. The matrix operation instruction controller of the data flow processor according to claim 1, characterized in that, It includes: The Matrix instruction execution state machine control module, which controls the states of state machine Ⅰ and state machine Ⅱ; The Matrix instruction read-write address generator, which receives the loop index output by the Matrix instruction execution state machine control module, and generates two source input addresses for each cycle of the inner loop, 1 ACC initialization address at the start of the uop loop, and multi-cycle output addresses after the uop ends according to the address calculation formula; The Matrix instruction cache module, which caches Matrix and parameter configuration instructions; Uop instruction cache module, which caches Uop instructions. The Uop instructions are used to initialize the starting address for reading / writing from the storage unit. The starting addresses of different Uop loops are different. At the start of each Uop loop, a Uop instruction is read from the Uop instruction cache module and the starting address for reading / writing is obtained. Matrix instruction decoder, which decodes the instructions read from the Matrix instruction cache module, extracts the calculation control information of the Matrix multiply-accumulate execution unit, and the output information includes: Uop loop end flag, instruction type, calculation mode, output data width, output mode, calculation data type, ACC initialization data width. Uop instruction decoder, which decodes the instructions read from the Uop instruction cache module and extracts the starting address for reading / writing from the storage unit.
4. The matrix operation instruction controller of the data flow processor according to claim 1, wherein It also includes a Matrix instruction dependency module, which decodes the dependency relationship information between Matrix instructions and other instructions and sends it to the data dependency resolution module for resolving the dependency relationships of different instructions.
5. The matrix operation instruction controller of the data flow processor according to claim 4, characterized in that When there is no dependency between the Matrix instruction and other instructions, it is immediately executed; when there is a dependency between the Matrix instruction and other instructions, the Matrix instruction dependency module sends the dependency information to the data dependency resolution module and waits to be executed after receiving the inform signal.