Convolutional neural network accelerator instruction set architecture and Faster R-CNN algorithm deployment method
Through the assistance of custom instruction set architecture and coprocessor, the problem of Faster R-CNN algorithm in fully quantized inference and complex control algorithm deployment on FPGA devices is solved, and efficient parallel computing and flexible adaptation of the algorithm are realized.
Patent Information
- Application Number
- CN202510290376.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
AI Technical Summary
The existing Faster R-CNN algorithm cannot implement full quantization inference on FPGA devices, and traditional convolutional neural network accelerators are difficult to deal with complex control algorithms such as non-maximum suppression algorithms, resulting in high algorithm deployment complexity.
It provides a method for deploying a convolutional neural network accelerator instruction set architecture and Faster R-CNN object detection algorithm. It assists the main processor in sharing computing pressure through a coprocessor, and optimizes data interaction and computing pipelines through a custom instruction set architecture, supporting multi-task parallel computing.
The fast deployment of Faster R-CNN algorithm on FPGA devices is realized, reducing the complexity of the algorithm, enhancing parallel computing capabilities, and adapting to the flexibility of different processor architectures.
Smart Images

Figure CN120297341A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of AI hardware accelerators, and relates to an instruction set architecture of a convolutional neural network accelerator and a method for deploying the Faster R-CNN object detection algorithm to an FPGA device. Background Art
[0002] At the present stage, the Faster R-CNN algorithm with high recognition accuracy still cannot achieve full quantization inference, and the existing Pytorch quantization methods only support fully convolutional neural networks, and cannot directly perform full quantization work on two-stage object detection algorithms. At the same time, the Faster R-CNN algorithm has high complexity, including not only the convolutional neural network operation layer but also other data processing algorithms, among which there are complex control algorithms such as non-maximum suppression algorithms. This is often difficult to achieve for traditional convolutional neural network accelerators, and such control algorithms must be implemented by the CPU. This architecture provides a special solution to this problem, which can realize the rapid deployment of the Faster R-CNN algorithm on the FPGA device, and at the same time change the circuit scale according to the FPGA board resources to achieve a compromise that meets the actual requirements. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide an instruction set architecture of a convolutional neural network accelerator and a deployment method of the Faster R-CNN object detection algorithm, which can realize the deployment of the Faster R-CNN algorithm from a higher level without deeply understanding the design method of the hardware circuit.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] An instruction set architecture of a convolutional neural network accelerator and a deployment method of the Faster R-CNN object detection algorithm, which includes an instruction set architecture and a series of methods for deploying the Faster R-CNN algorithm. Among them,
[0006] The present invention can use a coprocessor to help the main processor share the operating pressure of the object detection algorithm. Therefore, there are two ways to realize the data interaction between the main processor and the accelerator according to the different architectures of the main processor:
[0007] ① It is completed through the dedicated access interfaces at both the master and slave ends. For example, the RISC-V architecture itself can add a coprocessor interface according to design requirements, and access the coprocessor through specific instruction set extensions and other means.
[0008] ② The instructions, configurations and data in the registers at both ends of the master and slave are transferred through the coprocessor communication instructions to realize the master-slave interaction function. For example, the ARM architecture can realize the communication between the main processor and the coprocessor through the MCR and MRC instructions.
[0009] The two have great similarities in data interaction channels and synchronization mechanisms, and have similar features. The only difference is that the RISC-V architecture, as an open source instruction set architecture, gives developers the space to flexibly design coprocessors, while the ARM architecture requires developers to design coprocessors in accordance with relevant specifications. In addition to providing a more popular data interaction interface, the present invention can still adapt to the corresponding main processor requirements by adding an external compatible module.
[0010] The Faster R-CNN accelerator can usually continuously receive multiple instructions from the main processor, which act on configuration, calculation, memory access and routing functions respectively. After receiving a certain number and type of instructions, the accelerator transmits these instructions to the relevant processing unit, after which these instructions begin to take effect. These computing modules can perform parallel calculations under the guidance of corresponding instructions. For the circuit calculation module in the present invention, its parallel calculation granularity is relatively small, and it can usually complete two to four computing tasks at the same stage, and is generally configured as a pipeline operation mode.
[0011] According to the definition of Flynn computer classification, the convolutional neural network accelerator instruction set architecture provided by the present invention can control the accelerator to implement the MIMD (multiple instruction stream multiple data stream) execution process, that is, it can process multiple different instructions at the same time, and each operation instruction can calculate the data stream in the form of an array. Such a top-level architecture can help data calculations to achieve a combination of parallelization and pipeline.
[0012] In order to better adapt to the Faster R-CNN neural network target detection algorithm, the present invention also makes the following optimizations to the instruction flow and configuration flow architecture of the custom architecture compared to the traditional MIDI architecture:
[0013] ① This architecture deletes control instructions such as instruction jumps and branch challenges, and implements the complex control flow required by some operators in the Faster R-CNN algorithm with fixed control circuits to control the transmission direction of the data flow;
[0014] ② Its operation instructions are mainly coarse-grained operation instructions, which are dedicated to the processing work of the internal operation layer of the convolutional neural network;
[0015] ③ Its memory access instructions are mainly used for data access between the accelerator and external storage. The cache inside the accelerator does not access data through memory access instructions, and its memory access behavior is indirectly regulated by arithmetic instructions and arithmetic configurations.
[0016] ④ The configuration stream first has the ability to program the circuit operation scale. Secondly, since most of the operators in this algorithm are vector parallel operation types, the multi-processor routing configuration in the traditional configuration stream is simplified into the routing configuration between several commonly used operator operation arrays.
[0017] The custom instruction set architecture of the convolutional neural network accelerator proposed by the present invention uses a 32-bit format and can be adapted to mainstream processors in the form of soft instruction configuration. Since the configuration stream of the present invention is simplified, the configuration types and functions are relatively simple, and the configuration stream is tightly coupled with the instruction stream, so the instructions and configuration commands are uniformly included in the instruction set architecture below. This instruction set can be roughly divided into three behaviors: computing, configuration, and memory access and routing, including seven instruction types, namely: C1-type, C2-type, C3-type, A-type, U-type, F-type, and D-type.
[0018] All instruction types contain i / f and opcode fields. These two fields are the key indexes for the accelerator decoding system to distinguish instruction modes. Among them, the i / f field is used to indicate whether this instruction is used for configuration functions. When it is 0, it means that the current is an operation mode instruction, and when it is 1, it means that the current is a configuration mode instruction. The opcode field is an instruction type index used to indicate the detailed operation or configuration type. The instruction formats of some types also include fun3 field and other field. Since there are a certain number of variants under some instructions of the accelerator, the fun3 field is used to describe the variant type selected by the current specific instruction, and the other field usually serves two purposes: one is to load the information that is frequently changed when switching different variants, and the other is to provide it as an additional field for subsequent optimization.
[0019] Instructions of three types, C1-type, C2-type, and C3-type, are collectively referred to as Cx-type execution. These instructions share a basic instruction format, and only some formats are inconsistent. Cx-type is used to support various operations of sliding window types. Among them, C1-type is dedicated to convolutional operation types, C2-type is dedicated to pooling operation types, and C3-type is dedicated to fully connected layer operation types. The fields in the instruction format are used to indicate the specific instruction indexes under the current type, including fields indicating the sliding step of convolution, the size of the convolution kernel, and the output channels, etc. These common key features are sufficient to support any convolutional layer for specific calculations.
[0020] The general convolution operation adapts to different target algorithms with different convolution channel depths, convolution kernel sizes, and other operation characteristics. By using the fun3 bit for switching, different bit expansion modes are selected to increase the lengths of different fields within the instruction.
[0021] A-type instructions refer to various array operation instructions that need to be executed on the complex vector calculation processing module except for Cx-type instructions, including: A_ACC, A_CUM, A_MAC, A_ADD, A_MAC, etc., as well as their immediate operand version operation instructions. A-type instructions are designed for the RPN network of the Faster R-CNN algorithm. In addition to having three convolutional layers, the RPN network also includes other operations such as the Softmax layer, Loc2bbox (algorithm for calculating the position and size of the corrected bounding box), NMS algorithm (non-maximum suppression algorithm), and system screening algorithm. For example, Softmax in these operations can also utilize the multiply-accumulate array, such as the method for directly executing the Softmax algorithm proposed in the present invention. Therefore, compared with C-type which is used for the dedicated multiply-accumulate array convolution controller, A-type realizes the direct deployment of other sub-algorithms through instruction control.
[0022] U-type instructions are used to support other general operation modes except for array and sliding window type operations, such as the non-maximum suppression algorithm and bounding box correction algorithm mentioned above. This architecture requires these algorithms to be implemented using dedicated circuits, and the key variables of their interfaces are transmitted by the immediate operand field of U-type.
[0023] F-type instructions are operation-specific configuration instructions, and their main function is to guide the mapping of array algorithms, and also include some re-quantization parameters, etc.
[0024] D-type instructions are a combined type of memory access instructions and configuration instructions, which are used to guide the behavior of external memory access and for the specific allocation of data routing.
[0025] The instruction set architecture proposed by the present invention allows the accelerator to have the parallel computing ability of multiple computing modules working in the same stage, and each computing module has its own independent instruction decoder, controller, and arithmetic unit inside. However, these computing modules are different from general processing elements. Different computing modules inside the accelerator are dedicated to different types of operations, and their circuit structures are completely different, so there is only a local general processing function.
[0026] In order to enable object detection neural network algorithms such as Faster R-CNN to have a multi-task parallel computing mode, the instruction set architecture proposed by the present invention supports driving the accelerator with three working modes: single-task working mode, dual-task mode, and multi-task working mode. The single-task working mode is suitable for centralized processing of some adjacent structures in the convolution fusion layer of the CNN model. The dual-task working mode is usually suitable for structures with one layer of other operation types sandwiched between two convolution fusion layers. The multi-task working mode is suitable for structures such as the RPN network where a single convolution layer is followed by multiple other operations. The principle for determining the current task parallel working mode is to calculate the length of the non-repeating calculation types of multiple operation instructions fetched by the instruction fetch unit of the accelerator, and this length is the number of tasks that can be parallelly calculated in the current stage. In the circuit structure of task parallelism, the accelerator under this instruction set architecture builds a module pipeline structure with multiple computing modules to achieve its function. Since the calculation period of the convolutional neural network algorithm is relatively long, the load delay of the pipeline composed of multiple task modules can be ignored. The present invention proposes a single-task working mode, which requires the accelerator to have a three-stage pipeline circuit structure composed of convolutional operation, requantization operation, and Clamp operation. In addition to the basic pipeline structure, the pipeline structure of the dual-task computing mode also has a fourth-stage pipeline structure for auxiliary computing modules. Any computing unit of the auxiliary computing module can be used as the circuit that can act as the fourth-stage pipeline in the Faster R-CNN algorithm accelerator, such as an NMS computing module, an exponential operation module, or a pooling operation module. The addition of the subsequent pipeline will not significantly increase the overall computing time of the accelerator. When designing the auxiliary computing module of the subsequent pipeline, the key rules of pipeline design are fully considered to ensure that the data processing period of the auxiliary computing module is less than or equal to the data output period of the main computing module (generally speaking, the requantization algorithm and the auxiliary computing algorithm are mostly single-cycle algorithm types or the operation cycles are much less than the previous convolutional operation cycle). Therefore, the dual-task mode will not cause a reduction in the resource utilization rate of the main computing array.
[0027] The Faster RCNN quantization method proposed by the present invention performs quantization by dividing the schemes according to the principle of whether various operation functions in the Faster R-CNN algorithm belong to the linear mapping relationship. The present invention can divide the specific quantization strategies into two categories: one is for operation processes such as convolutional layers, fully connected layers, pooling layers, Relu activation functions, and BN layers. These operations are either linear mapping relationships themselves, or do not need to be re-quantized, or can be quantized in combination with other layers. For such objects, general static quantization methods are adopted. The other is for non-linear operations such as Softmax. Since using the direct quantization method will cause the quantized data to not be accurately restored to the original floating-point numbers, the present invention adopts the piecewise fitting function method for such objects. The quantization of the Faster R-CNN algorithm should first perform preparatory work: prepare the trained neural network, prepare the training dataset and the test dataset, set the quantization basic configuration, etc.; the second step is to perform model fusion on the neural network. The neural network operation layers that can be fused include the following categories:
[0028] ①[Conv,BN]
[0029] ②[Conv,BN,Relu]
[0030] ③[Conv,Relu]
[0031] ④[BN,Relu]
[0032] The third step is to classify the linear operations and linear operations in the entire algorithm; the third step will perform tasks such as marking, inserting quantization nodes, and additional processing on the two types of objects and the internal detailed operation layers (functions) of each object. The task of marking is to remind the quantization observer which operation layers (functions) do not need to be quantized, or can be quantized after model fusion, which operation layers (functions) need to be specially quantized, and which need to be normally quantized. Inserting quantization nodes is directly related to marking and is used to distinguish the scopes of different quantization strategies. Additional processing is an extra operation for special quantization; the next step is the process of specifically collecting parameters and quantizing data. By looping the training data into the neural network with quantization nodes inserted, the quantization observer can capture the maximum and minimum values of each channel / layer data. The data for which the maximum and minimum values need to be statistically calculated include the convolutional layer weights, convolutional layer biases, and activation values. Using the maximum and minimum values, the quantization scaling factors and zero-point offsets of the corresponding data can be calculated. Finally, the weights and bias data of the convolutional layer or fully connected layer can be quantized according to relevant strategies and the determined quantization parameters.
[0033] The hardware architecture pre-deployment processing method proposed by the present invention includes the following contents:
[0034] 1) Image data and weight data reshape: Deep convolutional neural network accelerators generally adopt design schemes that can specifically accelerate tensor calculations. However, due to the limitations of embedded devices, the tensor calculation unit only sequentially accesses each dimension of the operand tensor according to a fixed scale. When accessing a certain dimension, the tensor calculation unit needs to fetch all the weights and bias data for the current calculation from the data memory at once. Considering that the data involved in a single tensor calculation may involve multiple dimensions, the image data and convolution kernel parameters need to be reshaped before being stored in the data memory.
[0035] 2) Considering the requirements of FPGA tools and Verilog programs for externally read data, various types of data need to be normalized and exported from Python tools. The requirements include: during the FPGA simulation stage, the $readmemh function needs to be used to read in data; when actually programming the FPGA board, the weight data needs to be converted to Hex format and then read into the SD card; the ROM of the FPGA outputs a long vector. When the neural network has been quantized, the image data, convolution kernel weight parameters, convolution kernel bias, quantization scaling factor, and quantization zero-point offset all need to be converted into the Hex format of the corresponding signed (unsigned) numbers, and then data dimension adjustment and splicing and other processing work need to be carried out.
[0036] 3) Fixed-point processing of quantization parameters: For currently well-developed deep convolutional neural networks, the absolute values of the weights and bias data of the convolutional layer after training are generally less than 1. Therefore, its quantization scaling factor is much less than 1, and the zero-point offset is an 8-bit integer. In order to store the scaling factor on the FPGA and use it to complete quantization, it is necessary to pre-complete the fixed-point processing of the scaling factor-related parameters on the Python tool. The present invention completes this preprocessing according to the binary fixed-point mapping method.
[0037] 4) Neural network model fusion: Although neural network quantization can make neural network parameters become concise, there are still many calculation layers in the neural network model that will not be used or can be optimized during the inference stage. Therefore, it is necessary to adopt neural network model fusion, that is, the convolutional fusion layer introduced above. Description of the Drawings
[0038] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably below in conjunction with the drawings, where:
[0039] Figure 1 is the instruction set architecture of the convolutional neural network accelerator;
[0040] Figure 2 is the Cx-type instruction bit-width extension rule;
[0041] Figure 3 It is a schematic diagram of the operation pipeline in the dual-task working mode;
[0042] Figure 4 It is a flowchart of the static quantization acquisition parameters of the Faster R-CNN algorithm; Specific implementation manners
[0043] Next, the technical solutions in the embodiments of the present invention will be described in detail with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0044] The present invention can use a coprocessor to help the main processor share the operation pressure of the target detection algorithm. Therefore, there are two ways of data interaction between the main processor and the accelerator according to different architectures of the main processor.
[0045] The present invention can use a coprocessor to help the main processor share the operation pressure of the target detection algorithm. Therefore, according to different architectures of the main processor, there are two ways to realize the data interaction between the main processor and the accelerator:
[0046] ① It is completed through the dedicated access interfaces at both the master and slave ends. For example, the RISC-V architecture itself can add a coprocessor interface according to design requirements and realize the access to the coprocessor through specific instruction set extensions and other methods.
[0047] ② The instructions, configurations, and data in the registers at both the master and slave ends are transmitted through the coprocessor communication instructions to realize the function of master-slave interaction. For example, the ARM architecture can realize the communication between the main processor and the coprocessor through the MCR and MRC instructions.
[0048] These two have great similarities in the data interaction channel and the synchronization mechanism and have the same effect. The only difference is that the RISC-V architecture, as an open-source instruction set architecture, gives developers the space to flexibly design the coprocessor, while the ARM architecture requires developers to design the coprocessor in accordance with relevant specifications. In addition to the relatively popular data interaction interface given by the present invention, it can still adapt to the requirements of the corresponding main processor by adding an external compatible module.
[0049] The present invention realizes the access of the main processor through the following access channels:
[0050] The Inst-Response Channel (instruction response channel) includes interface signals such as Req_sys_valid, Req_sys_ready, and Instruction. This channel is used to realize the function of the main processor sending instruction request commands and related content to the coprocessor, and it ensures the correct communication between the main processor and the coprocessor through the handshake protocol.
[0051] The Imem-Request Channel (memory access request channel) includes interface signals such as App_wdf_rdy, App_wdf_wren, App_wdf_end, and App_wdf_data. This channel is used to implement the function of the coprocessor sending memory access request commands and related content to the off-chip memory. It ensures the correct communication between the off-chip memory and the coprocessor through a handshake protocol.
[0052] The Imem-Response Channel (memory access response channel) includes interface signals such as App_rd_data_valid, App_rd_data_end, and App_rd_data. This channel is used to implement the function of the off-chip memory sending memory access response commands and related content of read data to the coprocessor. It ensures the correct communication between the off-chip memory and the coprocessor through a handshake protocol.
[0053] The Additional Alu-Request Channel (additional arithmetic unit request channel) includes interface signals such as Req_alu_valid, Req_alu_ready, Cal_data, and Cal_itag. This channel is used to implement the function of the coprocessor sending memory access request commands and the data requested for transmission to the additional arithmetic unit. It ensures the correct communication between the off-chip memory and the coprocessor through a handshake protocol.
[0054] The present invention proposes a custom instruction set architecture based on the Faster R-CNN accelerator. It adopts a 32-bit format and can be adapted to mainstream processors in the form of soft instruction configuration. Since the configuration flow of this design is simplified, the configuration types and functions are relatively simple, and the configuration flow is tightly coupled with the instruction flow, the configuration flow and instruction flow situations are described together in this article, and their architectures are uniformly included in the instruction set architecture. This instruction set can be roughly divided into three behaviors: computing, configuration, and memory access and routing, including seven instruction types, namely: C1-type, C2-type, C3-type, A-type, U-type, F-type, and D-type. As Figure 1 shows the formats of different types of instructions adopted by the present invention.
[0055] Among all instruction types, both the i / f and opcode fields are crucial indexes for the accelerator decoding system to identify instruction patterns. The i / f field is used to indicate whether this instruction is for configuration functions. When it is 0, it means the current is an operation mode instruction, and when it is 1, it means the current is a configuration mode instruction. The opcode field is the instruction type index, used to indicate the detailed operation or configuration type. The instruction formats of some types also include the fun3 field and the other field. Since there are a certain number of variants for some instructions of the accelerator, the fun3 field is used to describe the variant type selected by the current specific instruction, and the other field usually serves two purposes: one is to load the information that is frequently changed when switching different variants, and the other is to be provided as an additional field for subsequent optimization.
[0056] Figure 1 Use Cx-type to represent three types of instructions: C1-type, C2-type, and C3-type. These instructions share a common basic instruction format, and only some formats are inconsistent. Cx-type is used to support various sliding window type operations. Among them, C1-type is dedicated to convolution operation types, C2-type is dedicated to pooling operation types, and C3-type is dedicated to fully connected layer operation types. The fun3 field in the instruction format is used to indicate the specific instruction index under the current type. The ksize field represents the length (width) of the sliding window, the pd field represents the zero element padding length (width) of the image, the imme field represents the immediate number, the std field represents the sliding step (width), and the other field is the switching extension bit. The Cx-type format shown in the figure is a unified demonstration of the overall formats of the three types of instructions. In fact, when the opcode field takes different values, different instructions should have different imme bits and other bits. When opcode = 10110, it correspondingly indicates that the current instruction is a general convolution operation instruction. The imme is used to transfer the number of output channels (C o ) of the convolution operation, that is, the oc field, and the other bit is used for expanding the bit width. The expansion situation is as Figure 2 shown.
[0057] In order to adapt to different target algorithms with different convolution channel depths, convolution kernel sizes, and other operation characteristics, the general convolution operation switches through the fun3 bit to select different bit expansion modes to increase the lengths of different fields within the instruction.
[0058] A-type instructions refer to various array operation instructions that need to be executed on the complex vector calculation processing module, excluding Cx-type instructions, including: A_ACC, A_CUM, A_MAC, A_ADD, A_MAC, etc., as well as their immediate operand version operation instructions. The rs1 field and rs2 field in the instruction format represent the memory numbers of the operand vectors for the current array calculation. The m1 and m2 fields are used to indicate whether the current instruction needs to access the rs1 and rs2 memories. The imme bit is still the immediate operand bit, providing data for immediate operand array operations. The size1 and size2 fields represent the lengths of continuously read data from the operand vector 1 memory and the operand vector 2 memory respectively for the current array operation.
[0059] U-type instructions are used to support other general operation modes except for array and sliding window type operations, such as exponential operations included in the RPN network. Its instruction format contains three immediate operand fields and is often used to pass special parameters within the algorithm.
[0060] F-type instructions are dedicated operation configuration instructions, and their main function is to guide the mapping of array algorithms. The scs, seq, and cb fields in this instruction format are used to select the memory-computation mode, the calculation order of sliding window type algorithms, and the data post-processing method during the array calculation stage respectively. The quan field indicates the type of quantization strategy of the current requantization unit, and zero is the zero offset parameter required for requantizing the current calculation result. The size0 and size1 fields are used to guide the two-dimensional mapping method of array algorithms.
[0061] D-type instructions are a combined type of memory access instructions and configuration instructions, which are used to guide the behavior of external memory access and for the specific allocation of data routing. Among them, the rs1 and rs2 fields represent the first-level and third-level pipeline data memory numbers of the accelerator system respectively, indicating the other end of data transmission during external memory access. The ff and fp fields are the first calculation indication bit and the ping-pong memory access indication bit respectively. The bank0 and bank1 fields are the bank numbers for writing and reading DDR during external memory access. Only when the current memory access behavior is not in the ping-pong memory access state, the external memory access controller of the accelerator can perform specified bank read and write according to the data in the bank1 and bank0 fields.
[0062] C1-type instructions include instructions such as UCV, C2-type instructions include MPL, APL, and RPL, etc., C3-type instructions include instructions such as FCL, A-type instructions include instructions such as A_ACC, A_CMU, A_MAC, A_ADD, A_MUL, A_ACCI, A_CMUI, A_MACI, A_ADDI, and A_MULI, F-type instructions only include the ACG instruction, U-type instructions include instructions such as SFT, LBB, and NMS, and D-type instructions include instructions such as FWD, DRD, WAR, and RAW.
[0063] The Faster R-CNN accelerator can usually continuously receive multiple instructions from the main processor, and these instructions act on the configuration, computing, memory access, and routing functions respectively. After receiving a certain number and type of instructions, the instruction control unit drives the instruction issue slot to fill with instructions. The instruction scheduling strategy of the Faster R-CNN accelerator allows the instruction issue slot to assemble multiple configuration instructions and computing instructions and issue them to different computing modules. These computing modules can perform parallel computing under the guidance of the corresponding instructions. For the circuit computing module in the present invention, its parallel computing granularity is relatively small, and it can usually complete two to four computing tasks in the same stage and is generally configured to operate in a pipeline mode.
[0064] The instruction set architecture proposed by the present invention allows the accelerator to have the parallel computing ability of multiple computing modules working in the same stage, and each computing module has its own independent instruction decoder, controller, and arithmetic unit inside. However, these computing modules are different from general processing elements. Different computing modules inside the accelerator are dedicated to different types of operations, and the circuit structures are completely different, so there is only a partial general processing function.
[0065] In order to enable object detection neural network algorithms such as Faster R-CNN to have a multi-task parallel computing mode, the instruction set architecture proposed by the present invention supports driving the accelerator with three working modes: single-task working mode, dual-task mode, and multi-task working mode. The single-task working mode is suitable for centrally processing some adjacent structures in the convolutional fusion layer of the CNN model. The dual-task working mode is usually suitable for structures with one layer of other operation types interspersed between two convolutional fusion layers. The multi-task working mode is suitable for structures such as the RPN network where a single convolutional layer is followed by multiple other operations. The principle for determining the current task parallel working mode is to calculate the length of non-repeating computational types of multiple operation instructions fetched by the instruction fetch unit in the current stage, and this length is the number of tasks that can be parallelly computed in the current stage. In the circuit structure of task parallelism, the accelerator under this instruction set architecture constructs a module pipeline structure with multiple computing modules to achieve its function. Since the computational cycle of the convolutional neural network algorithm is relatively long, the loading delay of the pipeline composed of multiple task modules can be ignored. As Figure 3 shown, the present invention proposes a single-task working mode, which requires the accelerator to have a three-stage pipeline circuit structure composed of convolutional operation, requantization operation, and Clamp operation. In addition to the basic pipeline structure, the pipeline structure of the dual-task computing mode also has a fourth-stage pipeline structure for auxiliary computing modules. Operations that can act as the fourth-stage pipeline within the Faster R-CNN algorithm can be the NMS algorithm or exponential operation, etc. The addition of the subsequent pipeline will not significantly increase the computational duration of the sub-algorithm running in the current stage.
[0066] As Figure 4 shown, according to whether various operation functions in the Faster R-CNN algorithm are linear mapping relationships, the present invention can divide specific quantization strategies into two categories: one is for operation processes such as convolutional layers, fully connected layers, pooling layers, Relu activation functions, and BN layers. These operations are either linear mapping relationships themselves, or do not require requantization, or can be quantized in combination with other layers. For such objects, general static quantization methods are adopted. The other is for non-linear operations such as Softmax. Since using the direct quantization method will result in the inability to accurately restore the quantized data back to the original floating-point numbers, in this paper, the method of using piecewise algebraic function to fit the original function is adopted to achieve quantization for such objects.
[0067] According to Figure 4 shown, the quantization of the Faster R-CNN algorithm should first carry out preparatory work: prepare the trained neural network, prepare the training dataset and test dataset, set the quantization basic configuration, etc.; the second step is to fuse the neural network. The neural network operation layers that can be fused include the following categories:
[0068] ① [Conv,BN]
[0069] ②[Conv, BN, Relu]
[0070] ③[Conv, Relu]
[0071] ④[BN, Relu]
[0072] The third step is to classify the linear operations and linear operations in the entire algorithm; the third step will perform tasks such as marking, inserting quantization nodes, and additional processing on two types of objects and the internal detailed operation layers (functions) of each object. The task of marking is to remind the quantization observer which operation layers (functions) do not need to be quantized, or can be quantized after model fusion, which operation layers (functions) need to be specially quantized, and which need to be normally quantized. Inserting quantization nodes is directly related to marking and is used to distinguish the ranges of different quantization strategies. Additional processing is an additional operation for special quantization; the next step is the process of specifically collecting parameters and quantization data. By looping the training data into the neural network with quantization nodes inserted, the quantization observer can capture the maximum and minimum values of each channel / layer data. The data for which the maximum and minimum values need to be statistically calculated include the weights of the convolutional layer, the biases of the convolutional layer, and the activation values. Using the maximum and minimum values, the quantization scaling factors and zero-point offsets of the corresponding data can be calculated. Finally, the weights and bias data of the convolutional layer or fully connected layer can be quantized according to relevant strategies and the determined quantization parameters.
[0073] It is determined that operations that do not change the value range do not need to collect quantization parameters again and can be directly inherited. For example, non-linear operations such as Softmax that change the value range need to be implemented by fitting with a piecewise algebraic function. Since Softmax uses the exponential function as the key operator, taking the exponential algorithm of the present invention as an example, if it is determined that the current calculation satisfies the exponential function, it needs to be processed by an identity transformation first, as shown in the following formula:
[0074]
[0075] This is to limit the value range of the exponential function within a finite interval, such as within (0, 1], and then scale it up to the interval [0, 255]. Because the independent variable range of the exponential function in the Faster R-CNN algorithm is normalized and generally within the range of [-1, 1], and in extreme cases, it will not exceed much.
[0076] After the identity transformation, only \(e\) needs to be considered z How to implement it in the integer calculation unit on the hardware, the present invention uses a piecewise function fitting method to replace the original exponential function. The expression of this piecewise function is as follows:
[0077]
[0078] The mean square error after fitting this piecewise function is less than 0.0048, and the average absolute error is less than 0.038. Using this function to replace the exponential function and combining with the maximum and minimum values obtained from statistics to calculate the quantization scaling parameter, the quantized value of \(e\) can be calculated. z Finally, multiplying it by \(e\) ramx can obtain the quantized value of the exponent. In the present invention, \(e\) ramx and the coefficients of the piecewise function can all be replaced by fixed-point numbers through the fixed-point processing of quantization parameters.
[0079] The preprocessing scheme for the deep convolutional neural network according to the hardware architecture includes: reshape of image data and weight data, data normalization export, fixed-point processing of quantization parameters, and neural network model fusion, etc.
[0080] 1) Reshape of image data and weight data: The deep convolutional neural network accelerator generally adopts a design scheme that can specifically accelerate tensor calculations. However, due to the limitations of embedded devices, the tensor calculation unit will only sequentially access each dimension of the operand tensor according to a fixed scale. When accessing a certain dimension, the tensor calculation unit needs to fetch all the weights and bias data for the current calculation from the data memory at one time. Considering that the data involved in a single tensor calculation may involve multiple dimensions, the image data and convolution kernel parameters need to be reshaped before being stored in the data memory. The conversion rules are shown in the following formula. In this article, a 16×16 MAC array is used for tensor calculations, including 16 input channel dimensions and 16 output channel dimensions. Therefore, the image tensor needs to be reshaped into data with the shape of [A, 16]. When the input channel size cannot be divided evenly by 16, the method of padding 0 is adopted to make up for this defect, and the convolution kernel weight tensor needs to be reshaped into the shape of [M, N, O, 16, 16]. When non-divisibility occurs, the method of pre-padding 0 values is also adopted to solve it. In order to import the data into the FPGA tool for analysis and simulation operations, finally, the image data and convolution kernel parameters also need to be turned into one-dimensional data, and at the same time, the previously padded 0 values are removed.
[0081]
[0082]
[0083] N = Hk × Wk
[0084]
[0085] Where Hin, Win, and Ci respectively represent the height, width, and number of input channels of the image data, and Co, Hk, Wk, and Ci respectively represent the number of output channels, width, width, and number of input channels of the weight data.
[0086] 2) Considering the requirements of FPGA tools and Verilog programs for externally read data, it is necessary to normalize and export various types of data from the Python tool. The requirements include: during the FPGA simulation phase, the $readmemh function is required to read in data; while when actually programming the FPGA board, the weight data needs to be converted to Hex format and then read into the SD card; the ROM output of the FPGA is a long vector. When the neural network has been quantized, the image data, convolutional kernel weight parameters, convolutional kernel biases, quantization scaling factors, and quantization zero-point offsets all need to be converted into the Hex format of the corresponding signed (unsigned) numbers, and then data dimension adjustment and splicing and other processing work need to be carried out.
[0087] 3) Fixed-point processing of quantization parameters: For the currently well-developed deep convolutional neural network, the absolute values of the weights and biases of the convolutional layer after training are generally less than 1. Therefore, its quantization scaling factor is much less than 1, and the zero-point offset is an 8-bit integer. In order to store the scaling factor on the FPGA and use it to complete quantization, it is necessary to pre-process the fixed-point of the parameters related to the scaling factor on the Python tool. This design uses a binary fixed-point mapping method to complete this pre-processing, and the mapping relationship is as follows.
[0088]
[0089] Where M represents the parameter related to the scaling factor - the requantization scaling factor, M0 represents an integer in the range [0, 256], and n is an integer in the range [0, 32].
[0090] Specifically deploy the Faster R-CNN algorithm using the custom instruction set architecture proposed by the present invention. First, decompose the Faster R-CNN algorithm sequentially. The decomposition rule depends on the instruction type. The instruction after translating each sub-algorithm must start with a D-type instruction, and its ff bit is 1. Generally, there is no need to configure the bank0 or bank1 bit, and the fp bit can be directly set to 1, that is, turn on the ping-pong mode of data storage and calculation. The D-type instruction is generally followed by an operation type instruction. If the current sub-algorithm can be configured into a multi-task or dual-task mode, then multiple operation instructions can be added, and the appropriate routing type can be selected through the route bit of the D-type to ensure switching to the specific mode type. As for the configuration instruction U-type instruction, it is not necessary. It can be added when the array operation scale needs to be changed, otherwise, it is carried out according to the default situation. For the sub-algorithms within the Faster R-CNN algorithm that cannot be directly translated, such as the adaptive average pooling algorithm, it can also be written through the Cx-type instruction. Although this algorithm does not have such keywords, it can also be indirectly calculated. The sliding feature of the adaptive average pooling algorithm is determined by the input image and output image sizes of this operation layer. Therefore, the pooling kernel size, sliding step size and other features of the adaptive average pooling algorithm can be indirectly obtained through calculation. The calculation formula is as follows:
[0091]
[0092] ksizeh = Hin - (Hout - 1) × stdh
[0093] ksizew = Win - (Wout - 1) × stdw
[0094] pdh = pdh = 0
[0095] Among them, stdh, stdw, ksizeh, ksizew, pdh, and pdh respectively represent the vertical sliding step size, horizontal sliding step size, vertical size, horizontal size, vertical padding length, and horizontal and vertical padding lengths of the pooling kernel.
Claims
1. A convolutional neural network accelerator instruction set architecture and a method for deploying the Faster R-CNN algorithm, characterized in that It can adapt to common convolutional neural network target detection algorithms including the Faster R-CNN algorithm, and complete full quantization reasoning on the edge side. The convolutional neural network accelerator instruction set architecture includes an accelerator instruction access interface, a custom instruction set, and an accelerator parallel computing mode that can be selected by the user. The instruction set of the convolutional neural network accelerator includes two control flow commands, instructions and configurations, with instruction flow as the main and configuration flow as the auxiliary. Because the configuration type and function are relatively simple, and the configuration flow is closely coupled with the instruction flow, the present invention will focus on the configuration flow and instruction flow, and its architecture is uniformly included in the instruction set architecture. Because it focuses on the convolutional neural network target detection algorithm, its instructions and configurations are dedicated to driving the operation units of coarse-grained operators, and its instruction flow and configuration flow architecture are also specifically optimized accordingly. The accelerator interface refers to the accelerator as a coprocessor for accelerated computing, which has a channel for receiving instructions issued by the main processor. The neural network accelerator can usually continuously receive multiple instructions from the main processor, which act on configuration, calculation, memory access and routing functions respectively. At the same time, the instruction set architecture requires the accelerator to have a variety of different parallel computing working modes. Users can choose a suitable algorithm deployment method according to the specific characteristics of the trained neural network algorithm and the parallel computing mode, improve the parallelism of the algorithm reasoning stage, and thus improve the throughput. The Faster R-CNN algorithm deployment method includes the quantization technology of the Faster R-CNN algorithm, the pre-deployment processing of the neural network parameters, and the method of translating the Python algorithm into assembly instructions during the algorithm deployment stage. The post-training static quantization technology of the Faster R-CNN algorithm is divided into two stages: offline data collection and online quantization. The offline acquisition parameters refer to relying on a large amount of data to determine the data distribution, and on this basis to achieve the static quantization of the two-stage neural network. The online quantization requires that it must be able to reason with a fully quantized data flow structure, which not only reduces the consumption of most storage and computing resources, but also does not require CPU-assisted calculations (transmitting unquantifiable operations back to the CPU, calculating in its floating-point arithmetic unit, and then sending them to the accelerator for subsequent reasoning). Neural network parameters and deployment processing refer to the targeted processing of the convolutional neural network structure and parameters according to the hardware acceleration circuit architecture. The translation of the neural network algorithm into assembly instructions is to enable the algorithm to better adapt to the accelerator parallel computing mode. The targeted compilation method proposed in the present invention can be executed as a practical example.
2. The convolutional neural network accelerator instruction set architecture according to claim 1, wherein This instruction set architecture is designed in a custom way, but it can interact with the processors of the current mainstream architecture through the accelerator instruction access interface. It can not only adapt to the Faster R-CNN algorithm, but also other convolutional neural network algorithms, and reserves an external computing unit access interface to meet the needs of new special computing modules. This instruction set architecture solves the problem of fast update and iteration of neural network target detection algorithms.
3. The convolutional neural network accelerator instruction set architecture according to claim 1 includes seven types of instructions, which can support various functions such as users deploying algorithms, changing routing, making programmable changes to arithmetic circuits, and configuring parallel computing modes. The accelerator based on this instruction set architecture can be deployed on many common FPGA boards, and can make a trade-off between resources and speed according to requirements. At the same time, this architecture supports static parameter files to configure the accelerator circuit scale and related control functions.
4. The Faster R-CNN algorithm deployment method according to claim 1, wherein It can achieve full quantization inference for two-stage object detection neural network algorithms such as Faster R-CNN. By using a linear mapping method, many neural network operation layers such as convolutional layers, Relu layers, exponential operations, and Softmax layers are mapped to the integer domain to complete full quantization inference. This solves the problem that the traditional Faster R-CNN algorithm cannot perform full quantization inference on the edge side, and realizes the exponential operation in the form of a polynomial without complex floating-point calculations.
5. The method for deploying the Faster R-CNN algorithm according to claim 1, wherein Perform pre-deployment processing operations on convolutional neural network parameters and compilation processing during the deployment phase of neural network algorithms. This operation can help connect any FPGA device or ASIC circuit to Pytorch, enabling the Faster R-CNN algorithm to have a wider deployment space and application. The processed parameters and the compiled instructions can simultaneously meet the requirements of the accelerator architecture and FPGA tools to achieve the main work of deploying the Faster R-CNN algorithm.