Convolutional neural network model training method and device, electronic equipment and medium

By deploying the target lightweight network model in the field programmable logic gate array and designing a deterministic random computing module, the problems of computational error accumulation and hardware overhead in the fusion scheme of random computing and convolutional neural networks are solved, and efficient training and deployment are achieved.

CN120012860AActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510486658.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing random computing and convolutional neural network fusion schemes face the problems of accumulation of computational errors caused by probability fluctuations, excessive hardware overhead of traditional random number generators, and lack of computing-storage collaborative optimization mechanisms for edge devices.

Method used

By deploying the target lightweight network model in a field programmable logic gate array, a high-level comprehensive method is used to design the deterministic random computing convolution module and pooling module, and a random sequence is generated through a linear feedback shift register to replace the traditional random number generator.

Benefits of technology

It effectively reduces training costs, improves the deployment efficiency of high-precision convolutional neural network models, and reduces the overhead of hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012860A_ABST
    Figure CN120012860A_ABST
Patent Text Reader

Abstract

The invention relates to a convolutional neural network model training method and apparatus, an electronic device and a medium. The method comprises the steps of obtaining a pre-trained network model; performing quantification processing and pruning processing on the pre-trained network model to obtain a target lightweight network model; designing a universal deterministic random computation convolution module and a deterministic random computation pooling module by adopting a high-level synthesis mode; and deploying the target lightweight network model in the field programmable gate array according to the deterministic stochastic computation convolution module and the deterministic stochastic computation pooling module. According to the invention, the combination of the convolutional neural network model and random calculation is realized through the linear feedback shift register, and the deployment of the convolutional neural network model in the field programmable gate array is realized through high-level integration, so that the deployment efficiency of the convolutional neural network model can be greatly improved; and the deployment cost of the convolutional neural network model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of neural network model training, and in particular to a training method, device, electronic device and medium for a convolutional neural network model. Background Art

[0002] In the context of the big data era, traditional CMOS circuits are limited by the bottleneck of Moore's Law. Stochastic computing (SC) has the characteristics of low hardware cost and low power consumption due to its probabilistic bit stream encoding mechanism, and it has a complementary advantage with neural networks in particular.

[0003] Existing random computing and convolutional neural network fusion solutions face three major challenges: the accumulation of computational errors caused by probability fluctuations, the excessive hardware overhead of traditional random number generators, and the lack of computing-storage collaborative optimization mechanisms for edge devices. Summary of the invention

[0004] Based on this, it is necessary to provide a training method, device, electronic device and storage medium for the convolutional neural network model that can effectively reduce the training cost and improve the high precision in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for training a convolutional neural network model, which is applied to a field programmable gate array, comprising: Get the pre-trained network model; The pre-trained network model is quantized and pruned to obtain a target lightweight network model; A general deterministic random computing convolution module and a deterministic random computing pooling module are designed by adopting a high-level synthesis method, wherein a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; According to the deterministic random computing convolution module and the deterministic random computing pooling module, the target lightweight network model is deployed on the field programmable logic gate array.

[0006] In one embodiment, the quantization and pruning of the pre-trained network model are performed respectively to obtain a target lightweight network model, including: Quantizing the pre-trained network model to obtain a quantized intermediate network model, wherein the quantization process adopts a method of rewriting floating-point data into integer data and then performing a right-shift quantization process; The intermediate network model is pruned to obtain a target lightweight network model after the pruned process.

[0007] In one embodiment, the quantizing the pre-trained network model to obtain a quantized intermediate network model includes: Receiving target input data, wherein the target input data has not been regularized and is within a preset range; Determining a target integer weight according to the distribution of the target input data; Right-shift quantization is performed on the target integer weight and the target input data to obtain quantization weight, quantization data and quantization-related parameters; Calling a preset update function for verification to determine whether the pre-trained network model has completed quantization processing; When the quantization-related parameters of each layer include a reduction factor, a zero offset and an offset value, and when the reduction factor has completed the offset processing, it is determined that the pre-trained network model has completed the quantization processing to obtain the intermediate network model.

[0008] In one embodiment, the pruning of the intermediate network model to obtain a target lightweight network model after pruning includes: For the fully connected layer, the weights less than a preset threshold are set to 0, wherein the preset threshold is calculated according to the pruning percentage corresponding to the fully connected layer; For the convolutional layer, the convolution kernel with the smallest L2 norm in all layers is continuously pruned until the preset pruning percentage corresponding to the convolutional layer is met.

[0009] In one embodiment, the deterministic random calculation unit includes an 8-bit linear feedback shift register, and uses an 8-bit polynomial to generate a random sequence. In a multi-bit technology, four new bits are generated per clock cycle, wherein the 8-bit polynomial is , where x is the input and y is the output.

[0010] In one embodiment, the field programmable logic gate array is connected to an external memory; the method further comprises: The data flow operation of the field programmable gate array is optimized, wherein the optimization includes partitioning and deforming the field programmable gate array to improve the data exchange rate and data exchange volume of the field programmable gate array; defining the loop operation in the field programmable gate array according to loop unrolling and pipeline processing; and interacting with the external memory during the data operation to cache the intermediate results in the external memory.

[0011] In one embodiment, deploying the target lightweight network model on the field programmable logic gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module includes: Performing simulation tests on the deterministic random calculation convolution module and the deterministic random calculation pooling module to obtain a target bitstream file; Packing the target bitstream file into a hardware overlay file, and automatically burning the code and connecting the circuit according to the hardware overlay file; Define the image file to be recognized and run the accelerated optimized convolution module and pooling module; The test set and the target lightweight network model are imported at the same time.

[0012] In a second aspect, the present application also provides a training device for a convolutional neural network model, which is applied to a field programmable logic gate array, comprising: Model pre-training module, used to obtain pre-trained network models; A quantization and pruning module, used to perform quantization and pruning processing on the pre-trained network model to obtain a target lightweight network model; A high-level synthesis design module, used for designing a general deterministic random computing convolution module and a deterministic random computing pooling module by adopting a high-level synthesis method, wherein the deterministic random computing convolution module and the deterministic random computing pooling module are integrated with a deterministic random computing unit, and the deterministic random computing unit is used for generating a random sequence through a linear feedback shift register; A model deployment module is used to deploy the target lightweight network model on the field programmable logic gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module.

[0013] In a third aspect, the present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the training method of the convolutional neural network model described in the first aspect.

[0014] In a fourth aspect, the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the training method for the convolutional neural network model described in the first aspect.

[0015] In summary, the present application proposes a training method, device, electronic device and medium for a convolutional neural network model, including: obtaining a pre-trained network model; quantizing and pruning the pre-trained network model to obtain a target lightweight network model; designing a general deterministic random computing convolution module and a deterministic random computing pooling module by a high-level synthesis method; deploying the target lightweight network model in a field programmable gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module. The present application realizes the combination of a convolutional neural network model and random computing through a linear feedback shift register, and realizes the deployment of a convolutional neural network model in a field programmable gate array through high-level synthesis, which can greatly improve the deployment efficiency of the convolutional neural network model and reduce the deployment cost of the convolutional neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of a flow chart of a training method for a convolutional neural network model in an embodiment; Figure 2 A schematic diagram of a flow chart of steps for performing quantization processing in an embodiment; Figure 3 A schematic diagram of the recognition effect of a CNN model before and after quantization in one embodiment; Figure 4 Schematic diagram of the recognition effect of a CNN model after pruning in one embodiment; Figure 5 A schematic diagram of resource usage for simulating a network structure of a high-level integrated design in one embodiment; Figure 6 A schematic diagram of the recognition effect after a convolutional neural network model is deployed in an FPGA in one embodiment; Figure 7 is a structural block diagram of a training device for a convolutional neural network model in one embodiment; Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0018] To facilitate understanding of the training method of the convolutional neural network model provided in this embodiment, some concepts are first introduced below: Convolutional Neural Network (CNN) is a deep learning model suitable for processing image and video data. Through structures such as convolutional layers, pooling layers, and fully connected layers, CNN can automatically extract features from images and perform tasks such as classification and recognition.

[0019] Field-Programmable Gate Array (FPGA) is a semi-custom circuit that combines the high performance of custom hardware with the flexibility of general-purpose processors. FPGAs are composed of a large number of basic logic units such as logic gates, triggers, multiplexers, and programmable interconnection resources. Users can configure these logic units and interconnections through specific programming languages ​​(such as VHDL, Verilog, etc.) or graphical tools (such as Xilinx's Vivado, Altera's Quartus, etc.) to achieve specific circuit functions.

[0020] Stochastic computing (SC) is a computing method based on probability calculation. It uses random bit streams to represent values ​​and performs computing tasks through probabilistic operations. The stochastic computing module uses bit streams instead of traditional binary codes for calculations, which can achieve high area efficiency arithmetic circuits and good fault tolerance, making it very suitable for neural network acceleration for edge computing.

[0021] High-Level Synthesis (HLS) is a process that automatically converts the logic structure described in a high-level language into a circuit model described in a low-level language. High-level synthesis involves automatically converting the algorithm or logic structure described in a high-level programming language (such as C, C++, SystemC, etc.) into a register transfer level (RTL) circuit model described in a low-level hardware description language (such as Verilog, VHDL, SystemVerilog, etc.) through specific tools or compilers.

[0022] In related technologies, the fusion solution of random computing and convolutional neural network faces three major challenges: the accumulation of computational errors caused by probability fluctuations, the excessive hardware overhead of traditional random number generators, and the lack of computing-storage collaborative optimization mechanism for edge devices.

[0023] This embodiment provides a training method for a CNN model that can be deployed in an FPGA, and effectively combines a random computing unit in the CNN model. By designing the structure of the random computing unit and optimizing the acceleration framework in the FPGA, the above-mentioned problems of the random computing and convolutional neural network fusion solution are systematically solved.

[0024] like Figure 1 As shown, a training method for a convolutional neural network model is provided, including: S101, obtaining a pre-trained network model.

[0025] In this embodiment, the pre-trained network model can be implemented based on Python language. The steps of obtaining the pre-trained network model include data pre-processing, forward propagation, loss calculation, back propagation and iterative training.

[0026] It should be noted that this embodiment does not limit the training steps of the pre-trained model, and can be implemented by constructing a basic LeNet-5 model, or by constructing other CNN models that can achieve the same image processing function.

[0027] S102, quantizing and pruning the pre-trained network model to obtain a target lightweight network model.

[0028] In this embodiment, quantization and pruning methods are used to achieve lightweight CNN model, so that the CNN model is more suitable for FPGA, convenient for compatibility with random computing and adaptation to subsequent network deployment.

[0029] In this embodiment, quantization processing is mainly used to convert floating-point data in CNN into integer data, such as INT8 integer data, thereby greatly reducing the memory usage of the CNN model and improving the reasoning speed. When implementing the quantized CNN model on FPGA, a large amount of resources can be saved to improve the parallelism of the acceleration design and improve the acceleration effect.

[0030] Pruning can simplify the convolutional neural network model to a greater extent based on model quantization, allowing CNN to be deployed with smaller memory usage, fewer hardware resources, and lighter processors.

[0031] In one embodiment, the pre-trained network model is quantized and pruned respectively to obtain a target lightweight network model, including: The pre-trained network model is quantized to obtain a quantized intermediate network model, wherein the quantization process is to rewrite floating-point data into integer data and then perform right-shift quantization.

[0032] The intermediate network model is pruned to obtain the target lightweight network model after pruning.

[0033] In this embodiment, the pre-trained network model can be quantized first, and then the quantized intermediate network model can be pruned. In the specific application process, when the model is quantized first and then pruned, the advantages of both can be further utilized to achieve more efficient model compression and acceleration.

[0034] Based on the above steps, quantization can reduce the storage and computational overhead of each parameter, while pruning can reduce the number of parameters of the model. The combination of the two can better reduce the storage requirements and computational workload of the model. Quantization before pruning can help reduce the accuracy loss that may be caused by the pruning process. Because quantization has reduced the complexity of the model, pruning has less impact on model performance. By reducing the accuracy of parameters through quantization and then removing unimportant parameters through pruning, the CNN model can be more efficient and compact while maintaining certain performance.

[0035] S103, a general deterministic random computing convolution module and a deterministic random computing pooling module are designed by adopting a high-level synthesis method, wherein a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register.

[0036] In this embodiment, the HLS method is used to implement the general convolution module and pooling module, and the designed deterministic random computing unit is added to the general convolution module and pooling module to obtain a general deterministic random computing convolution module and a deterministic random computing pooling module.

[0037] In the related art, a stochastic number generator (SNG) usually includes a random number source (RNS) and a probability conversion unit (PCU). The two parts work together to generate random numbers with a specific probability distribution. The random number generator unit generates a specific bit stream based on the input X and generates a random bit stream S at the same time, so that the probability is equal to the corresponding binary number.

[0038] In this embodiment, a linear feedback shift register (LFSR) is used as a pseudo-random number generator to form a deterministic random calculation unit.

[0039] In one embodiment, the deterministic random calculation unit includes an 8-bit linear feedback shift register, and uses an 8-bit polynomial to generate a random sequence. In a multi-bit technique, four new bits are generated per clock cycle, wherein the 8-bit polynomial is , where x is input and y is output. In a specific application, one bit is transferred from the register to the Move to and from the register arrive Perform an XOR operation.

[0040] It should be noted that this embodiment uses LFSR as a deterministic random unit, combined with the CNN model, which can greatly reduce the cost of combining random computing with CNN. Combining random computing with convolutional neural networks can solve the large amount of data accumulation and transmission that may exist between the processor and memory of the traditional convolutional neural network model. The collaborative architecture of random computing and convolutional neural networks can effectively alleviate the storage wall problem under the von Neumann architecture through the spatiotemporal local characteristics of the probabilistic bit stream, which is very beneficial to the hardware implementation of convolutional neural networks in edge computing applications, and facilitates the deployment of large-scale convolutional neural network models with large-scale big data parallelization.

[0041] In this embodiment, the neural network model is designed by high-level synthesis (HLS), taking advantage of its high hardware design and development efficiency, high software design system performance, and high readability of C language code. On the basis of designing the deterministic random computing module, the network is planned and designed as a whole.

[0042] This embodiment implements the general convolution and pooling modules through high-level synthesis, and adds the designed deterministic random calculation module for use. The high-level synthesis design tool in this embodiment can be used by language is converted to register transfer level implementation and integrated into In The massively parallel architecture provided outperforms traditional processors in terms of performance, cost, and power consumption.

[0043] S104, deploying a target lightweight network model on a field programmable logic gate array according to a deterministic random computing convolution module and a deterministic random computing pooling module.

[0044] In this embodiment, the network hardware accelerator framework design can be further completed based on the HLS implementation method or the underlying logic synthesis implementation method, so as to deploy the target lightweight network model on the field programmable logic gate array and complete the coverage of the deterministic random computing convolution module and the deterministic random computing pooling module.

[0045] In summary, this embodiment provides a training method for a convolutional neural network model. By quantizing and pruning the convolutional neural network model, the convolutional neural network model can be lightweight, which is more conducive to the hardware design of the convolutional neural network model on the FPGA. By forming a deterministic random calculation unit through a linear feedback shift register, replacing the traditional random number generator, and generating a corresponding random sequence, the problem of probability fluctuation leading to accumulation of calculation errors and excessive hardware overhead of the traditional random number generator in the fusion scheme of the random network and the convolutional neural network can be effectively overcome. And based on HLS, a general deterministic random calculation convolution module and a deterministic random calculation pooling module can be implemented, which can be combined with the acceleration framework in the FPGA to effectively improve the speed and efficiency of deploying the convolutional neural network model in the FPGA, and can further speed up the image processing speed of the CNN model in the FPGA.

[0046] In one embodiment, if Figure 2 As shown, the pre-trained network model is quantized to obtain a quantized intermediate network model, including: S201, receiving target input data, wherein the target input data has not been regularized and is in a preset range.

[0047] Since this embodiment is deployed on FPGA, the quantization scheme adopted is to rewrite the floating point type into the integer right shift quantization form. Therefore, during the model training process, there is no need to regularize the input data, so that the input data is in a preset range, for example, keeping the input data between 0 and 1, which can facilitate the quantization of the target input data.

[0048] In this embodiment, the target input data is the floating-point data in the pre-trained model.

[0049] In this embodiment, the purpose of the quantization processing task is mainly to convert floating-point data into NT8 integer data, that is, to map the weight parameters of each layer to between -127 and 127.

[0050] S202, determining a target integer weight according to the distribution of the target input data.

[0051] In this embodiment, it is necessary to observe the distribution of parameters first and select an appropriate mapping method. For example, when the absolute value of the maximum value of the weight of the first convolutional layer is , the single weight to be converted is , first divide the weight to be converted by the absolute value of the maximum value. As shown in the following formula: in, is the integer weight corresponding to NT8 quantization, is the absolute value of the maximum weight corresponding to the current convolutional layer, is the weight to be converted corresponding to the current convolutional layer.

[0052] S203, right-shift quantization is performed on the target integer weight and the target input data to obtain quantization weight, quantization data and quantization-related parameters.

[0053] In this embodiment, the specific formula for right shift quantization is: in, is the quantized target integer weight, is the bias, For gain, Indicates a shift operation. The shift value of the shift operation is . ,Right now , is the zero point offset parameter, ,Right now , which is the shifting process for the reduction factor After the reduction parameter, is the reduction factor, is the shift value, represents the target input data, Represents the kernel.

[0054] The calculation formulas for the reduction factor and zero offset parameters are: in, is the reduction factor, To initialize the zero point, , , The maximum value of floating point data. The minimum value of floating-point data.

[0055] S204, calling a preset update function for verification to determine whether the pre-trained network model has completed the quantization process.

[0056] In this embodiment, based on the floating-point reasoning of the traditional quantization scheme, an integer forward reasoning process is added, that is, the convolution layer, pooling layer and fully connected layer are rewritten. While changing the calculation method to integer, an update function is added to each layer to shift and adjust the trained integer weights.

[0057] During the quantization process, the network weights, biases, and and The parameters are saved and in the next step, the weights and biases are imported and called at the same time update The function verifies the quantized network.

[0058] S205, when the quantization-related parameters of each layer include a reduction factor, a zero offset and an offset value, and when the reduction factor has completed the offset processing, it is determined that the pre-trained network model has completed the quantization processing to obtain an intermediate network model.

[0059] During the verification process, each layer has three parameters: By determining Whether the offset processing has been completed, resulting in a certain loss of accuracy, further determines whether the pre-trained network model has completed the quantization processing.

[0060] The parameters of the model after quantization training are shown in Table 1: Table 1 in, , and Represents the convolutional layer. Represents a fully connected layer (FullyConnected Layer), represents the output layer, The parameter refers to the reduction factor Shift corresponds to the shift value The reduction factor parameter after The parameter indicates the initialization zero point. represents the reduction factor, Indicates the shift value.

[0061] This embodiment conducts a comparative test on the model before and after quantization. Figure 3 As shown in the figure, non-quantized represents the model before quantization, and quantized represents the model after quantization. The digital recognition test (predicted: "2") is performed on the models before and after quantization to obtain the time consumption and tensor situation. It can be found that the difference in the tensor output by the CNN model before and after quantization is not large, and the detection time is greatly reduced.

[0062] In one embodiment, pruning the intermediate network model to obtain a target lightweight network model after pruning includes: For the fully connected layer, the weights less than the preset threshold are set to 0, where the preset threshold is calculated based on the pruning percentage corresponding to the fully connected layer; For the convolutional layer, the convolution kernel with the smallest L2 norm in all layers is continuously pruned until the preset pruning percentage corresponding to the convolutional layer is met.

[0063] In this embodiment, on the basis of quantizing the model, pruning training is further performed to simplify the network model to a greater extent so that it can be deployed with smaller memory usage, fewer hardware resources and a lighter processor.

[0064] The pruning process in this embodiment mainly uses a mask matrix to prune the CNN model, and focuses on the fully connected layer and the convolutional layer. The fully connected layer uses a threshold screening method and prunes with a single weight within the layer, while the convolutional layer uses the L2 norm and prunes with the convolution kernel within the layer.

[0065] In this embodiment, the fully connected layer sets the weights less than a preset threshold to 0, and the preset threshold is calculated by the pruning percentage. The convolution layer continuously prunes the convolution kernel with the smallest L2 norm in all layers and determines whether the percentage is reached.

[0066] The side view result after pruning is as follows Figure 4 As shown in Figure 2, the detection time is greatly reduced after pruning. It should be noted that the convolution formula after quantization is as shown in the above formula, and the weight data is 0 after pruning, and the quantization result is still valid.

[0067] In one embodiment, the field programmable logic gate array is connected to the external memory. In this embodiment, the data operation of the convolutional neural network is However, In Resources are often very low, and intermediate results are compared with external The interactions are cached and data is transmitted through a high-speed data interface to speed up the operation process.

[0068] In this embodiment, in order to avoid the time delay caused by the processor and the low bandwidth effect of the interface part, a direct memory acquisition method is used in the data stream acquisition operation. Among them, direct memory acquisition (Direct Memory Access, referred to as DMA) technology is a technology that allows external devices (such as I / O devices) to exchange data directly with the memory without the intervention of the CPU. This method can significantly improve the data transmission efficiency when processing large amounts of data or high-speed I / O devices.

[0069] The overall design of accelerating convolutional neural networks on FPGAs implements the data flow operation mode on FPGAs to reduce repeated data reading through register buffers during the convolution process, which can effectively avoid resource waste and time delays caused by multiple reading and storage of data. While making reasonable use of FPGA resources, efficiency optimization is achieved as much as possible.

[0070] In one embodiment, the training method of the convolutional neural network model further includes: The data flow operation of the field programmable logic gate array is optimized, wherein the optimization includes partitioning and deforming the field programmable logic gate array array to improve the data exchange rate and data exchange volume of the field programmable logic gate array; loop operations in the field programmable logic gate array are defined according to loop unrolling and pipeline processing; and during the data operation process, an external memory is interacted with to cache intermediate results in the external memory.

[0071] In this embodiment, the steps of optimizing the data flow operation of the field programmable logic gate array include: Array partitioning can change the order of arrays in memory and can also increase the data exchange rate by changing the number of ports.

[0072] Perform array deformation processing. Array deformation can change the bit width of the memory, and by changing the bit width, it can transmit more data in one input or output.

[0073] By defining loop operations through loop unrolling, consecutive loops can be expanded through loop unrolling operations, increasing read and write parallelism while reducing latency.

[0074] Define loop operations with pipelining, and pipeline loops so that the next command operation starts before the previous one ends.

[0075] For example, when this embodiment implements the accelerated optimized convolution module and pooling module through the aforementioned optimization processing, the specific steps are as follows: First, extract the overall feature layer of the CNN network, that is, the convolution layer (fully connected layer) and pooling layer of LeNet-5, and control and store data of the CNN network through the CPU and memory, which greatly improves the versatility and parallelism of the designed circuit. Then use the CPU to control the convolution layer and pooling layer to collect images and store the read images in the memory. The CPU controls the convolution module to perform the first convolution operation and stores the calculation results in the memory. After completion, the CPU controls the pooling module to read and store the results of the previous convolution according to the CNN network model; repeat the above process until the network operation is completed.

[0076] This embodiment uses the accelerator in the FPGA to implement the data operation of the convolutional neural network, and combines the high-level synthesis design tool to complete the design of the convolutional neural network model, which can effectively improve the development efficiency of hardware design and improve the system performance of software design. language environment for algorithm development and verification, using optimized instructions to complete Language to A comprehensive implementation of Language code.

[0077] In one embodiment, a target lightweight network model is deployed on a field programmable logic gate array according to a deterministic random computing convolution module and a deterministic random computing pooling module, including: Perform simulation tests on the deterministic random computing convolution module and the deterministic random computing pooling module to obtain the target bitstream file; Package the target bitstream file into a hardware overlay file, and automatically burn the code and connect the circuit according to the hardware overlay file; Define the image file to be recognized and run the accelerated optimized convolution module and pooling module; Import the test set and target lightweight network model at the same time.

[0078] In this embodiment, network deployment is performed on FPGA. First, the convolutional neural network model of the high-level synthesis design needs to be simulated and tested, that is, the deterministic random computing convolution module and the deterministic random computing pooling module need to be simulated and tested to obtain the target bitstream file, wherein the target bitstream file should at least include a (.tcl) file and a (.bit) file.

[0079] By packaging the target bitstream files, i.e., (.tcl) files and (.bit) files, we can get the hardware overlay file. We can call the hardware overlay file through Python statements to realize automatic burning of the codes of the deterministic random computing convolution module and the deterministic random computing pooling module, and automatic connection with the FPGA circuit.

[0080] Define the image file to be recognized and run the accelerated optimized convolution module and pooling module to complete the FPGA hardware deployment.

[0081] Finally, the test set and the target lightweight network model are imported at the same time, completing the training and deployment of the convolutional neural network model on the FPGA.

[0082] In summary, this embodiment provides a training method for a convolutional neural network model. By quantizing and pruning the convolutional neural network model, the convolutional neural network model can be lightweight, which is more conducive to the hardware design of the convolutional neural network model on the FPGA. By forming a deterministic random calculation unit through a linear feedback shift register, replacing the traditional random number generator, and generating a corresponding random sequence, the problem of probability fluctuation leading to accumulation of calculation errors and excessive hardware overhead of the traditional random number generator in the fusion scheme of the random network and the convolutional neural network can be effectively overcome. And based on HLS, a general deterministic random calculation convolution module and a deterministic random calculation pooling module can be implemented, which can be combined with the acceleration framework in the FPGA to effectively improve the speed and efficiency of deploying the convolutional neural network model in the FPGA, and can further speed up the image processing speed of the CNN model in the FPGA.

[0083] In a more detailed embodiment, in order to simulate, implement and verify the convolutional neural network, the experiment intends to test the network on three different platforms: network module simulation and verification test based on Vivado software, network model quantization and pruning training based on Pytorch, and network model quantization and pruning training based on Pytorch. Network deployment and handwritten digit recognition.

[0084] In the simulation test process, in order to ensure that the performance parameters of the convolutional neural network meet the task book indicators, Vivado is used to simulate the high-level integrated design. First, constraints are added to it: constrain the input and output ports of the circuit, constrain the clock cycle, and constrain the function name. Secondly, the CPU control is set to AXI bus control. Finally, the convolution and pooling module IP cores of the designed general deterministic random calculation are added and connected. After comprehensive simulation of the network, the resource usage can be obtained as shown in Figure 5. Among them, the resource usage in Vivado meets the basic requirements of LUT resource usage within 100,000 and FF resource usage within 50,000.

[0085] Network deployment requires the bitstream files generated in the simulation test: .tcl, .bit, which are read on the pynq board and packaged as Overlay, to automatically burn the code and connect the circuit. Finally, the network is called directly with Python statements on JupyterNotebook.

[0086] To deploy the network using Jupyter, first call overlay to automatically burn and connect the circuit: Then define the image file to be recognized and run the accelerated optimized convolution module and pooling module: Finally, import the test set and then import the LeNet-5 network after quantization and pruning training: The recognition result is as follows Figure 6 As shown, it is obvious that the convolutional neural network model deployed on FPGA provided in this embodiment can effectively run and quickly complete image recognition.

[0087] When deployed on FPGA through different methods, the corresponding logic resource usage of different implementation methods is shown in Table 2: Table 2 Among them, the high-level synthesis implementation provided in this embodiment allows the use of C language to implement the accelerator and export the RTL of the IP core, so that the network can be parallelized by adding the compilation wizard defined by the high-level synthesis implementation to design the C code, and the parallel version is verified by the time series analysis tool, thereby achieving fast pre-synthesis simulation. The resource usage of other implementations is higher than that of the high-level synthesis implementation provided in this embodiment, which can further prove the advantages of the training method of the convolutional neural network model provided in this embodiment.

[0088] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0089] Based on the same inventive concept, the embodiment of the present application also provides a training device for a convolutional neural network model for implementing the training method of the convolutional neural network model involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of the training device for one or more convolutional neural network models provided below can refer to the limitations of the training method for the convolutional neural network model above, and will not be repeated here.

[0090] In one embodiment, Figure 7As shown, a training device 700 for a convolutional neural network model is provided, comprising: a model pre-training module 710, a quantization pruning module 720, a high-level synthesis design module 730 and a model deployment module 740, wherein: Model pre-training module 710, used to obtain a pre-trained network model; A quantization and pruning module 720 is used to perform quantization and pruning processing on the pre-trained network model to obtain a target lightweight network model; A high-level synthesis design module 730, which is used to design a general deterministic random computing convolution module and a deterministic random computing pooling module by adopting a high-level synthesis method, wherein the deterministic random computing convolution module and the deterministic random computing pooling module are integrated with a deterministic random computing unit, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; The model deployment module 740 is used to deploy the target lightweight network model on the field programmable logic gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module.

[0091] In summary, this embodiment provides a training device for a convolutional neural network model. By quantizing and pruning the convolutional neural network model, the convolutional neural network model can be lightweight, which is more conducive to the hardware design of the convolutional neural network model on the FPGA. By forming a deterministic random calculation unit through a linear feedback shift register, replacing the traditional random number generator, and generating a corresponding random sequence, it can effectively overcome the problems existing in the fusion scheme of random networks and convolutional neural networks, such as probability fluctuations leading to accumulation of calculation errors and excessive hardware overhead of traditional random number generators. And based on HLS, a general deterministic random calculation convolution module and a deterministic random calculation pooling module can be implemented, which can be combined with the acceleration framework in the FPGA to effectively improve the speed and efficiency of deploying the convolutional neural network model in the FPGA, and can further speed up the image processing speed of the CNN model in the FPGA.

[0092] Each module in the training device of the above-mentioned convolutional neural network model can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0093] In one embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 8As shown. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and the external device. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a training method for a convolutional neural network model is implemented. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the electronic device casing, or an external keyboard, touchpad or mouse.

[0094] Those skilled in the art will understand that Figure 8 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0095] In one embodiment, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented: Get the pre-trained network model; The pre-trained network model is quantized and pruned respectively to obtain the target lightweight network model; A general deterministic random computing convolution module and a deterministic random computing pooling module are designed by adopting a high-level synthesis method, wherein a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; According to the deterministic random computing convolution module and the deterministic random computing pooling module, the target lightweight network model is deployed on the field programmable logic gate array.

[0096] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: Get the pre-trained network model; The pre-trained network model is quantized and pruned respectively to obtain the target lightweight network model; A general deterministic random computing convolution module and a deterministic random computing pooling module are designed by adopting a high-level synthesis method, wherein a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; According to the deterministic random computing convolution module and the deterministic random computing pooling module, the target lightweight network model is deployed on the field programmable logic gate array.

[0097] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps: Get the pre-trained network model; The pre-trained network model is quantized and pruned respectively to obtain the target lightweight network model; A general deterministic random computing convolution module and a deterministic random computing pooling module are designed by adopting a high-level synthesis method, wherein a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; According to the deterministic random computing convolution module and the deterministic random computing pooling module, the target lightweight network model is deployed on the field programmable logic gate array.

[0098] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0099] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0100] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A training method for a convolutional neural network model, characterized in that: Applications in field programmable gate arrays include: Get the pre-trained network model; The pre-trained network model is quantized and pruned to obtain a target lightweight network model; A general deterministic random computing convolution module and a deterministic random computing pooling module are designed by adopting a high-level synthesis method, wherein a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; According to the deterministic random computing convolution module and the deterministic random computing pooling module, the target lightweight network model is deployed on the field programmable logic gate array.

2. The method according to claim 1, characterized in that The step of performing quantization and pruning on the pre-trained network model to obtain a target lightweight network model includes: Quantizing the pre-trained network model to obtain a quantized intermediate network model, wherein the quantization process adopts a method of rewriting floating-point data into integer data and then performing a right-shift quantization process; The intermediate network model is pruned to obtain a target lightweight network model after the pruned process.

3. The method according to claim 2, characterized in that The quantizing the pre-trained network model to obtain a quantized intermediate network model includes: Receiving target input data, wherein the target input data has not been regularized and is within a preset range; Determining a target integer weight according to the distribution of the target input data; Right-shift quantization is performed on the target integer weight and the target input data to obtain quantization weight, quantization data and quantization-related parameters; Calling a preset update function for verification to determine whether the pre-trained network model has completed quantization processing; When the quantization-related parameters of each layer include a reduction factor, a zero offset and an offset value, and when the reduction factor has completed the offset processing, it is determined that the pre-trained network model has completed the quantization processing to obtain the intermediate network model.

4. The method according to claim 2, characterized in that: The pruning of the intermediate network model to obtain a target lightweight network model after pruning includes: For the fully connected layer, the weights less than a preset threshold are set to 0, wherein the preset threshold is calculated according to the pruning percentage corresponding to the fully connected layer; For the convolutional layer, the convolution kernel with the smallest L2 norm in all layers is continuously pruned until the preset pruning percentage corresponding to the convolutional layer is met.

5. The method according to claim 1, characterized in that The deterministic random calculation unit includes an 8-bit linear feedback shift register, and uses an 8-bit polynomial to generate a random sequence. In the multi-bit technology, four new bits are generated per clock cycle, wherein the 8-bit polynomial is , where x is the input and y is the output.

6. The method according to claim 1, characterized in that The field programmable logic gate array is connected to an external memory; the method further comprises: The data flow operation of the field programmable gate array is optimized, wherein the optimization includes partitioning and deforming the field programmable gate array to improve the data exchange rate and data exchange volume of the field programmable gate array; defining the loop operation in the field programmable gate array according to loop unrolling and pipeline processing; and interacting with the external memory during the data operation to cache the intermediate results in the external memory.

7. The method according to claim 1, characterized in that The method of deploying the target lightweight network model on the field programmable logic gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module includes: Performing simulation tests on the deterministic random calculation convolution module and the deterministic random calculation pooling module to obtain a target bitstream file; Packing the target bitstream file into a hardware overlay file, and automatically burning the code and connecting the circuit according to the hardware overlay file; Define the image file to be recognized and run the accelerated optimized convolution module and pooling module; The test set and the target lightweight network model are imported at the same time.

8. A training device for a convolutional neural network model, characterized in that: Applications in field programmable gate arrays include: Model pre-training module, used to obtain pre-trained network models; A quantization and pruning module, used to perform quantization and pruning processing on the pre-trained network model to obtain a target lightweight network model; A high-level synthesis design module, used for designing a general deterministic random computing convolution module and a deterministic random computing pooling module by adopting a high-level synthesis method, wherein the deterministic random computing convolution module and the deterministic random computing pooling module are integrated with a deterministic random computing unit, and the deterministic random computing unit is used for generating a random sequence through a linear feedback shift register; A model deployment module is used to deploy the target lightweight network model on the field programmable logic gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the training method of the convolutional neural network model described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the training method of the convolutional neural network model described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data quantification method and device, computer equipment and storage medium

    CN111176853A

  • Intelligent inference network system and addition unit and pooling unit circuitry

    CN112949830A

  • Neural network quantification processing method, apparatus and device, and readable storage medium

    CN114781618A

  • Data processing method and device, electronic equipment and storage medium

    CN117472326A

  • FPGA-based neural network acceleration method and system

    CN117521752A