Training Method, Device, Electronic Device and Medium of Convolutional Neural Network Model

By quantizing and pruning the pretrained network model, combined with a high-level comprehensive design of the deterministic random computing module, the calculation error accumulation and hardware overhead problems in the fusion of random computing and convolutional neural networks are solved, and efficient and low-cost convolutional neural network model deployment is achieved.

CN120012860BActive Publication Date: 2025-07-25NAT UNIV OF DEFENSE TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510486658.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing random computing and convolutional neural network fusion schemes face the accumulation of computational errors caused by probability fluctuations, the excessive hardware overhead of traditional random number generators, and the lack of computing-storage collaborative optimization mechanism for edge devices.

Method used

By quantizing and pruning the pretrained network model, combining high-level comprehensive design of deterministic random computing convolution modules and pooling modules, and deploying in a field programmable logic gate array, random sequences are generated using linear feedback shift registers to replace traditional random number generators.

Benefits of technology

It effectively reduces the training cost of convolutional neural network models, improves high accuracy and deployment efficiency, solves the calculation error accumulation and hardware overhead problems in random computing and convolutional neural network fusion schemes, and is suitable for edge computing devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012860B_ABST
    Figure CN120012860B_ABST
Patent Text Reader

Abstract

The present application relates to a training method, device, electronic device and medium of a convolutional neural network model, including: obtaining a pre-trained network model; respectively performing quantization processing and pruning processing on the pre-trained network model to obtain a target lightweight network model; designing a general deterministic stochastic computing convolutional module and a deterministic stochastic computing pooling module by means of high-level synthesis; and deploying the target lightweight network model on a field-programmable gate array according to the deterministic stochastic computing convolutional module and the deterministic stochastic computing pooling module. The present application realizes the combination of the convolutional neural network model and stochastic computing through a linear feedback shift register, and realizes the deployment of the convolutional neural network model in a field-programmable gate array through high-level synthesis, which can greatly improve the deployment efficiency of the convolutional neural network model and reduce the deployment cost of the convolutional neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of neural network model training, and particularly to a training method, device, electronic device and medium for a convolutional neural network model. Background Art

[0002] In the context of the big data era, traditional CMOS circuits are limited by the Moore's law bottleneck. Stochastic computing (SC) exhibits the characteristics of low hardware cost and excellent power consumption due to its probabilistic bitstream encoding mechanism, and particularly forms a complementary advantage with neural networks.

[0003] Existing stochastic computing and convolutional neural network fusion solutions face three major challenges: cumulative computational errors caused by probability fluctuations, excessive hardware overhead of traditional random number generators, and lack of a compute-storage co-optimization mechanism for edge devices. Summary of the Invention

[0004] Based on this, it is necessary to provide a training method, device, electronic device and storage medium for a convolutional neural network model that can effectively reduce training costs and improve the accuracy of the model in view of the above technical problems.

[0005] In a first aspect, the present application provides a training method for a convolutional neural network model, which is applied to a field programmable gate array and includes:

[0006] Obtain a pre-trained network model;

[0007] Perform quantization processing and pruning processing on the pre-trained network model respectively to obtain a target lightweight network model;

[0008] Design a general deterministic stochastic computing convolution module and a deterministic stochastic computing pooling module in a high-level synthesis manner, wherein the deterministic stochastic computing convolution module and the deterministic stochastic computing pooling module integrate deterministic stochastic computing units, and the deterministic stochastic computing units are used to generate random sequences through linear feedback shift registers;

[0009] Deploy the target lightweight network model on the field programmable gate array according to the deterministic stochastic computing convolution module and the deterministic stochastic computing pooling module.

[0010] In one embodiment, the performing quantization processing and pruning processing on the pre-trained network model respectively to obtain a target lightweight network model includes:

[0011] Perform quantization processing on the pre-trained network model to obtain a quantized intermediate network model, wherein the quantization processing adopts a method of rewriting floating-point data into integer data and then performing right-shift quantization processing;

[0012] Prune the intermediate network model to obtain a target lightweight network model after pruning.

[0013] In one embodiment, the quantizing the pre-trained network model to obtain a quantized intermediate network model includes:

[0014] Receive target input data, where the target input data is not regularized and is within a preset range;

[0015] Determine target integer weights according to the distribution of the target input data;

[0016] Perform right shift quantization on the target integer weights and the target input data to obtain quantized weights, quantized data, and quantization related parameters;

[0017] Call a preset update function for verification to determine whether the pre-trained network model has completed quantization processing;

[0018] When the quantization related parameters of each layer all include a reduction factor, a zero point offset, and an offset value, and the reduction factor has completed offset processing, determine that the pre-trained network model has completed quantization processing to obtain the intermediate network model.

[0019] In one embodiment, the pruning the intermediate network model to obtain a target lightweight network model after pruning includes:

[0020] For the fully connected layer, set the weights less than a preset threshold to 0, where the preset threshold is calculated according to the pruning percentage corresponding to the fully connected layer;

[0021] For the convolutional layer, continuously prune the convolutional kernels with the smallest L2 norm in all layers until the preset pruning percentage corresponding to the convolutional layer is satisfied.

[0022] In one embodiment, the deterministic random calculation unit includes an 8-bit linear feedback shift register, uses an 8-bit polynomial to generate a random sequence, and in multi-bit technology, four new bits are generated in each clock cycle, where the 8-bit polynomial is , where x is the input and y is the output.

[0023] In one embodiment, the field programmable logic gate array is connected to an external memory; the method further includes:

[0024] Optimize the data stream operation of the field programmable gate array, where the optimization process includes partitioning and deforming the field programmable gate array array to improve the data exchange rate and data exchange volume of the field programmable gate array; define loop operations in the field programmable gate array according to loop unrolling and pipelining processing; during data operation, interact with the external memory to cache intermediate results in the external memory.

[0025] In one embodiment, deploying the target lightweight network model on the field programmable gate array according to the deterministic random calculation convolution module and the deterministic random calculation pooling module includes:

[0026] Perform simulation tests on the deterministic random calculation convolution module and the deterministic random calculation pooling module to obtain a target bitstream file;

[0027] Package the target bitstream file into a hardware overlay file, and automatically burn the code and connect the circuit according to the hardware overlay file;

[0028] Define the image file to be recognized, and run the optimized convolution module and pooling module;

[0029] Import the test set and the target lightweight network model simultaneously.

[0030] In a second aspect, the present application also provides a training device for a convolutional neural network model, which is applied to a field programmable gate array and includes:

[0031] A model pre-training module for obtaining a pre-trained network model;

[0032] A quantization pruning module for respectively performing quantization processing and pruning processing on the pre-trained network model to obtain a target lightweight network model;

[0033] A high-level synthesis design module for designing a general deterministic random calculation convolution module and a deterministic random calculation pooling module in a high-level synthesis manner, where a deterministic random calculation unit is integrated in the deterministic random calculation convolution module and the deterministic random calculation pooling module, and the deterministic random calculation unit is used to generate a random sequence through a linear feedback shift register;

[0034] A model deployment module for deploying the target lightweight network model on the field programmable gate array according to the deterministic random calculation convolution module and the deterministic random calculation pooling module.

[0035] In a third aspect, the present application also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the training method of the convolutional neural network model described in the first aspect are implemented.

[0036] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the training method of the convolutional neural network model described in the first aspect are implemented.

[0037] In summary, the present application proposes a training method, device, electronic device and medium for a convolutional neural network model, including: obtaining a pre-trained network model; respectively performing quantization processing and pruning processing on the pre-trained network model to obtain a target lightweight network model; designing a general deterministic stochastic computing convolution module and a deterministic stochastic computing pooling module by means of high-level synthesis; and deploying the target lightweight network model on a field-programmable gate array according to the deterministic stochastic computing convolution module and the deterministic stochastic computing pooling module. The present application realizes the combination of the convolutional neural network model and stochastic computing through a linear feedback shift register, and realizes the deployment of the convolutional neural network model on a field-programmable gate array through high-level synthesis, which can greatly improve the deployment efficiency of the convolutional neural network model and reduce the deployment cost of the convolutional neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic flowchart of the training method of the convolutional neural network model in an embodiment;

[0039] Figure 2 It is a schematic flowchart of the steps of quantization processing in an embodiment;

[0040] Figure 3 It is a schematic diagram of the recognition effect of the CNN model before and after quantization in an embodiment;

[0041] Figure 4 It is a schematic diagram of the recognition effect of the CNN model after pruning in an embodiment;

[0042] Figure 5 It is a schematic diagram of the resource occupancy of simulating the network structure designed by high-level synthesis in an embodiment;

[0043] Figure 6 It is a schematic diagram of the recognition effect after deploying the convolutional neural network model in an FPGA in an embodiment;

[0044] Figure 7 It is a structural block diagram of the training device of the convolutional neural network model in an embodiment;

[0045] Figure 8 The internal structure diagram of a computer device in an embodiment. Detailed implementation

[0046] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] To facilitate the understanding of the training method of the convolutional neural network model provided in this embodiment, some concepts will be introduced first as follows:

[0048] A convolutional neural network (CNN for short) is a deep learning model suitable for processing image and video data. Through structures such as convolutional layers, pooling layers, and fully connected layers, a CNN can automatically extract features in images and perform tasks such as classification and recognition.

[0049] A field-programmable gate array (FPGA for short) is a semi-custom circuit that combines the high performance of custom hardware and the flexibility of general-purpose processors. An FPGA is internally composed of a large number of basic logic units such as logic gates, flip-flops, and multiplexers, as well as programmable interconnect resources. Users can configure these logic units and interconnects through specific programming languages (such as VHDL, Verilog, etc.) or graphical tools (such as Vivado from Xilinx, Quartus from Altera, etc.) to implement specific circuit functions.

[0050] Stochastic Computing (SC for short), as a computing method based on probabilistic computing, uses a random bit stream to represent numerical values and performs computing tasks through probabilistic operations. By using a bit stream instead of traditional binary encoding for computing, a stochastic computing module can achieve arithmetic circuits with high area efficiency and good fault tolerance, and is very suitable for neural network acceleration in edge computing.

[0051] High-Level Synthesis (HLS for short) is a process of automatically converting a logical structure described in a high-level language into a circuit model described in a low-level abstraction language. High-level synthesis involves automatically converting an algorithm or logical structure described in a high-level programming language (such as C, C++, SystemC, etc.) into a register transfer level (RTL) circuit model described in a low-level hardware description language (such as Verilog, VHDL, SystemVerilog, etc.) through specific tools or compilers.

[0052] In the related art, the fusion scheme of stochastic computing and convolutional neural network faces three major challenges: the accumulation of computational errors caused by probability fluctuations, the excessive hardware overhead of traditional random number generators, and the lack of a computational-storage co-optimization mechanism for edge devices.

[0053] This embodiment provides a training method for a CNN model deployable in an FPGA, effectively combines a stochastic computing unit in the CNN model, and optimizes the acceleration framework in the FPGA by designing the structure of the stochastic computing unit, systematically solving the above problems of the fusion scheme of stochastic computing and convolutional neural network.

[0054] As Figure 1 shown, a training method for a convolutional neural network model is provided, including:

[0055] S101, obtaining a pre-trained network model.

[0056] In this embodiment, the pre-trained network model can be implemented based on the Python language. The steps of obtaining the pre-trained network model include data preprocessing, forward propagation, loss calculation, backpropagation, and iterative training, etc.

[0057] It should be noted that this embodiment does not limit the training steps of the pre-trained model, and it can be implemented by constructing a basic LeNet-5 model, or by constructing other CNN models that can achieve the same image processing function.

[0058] S102, respectively performing quantization processing and pruning processing on the pre-trained network model to obtain a target lightweight network model.

[0059] In this embodiment, by using quantization and pruning means, the lightweight of the CNN model can be realized, so that the CNN model is more suitable for the FPGA, facilitating compatibility with stochastic computing and adapting to subsequent network deployment.

[0060] In this embodiment, quantization processing is mainly used to convert floating-point data in the CNN into integer data, such as INT8 integer data, thereby greatly reducing the memory occupancy of the CNN model and improving the inference speed. And when implementing the quantized CNN model on the FPGA, a large amount of resources can be saved to improve the parallelism of the acceleration design and enhance the acceleration effect.

[0061] Pruning processing can further simplify the convolutional neural network model to a greater extent on the basis of model quantization processing, enabling the CNN to be deployed with less memory occupancy, fewer hardware resources, and a more lightweight processor.

[0062] In one embodiment, the pre-trained network model is respectively quantized and pruned to obtain a target lightweight network model, including:

[0063] The pre-trained network model is quantized to obtain a quantized intermediate network model. Among them, the quantization process adopts the method of rewriting floating-point data into integer data and then performing right-shift quantization.

[0064] The intermediate network model is pruned to obtain the target lightweight network model after pruning.

[0065] In this embodiment, the pre-trained network model can be quantized first, and then the quantized intermediate network model can be pruned. In the specific application process, when the model is quantized first and then pruned, the advantages of both can be further exploited to achieve more efficient model compression and acceleration.

[0066] Based on the above steps, quantization can reduce the storage occupancy and computational overhead of each parameter, while pruning can reduce the number of parameters in the model. Using both in combination can better reduce the storage requirements and computational amount of the model. Quantization before pruning can help reduce the accuracy loss that may be caused during the pruning process. Because quantization has already reduced the complexity of the model, making the impact of pruning on the model performance smaller. By reducing the parameter accuracy through quantization and then removing unimportant parameters through pruning, the CNN model can be made more efficient and compact while maintaining a certain performance.

[0067] S103, design a general deterministic random computing convolution module and a deterministic random computing pooling module in a high-level synthesis manner. Among them, a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register.

[0068] In this embodiment, a general convolution module and a pooling module are implemented in an HLS manner, and the designed deterministic random computing unit is added to the general convolution module and the pooling module to obtain a general deterministic random computing convolution module and a deterministic random computing pooling module.

[0069] In the related art, a Stochastic Number Generator (SNG) generally includes a Random Number Source (RNS) and a Probability Conversion Unit (PCU). The RNS and the PCU cooperate together to generate random numbers with a specific probability distribution. While generating a specific bit stream according to the input X, the random number generator unit generates a random bit stream S such that the probability is equal to the corresponding binary number.

[0070] In this embodiment, a Linear Feedback Shift Register (LFSR) is used as a pseudo-random number generator to form a deterministic random calculation unit.

[0071] In one embodiment, the deterministic random calculation unit includes an 8-bit linear feedback shift register, which generates a random sequence using an 8-bit polynomial. In the multi-bit technology, four new bits are generated in each clock cycle. Among them, the 8-bit polynomial is , where x is the input and y is the output. In a specific application, within each clock cycle, one bit moves from to in the register, and an exclusive OR operation is performed from the register to .

[0072] It should be noted that in this embodiment, using the LFSR as the deterministic random unit and combining it with the CNN model can greatly reduce the cost of combining random calculation and CNN. Combining random calculation with the convolutional neural network can solve the possible large amount of data accumulation and transmission between the traditional convolutional neural network model processor and memory. The collaborative architecture of random calculation and the convolutional neural network can effectively alleviate the memory wall problem under the von Neumann architecture through the spatio-temporal locality characteristics of the probability bit stream, which is very beneficial to the hardware implementation of the convolutional neural network in edge computing applications and facilitates the deployment of large-scale large data parallel large convolutional neural network models.

[0073] In this embodiment, the design of the neural network model is completed through the High-Level Synthesis (HLS) method, taking advantage of its high hardware design and development efficiency, high software design system performance, and highly readable C language code. Based on the design of the deterministic random calculation module, an overall planning and design of the network is carried out.

[0074] This embodiment realizes general convolutional and pooling modules through high-level synthesis and adds and uses the designed deterministic random calculation module. The high-level synthesis design tool in this embodiment can The language is converted into a register transfer level implementation and integrated into while the provided large-scale parallel architecture is superior to traditional processors in terms of performance, cost, and power consumption.

[0075] S104. Deploy the target lightweight network model on the field-programmable gate array according to the deterministic random calculation convolution module and the deterministic random calculation pooling module.

[0076] In this embodiment, the network hardware accelerator framework design can be further completed based on the HLS implementation method or the underlying logic synthesis implementation method, so as to deploy the target lightweight network model on the field-programmable gate array and complete the coverage of the deterministic random calculation convolution module and the deterministic random calculation pooling module.

[0077] In summary, this embodiment provides a training method for a convolutional neural network model. By performing quantization processing and pruning processing on the convolutional neural network model, the lightweight of the convolutional neural network model can be realized, which is more conducive to the hardware design of the convolutional neural network model on the FPGA. By constructing a deterministic random calculation unit with a linear feedback shift register to replace the traditional random number generator and generate the corresponding random sequence, the problems existing in the fusion scheme of the random network and the convolutional neural network, such as the accumulation of calculation errors caused by probability fluctuations and the excessive hardware overhead of the traditional random number generator, can be effectively overcome. And based on the HLS to implement the general deterministic random calculation convolution module and the deterministic random calculation pooling module, it can be combined with the acceleration framework in the FPGA to effectively improve the speed and efficiency of deploying the convolutional neural network model in the FPGA, and can further accelerate the image processing speed of the CNN model in the FPGA.

[0078] In one of the embodiments, as Figure 2 shown, perform quantization processing on the pre-trained network model to obtain the quantized intermediate network model, including:

[0079] S201. Receive the target input data, where the target input data is not regularized and is within a preset interval.

[0080] Since this embodiment is deployed on the FPGA and the quantization scheme adopted is to rewrite the floating-point type into the form of integer right-shift quantization. Therefore, during the model training process, there is no need to regularize the input data to make the input data within the preset interval. For example, keeping the input data between 0 and 1 can facilitate the quantization processing of the target input data.

[0081] In this embodiment, the target input data is the floating-point type data in the pre-trained model.

[0082] In this embodiment, the purpose of the quantization process is mainly to convert floating-point data into NT8 integer data, that is, to map the weight parameters of each layer to the range between -127 and 127.

[0083] S202. Determine the target integer weight according to the distribution of the target input data.

[0084] In this embodiment, it is necessary to first observe the distribution of the parameters and select an appropriate mapping method. For example, when the absolute value of the maximum value of the weights in the first convolutional layer is and the single weight to be converted is , first divide the weight to be converted by the absolute value of the maximum value. As shown in the following formula:

[0085]

[0086] where is the integer weight corresponding to NT8 quantization, is the absolute value of the maximum weight corresponding to the current convolutional layer, and is the weight to be converted corresponding to the current convolutional layer.

[0087] S203. Perform right-shift quantization on the target integer weight and the target input data to obtain the quantized weight, quantized data, and quantization-related parameters.

[0088] In this embodiment, the specific formula for right-shift quantization is:

[0089]

[0090] where is the quantized target integer weight, is the bias, is the gain, represents the shift operation, and the shift value of the shift operation is . , that is, is the zero-point offset parameter, , that is, is the reduction parameter after shifting the reduction factor , is the reduction factor, is the shift value, represents the target input data, and represents the kernel.

[0091] where the calculation formulas for the reduction factor and the zero-point offset parameter are:

[0092]

[0093]

[0094] Among them, is the reduction factor, is the initialization zero point, , , is the maximum value of floating-point data, is the minimum value of floating-point data.

[0095] S204, call the preset update function for verification to determine whether the pre-trained network model has completed quantization processing.

[0096] In this embodiment, based on the floating-point inference of the traditional quantization scheme, an integer forward inference process is added, that is, the convolutional layer, pooling layer, and fully connected layer are rewritten. While changing its calculation method to integer, an update function is added to each layer to perform shift adjustment on the trained integer weights.

[0097] During the quantization process, the network weights, biases, and and parameters generated during the training model process are saved, and in the next step, by importing the weights and biases, and at the same time calling the update function to verify the quantized network.

[0098] S205, when the quantization-related parameters of each layer all include a reduction factor, zero-point offset, and offset value, and the reduction factor has completed the offset processing, it is determined that the pre-trained network model has completed quantization processing to obtain an intermediate network model.

[0099] During the verification process, each layer of parameters has three: . By determining whether has completed the offset processing, causing a certain accuracy loss, and then determining whether the pre-trained network model has completed quantization processing.

[0100] The parameters of the model after quantization training are shown in Table 1 below:

[0101] Table 1

[0102]

[0103] Among them, , and represent the convolutional layer, represents the fully connected layer, represents the output layer, The parameter refers to shifting the reduction factor by the corresponding shift value The reduction factor parameter after The parameter represents the initialization zero point, represents the reduction factor, represents the shift value.

[0104] In this embodiment, the models before and after quantization are compared and tested. As Figure 3 shown, non - quantized represents the model before quantization, and quantized represents the model after quantization. Digital recognition tests (predicted: "2") are performed on the models before and after quantization respectively to obtain the time consumption and tensor conditions. It can be obtained that the tensors output by the CNN models before and after quantization have little difference, and at the same time, the detection time is greatly reduced.

[0105] In one of the embodiments, pruning is performed on the intermediate network model to obtain the target lightweight network model after pruning, including:

[0106] For the fully - connected layer, the weights less than the preset threshold are set to 0, where the preset threshold is calculated according to the pruning percentage corresponding to the fully - connected layer;

[0107] For the convolutional layer, continuously prune the convolutional kernels with the smallest L2 norm in all layers until the preset pruning percentage corresponding to the convolutional layer is met.

[0108] In this embodiment, on the basis of quantizing the model, pruning training is further performed to simplify the network model to a greater extent, so that it can be deployed with less memory occupancy, fewer hardware resources, and a more lightweight processor.

[0109] The pruning process in this embodiment mainly uses a mask matrix (mask) to prune the CNN model, and focuses on writing for the fully - connected layer and the convolutional layer. Among them, the fully - connected layer uses a threshold screening method and prunes with individual weights within the layer, while the convolutional layer uses the L2 norm and prunes with convolutional kernels within the layer.

[0110] In this embodiment, for the fully - connected layer, the weights less than the preset threshold are set to 0, and the preset threshold is calculated from the pruning percentage. For the convolutional layer, continuously prune the convolutional kernels with the smallest L2 norm in all layers and judge whether the percentage is reached.

[0111] The side - view result after pruning is as Figure 4 shown. The detection time is greatly reduced after pruning. It should be noted that the convolutional formula after quantization is as shown in the foregoing formula, and the weight data is 0 after being pruned, and the quantization result is still valid.

[0112] In one embodiment, a field-programmable gate array is connected to an external memory. In this embodiment, data operations of the convolutional neural network are performed in but in the resources are often very low. Caching the intermediate results for interaction with the external cache and performing data transmission through a high-speed data interface to accelerate the operation process.

[0113] In this embodiment, to avoid the time delay caused by the processor and the low-bandwidth effect of the interface part, the direct memory acquisition method is adopted in the data stream acquisition operation. Among them, the direct memory access (DMA) technology is a technology that allows direct data exchange between external devices (such as I / O devices) and memory without the intervention of the CPU. This method can significantly improve the data transmission efficiency when processing a large amount of data or high-speed I / O devices.

[0114] For the overall design of accelerating the convolutional neural network on the FPGA, by implementing the operation mode of the data stream on the FPGA to reduce the repeated reading of data through the register buffer during the convolution process, it can effectively avoid the resource waste and time delay caused by multiple data readings and multiple data storages. While reasonably utilizing the FPGA resources, it meets the efficiency optimization as much as possible.

[0115] In one embodiment, the training method of the convolutional neural network model further includes:

[0116] Optimizing the data stream operation of the field-programmable gate array, where the optimization process includes partitioning and deforming the field-programmable gate array array to improve the data exchange rate and data exchange volume of the field-programmable gate array; defining loop operations in the field-programmable gate array according to loop unrolling and pipelining processing; during the data operation process, interacting with the external memory to cache the intermediate results in the external memory.

[0117] In this embodiment, the steps of optimizing the data stream operation of the field-programmable gate array include:

[0118] Performing array partitioning, and array partitioning can change its order in memory or change the number of ports, thereby improving the data exchange rate.

[0119] Performing array deformation, and the deformation of the array can change the bit width of the memory, and by changing the bit width, more data can be transmitted in one input or output.

[0120] Define loop operations through loop unrolling. Consecutive loops can be expanded through loop unrolling operations, reducing latency while increasing read-write parallelism.

[0121] Define loop operations through pipelining. Pipeline the loop so that the next one starts running before the previous command operation ends.

[0122] For example, when implementing an accelerated and optimized convolution module and pooling module through the aforementioned optimization process in this embodiment, the specific steps are as follows:

[0123] First, extract the overall feature layers of the CNN network, that is, the convolutional layer (fully connected layer) and pooling layer of LeNet-5, and control and store data of the CNN network through the CPU and memory, greatly improving the generality and parallelism of the designed circuit. Then, control the convolutional layer and pooling layer to perform image acquisition through the CPU, and store the read pictures in the memory. The CPU controls the convolutional module to perform the first convolution operation and stores the calculation result in the memory. After completion, the CPU controls the pooling module to read the result of the previous convolution and store it according to the CNN network model; continuously repeat the above process until the network operation ends.

[0124] This embodiment realizes data operations of the convolutional neural network through the accelerator in the FPGA, and completes the design of the convolutional neural network model in combination with the high-level synthesis design tool, which can effectively improve the development efficiency of the hardware design and the system performance of the software design. It conducts algorithm development and verification in the language environment, and uses optimized instructions to complete language to synthesis implementation, creating readable and portable language code.

[0125] In one of the embodiments, according to the deterministic random computing convolution module and the deterministic random computing pooling module, deploy the target lightweight network model on the field programmable logic gate array, including:

[0126] Perform simulation tests on the deterministic random computing convolution module and the deterministic random computing pooling module to obtain the target bitstream file;

[0127] Pack the target bitstream file into a hardware overlay file, and automatically burn the code and connect the circuit according to the hardware overlay file;

[0128] Define the image file to be recognized, and run the accelerated and optimized convolution module and pooling module;

[0129] Import the test set and the target lightweight network model at the same time.

[0130] In this embodiment, for network deployment on the FPGA, it is first necessary to perform simulation tests on the convolutional neural network model designed by high-level synthesis, that is, to perform simulation tests on the deterministic random computing convolutional module and the deterministic random computing pooling module to obtain the target bitstream file. Among them, the target bitstream file should at least include the (.tcl) file and the (.bit) file.

[0131] By packing the target bitstream file, that is, the (.tcl) file and the (.bit) file, a hardware overlay file can be obtained. The code of the deterministic random computing convolutional module and the deterministic random computing pooling module can be automatically burned, and the FPGA circuit can be automatically connected by calling the hardware overlay file through Python statements.

[0132] Define the image file to be recognized, and run the optimized convolutional module and pooling module for acceleration to complete the hardware deployment of the FPGA.

[0133] Finally, import the test set and the target lightweight network model at the same time to complete the training and deployment of the convolutional neural network model on the FPGA.

[0134] In summary, this embodiment provides a training method for a convolutional neural network model. By quantizing and pruning the convolutional neural network model, the lightweight of the convolutional neural network model can be realized, which is more conducive to the hardware design of the convolutional neural network model on the FPGA. By using a linear feedback shift register to form a deterministic random computing unit to replace the traditional random number generator and generate the corresponding random sequence, the problems existing in the fusion scheme of the random network and the convolutional neural network can be effectively overcome, such as the accumulation of calculation errors caused by probability fluctuations and the excessive hardware overhead of the traditional random number generator. And based on HLS, a general deterministic random computing convolutional module and a deterministic random computing pooling module are implemented, which can be combined with the acceleration framework in the FPGA to effectively improve the speed and efficiency of deploying the convolutional neural network model in the FPGA, and can further accelerate the image processing speed of the CNN model in the FPGA.

[0135] In a more detailed embodiment, in order to perform the three steps of simulation, implementation, and verification on the convolutional neural network, the experiment is planned to test the network on three different platforms: network module simulation and verification test based on Vivado software, network model quantization and pruning training based on Pytorch, and network deployment and handwritten digit recognition based on

[0136] During the simulation test, in order to ensure that the performance parameters of the convolutional neural network meet the requirements of the task book, Vivado is intended to be used to simulate and test the high-level synthesis design. First, constraints are added to it: the input and output ports of the circuit are constrained, the clock cycle is constrained, and the function name is constrained. Secondly, the CPU control is set to AXI bus control. Finally, the IP cores of the convolution and pooling modules for general deterministic random calculation designed are added and connected. By comprehensively simulating the network, the resource occupancy can be obtained as shown in Figure 5. Among them, the resource occupancy in Vivado meets the basic requirements that the LUT resource occupancy is within 100,000 and the FF resource occupancy is within 50,000.

[0137] For network deployment, the bitstream files generated during the simulation test are required: .tcl and .bit. Read them on the Pynq board and package them as Overlay, then the code can be automatically burned and the circuit can be connected. Finally, the network can be called directly using Python statements on Jupyter Notebook.

[0138] When using Jupyter to deploy the network, first call overlay to automatically burn and connect the circuit:

[0139]

[0140] Then define the image file to be recognized, and run the optimized convolution module and pooling module after acceleration:

[0141]

[0142] Finally, while importing the test set, import the LeNet-5 network after quantization and pruning training:

[0143]

[0144] The recognition result is as Figure 6 shown. Obviously, the convolutional neural network model deployed on the FPGA provided in this embodiment can effectively run and quickly complete image recognition.

[0145] When deploying on the FPGA through different methods, the corresponding logic resource occupancies of different implementation methods are shown in Table 2:

[0146] Table 2

[0147]

[0148] Among them, the high-level synthesis implementation method provided in this embodiment allows the use of C language to implement the accelerator and export the RTL of the IP core, enabling the network to design C code by adding a compilation wizard defined by high-level synthesis implementation to achieve parallelization. The parallel version is verified by a time series analysis tool, thereby achieving fast pre-synthesis simulation. The resource occupancy of other implementation methods is higher than that of the high-level synthesis implementation method provided in this embodiment, which can further prove the advantages of the convolutional neural network model training method provided in this embodiment.

[0149] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0150] Based on the same inventive concept, the embodiment of the present application also provides a training device for a convolutional neural network model for implementing the training method of the convolutional neural network model involved above. The implementation solution for solving problems provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the training device for the convolutional neural network model provided below can refer to the limitations on the training method of the convolutional neural network model in the above text, and will not be repeated here.

[0151] In one embodiment, as Figure 7 shown, a training device 700 for a convolutional neural network model is provided, including: a model pre-training module 710, a quantization pruning module 720, a high-level synthesis design module 730, and a model deployment module 740, where:

[0152] The model pre-training module 710 is used to obtain a pre-trained network model;

[0153] The quantization pruning module 720 is used to perform quantization processing and pruning processing on the pre-trained network model respectively to obtain a target lightweight network model;

[0154] A high-level synthesis design module 730 is used to design a general deterministic random computing convolution module and a deterministic random computing pooling module in a high-level synthesis manner. Among them, a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register;

[0155] A model deployment module 740 is used to deploy the target lightweight network model on the field programmable gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module.

[0156] In summary, this embodiment provides a training device for a convolutional neural network model. By performing quantization processing and pruning processing on the convolutional neural network model, the lightweight of the convolutional neural network model can be realized, which is more conducive to the hardware design of the convolutional neural network model on the FPGA. By constructing a deterministic random computing unit through a linear feedback shift register to replace the traditional random number generator and generate the corresponding random sequence, the problems existing in the fusion scheme of the random network and the convolutional neural network, such as the accumulation of calculation errors caused by probability fluctuations and the excessive hardware overhead of the traditional random number generator, can be effectively overcome. And based on HLS, a general deterministic random computing convolution module and a deterministic random computing pooling module are implemented, which can combine with the acceleration framework in the FPGA to effectively improve the speed and efficiency of deploying the convolutional neural network model in the FPGA, and can further accelerate the image processing speed of the CNN model in the FPGA.

[0157] Each module in the above-mentioned training device for a convolutional neural network model can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0158] In one embodiment, an electronic device is provided. The electronic device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and external devices. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a training method for a convolutional neural network model. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.

[0159] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0160] In one embodiment, an electronic device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0161] Obtain a pre-trained network model;

[0162] Perform quantization processing and pruning processing on the pre-trained network model respectively to obtain a target lightweight network model;

[0163] Design a general deterministic random computing convolution module and a deterministic random computing pooling module in a high-level synthesis manner. Among them, a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module. The deterministic random computing unit is used to generate a random sequence through a linear feedback shift register;

[0164] Deploy the target lightweight network model on a field programmable gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module.

[0165] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0166] Obtain a pre-trained network model;

[0167] Perform quantization processing and pruning processing on the pre-trained network model respectively to obtain the target lightweight network model;

[0168] Design a general deterministic random computing convolution module and a deterministic random computing pooling module in a high-level synthesis manner. Among them, a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register;

[0169] Deploy the target lightweight network model on a field programmable gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module.

[0170] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0171] Obtain a pre-trained network model;

[0172] Perform quantization processing and pruning processing on the pre-trained network model respectively to obtain the target lightweight network model;

[0173] Design a general deterministic random computing convolution module and a deterministic random computing pooling module in a high-level synthesis manner. Among them, a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register;

[0174] Deploy the target lightweight network model on a field programmable gate array according to the deterministic random computing convolution module and the deterministic random computing pooling module.

[0175] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0176] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0177] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A training method for a convolutional neural network model, characterized in that Applied to a field programmable gate array, including: Obtain a pre-trained network model; Perform quantization processing and pruning processing on the pre-trained network model respectively to obtain a target lightweight network model; Design a general deterministic random computing convolution module and a deterministic random computing pooling module in a high-level synthesis manner. Among them, a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module, and the deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; the deterministic random computing unit includes an 8-bit linear feedback shift register, and a random sequence is generated by using an 8-bit polynomial. In the multi-bit technology, four new bits are generated in each clock cycle, where the 8-bit polynomial is , where x is the input and y is the output; Deploy the target lightweight network model on the field programmable gate array according to the deterministic stochastic computing convolution module and the deterministic stochastic computing pooling module; The performing quantization processing and pruning processing on the pre-trained network model respectively to obtain a target lightweight network model includes: Perform quantization processing on the pre-trained network model to obtain a quantized intermediate network model, wherein the quantization processing adopts a method of rewriting floating-point data into integer data and then performing right shift quantization processing; Perform pruning processing on the intermediate network model to obtain a target lightweight network model after pruning processing; The performing quantization processing on the pre-trained network model to obtain a quantized intermediate network model includes: Receive target input data, wherein the target input data is not regularized and is within a preset range; Determine a target integer weight according to the distribution of the target input data; Perform right shift quantization on the target integer weight and the target input data to obtain a quantized weight, quantized data and quantization related parameters; Call a preset update function for verification to determine whether the pre-trained network model has completed quantization processing; When the quantization related parameters of each layer all include a reduction factor, a zero point offset and an offset value, and the reduction factor has completed offset processing, determine that the pre-trained network model has completed quantization processing to obtain the intermediate network model; Wherein, the right shift quantization includes: Among them, is the output value of the right shift quantization, is the target integer weight after quantization, is the bias, is the gain, represents a shift operation, and the shift value of the shift operation is , , that is , which is the zero-point offset parameter, , that is , and is the reduced parameter after the reduction factor is subjected to a shift process after is the reduction factor, is the shift value, represents the target input data, represents the kernel.

2. The method according to claim 1, wherein The performing pruning processing on the intermediate network model to obtain a target lightweight network model after pruning processing includes: For a fully connected layer, set the weights less than a preset threshold to 0, wherein the preset threshold is calculated according to the pruning percentage corresponding to the fully connected layer; For a convolutional layer, continuously prune the convolutional kernel with the smallest L2 norm in all layers until the preset pruning percentage corresponding to the convolutional layer is satisfied.

3. The method according to claim 1, characterized in that The field programmable gate array is connected to an external memory; the method further includes: Optimize the data stream operation of the field programmable gate array, wherein the optimization processing includes partitioning processing and deformation processing on the field programmable gate array array to improve the data exchange rate and data exchange volume of the field programmable gate array; define the loop operation in the field programmable gate array according to loop unrolling and pipeline processing; during the data operation process, interact with the external memory to cache the intermediate results in the external memory.

4. The method according to claim 1, characterized in that, The deploying the target lightweight network model on the field programmable gate array according to the deterministic stochastic computing convolution module and the deterministic stochastic computing pooling module includes: Perform simulation testing on the deterministic stochastic computing convolution module and the deterministic stochastic computing pooling module to obtain a target bitstream file; Pack the target bitstream file into a hardware overlay file, and automatically burn the code and connect the circuit according to the hardware overlay file; Define the image file to be recognized, and run the optimized convolutional module and pooling module for acceleration; Import the test set and the target lightweight network model at the same time.

5. A training device for a convolutional neural network model, characterized in that, Applied to a field programmable gate array, including: A model pre-training module for obtaining a pre-trained network model; A quantization pruning module for respectively performing quantization processing and pruning processing on the pre-trained network model to obtain a target lightweight network model; A high-level synthesis design module is used to design a general deterministic random computing convolution module and a deterministic random computing pooling module in a high-level synthesis manner. Among them, a deterministic random computing unit is integrated in the deterministic random computing convolution module and the deterministic random computing pooling module. The deterministic random computing unit is used to generate a random sequence through a linear feedback shift register; the deterministic random computing unit includes an 8-bit linear feedback shift register, and uses an 8-bit polynomial to generate a random sequence. In the multi-bit technology, four new bits are generated in each clock cycle. Among them, the 8-bit polynomial is , where x is the input and y is the output; A model deployment module for deploying the target lightweight network model on the field programmable gate array according to the deterministic random calculation convolutional module and the deterministic random calculation pooling module; The quantization pruning module is specifically used for performing quantization processing on the pre-trained network model to obtain a quantized intermediate network model, wherein the quantization processing adopts a method of rewriting floating-point data into integer data and then performing right-shift quantization processing; performing pruning processing on the intermediate network model to obtain the target lightweight network model after pruning processing; the quantization pruning module is specifically used for receiving target input data, wherein the target input data is not normalized and is within a preset interval; determining target integer weights according to the distribution of the target input data; performing right-shift quantization on the target integer weights and the target input data to obtain quantized weights, quantized data and quantization-related parameters; calling a preset update function for verification to determine whether the pre-trained network model has completed quantization processing; when the quantization-related parameters of each layer include a reduction factor, a zero-point offset and an offset value, and the reduction factor has completed the offset processing, determining that the pre-trained network model has completed quantization processing to obtain the intermediate network model; Among them, the right-shift quantization includes: Among them, is the output value of the right shift quantization, is the target integer weight after quantization, is the bias, is the gain, represents a shift operation, and the shift value of the shift operation is , , that is , is the zero-point offset parameter, , that is , is the reduction parameter after the reduction factor is subjected to a shift process after is the reduction factor, is the shift value, represents the target input data, represents the kernel.

6. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the training method of the convolutional neural network model according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the training method of the convolutional neural network model according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Data quantification method and device, computer equipment and storage medium

    CN111176853A

  • Intelligent inference network system and addition unit and pooling unit circuitry

    CN112949830A

  • FPGA-based neural network acceleration method and system

    CN117521752A