A Neural Network Inference Acceleration Method Based on Heterogeneous Platforms

By designing neural network accelerators on heterogeneous platforms, the problem that the existing technology cannot effectively support lightweight convolutional neural networks is solved, and efficient acceleration and performance improvement of convolutional neural networks are achieved, which is suitable for AI applications on edge devices.

CN114742225BActive Publication Date: 2025-05-27HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210361419.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-05-27
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

The design of existing accelerated convolutional neural network inference processes cannot effectively support lightweight convolutional neural networks containing deep separable convolutions, making these models difficult to apply on embedded edge devices.

Method used

The neural network inference acceleration method based on heterogeneous platform is adopted, and the neural network accelerator is designed using a system-on-chip system of processors and FPGAs, including ordinary convolution modules, deep separable convolution modules, fully connected modules, pooling modules, batch normalization modules and activation function modules. The forward inference process of neural networks is accelerated through cache optimization, flow and data flow optimization.

Benefits of technology

It realizes efficient acceleration of convolutional neural networks, supports ordinary convolutional neural networks at different depths and lightweight convolutional neural networks, improves the performance and resource utilization of the system, and is suitable for AI application deployment on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114742225B_ABST
    Figure CN114742225B_ABST
Patent Text Reader

Abstract

A neural network inference acceleration method based on a heterogeneous platform of the present invention includes building a hardware system for neural network accelerator inference using a heterogeneous platform of a processor + FPGA. The processor is responsible for logical control, and the FPGA is responsible for parallel acceleration of computationally intensive tasks, giving full play to the advantages of the heterogeneous platform. First, a neural network accelerator is designed in the FPGA. The neural network accelerator includes a general convolution module, a depthwise separable convolution module, a fully connected module, a pooling module, a batch normalization module, and an activation function module to complete the convolution calculation of the neural network and the processing of output data. Then, effective acceleration is carried out by means of convolutional block division, parallel convolutional calculation, setting and optimization of caching, data flow optimization, and pipelining, improving the operation speed and resource utilization rate of the convolutional neural network accelerator. The present invention can be used to accelerate the forward inference of a convolutional neural network including general convolution, depthwise separable convolution, batch normalization, activation function, pooling, and fully connected operators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a neural network inference acceleration method based on a heterogeneous platform. Background Art

[0002] Convolutional neural network is the most important model in the current field of artificial intelligence deep learning. It is widely used in scenarios such as image recognition and object detection, and has achieved high accuracy. The application scenarios of deep learning mainly include: intelligent driving, intelligent vehicle damage assessment, goods sorting, crop identification, defect detection of parts in industrial manufacturing, face recognition in the security field, etc. The development of convolutional neural networks has also brought us more challenges. The number of weight parameters is increasing, and the amount of computation is also increasing. As a result, complex models are difficult to be transplanted to mobile and embedded devices. Convolutional neural network models usually have hundreds of layers of networks and millions of weight parameters. Saving a large number of weight parameters requires high memory for edge devices, while the storage capacity of most edge devices is very limited. Therefore, it is important to deploy lightweight convolutional neural networks on edge devices. The main idea of lightweight model design lies in designing a more efficient network computing method, so that while reducing network parameters, the network performance is not lost. Currently, lightweight neural networks such as SqueezeNet, MobileNet, ShuffleNet, Xception, etc. mainly use depthwise separable convolution to reduce the number of parameters and the amount of computation.

[0003] At the same time, with the continuous popularization of intelligent devices and mobile terminals, when AI applications are deployed on embedded devices, there are high requirements for their speed, performance, and power consumption. Currently, there are various hardware platform design methods for convolutional neural network acceleration algorithms: acceleration systems designed using graphics processing units (GPUs), application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs). GPU acceleration is a commonly used acceleration method, but it has a large volume, high power consumption, and high cost, and is not suitable for edge applications. ASIC acceleration has good performance, and the volume and power consumption are controllable, but its design is difficult, the development cycle is long, and the flexibility is poor. FPGA is a platform on which hardware circuits can be built according to different algorithms. With its powerful parallel capabilities, flexible design methods, and high performance-to-power ratio, FPGA has become one of the most attractive implementation platforms for hardware acceleration of convolutional neural networks in embedded devices.

[0004] At present, most of the accelerator designs for accelerating the inference process of convolutional neural networks only include ordinary convolution, pooling, and fully connected parts, so they can only support simple ordinary convolutional neural networks and cannot well support lightweight convolutional neural networks containing depthwise separable convolutions, resulting in many models with complex networks, high performance, and lightweight characteristics being unable to be well applied on embedded edge devices; most of the designs for accelerating the inference process of convolutional neural networks only use a small number of methods to optimize the acceleration process and do not combine various methods simultaneously. The convolution calculation and data processing processes are not efficient enough to fully utilize the parallelism of the FPGA; and most of them focus on hardware acceleration on the FPGA side, while the acceleration based on heterogeneous platforms can fully combine the processor's control of logic and the FPGA's parallel acceleration of computationally intensive tasks, giving full play to the advantages of heterogeneous platforms and thus improving the overall performance of the system. Summary of the Invention

[0005] A neural network inference acceleration method based on a heterogeneous platform proposed by the present invention can solve the above technical problems and is applicable to the application of convolutional neural networks on edge devices.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A neural network inference acceleration method based on a heterogeneous platform, based on a heterogeneous platform, the heterogeneous platform includes a processor and an FPGA on-chip system, and includes the following steps:

[0008] Design a neural network accelerator in the FPGA. The neural network accelerator includes a computing module and a data processing module to complete the convolution calculation of the neural network and the processing of output data; then, methods such as cache setting and optimization, pipelining, and data flow optimization are used to accelerate the forward inference process of the neural network;

[0009] Among them, the computing module includes an ordinary convolution module, a depthwise separable convolution module, and a fully connected module; the data processing module includes a pooling module, a batch normalization module, and an activation function module;

[0010] The ordinary convolution module performs block calculation on the convolution, and only performs convolution calculation between the input of a fixed block size and the convolution kernel weights each time, and then completes the calculation of pixel points on all output feature maps in turn through the sliding window method;

[0011] The depthwise separable convolution module includes a depth convolution module and a pointwise convolution module; a convolution kernel of the depth convolution has only one channel, and one convolution kernel is only responsible for one channel of the input feature map of the convolution, and the number of channels of the generated output feature map is the same as the number of input channels; the convolution kernel size of the pointwise convolution module is 1x1xD, and D is the number of channels of the output of the previous layer of convolution;

[0012] The fully connected module reuses the ordinary convolution part, and sets the input scale to 1x1xC through the AXI_Lite bus, where C is the number of channels;

[0013] The pooling module includes max pooling and average pooling, which are selected by configuring the values of the corresponding registers; max pooling is only a logical operation, and the input data is compared by a comparator to output a maximum value; average pooling uses an adder to sum the input data and then uses a shift register to implement division calculation to obtain the average value;

[0014] The batch normalization module performs data normalization on the outputs of the ordinary convolution module and the depthwise separable convolution module.

[0015] Furthermore, the method of using cache settings and optimization, pipelining, and data flow optimization is adopted to accelerate the forward inference process of the neural network, specifically including:

[0016] Cache settings and optimization: In the on-chip BRAM memory of the FPGA, input cache: IN[Tn][Tic][Tir], weight cache: W[Tm][Tn][Tkc][Tkr], and output cache: OUT[Tm][Toc][Tor] are respectively set. The size of the cache is determined according to the size of the convolution block variables; the cache splitting method is used to split the channel dimension of the input and output caches, and the input and output channel dimensions of the weight cache are split so that they are distributed in different BRAM blocks, increasing the number of its input and output ports and enabling simultaneous read and write operations; the input cache is divided into Tn independent cache blocks, the weight cache is divided into Tm*Tn independent cache blocks, and the output cache is divided into Tm independent cache blocks;

[0017] Data flow optimization: The dual-buffer + ping-pong operation method is adopted for parallel optimization of the task-level data flow, that is, two input caches, weight caches, and output caches of the same size are set in the on-chip BRAM of the FPGA, and the "ping-pong" data transfer mechanism is used to perform data reading, convolution calculation, and result writing back simultaneously.

[0018] Furthermore, it also includes the training of the neural network, specifically by building a neural network model on the server side, importing a dataset for training, obtaining the model parameters of each layer of the neural network after training, including the weights and bias parameters of the convolution layer and the batch normalization parameters, and saving the parameters as a binary file and putting it into the SD card.

[0019] Furthermore, the forward inference process of the neural network includes:

[0020] The application allocates a continuous array space ARRAY_IMAGE in the DDR memory, reads the input image, preprocesses it, and then places it in this array space;

[0021] The application allocates a continuous array space ARRAYi in the DDR memory, reads the neural network convolutional layer parameters and batch normalization layer parameters into this array space;

[0022] The application loads the entire accelerator into the FPGA in the form of a binary bitstream file;

[0023] The application configures the registers of each module inside the FPGA accelerator according to the structure of the neural network model, and adjusts the scales of convolution, pooling, and fully connected operations;

[0024] The application respectively calls the ordinary convolution module, depthwise separable convolution module, fully connected module, pooling module, batch normalization module, and activation function module in the FPGA accelerator according to the structure of the neural network model, and passes the input data, convolutional layer parameters, and batch normalization layer parameters in the memory DDR into the FPGA accelerator for accelerated calculation;

[0025] After the FPGA accelerator completes the inference calculations of all levels of the neural network, it returns the inference results to the memory DDR for the application to access.

[0026] Furthermore, the ordinary convolution module performs block-based convolution calculations. Each time, it only performs convolution calculations between the input of a fixed block size and the convolutional kernel weights, and then sequentially completes the calculations of all pixel points on the output feature map through the sliding window method. Specifically, it includes:

[0027] First, adopt the convolution block strategy to complete the entire convolution calculation by time-division multiplexing the convolution block unit. The block variables are the number of channels of the output block: Tm, width: Toc, height: Tor, the width of the convolutional kernel block: Tkc, height: Tkr, the number of channels of the input block: Tn, width: Tic, height: Tir. The input block variables and output block variables satisfy the following relationships: Tir = (Tor - 1) * S + Tkr, Tic = (Toc - 1) * S + Tkc, where S is the convolutional kernel moving step;

[0028] The block variables are determined by the size of the output feature map of the convolutional layer and the size of the convolutional kernel, and are restricted by the on-chip resources;

[0029] The convolution calculation consists of a nested loop structure in four dimensions: within the output feature map, output channels, input channels, and within the convolution kernel. There is parallelism in each dimension. By adjusting the loop order of the nested loops within the convolution block, the loops of the output channels and input channels are placed in the innermost layer and parallel convolution calculations are performed on these two dimensions, enabling the outer loops in the convolution block to continuously reuse the internal parallel computing cores to complete the calculation of the entire convolution block unit.

[0030] Furthermore, the parallel computing includes: Tm convolution computing engines PE, and each PE consists of Tn multipliers, an adder tree with a depth of and an adder. It performs the multiplication operations of Tn input feature map pixels and Tm*Tn convolution kernel weights in parallel, then accumulates the results through the adder tree, and finally adds them to the previously saved partial sums.

[0031] Furthermore, the design of the batch normalization module includes:

[0032] Let the four parameters of the output channels of the convolution module be u, σ, γ, and β respectively, where u is the mean value; σ is the variance; γ is the scaling coefficient, and β is the offset coefficient; the input data of the batch normalization part is x, and the output data is:

[0033]

[0034] where ∈ is a hyperparameter;

[0035] The design of the batch normalization module is to read out the corresponding parameters u, σ, γ, and β in the memory DDR. Each output channel corresponds to a set of parameters. According to the above formula, a specific arithmetic circuit module is built using the hardware resources in the FPGA, and the input is the result output after convolution calculation.

[0036] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the above method.

[0037] As can be seen from the above technical solutions, the neural network inference acceleration method based on a heterogeneous platform of the present invention uses a heterogeneous platform of a processor + FPGA to build a hardware system for neural network accelerator inference. The processor is responsible for logical control, and the FPGA is responsible for parallel acceleration of computationally intensive tasks, giving full play to the advantages of the heterogeneous platform. First, a neural network accelerator is designed in the FPGA. The neural network accelerator includes a general convolution module, a depthwise separable convolution module, a fully connected module, a pooling module, a batch normalization module, and an activation function module to complete the convolution calculation of the neural network and the processing of output data. Then, effective acceleration is achieved by means of convolutional block division, parallel convolution calculation, cache setting and optimization, data flow optimization, and pipelining, improving the operation speed and resource utilization rate of the convolutional neural network accelerator. The present invention can be used to accelerate the forward inference of convolutional neural networks including general convolution, depthwise separable convolution, batch normalization, activation function, pooling, and fully connected operators.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] 1. The present invention builds a general convolutional neural network accelerator, which is applicable to general convolutional neural networks and lightweight convolutional neural networks with different depths;

[0040] 2. The present invention configures the registers of each module in the accelerator through the parameters of each layer of the network, enabling the reuse of each module in the accelerator; the implementation of the fully connected operator in the present invention reuses the convolutional operator, further reducing the hardware resource overhead and improving the utilization rate of hardware resources;

[0041] 3. The present invention improves the parallelism of convolutional calculation, reduces the data transmission time, and accelerates the network inference process by means of convolutional block division, parallel convolution calculation, cache setting and optimization, data flow optimization, and pipelining in the FPGA;

[0042] 4. The present invention uses the processor for logical control and uses the FPGA for parallel acceleration of computationally intensive tasks, giving full play to the advantages of the heterogeneous platform. The method proposed by the present invention will be able to deploy AI applications on intelligent devices and mobile terminals at a relatively fast speed and with high performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the network structure diagram of the lightweight convolutional neural network MobileNetV2 of the present invention;

[0044] Figure 2 is the schematic diagram of convolutional block division calculation of the present invention;

[0045] Figure 3 is the schematic diagram of the hardware structure of the parallel convolution calculation core of the present invention;

[0046] Figure 4 This is the cache setting and optimization diagram of the present invention;

[0047] Figure 5 This is the data flow optimization diagram of the present invention;

[0048] Figure 6 This is the pipelining diagram of the present invention;

[0049] Figure 7 This is the hardware architecture diagram for building a neural network forward inference system based on a heterogeneous platform on the SoC in this embodiment. Detailed implementation manners

[0050] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0051] The acceleration method for neural network inference based on a heterogeneous platform proposed by the present invention is applicable to ordinary convolutional neural networks and lightweight convolutional neural networks with different depths, and is applied to a neural network inference acceleration system composed of an SoC system-on-chip, DDR memory, SD memory card, and application programs; the SoC system-on-chip is a heterogeneous system composed of a processor and an FPGA; it includes:

[0052] Step 1: Design of a neural network accelerator;

[0053] Design a neural network forward inference accelerator in the FPGA, including a calculation module and a data processing module to complete the convolution calculation and the processing of output data. The calculation module includes: an ordinary convolution module, a depthwise separable convolution module, and a fully connected module, and the data processing module includes: a pooling module, a batch normalization module, and an activation function module.

[0054] Step 2: Optimization method of the neural network accelerator;

[0055] Accelerate and optimize the ordinary convolution and depthwise separable convolution in Step 1, and use the methods of cache setting and optimization, pipelining, and data flow optimization to speed up the neural network forward inference process.

[0056] Step 3: Construction of a heterogeneous platform;

[0057] Utilize the neural network accelerator designed in Step 2, combined with the processor part, to build a complete hardware system for neural network inference;

[0058] Step 4: Training of the model;

[0059] Build a neural network model on the server side, import the dataset for training, and import the model parameters of each trained layer, including weights, biases, and batch normalization parameters (β, γ, u, σ), into the SD card.

[0060] Step 5: Execute neural network inference;

[0061] Configure the registers inside the accelerator according to the scale of each layer of the specific neural network, transfer the model parameters obtained from the training in Step 4 to the accelerator through the AXI bus, and call different modules in the accelerator according to the specific network structure to complete the forward inference of the entire network.

[0062] Specifically, the neural network accelerator in Step 1 integrates the basic operators required to build a convolutional neural network, especially the depthwise separable convolution operator for building lightweight neural networks. In the ordinary convolution part, a convolution block strategy is adopted to calculate the convolution in blocks. Only the convolution calculation between the input of a fixed block size and the convolution kernel weights is performed each time. Then, by adjusting the loop order of the nested loops within the convolution block, the loops of the output channels and the input channels are placed in the innermost layer and parallel convolution calculations are performed on these two dimensions, enabling the outer loop in the convolution block to continuously reuse the internal parallel computing cores, thereby completing the calculation of the entire convolution block unit. Then, the calculation of all pixel points on the output feature map is completed sequentially through the sliding window method;

[0063] The depthwise separable convolution part includes depthwise convolution and pointwise convolution. The design of the depthwise convolution part is basically the same as the implementation process of ordinary convolution, except for the calculation method. A convolution kernel in depthwise convolution has only one channel, and one convolution kernel is only responsible for convolving one channel of the input feature map, and the number of output feature map channels is the same as the number of input channels. Therefore, the depthwise separable convolution only requires 5 nested for loops and needs to change the data reading order and the convolution calculation order. The calculation of pointwise convolution is very similar to that of ordinary convolution, except that its convolution kernel size is 1x1x D, where D is the number of channels in the previous layer. Pointwise convolution will perform weighted combination on the output feature map of depthwise convolution in the depth direction to generate a new feature map. Therefore, the implementation of the pointwise convolution part can reuse the ordinary convolution part. Since the fully connected operation is a special convolution operation, the implementation of the fully connected part reuses the ordinary convolution part, further reducing the hardware resource overhead. The design of the pooling part includes max pooling and average pooling, which are selected by configuring the values of the corresponding registers. Batch normalization part: Each output channel of each convolution layer has four parameters β, γ, u, σ, and each pixel point of the output feature map of each channel is processed using the formula. The specific formula is as follows:

[0064]

[0065] where: u is the mean value; σ is the variance; γ is the scaling coefficient, and β is the offset coefficient; ∈ is the hyperparameter.

[0066] Specifically, the optimization method in Step 2: Cache setting and optimization: Set input cache, weight cache, and output cache in the on-chip BRAM memory in the FPGA. The size of the cache is determined according to the size of the block parameters. Input cache: IN[Tn][Tic][Tir], weight cache: W[Tm][Tn][Tkc][Tkr], output cache: OUT[Tm][Toc][Tor]. Due to the limitation of the number of input and output ports, the data throughput rate of the accelerator is affected by the read and write rates of the cache data. To ensure the parallelism of data processing in the accelerator, it is necessary to increase the throughput rate of data reading in the accelerator. Therefore, it is necessary to partition the input cache, weight cache, and output cache so that they are distributed in different BRAM blocks, increasing the number of input and output ports of each cache area. The input cache is divided into Tn independent caches, the weight cache is divided into Tm*Tn independent caches, and the output cache is divided into Tm independent caches. Data flow optimization method: Adopt the double buffer + ping-pong operation method to perform parallel optimization of the task-level data flow. That is, set two input caches, weight caches, and output caches of the same size in the on-chip BRAM of the FPGA. Using the "ping-pong" data transfer mechanism, data reading, convolution calculation, and result writing back can be carried out simultaneously, effectively overlapping the data transfer time and the calculation time, thereby masking the data transfer time and improving the throughput of RTL. Pipelining: Divide the original larger calculation process into multiple smaller steps. Each step can independently and parallelly process different data, which is a common circuit optimization method that exchanges area for speed. Its throughput is between serial processing and parallel processing.

[0067] Specifically, the processor system part in Step 3 is a Cortex-A9 processor based on the ARM core, mainly used for logic control. According to the specific neural network model, configure the registers of each module inside the neural network accelerator, adjust the scale of convolution, pooling, and fully connected operations, and transfer the input data, weights, and biases in the memory DDR to the neural network accelerator in Step 2. The AXI_HP slave interface of the processor is connected to the AXI_Master master interface of the accelerator, and the AXI_GP master interface of the processor is connected to the AXI_GP slave interface of the accelerator. Data flow and control flow are transmitted through the AXI4 bus and the AXI_Lite bus. Call different modules in the accelerator according to the hierarchical structure of the specific model to complete the inference of the entire neural network and return the final result to the DDR.

[0068] The following specific description includes:

[0069] 1.1 The 1.1 processor is responsible for the logical control of neural network inference, and the FPGA is responsible for the computational acceleration of neural network inference; in terms of the connection and communication method between the processor and the FPGA, the AXI_HP slave interface of the processor is connected to the AXI_Master master interface of the FPGA accelerator, and the AXI_GP master interface of the processor is connected to the AXI_GP slave interface of the FPGA accelerator; the processor and the FPGA accelerator transmit data through the AXI4 bus; the processor configures the internal registers of the accelerator through the AXI_Lite bus;

[0070] 1.2 The FPGA is a neural network inference accelerator, including a computing module, a data processing module, and an AXI bus interface, providing convolutional calculation of the neural network and data processing;

[0071] 1.3 The application program calls the modules in the accelerator according to the neural network structure to complete the neural network inference work, including loading the weight parameters, biases, and batch normalization parameters of the model from the SD card to the DDR memory; transmitting the configuration parameters of each layer of the neural network, such as the size of the input feature map, the size of the convolutional kernel, the stride, etc. to the accelerator;

[0072] Step 2.1, Design of the neural network accelerator;

[0073] Design a neural network inference accelerator on the FPGA, including a computing module and a data processing module; among them:

[0074] The computing module includes: a general convolution module, a depthwise separable convolution module, and a fully connected module;

[0075] The data processing module includes: a pooling module, a batch normalization module, and an activation function module;

[0076] 2.1.1 Design of the general convolution module: Due to the different scales of each layer of the convolutional neural network and the limitation of the FPGA's own resources, it is impossible to load the feature map data and weights of each layer onto the FPGA. Therefore, it is necessary to design a general convolution acceleration module with a fixed scale size, that is, perform block calculation on the convolution, and only perform the convolution calculation between the input of a fixed block size and the convolution kernel weights each time, and then complete the calculation of all pixel points on the output feature map in turn through the sliding window method.

[0077] Therefore, first, the convolution block strategy is adopted to complete the entire convolution calculation by time-division multiplexing convolution block units. The block variables are the number of channels of the output block: Tm, width: Toc, height: Tor, the width of the convolution kernel block: Tkc, height: Tkr, the number of channels of the input block: Tn, width: Tic, height: Tir. According to the principle of the convolutional neural network, the input block variables and output block variables satisfy the following relationship: Tir = (Tor - 1) * S + Tkr, Tic = (Toc - 1) * S + Tkc, where S is the convolution kernel moving step. The block variables are determined by the size of the output feature map of the convolution layer and the size of the convolution kernel, and are restricted by on-chip resources.

[0078] The convolution calculation mainly consists of nested loop structures in four dimensions: inside the output feature map, output channels, input channels, and inside the convolution kernel. And there is a certain degree of parallelism in each dimension. By adjusting the loop order of the nested loops within the convolution block, the loops of the output channels and input channels are placed in the innermost layer and parallel convolution calculations are performed on these two dimensions, enabling the outer loops in the convolution block to continuously reuse the internal parallel computing cores, thus completing the calculation of the entire convolution block unit. Among them, the parallel computing core includes: Tm convolution computing engines PE, and each PE consists of Tn multipliers, an adder tree with a depth of and an adder, which perform parallel multiplication operations on Tn input feature map pixels and Tm * Tn convolution kernel weights. Then, the results are accumulated through the adder tree and finally added to the previously saved partial sums.

[0079] 2.1.2 The depthwise separable convolution module includes a depthwise convolution module and a pointwise convolution module;

[0080] Design of the depthwise convolution module: Since the calculation method is different from that of ordinary convolution, a convolution kernel of depthwise convolution has only one channel, and one convolution kernel is only responsible for convolving one channel of the input feature map, and the number of output feature map channels is the same as the number of input channels. Therefore, depthwise convolution only requires 5 nested for loops and needs to change the data reading order and convolution calculation order, and other implementation processes are basically the same as those of ordinary convolution.

[0081] The implementation of the pointwise convolution module reuses the ordinary convolution module and sets the convolution kernel size to 1x1xD, where D is the number of channels of the output of the previous layer of convolution;

[0082] 2.1.3 Since the fully connected operation is a special convolution operation, the implementation of the fully connected module reuses the ordinary convolution part, thus reducing the overhead of hardware resources. Only the input scale needs to be set to 1x1xC through the AXI_Lite bus, where C is the number of channels.

[0083] 2.1.4 Design of the pooling module: The hardware circuit is mainly built using resources such as adders, comparators, selectors, and shift registers. The pooling module includes max pooling and average pooling, which can be selected by configuring the values of the corresponding registers. Among them, max pooling is very suitable for FPGA operations. It is only a logical operation. The input data is compared by a comparator, and a maximum value is output. Since the size of the pooling is not large, the pooling operation can be completed within one clock using combinational logic. Average pooling uses an adder to sum the input data and then uses a shift register to perform division calculations to obtain the average value. Since the pooling calculations between channels are independent and have the characteristic of parallel calculation in the channel direction, multiple groups of the same calculation resources can be used for parallel calculation to accelerate the pooling calculation process.

[0084] 2.1.5 The batch normalization module performs data normalization processing on the outputs of the ordinary convolution module and the depthwise separable convolution module. Let the four parameters of the output channels of the convolution module be u, σ, γ, and β respectively, where u is the mean value; σ is the variance; γ is the scaling coefficient, and β is the offset coefficient; the input data of the batch normalization part is x, and the output data is:

[0085]

[0086] where ∈ is a hyperparameter;

[0087] Design of the batch normalization module: Read out the corresponding parameters u, σ, γ, and β in the memory DDR. Each output channel corresponds to a set of parameters. According to the above formula, a specific operation circuit module is built using the hardware resources in the FPGA, and the input is the result output after convolution calculation.

[0088] Step 2.2, Optimization of the neural network accelerator:

[0089] Accelerate and optimize the ordinary convolution and depthwise separable convolution in step 2.1. It is characterized in that the methods of cache setting and optimization, pipelining, and data flow optimization are adopted to accelerate the forward inference process of the neural network.

[0090] 2.2.1 Cache Settings and Optimization. In the on-chip BRAM memory of the FPGA, the input cache: IN[Tn][Tic][Tir], weight cache: W[Tm][Tn][Tkc][Tkr], and output cache: OUT[Tm][Toc][Tor] are respectively set. The size of the cache is determined according to the size of the convolution block variables. Since parallel computing is performed on the input and output channels, the read and write parallelism of the input feature map, weights, and output feature map within the convolution block are Tn, Tm*Tn, and Tm respectively. Therefore, it is necessary to complete the simultaneous reading and writing of multiple data in the cache within one cycle to meet the requirements of parallel computing. However, the cache is continuously stored in the BRAM and is implemented as a BRAM with at most 2 data ports. Therefore, it is difficult to complete the simultaneous reading and writing of multiple data in a cache within one cycle. To improve the data throughput rate, the cache splitting method is used to split the channel dimension of the input and output caches and split the input and output channel dimensions of the weight cache. Make it distributed in different BRAM blocks, increase the number of its input and output ports, and enable simultaneous read and write operations. The input cache is divided into Tn independent cache blocks, the weight cache is divided into Tm*Tn independent cache blocks, and the output cache is divided into Tm independent cache blocks;

[0091] 2.2.2 Data Flow Optimization. To mask the data transmission time, the double buffer + ping-pong operation method is used for parallel optimization of the task-level data flow. That is, two input caches, weight caches, and output caches of the same size are set in the on-chip BRAM of the FPGA. Using the "ping-pong" data transmission mechanism, the data reading, convolution calculation, and result writing are performed simultaneously, which can effectively overlap the data transmission time and the calculation time, thereby reducing the delay and improving the throughput of the RTL, and further improving the performance of the accelerator. The specific ping-pong cache algorithm can be implemented through the if-else structure + ping-pong flag.

[0092] 2.2.3 Pipelining divides the original larger calculation process into multiple smaller steps. Each step can independently and parallelly process different data. It is a commonly used circuit optimization method that exchanges area for speed. Its throughput is between serial processing and parallel processing. It can make the register read data, calculation, and write data in the convolution calculation process flow, increasing the data throughput rate.

[0093] Step 2.3, Training of the neural network;

[0094] 2.3.1 Build a neural network model on the server side and import the dataset for training;

[0095] 2.3.2 After the training is completed, the model parameters of each layer of the neural network are obtained, including the weights and bias parameters of the convolutional layer and the batch normalization parameters, and the parameters are saved as a binary file and placed in the SD card;

[0096] Step 2.4, Execution of neural network inference;

[0097] 2.4.1 The application program allocates a continuous physical address array space ARRAY_IMAGE in the DDR memory, and reads the input image, preprocesses it, and then places it in this array space;

[0098] 2.4.2 The application program allocates a continuous physical address array space ARRAYi in the DDR memory, and reads the neural network convolutional layer parameters and batch normalization layer parameters into this array space;

[0099] 2.4.3 The application program loads the entire accelerator into the FPGA in the form of a binary bitstream file;

[0100] 2.4.4 The application program configures the registers of each module inside the FPGA accelerator according to the structure of the neural network model, and adjusts the scale of convolution, pooling, and fully connected operations;

[0101] 2.4.5 The application program calls the ordinary convolution module, depthwise separable convolution module, fully connected module, pooling module, batch normalization module, and activation function module in the FPGA accelerator respectively according to the structure of the neural network model, and passes the input data in the memory DDR, convolutional layer parameters, and batch normalization layer parameters into the FPGA accelerator for accelerated calculation;

[0102] 2.4.6 After the FPGA accelerator completes the inference calculation of all levels of the neural network, it returns the inference result to the memory DDR for the application program to access.

[0103] The following will be specifically described in conjunction with the accompanying drawings:

[0104] Figure 1 As shown, in this embodiment, the lightweight convolutional neural network MobileNetV2 is selected for the neural network. The network structure is specifically described as follows: The size of the input image of the network is (224, 224, 3). The main structure is based on depthwise separable convolution, and a reverse residual structure (block) with a linear bottleneck is adopted. The first layer of the network uses ordinary convolution, and then 16 reverse residual structures with linear bottlenecks are used to extract features. Finally, average pooling and full connection are used for classification. The network has a total of 3.4 million parameters. Specifically, one reverse residual structure with a linear bottleneck is used for illustration, and others are similar.

[0105] The acceleration method for neural network inference based on a heterogeneous platform proposed by the present invention includes the following steps.

[0106] Step 1) Design of a neural network accelerator: Design a forward inference accelerator for a neural network in an FPGA, including a computing module and a data processing module to complete convolution calculation and output data processing. The computing module includes: a general convolution module, a depthwise separable convolution module, and a fully connected module. The data processing module includes: a pooling module, a batch normalization module, and an activation function module. These parts are all general. The neural network accelerator integrates the basic operators required to build a neural network, especially the depthwise separable convolution operator used to build the lightweight convolutional neural network MobileNetV2. For the general convolution part: First, adopt a convolution block strategy to perform block calculation on the convolution. Then, by adjusting the loop order of the nested loops within the convolution block, place the loops of the output channels and input channels in the innermost layer and perform parallel convolution calculation on these two dimensions, enabling the outer loop in the convolution block to continuously reuse the internal parallel computing cores, thereby completing the calculation of the entire convolution block unit. Then, calculate the pixels on all output feature maps in sequence through the sliding window method; Figure 2 is a schematic diagram of the convolution block calculation in this embodiment. When performing block calculation on the convolution, only perform the convolution calculation between the input of a fixed block size and the convolution kernel weights each time. The block variables are: the number of channels of the output block: Tm, width: Toc, height: Tor; the width of the convolution kernel block: Tkc, height: Tkr; the number of channels of the input block: Tn, width: Tic, height: Tir. According to the principle of the convolutional neural network, the input block variables and output block variables satisfy the following relationship: Tir = (Tor - 1) * S + Tkr, Tic = (Toc - 1) * S + Tkc, where S is the convolution kernel movement step size. Figure 3 is a schematic diagram of the hardware structure of the parallel convolution calculation core in this embodiment, which includes: Tm convolution calculation engines PE, and each PE consists of Tn multipliers, an adder tree with a depth of and an adder, and performs parallel multiplication operations on the pixels of Tn input feature images and the weights of Tm * Tn convolution kernels. Then, accumulate the results through the adder tree, and finally add them to the previously saved partial sum.

[0107] The depthwise separable convolution part includes depthwise convolution and pointwise convolution. The design of the depthwise convolution part is basically the same as the implementation process of ordinary convolution, except for the calculation method. A convolution kernel in depthwise convolution has only one channel, and one convolution kernel is only responsible for convolving one channel of the input feature map, and the number of output feature map channels is the same as the number of input channels. Therefore, the depthwise separable convolution only requires 5 nested for loops and needs to change the order of data reading and the order of convolution calculation. The calculation of pointwise convolution is very similar to that of ordinary convolution, except that its convolution kernel size is 1x1x D, where D is the number of channels in the previous layer. Pointwise convolution will perform weighted combination on the output feature map of depthwise convolution in the depth direction to generate a new feature map. Therefore, the implementation of the pointwise convolution part can reuse the ordinary convolution part. Fully connected part: Since the fully connected operation is a special convolution operation, the implementation of the fully connected part reuses the ordinary convolution part, further reducing the hardware resource overhead. The design of the pooling part includes max pooling and average pooling, which are selected by configuring the values of the corresponding registers. Batch normalization part: Each output channel of the convolution layer has four parameters β, γ, u, σ, and each pixel point of the input feature map of each channel is processed using the formula. The specific formula is as follows:

[0108]

[0109] In the formula: u is the mean value; σ is the variance; γ is the scaling coefficient, β is the offset coefficient; ∈ is the hyperparameter.

[0110] Step 2) Optimization of the neural network accelerator, Figure 4 This is the setup and optimization diagram of the cache in this embodiment. An input cache, a weight cache, and an output cache are set in the on-chip BRAM memory in the FPGA. The size of the cache is determined according to the size of the block parameters. Input cache: IN[Tn][Tic][Tir], weight cache: W[Tm][Tn][Tkc][Tkr], output cache: OUT[Tm][Toc][Tor]. Due to the limitation of the number of input and output ports, the data throughput rate of the accelerator is affected by the cache data reading and writing rates. To ensure the parallelism of data processing in the accelerator, it is necessary to increase the throughput rate of data reading in the accelerator. Therefore, it is necessary to partition the input cache, weight cache, and output cache so that they are distributed in different BRAM blocks, increasing the number of input and output ports of each cache area. The input cache is divided into Tn independent caches, the weight cache is divided into Tm*Tn independent caches, and the output cache is divided into Tm independent caches.

[0111] Figure 5It is the data stream optimization diagram in this embodiment. When data stream optimization is not adopted, data reading, convolution calculation, and output data processing are performed serially. Assume that it takes 3 cycles to read data, 2 cycles for convolution calculation, and 3 cycles for output data processing. Then a total of 8 cycles are required. By using the data stream optimization method, data reading, convolution calculation, and output data processing are performed simultaneously, and a total of 5 cycles are required, further improving the throughput of the accelerator. Figure 6 It is the pipelining diagram in the embodiment. When pipelining is not adopted, it takes 3 cycles to read data, perform operations, and write back from the register each time, and a total of 9 cycles are required for 3 times. By using the pipelining method, data reading, operations, and write-back are pipelined, and a total of 4 cycles are required for 3 times, increasing the data throughput.

[0112] Step 3) Construction of the heterogeneous platform: Using the neural network accelerator designed in the second step, combined with the processor part, a complete hardware system for neural network inference is constructed. The processor system part is a Cortex-A9 processor based on the ARM core, mainly used for logical control. The AXI_HP slave interface of the processor is connected to the AXI_Master master interface of the accelerator, and the AXI_GP master interface of the processor is connected to the AXI_GP slave interface of the accelerator. Data stream and control stream are transmitted through the AXI4 bus and the AXI_Lite bus.

[0113] Figure 7This is the hardware architecture diagram of building a neural network forward inference system based on a heterogeneous platform on the SoC in this embodiment. The architecture details are described as follows: In this embodiment, a top-level accelerator module is designed in the FPGA, which includes a computing sub-module and a data processing sub-module. The computing sub-module includes: a general ordinary convolution module, a depthwise separable convolution module, and a fully connected partially multiplexed ordinary convolution module. The data processing module includes: a pooling module, a batch normalization module, and an activation function module. Then the accelerator is integrated into the SoC heterogeneous system. The AXI_Master master interface of the accelerator is connected to the slave interface of the AXI_HP of the ARM processor, and the slave interface of the AXI_GP of the accelerator is connected to the master interface of the AXI_GP of the ARM processor. The accelerator and the ARM processor transmit data through the AXI4 bus, and the ARM processor configures the internal registers of the accelerator through the AXI_Lite bus to form a complete convolutional neural network inference system hardware architecture. To reduce the overhead of data transmission and accelerate the operation speed, input buffers, weight buffers, and output buffers are set in the BRAM in the FPGA, and the buffers are segmented so that they are distributed in different BRAM blocks, increasing the number of their input and output ports, enabling simultaneous read and write operations, and improving the data throughput rate. The application mainly completes the preparatory work before neural network forward inference, the initialization of the system, and calls the modules in the accelerator according to the specific network structure to complete the inference work, specifically including: loading the weight parameters, biases, and batch normalization parameters of the model from the SD card into the DDR memory, and passing the configuration parameters of each layer of the neural network, such as the size of the input feature map, the size of the convolution kernel, and the stride, to the specific modules in the accelerator.

[0114] Step 4) Training of the model: Based on the tensorflow 2.0 deep learning framework on the server side, build a MobileNetV2 neural network image classification model, and use the ImageNet image dataset for training. Import the trained model parameters of each layer, including weights, biases, and batch normalization parameters (β, γ, u, σ), into the SD card.

[0115] Step 5) Execution of neural network inference: According to the configuration parameters of each layer of the MobileNetV2 neural network, such as the size of the input feature map, the size of the convolution kernel, the stride, and the pooling type, configure the registers of each module inside the neural network accelerator through the AXI_Lite bus, adjust the scale of convolution, pooling, and fully connected operations, and pass the input data, weights, biases, and batch normalization parameters in the DDR memory into the neural network accelerator in Step 2 through the AXI4 bus. Call different modules in the accelerator according to the hierarchical structure of MobileNetV2 to complete the inference of the entire neural network, and return the final result to the DDR.

[0116] The input image size of the MobileNetV2 network is (224, 224, 3). After padding, it becomes (225, 225, 3), which serves as the input to the first layer of ordinary convolution. The first layer of convolution operation consists of 32 kernels (convolution kernels), where the size of each kernel is (3, 3, 3). The stride of the convolution operation is 2. Configure the registers of the ordinary convolution part inside the neural network accelerator through the AXI_Lite bus, then input the weight parameters and input data, and call the ordinary convolution part in the accelerator. By means of convolution block division, parallel convolution calculation, cache setting and optimization, data flow optimization, and pipelining, improve the parallelism and throughput rate of the convolution operation, and obtain a feature map with a size of (112, 112, 32). Then, after data processing by the batch normalization part and the activation function part in the data processing module, the size of the feature map remains unchanged. For example, in the first layer of batch normalization, there are 32 channels, and each channel has four parameters β, γ, u, σ. Use formula 1 to perform operations on each pixel point of the input feature map for each channel, and pass the results to the activation function part. The activation function part uses ReLu6, that is, when the value output by batch normalization is less than 0, the output is 0; when it is greater than 0 and less than 6, it is output as it is; when it is greater than 6, the output is 6. The input of the first layer of depth convolution is (112, 112, 32), and there are 32 convolution kernels with a size of (3, 3). One convolution kernel is only responsible for convolving one channel of the input feature map, so the output feature map also has 32 channels. After data processing by the batch normalization layer and the RuLe6 layer, the size of the feature map remains unchanged and is input to the first layer of point convolution. The calculation of the first layer of point convolution is very similar to that of ordinary convolution, except that the size of its convolution kernel is (1, 1, 32, 16), where 32 is the number of output channels of the previous layer. The point convolution will perform weighted combination on the output feature map of the depth convolution in the depth direction to generate a feature map with a size of (112, 112, 16). The implementation of the point convolution part reuses the ordinary convolution part. The design of the pooling part includes max pooling and average pooling. According to the network requirements, configure the value of the pooling register as 1, that is, select max pooling. The input of the fully connected part is 1280, and the output is the probabilities of 1000 categories, which is implemented by reusing the ordinary convolution part. Finally, return the classification result to the DDR.

[0117] The above embodiments have fully illustrated the necessary technical content of the present invention. Those of ordinary skill in the art can implement it based on the description, so other technical details will not be elaborated.

[0118] In summary, in this embodiment, a general neural network forward inference accelerator is first designed in the FPGA, including a general convolution module, a depthwise separable convolution module, a fully connected module, a pooling module, a batch normalization module, and an activation function module. And methods such as convolutional block division, parallel convolution calculation, cache setting and optimization, data flow optimization, and pipelining are adopted to accelerate the neural network forward inference process. Then, the neural network accelerator is combined with the processor part to build a complete hardware system for neural network inference. Finally, a specific neural network model is trained on the server, the registers inside the accelerator are configured according to the scale of each layer of the neural network, the model parameters obtained from training are transmitted to the accelerator, and different modules in the accelerator are called respectively according to the network structure, realizing the forward inference of the entire network.

[0119] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to perform the steps of any of the above methods.

[0120] In yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to perform the steps of any of the above methods.

[0121] In yet another embodiment provided by the present application, a computer program product containing instructions is also provided, which when running on a computer causes the computer to perform the steps of any of the above embodiments.

[0122] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples, and beneficial effects of related content, reference can be made to the corresponding parts in the above methods.

[0123] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0124] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0125] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A neural network inference acceleration method based on a heterogeneous platform. The heterogeneous platform includes a processor and an FPGA on-chip system. Characterized in that, It includes the following steps: Design a neural network accelerator in the FPGA. The neural network accelerator includes a computing module and a data processing module to complete the neural network convolution calculation and the processing of output data. Then, methods such as cache setting and optimization, pipelining, and data flow optimization are used to accelerate the neural network forward inference process. Among them, the computing module includes a general convolution module, a depthwise separable convolution module, and a fully connected module; the data processing module includes a pooling module, a batch normalization module, and an activation function module. The general convolution module performs block-based convolution calculation. Each time, only the convolution calculation between the input of a fixed block size and the convolution kernel weights is performed, and then the calculation of pixel points on all output feature maps is completed in turn through the sliding window method. The depthwise separable convolution module includes a depthwise convolution module and a pointwise convolution module. A convolution kernel of the depthwise convolution has only one channel, and one convolution kernel is only responsible for one channel of the input feature map, and the number of channels of the generated output feature map is the same as the number of input channels. The convolution kernel size of the pointwise convolution module is 1x1xD, where D is the number of channels of the output of the previous layer of convolution. The fully connected module reuses the general convolution part, and sets the input scale to 1x1xC through the AXI_Lite bus, where C is the number of channels. The pooling module includes max pooling and average pooling, which are selected by configuring the values of the corresponding registers. Max pooling is only a logical operation, and the input data is compared by a comparator to output a maximum value. Average pooling uses an adder to sum the input data and then uses a shift register to implement division calculation to obtain the average value. The batch normalization module performs data normalization processing on the outputs of the general convolution module and the depthwise separable convolution module. The method of using cache setting and optimization, pipelining, and data flow optimization to accelerate the neural network forward inference process specifically includes: Cache setting and optimization. In the on-chip BRAM memory in the FPGA, input cache: IN[Tn][Tic][Tir], weight cache: W[Tm][Tn][Tkc][Tkr], and output cache: OUT[Tm][Toc][Tor] are respectively set. The size of the cache is determined according to the size of the convolution block variable. The cache splitting method is used to split the channel dimension of the input and output caches, and split the two dimensions of the input and output channels of the weight cache, so that they are distributed in different BRAM blocks, increasing the number of its input and output ports, and performing read and write operations at the same time. The input cache is divided into Tn independent cache blocks, the weight cache is divided into Tm*Tn independent cache blocks, and the output cache is divided into Tm independent cache blocks. Data flow optimization is to perform parallel optimization of task-level data flow using the double-buffer + ping-pong operation method, that is, two input caches, weight caches, and output caches of the same size are set in the on-chip BRAM of the FPGA, and the "ping-pong" data transfer mechanism is used to perform data reading, convolution calculation, and result writing back simultaneously; The ordinary convolution module performs block-based convolution calculation, only performing convolution calculation between the input and convolution kernel weights of a fixed block size each time, and then sequentially completing the calculation of all pixel points on the output feature map through the sliding window method. Specifically, it includes: First, adopt the convolution block strategy, and complete the calculation of the entire convolution by time-sharing multiplexing the convolution block unit. The block variables are the number of channels of the output block: Tm, width: Toc, height: Tor, the width of the convolution kernel block: Tkc, height: Tkr, the number of channels of the input block: Tn, width: Tic, height: Tir. The input block variables and output block variables satisfy the following relationship: Tir = (Tor - 1) * S + Tkr, Tic = (Toc - 1) * S + Tkc, where S is the convolution kernel movement step; The block variables are determined by the size of the output feature map of the convolution layer and the size of the convolution kernel, and are restricted by on-chip resources; The convolution calculation consists of a nested loop structure in four dimensions: inside the output feature map, output channels, input channels, and inside the convolution kernel. And there is parallelism in each dimension. By adjusting the loop order of the nested loops inside the convolution block, the loops of the output channels and input channels are placed in the innermost layer and parallel convolution calculation is performed on these two dimensions, so that the outer loop in the convolution block continuously reuses the internal parallel calculation cores to complete the calculation of the entire convolution block unit.

2. The neural network inference acceleration method based on a heterogeneous platform according to claim 1, characterized in that: It further includes the training of the neural network, specifically by building a neural network model on the server side, importing a data set for training, and obtaining the model parameters of each layer of the neural network after training, including the weights and bias parameters of the convolution layer, batch normalization parameters, and saving the parameters as a binary file and putting it into the SD card.

3. The neural network inference acceleration method based on a heterogeneous platform according to claim 1, characterized in that: The forward inference process of the neural network includes: Allocate a continuous physical address array space ARRAY_IMAGE in the DDR memory, and read the input picture and preprocess it and put it into this array space; Allocate a continuous physical address array space ARRAYi in the DDR memory, and read the neural network convolution layer parameters and batch normalization layer parameters into this array space; Load the entire accelerator into the FPGA in the form of a binary bitstream file; According to the structure of the neural network model, configure the registers of each module inside the FPGA accelerator, and adjust the scale of convolution, pooling, and fully connected operations; According to the structure of the neural network model, the general convolution module, depthwise separable convolution module, fully connected module, pooling module, batch normalization module, and activation function module in the FPGA accelerator are called respectively, and the input data of the memory DDR, convolution layer parameters, and batch normalization layer parameters are passed into the FPGA accelerator for accelerated calculation; After the FPGA accelerator completes the inference calculation of all levels of the neural network, the inference result is returned to the memory DDR and provided for the application program to access.

4. The neural network inference acceleration method based on a heterogeneous platform according to claim 3, characterized in that: Parallel computing includes: Tm convolutional computing engines PE, and each PE consists of Tn multipliers, an adder tree with a depth of , and an adder. It performs the multiplication operation of Tn input feature image pixels and Tm*Tn convolutional kernel weights in parallel. Then, the results are accumulated through the adder tree, and finally added to the previously saved partial sum.

5. The neural network inference acceleration method based on a heterogeneous platform according to claim 1, characterized in that: The design of the batch normalization module includes: Let the four parameters of the output channels of the convolution module be u, σ, γ, and β respectively, where u is the mean value; σ is the variance; γ is the scaling coefficient, and β is the offset coefficient; the input data of the batch normalization part is x, and the output data is: where ∈ is a hyperparameter; The design of the batch normalization module is to read the corresponding parameters u, σ, γ, and β in the memory DDR. Each output channel corresponds to a set of parameters. According to formula (1), specific arithmetic circuit modules are built using the hardware resources in the FPGA, and the input is the result output after convolution calculation.

6. A computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the method according to any one of claims 1 to 5.