A lightweight convolutional neural network acceleration system and acceleration method

By adopting high-parallel calculation strategies and dedicated module design in the lightweight convolutional neural network acceleration system, the acceleration processing of different types of convolutional calculations is performed, and intermediate results are stored on chip, which solves the problems of insufficient resource utilization and low computing efficiency in the existing technology, and realizes efficient convolutional neural network calculation.

CN116882454BActive Publication Date: 2025-05-13XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310713692.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2025-05-13
Estimated Expiration
2043-06-15

AI Technical Summary

Technical Problem

When designing convolutional neural network hardware, the resource utilization is insufficient, the computing efficiency is low, and there is no special computing engine design for the structural characteristics of lightweight convolutional neural networks, which leads to huge challenges in the actual deployment on edge devices.

Method used

It provides a lightweight convolutional neural network acceleration system, including control module, parameter storage module, acceleration fill zero-filling module, dedicated convolution calculation module, post-processing module, special on-chip storage module and result outgoing module. It adopts high-parallel in-layer parallel computing strategy and inter-layer parallel computing strategy to perform special acceleration processing for different types of convolutional computing, and store intermediate results on chip to reduce data transmission.

Benefits of technology

By accelerating the fill-up module to reduce data preprocessing time, dedicated convolutional computing modules improve computing throughput, dedicated on-chip storage modules reduce power consumption, improve computing efficiency, reduce resource waste and idle time, and improve the overall computing speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116882454B_ABST
    Figure CN116882454B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight convolutional neural network acceleration system and acceleration method, the system includes a control module, a parameter storage module, an acceleration filling and zero-padding module, a dedicated convolution calculation module, a dedicated on-chip storage module, a post-processing module and a result output module. The system eliminates useless windows through the acceleration filling and zero-padding module, reducing the data preprocessing time; the dedicated convolution calculation module specifically calculates different types of convolutions, saving and improving resource utilization; the dedicated on-chip storage module stores intermediate results, reducing data transmission with the outside of the chip, and reducing the power consumption caused by data transmission; at the same time, the parallelism of the calculation within the convolution layer is improved by cooperating with high-parallelism intra-layer parallel calculation, accelerating the calculation efficiency, and calculating different types of convolution layers in parallel through the inter-layer parallel calculation strategy, reducing resource waste and idle time, and further improving the calculation speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of chip design, and in particular relates to a lightweight convolutional neural network acceleration system and acceleration method. Background Art

[0002] With the introduction of deep learning theory and the improvement of numerical computing equipment, convolutional neural networks have developed rapidly and have been applied to fields such as computer vision and natural language processing, and their application scenarios are still expanding. However, as the number of network layers increases, the amount of computation and the number of parameters also increase. In actual application deployment, cloud computing is currently used to complete the calculation of convolutional neural networks. However, for edge application scenarios such as unmanned driving that have high real-time requirements, the delay and stability brought by cloud computing are unacceptable. The huge amount of computation and the number of parameters pose a huge challenge to the actual deployment of convolutional neural networks on edge devices.

[0003] At present, many studies have been conducted on the hardware design and implementation of convolutional neural networks at home and abroad. However, most of the current work only designs a separate general-purpose convolution engine to complete all convolution calculations. When calculating convolution types with small parameters and calculation amount, such as deep convolution, it will cause a lot of resource waste and idle time. In addition, there is no dedicated computing engine design and appropriate parallel computing strategy designed specifically for the structural characteristics of lightweight convolutional neural networks, resulting in insufficient utilization of hardware resources and low computing efficiency. In addition, there is no suitable on-chip storage designed for the inverse residual structure, and the intermediate results of the calculation are transmitted to the outside of the chip, resulting in higher power consumption. Or the storage structure is designed unreasonably, resulting in a waste of on-chip storage resources. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a lightweight convolutional neural network acceleration system, namely, an acceleration method. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0005] In a first aspect, the present invention provides a lightweight convolutional neural network acceleration system, comprising a control module, a parameter storage module, an acceleration filling and zeroing module, a dedicated convolution calculation module, a post-processing module, a dedicated on-chip storage module, and a result transmission module; wherein,

[0006] The control module is used to control the flow of data in the entire system and the calling of each module;

[0007] The parameter storage module is used to store the transmitted weight data and bias, and is also used to store the input feature data required for the calculation of the first convolution layer;

[0008] The accelerated zero-filling module is used to rearrange the transmitted input data and perform zero-filling operation based on the improved Line-buffer technology;

[0009] The dedicated convolution calculation module is used to perform special convolution acceleration calculation on input data of different convolution types, and transmit the calculation results to the post-processing module according to the configuration of the control module;

[0010] The post-processing module is used to implement function activation, requantization and inverse residual connection operations on the data transmitted by the dedicated convolution calculation module; and transmit the obtained calculation results to the dedicated on-chip storage module as input data of the subsequent convolution layer, or to other calculation modules according to the configuration of the control module;

[0011] The dedicated on-chip storage module is used to store the input feature data required for the calculation of all layers except the first convolution layer, and is also used to store the intermediate data passed in by the post-processing module;

[0012] The other computing modules are used to perform pooling calculations and full-connection calculations on the output data of the post-processing module, and transmit the calculation results to off-chip storage through the result output module after the calculations are completed.

[0013] In a second aspect, the present invention provides a lightweight convolutional neural network acceleration method, which is applied to the acceleration system provided in the above embodiment, and includes the following steps:

[0014] Step 1: The control module receives the command on the transmission bus and the layer number information of the network layer, and transmits the enable signal in sequence and schedules other modules based on the layer number information and the command, and configures the transmission order of the data;

[0015] Step 2: The parameter storage module receives the weight data related to the current layer calculation and the input feature data of the first convolutional layer transmitted from outside the chip;

[0016] Step 3: The accelerated zero-filling module rearranges the transmitted input data based on the improved Line-buffer technology and performs zero-filling operations to form an input calculation window;

[0017] Step 4: The dedicated convolution calculation module first receives the weight data input by the parameter storage module and the input calculation window input by the acceleration filling zero module, and then schedules different convolution calculation units to perform convolution calculation according to the enable signal transmitted by the control module, and transmits the data to the post-processing module after the calculation is completed;

[0018] Step 5: After the post-processing module processes the incoming data, it selects to transfer the obtained calculation results to the dedicated on-chip storage module as the input data of the subsequent convolutional layer under the enablement of the control module; or transfer them to other calculation modules;

[0019] Step 6: The control module enables other computing modules to perform pooling calculations and fully connected calculations on the input data, and transmits the calculation results to the off-chip storage through the calculation result output module based on the bus protocol.

[0020] Beneficial effects of the present invention:

[0021] 1. The lightweight convolutional neural network acceleration system provided by the present invention eliminates useless windows and reduces data preprocessing time by accelerating the filling and zero-padding module; uses a dedicated convolution calculation module to specifically calculate different types of convolution calculations, reduces useless calculation time, improves calculation throughput, and improves the overall calculation efficiency of the system; uses a dedicated on-chip storage module to store intermediate results, reduces data transmission with the outside of the chip, reduces power consumption caused by data transmission, and speeds up calculation efficiency;

[0022] 2. The lightweight convolutional neural network acceleration system provided by the present invention adopts a high-parallel intra-layer parallel computing strategy to improve the parallelism of calculations within the convolution layer, and at the same time cooperates with the inter-layer parallel computing strategy to parallelly calculate different types of convolution layers, reducing resource waste and idle time, and further improving the calculation speed.

[0023] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a structural schematic diagram of a lightweight convolutional neural network acceleration system provided by an embodiment of the present invention;

[0025] Figure 2 It is the existing ordinary filling and zero-filling effect diagram;

[0026] Figure 3 This is a schematic diagram of an invalid window and a valid window formed by existing common zero padding;

[0027] Figure 4 This is a diagram showing the effect of accelerated zero filling provided by an embodiment of the present invention;

[0028] Figure 5 Schematic diagram of an invalid window and a valid window formed by accelerating zero filling provided by an embodiment of the present invention;

[0029] Figure 6 It is a schematic diagram of the parallel computing strategy within the inner layer of the convolution kernel and the input and output channels provided by an embodiment of the present invention;

[0030] Figure 7 Schematic diagram of a parallel strategy for point-by-point convolution calculation input data provided by an embodiment of the present invention;

[0031] Figure 8 is a schematic diagram of an inter-layer parallel strategy provided by an embodiment of the present invention;

[0032] Fig. 9 It is a schematic diagram of the internal structure and related data flow of a dedicated on-chip storage module provided by an embodiment of the present invention;

[0033] Fig.10 It is a flow chart of a lightweight convolutional neural network acceleration method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0035] Embodiment 1

[0036] See also Figure 1 , Figure 1 1 is a schematic diagram of the structure of a lightweight convolutional neural network acceleration system provided by an embodiment of the present invention, the system includes a control module, a parameter storage module, an acceleration filling and zero-padding module, a dedicated convolution calculation module, a post-processing module, a dedicated on-chip storage module, and a result transmission module; wherein,

[0037] The control module is used to control the flow of data in the entire system and the calling of each module;

[0038] The parameter storage module is used to store the transmitted weight data and bias, and is also used to store the input feature data required for the calculation of the first convolutional layer;

[0039] The accelerated zero-filling module is used to rearrange the transmitted input data and perform zero-filling operations based on the improved Line-buffer technology;

[0040] The dedicated convolution calculation module is used to perform special convolution acceleration calculations on input data of different convolution types, and transmit the calculation results to the post-processing module according to the configuration of the control module;

[0041] The post-processing module is used to implement function activation, requantization and inverse residual connection operations on the data passed in by the dedicated convolution calculation module; and according to the configuration of the control module, the obtained calculation results are passed to the dedicated on-chip storage module as input data of the subsequent convolution layer, or passed to other calculation modules;

[0042] The dedicated on-chip storage module is used to store the input feature data required for the calculation of all layers except the first convolutional layer, and is also used to store the intermediate data passed in by the post-processing module;

[0043] Other computing modules are used to perform pooling and full-connection calculations on the output data of the post-processing module, and after the calculation is completed, the calculation results are transmitted to off-chip storage through the result output module.

[0044] Specifically, under the scheduling of the control module, the input data and weights to be calculated are respectively transmitted to the accelerated zero-filling module and the parameter storage module, and then the rearranged data is sent to the dedicated convolution calculation module for different convolution types for convolution calculation, and then transmitted to the post-processing module and other calculation modules according to the configuration of the control module. When all the calculations of this lightweight convolutional neural network are completed, it is transmitted out of the chip.

[0045] The following is a detailed introduction to each module.

[0046] Control Module:

[0047] The control module is mainly used to receive the command on the transmission bus and the layer number information of the network layer. According to the layer number information and the command, the control module transmits the enable signal in sequence and schedules other modules, thereby defining the transmission order of the data.

[0048] Parameter storage module:

[0049] The parameter storage module receives the weight data related to the current layer calculation and the input features of the first layer transmitted from the chip. It is composed of ping-pong RAM. While one RAM is writing the transmission weight data, the other RAM is reading the required calculation data by the subsequent calculation module.

[0050] Accelerated filling zero module:

[0051] The input data is rearranged using the improved Line-buffer technology. When zero padding is needed, it is only done at the top and bottom of the data, rather than at the two edges of each line of data, thus reducing invalid data windows.

[0052] See also Figure 2-3 , Figure 2 This is the existing normal filling and zero-filling effect diagram. Figure 3 It is a schematic diagram of invalid windows and valid windows formed by the existing common zero-filling. It can be seen that the common zero-filling module will inevitably generate invalid windows during the conversion process of every two rows, which increases the preprocessing time and wastes computing resources.

[0053] The acceleration effect of the accelerated filling zero-padding module provided in this embodiment is as follows: Figure 4-5 As shown, Figure 4 This is a diagram showing the effect of accelerated zero filling provided by an embodiment of the present invention. Figure 5Schematic diagram of invalid windows and valid windows formed by the accelerated zero-filling provided by an embodiment of the present invention. It can be seen that the accelerated zero-filling module of this embodiment is mainly reflected in the fact that no zero padding is required at the beginning and end of each row compared with the ordinary zero-filling module. When the output window moves to the end of each row and is about to switch to the next row to generate an unnecessary error window, it is set to zero and converted into the required zero-filling window. The accelerated zero-filling module avoids the invalid windows caused by the zero-filling at the beginning and end, and only needs to fill zeros at the top and the tail. Taking the m×n input data with a filling number of 1 as an example, due to the need to fill zeros at the boundary, the ordinary zero-filling module needs to circulate (m+2)×(n+2) times on each input channel, while the accelerated zero-filling module only circulates (m+2)×n times on each input channel.

[0054] This embodiment eliminates useless windows and reduces data preprocessing time by accelerating the filling of the zero-padding module.

[0055] Dedicated convolution computation module:

[0056] Please continue to see Figure 1 , wherein the dedicated convolution calculation module includes a common convolution calculation unit, a depth convolution calculation unit and a point-by-point convolution calculation unit;

[0057] When enabled by the control module, the ordinary convolution calculation unit, the depth convolution calculation unit and the point-by-point convolution calculation unit are used to perform special convolution calculation acceleration on the data after rearranging and zero-padding in the accelerated filling and zero-padding module and the weight data in the parameter storage module based on the three different convolution types in the network.

[0058] In addition, the dedicated convolution calculation module also includes a universal adder tree unit for performing accumulation calculation on the data output by different convolution calculation units.

[0059] Specifically, the dedicated convolution calculation module first receives the weight data input by the parameter storage module and the input calculation window input by the acceleration filling and zero-padding module, and then schedules different convolution calculation modules to perform related convolution calculations according to the enable signal transmitted by the control module.

[0060] After each convolution calculation module completes the relevant convolution calculation, the universal adder tree module completes the parallel accumulation within the convolution kernel or the accumulation part at the input and output channel level. The universal adder tree is completed by an adder tree that can complete 8 or 9 inputs at the same time. According to different convolution types, relevant configurations can be performed to be compatible with the accumulation and summation types required by various convolutions. For the parallel accumulation and summation within the convolution kernel, the adder tree is configured to be composed of 9-input adder trees. For channel-level accumulation and summation, the adder tree is configured to be composed of 8-input adder trees. For various types of convolution, it is only necessary to perform relevant configuration and scheduling according to the parallel accumulation within the convolution kernel or at the channel level to meet the relevant needs.

[0061] Furthermore, the dedicated convolution calculation module cooperates with the high-parallelism intra-layer parallel calculation strategy and inter-layer parallel calculation strategy during convolution calculation to achieve and accelerate the convolution calculation that accounts for the main part of the entire convolutional neural network operation. Among them, the intra-layer parallel calculation strategy is applied to the convolution kernel, between input and output channels, and between input data.

[0062] Specifically, in addition to the common convolution kernel intra-parallelism and input and output channel parallelism, the high-parallelism intra-layer parallel computing strategy also adopts parallelism between input data. Since point-by-point convolution is a 1×1 convolution, there is no computational correlation between adjacent input data. Therefore, multiple input data can be calculated in parallel with the 1*1 convolution kernel at the same time to improve the parallelism. The details are as follows:

[0063] 1. Parallelism within the convolution kernel: In convolution calculation, the parallel calculation of k*k data within the convolution kernel is performed by the convolution kernel of size k*k and the k*k data in the input data in parallel. This is the parallelism within the convolution kernel of size k*k. The process is as follows: Figure 6 shown.

[0064] 2. Parallelism between input and output channels: k*k m-dimensional input data and n k*k m-dimensional convolution kernels are calculated in parallel, and finally the corresponding data on n output channels are obtained after summing up. This is the parallelism between input channels with a parallelism of m and the parallelism between output channels with a parallelism of n. The process is as follows: Figure 6 shown.

[0065] 3. Parallelism between input data: For point-by-point convolution, v m-dimensional data on the input data are convolved with n 1*1 m-dimensional convolution kernels. Since these v m-dimensional data have no correlation with each other in point-by-point convolution, after convolution and accumulation, v n-dimensional output data can be obtained. This is parallelism between input data with a parallelism of v. The process is as follows: Figure 7 shown.

[0066] Furthermore, in lightweight convolutional neural networks, the computational complexity and parameter count of deep convolution and traditional convolution are greatly reduced compared to point-by-point convolution, so the workload between them is not balanced, and an inter-layer parallel computing strategy is adopted.

[0067] Specifically, inter-layer parallelism is mainly the parallel calculation between depthwise convolution and pointwise convolution. Since the output of the previous convolution layer can be used as the input of the next layer, pipeline technology is used to parallelize the calculation of different adjacent convolution layers. The two layers can be fused into one layer in the order of depthwise convolution and pointwise convolution as needed. The three layers can also be fused into one layer in the order of pointwise convolution, depthwise convolution, and pointwise convolution.

[0068] See also Figure 8 , Figure 8 It is a schematic diagram of the inter-layer parallel strategy provided by an embodiment of the present invention. Specifically, different adjacent convolutional layers are calculated in parallel using pipeline technology, and the output of the previous layer is used as the input of the next layer. In view of the depth-separable convolution characteristics of lightweight convolutional neural networks, inter-layer parallelism can fuse two layers into one layer in the order of depth convolution and point-by-point convolution. These three layers can also be fused into one layer in the order of point-by-point convolution, depth convolution, and point-by-point convolution. It is worth noting that even with such division, due to the interdependence of input and output data of different convolutional layers, inter-layer parallelism will only calculate one layer of depth convolution and one layer of point-by-point convolution at the same time.

[0069] The present invention adopts a high-parallelism intra-layer parallel computing strategy to improve the parallelism of calculations within the convolution layer, and at the same time cooperates with the inter-layer parallel computing strategy to parallelly calculate different types of convolution layers, reducing resource waste and idle time, and further improving the calculation speed.

[0070] Post-processing module:

[0071] Please continue to see Figure 1 , the post-processing module includes an activation function unit, a requantization unit and a residual connection unit; wherein,

[0072] The activation function unit mainly implements the Relu activation function;

[0073] The requantization unit remaps the accumulated 32-bit result to 8 bits;

[0074] The residual connection unit implements the residual connection of the inverse residual structure.

[0075] Specifically, the activation function mainly implements the Relu activation function. Requantization implements remapping of the accumulated 32-bit result to 8 bits. The inverse residual connection step reads the initial data entering the inverse residual structure from the dedicated on-chip storage module of the inverse residual structure, waits for the entire inverse residual structure to be calculated, and then, according to the instructions, accumulates the calculated output data with the input initial data. At the same time, according to the idea of ​​requantization, in conjunction with the quantization coefficient and the number of shifts, ensure that the accumulated data is still within the 8-bit range. According to the enable signal of the control module and the number of layers transmitted, choose to store the calculated results in the dedicated on-chip storage module as the input data of the subsequent convolutional layer; or transmit the calculation results to other calculation modules to complete the final number of layers. Calculation calculation.

[0076] Dedicated on-chip memory modules:

[0077] In this embodiment, a plurality of true dual-port RAMs or pseudo dual-port RAMs are provided inside the dedicated on-chip storage module, and are configured into different areas for storing calculation results of different convolutional layers respectively.

[0078] Specifically, the dedicated on-chip storage module cooperates with the inter-layer parallel strategy as a dedicated on-chip storage module for the inverse residual structure, and its internal area is divided into three areas to store the calculation results of the three groups of convolutional layers of the inverse residual structure and complete the residual accumulation. Fig. 9 , Fig. 9 It is a schematic diagram of the internal structure and related data flow of a dedicated on-chip storage module provided by an embodiment of the present invention; wherein, the input data of the point-by-point convolution layer 1 is read from storage unit 1, and the calculation result is stored in storage unit 2. The input data of the depth convolution layer is read from storage unit 2, and the calculation result is stored in storage unit 3. The input data of the point-by-point convolution layer 2 is read from storage unit 3. When a residual connection is required, the initial input data is read from storage unit 1, and the residual connection with the current calculation result is stored in storage unit 1. When a residual connection is not required, it is directly stored in storage unit 1. In conjunction with the inter-layer parallel computing strategy, since only one point-by-point convolution layer and the depth convolution layer are being calculated in parallel at the same time. Therefore, only storage unit 2 or storage unit 3 will be read and written at the same time. This meets the simultaneous reading and writing requirements of dual-port RAM.

[0079] This embodiment uses a dedicated on-chip storage module to store intermediate results on-chip, which reduces data transmission outside the chip, reduces power consumption caused by data transmission, and improves computing efficiency.

[0080] Other computing modules:

[0081] like Figure 1As shown in the figure, other computing modules include pooling units and fully connected units. In lightweight convolutional neural networks, pooling units and fully connected units generally appear as the last few layers of the entire neural network. Therefore, when the control module enables other computing modules, the output results of their calculations can be directly passed to the result output module and transmitted to the off-chip storage as the final output results, thereby realizing hardware acceleration of lightweight convolutional neural networks.

[0082] The lightweight convolutional neural network acceleration system provided by the present invention first reads the specific parameters of the network, sends instructions and selects to enable related calculation modules through the control module, and reads the weights and input features in the parameter storage module and the dedicated on-chip storage module of the inverse residual structure. The parameter storage module achieves read-write separation through the ping-pong structure, and uses the calculation time to cover the waiting time of data transmission. The dedicated on-chip storage module of the inverse residual structure only stores all input and output data of the currently calculated inverse residual structure and related intermediate calculation data, avoiding the input data of the first layer of the inverse residual structure from being overwritten while reducing the on-chip storage pressure.

[0083] After that, the weights and input features are passed to the accelerated zero-filling module, the dedicated convolution calculation module and the general adder tree module in sequence. The data is rearranged, the convolution calculation is performed, and the cumulative sum is performed to obtain the final output data. The accelerated zero-filling module improves the computing efficiency by reducing the waiting time for data transmission and the calculation time of the invalid zero-filling window. The dedicated convolution calculation module can perform specialized convolution calculations for different convolution types, making full use of computing resources to speed up the calculation speed and improve the throughput. The general adder tree module saves more computing resources by calling and configuring the same module.

[0084] Furthermore, the parallelism of the calculation is improved by coordinating the intra-layer parallel strategy with high parallelism, and the computing resources are further fully utilized to speed up the calculation and improve the computing efficiency. In addition, it is understandable that if the inter-layer calculations are only performed sequentially between different convolutional layer calculations, a large amount of computing resources will be idle and wasted when calculating the depth convolution. Therefore, the inter-layer parallel computing strategy is adopted to make full use of computing resources and improve efficiency. It is conducive to realizing pipeline scheduling and more efficient parallel processing of calculations between different types of convolutional layers.

[0085] Embodiment 2

[0086] Based on the above-mentioned embodiment 1, this embodiment provides a lightweight convolutional neural network acceleration method. Fig.10 , the method comprises the following steps:

[0087] Step 1: The control module receives the command on the transmission bus and the layer number information of the network layer, and transmits the enable signal in sequence and schedules other modules based on the layer number information and the command, and configures the transmission order of the data;

[0088] Step 2: The parameter storage module receives the weight data related to the current layer calculation and the input feature data of the first convolutional layer transmitted from outside the chip;

[0089] Step 3: The accelerated zero-filling module rearranges the transmitted input data based on the improved Line-buffer technology and performs zero-filling operations to form an input calculation window;

[0090] Step 4: The dedicated convolution calculation module first receives the weight data input by the parameter storage module and the input calculation window input by the acceleration filling zero module, and then schedules different convolution calculation units to perform convolution calculation according to the enable signal transmitted by the control module, and transmits the data to the post-processing module after the calculation is completed;

[0091] Step 5: After the post-processing module processes the incoming data, it selects to transfer the obtained calculation results to the dedicated on-chip storage module as the input data of the subsequent convolutional layer under the enablement of the control module; or transfer them to other calculation modules;

[0092] Step 6: The control module enables other computing modules to perform pooling calculations and fully connected calculations on the input data, and transmits the calculation results to the off-chip storage through the calculation result output module based on the bus protocol.

[0093] Specifically, in step 4, when performing different convolution calculations, the dedicated convolution calculation module adopts a high degree of parallelism intra-layer parallel calculation strategy and an inter-layer parallel calculation strategy to accelerate the convolution calculation.

[0094] The method provided in this embodiment can be applied to the system provided in the above embodiment 1, and the specific process can refer to the description of the above embodiment 1. Therefore, the method can also reduce the power consumption of the lightweight convolutional neural network, reduce resource waste, and improve the calculation speed.

[0095] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A lightweight convolutional neural network acceleration system, characterized in that: It includes a control module, a parameter storage module, an accelerated zero-filling module, a dedicated convolution calculation module, a post-processing module, a dedicated on-chip storage module, and a result output module; wherein, The control module is used to control the flow of data in the entire system and the calling of each module; The parameter storage module is used to store the transmitted weight data and bias, and is also used to store the input feature data required for the calculation of the first convolution layer; The accelerated zero-filling module is used to rearrange the input data transmitted and perform zero-filling operation based on the improved Line-buffer technology; wherein, when the accelerated zero-filling module performs zero-filling operation on the input data, it only fills zero at the top and bottom of the whole; The dedicated convolution calculation module is used to perform special convolution acceleration calculation on input data of different convolution types, and transmit the calculation results to the post-processing module according to the configuration of the control module; The post-processing module is used to implement function activation, requantization and inverse residual connection operations on the data transmitted by the dedicated convolution calculation module; and transmit the obtained calculation results to the dedicated on-chip storage module as input data of the subsequent convolution layer, or to other calculation modules according to the configuration of the control module; The dedicated on-chip storage module is used to store the input feature data required for the calculation of all layers except the first convolution layer, and is also used to store the intermediate data passed in by the post-processing module; The other computing modules are used to perform pooling calculations and full-connection calculations on the output data of the post-processing module, and transmit the calculation results to off-chip storage through the result output module after the calculations are completed.

2. A lightweight convolutional neural network acceleration system according to claim 1, characterized in that: The dedicated convolution calculation module includes a common convolution calculation unit, a depth convolution calculation unit and a point-by-point convolution calculation unit; When enabled by the control module, the ordinary convolution calculation unit, the depth convolution calculation unit and the point-by-point convolution calculation unit respectively perform special convolution calculation acceleration on the data rearranged and zero-filled by the accelerated filling and zero-filling module and the weight data in the parameter storage module based on the three different convolution types in the network.

3. A lightweight convolutional neural network acceleration system according to claim 2, characterized in that: The dedicated convolution calculation module also includes a universal adder tree unit; the universal adder tree unit is used to perform accumulation calculation on data output by different convolution calculation units.

4. A lightweight convolutional neural network acceleration system according to claim 2, characterized in that: The dedicated convolution calculation module uses a high degree of parallelism intra-layer parallel calculation strategy and an inter-layer parallel calculation strategy to process data; The intra-layer parallel computing strategy is applied within the convolution kernel, between input and output channels, and between input data.

5. The lightweight convolutional neural network acceleration system according to claim 1, characterized in that: The dedicated on-chip storage module is internally provided with a plurality of true dual-port RAMs or pseudo dual-port RAMs.

6. A lightweight convolutional neural network acceleration system according to claim 5, characterized in that: The dedicated on-chip storage module is internally configured into different areas for storing calculation results of different convolutional layers.

7. The lightweight convolutional neural network acceleration system according to claim 1, characterized in that: The post-processing module includes an activation function unit, a requantization unit and a residual connection unit; wherein, The activation function unit mainly implements the Relu activation function; The requantization unit is used to remap the accumulated 32-bit result to 8 bits; The residual connection unit is used to implement the residual connection of the inverse residual structure.

8. A lightweight convolutional neural network acceleration method, characterized in that: The acceleration system according to any one of claims 1 to 7 comprises the following steps: Step 1: The control module receives the command on the transmission bus and the layer number information of the network layer, and transmits the enable signal in sequence and schedules other modules based on the layer number information and the command, and configures the transmission order of the data; Step 2: The parameter storage module receives the weight data related to the current layer calculation and the input feature data of the first convolutional layer transmitted from outside the chip; Step 3: The accelerated zero-filling module rearranges the input data transmitted and performs zero-filling operation based on the improved Line-buffer technology to form an input calculation window; wherein, when the accelerated zero-filling module performs zero-filling operation on the input data, it only fills zero at the top and bottom of the whole; Step 4: The dedicated convolution calculation module first receives the weight data input by the parameter storage module and the input calculation window input by the acceleration filling zero module, and then schedules different convolution calculation units to perform convolution calculation according to the enable signal transmitted by the control module, and transmits the data to the post-processing module after the calculation is completed; Step 5: After the post-processing module processes the incoming data, it selects to transfer the obtained calculation results to the dedicated on-chip storage module as the input data of the subsequent convolution layer under the enablement of the control module; or transfer them to other calculation modules; Step 6: The control module enables other computing modules to perform pooling calculations and fully connected calculations on the input data, and transmits the calculation results to the off-chip storage through the calculation result output module based on the bus protocol.

9. A lightweight convolutional neural network acceleration method according to claim 8, characterized in that: In step 4, when performing different convolution calculations, the dedicated convolution calculation module adopts a high degree of parallelism intra-layer parallel calculation strategy and an inter-layer parallel calculation strategy to accelerate the convolution calculation.

Citation Information

Patent Citations

  • Audio data processing method and device, equipment and storage medium

    CN113377331A

  • Integrated machine olfaction chip using convolutional neural network for calculation

    CN114113491A