Convolutional Neural Network Deployment Method, Device, Computer Equipment and Storage Medium
By modifying the instruction set and data structure, using PCIe interface and DDR memory, rapid deployment and structural modification of convolutional neural networks in different structures are achieved, solving the problem of insufficient universality of traditional deployment methods and improving hardware resource utilization efficiency.
Patent Information
- Application Number
- CN202111290658.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-02
AI Technical Summary
Traditional neural network deployment methods have shortcomings in terms of universality, and it is difficult to quickly deploy convolutional neural networks of different structures or modify the deployed networks.
By modifying the instruction set and data structure, using the PCIe interface to transmit model parameters and instructions, using DDR memory to read and write data, and performing convolution operations through the convolution calculation module, the rapid deployment and structural modification of convolution neural networks in different structures are achieved.
Without hardware knowledge, it is possible to quickly deploy convolutional neural networks of different structures at different times, reducing hardware resource consumption, and supporting the calculation of convolutions of almost any feature map size, number of layers, kernel size, and convolution step size.
Smart Images

Figure CN114021711B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of mobile deployment in artificial intelligence, and particularly relates to a convolutional neural network deployment method, device, computer device, and storage medium applied to a mobile terminal. Background Art
[0002] Currently, in the AI technology for image processing, such as tasks like object recognition or object detection, the performance of convolutional neural networks is generally superior to other types of neural networks, so convolutional neural networks are used more frequently. The implementation methods of AI technology include CPU, GPU, ASIC, or FPGA. As a general-purpose processor, the CPU not only has to meet the computing requirements but also has to satisfy the applications for responding to human-computer interaction and handle the synchronization and coordination between tasks. Such a hardware structure cannot handle well the convolutional neural network model where the data computing demand is greater than the control instruction demand. The GPU is designed to maximize the computing output, and almost all of its space is given to the ALU, so the computing power advantage of the GPU is very obvious. However, due to its area, power consumption, and energy efficiency ratio, the GPU cannot be deployed on small-sized and low-power mobile terminals. The ASIC chip is customized for the algorithm, without redundancy, with low power consumption and high energy efficiency ratio. However, its development cycle is long, and due to its customization, its internal structure cannot be changed after production, lacking flexibility, which is not conducive to accelerating the currently rapidly developing convolutional neural networks, and in small-scale deployment applications, the cost is huge. Compared with the GPU, the FPGA is smaller in volume, lower in power consumption, and has stronger pipeline parallel and data parallel capabilities. Compared with the CPU, the FPGA has stronger computing power; compared with the GPU, the FPGA is smaller in volume and lower in power consumption; compared with the ASCI, the FPGA has lower costs in small-scale deployment.
[0003] Due to different specific tasks, the structures of convolutional neural networks are also different, for example: convolutional calculation methods, feature map sizes, number of layers, kernel sizes, convolutional strides, activation functions, pooling, etc. Generally, the deployment method of neural networks on the FPGA is mostly to design the hardware structure according to its network structure, and after deployment, it can only accelerate the network of this structure. If you want to accelerate another network with a different structure, you can only re-design the hardware. This process is relatively easy for FPGA designers, but in most cases, the users of this FPGA board are not its designers, and it is very difficult for users to complete the hardware design process. Thus, it can be seen that the traditional neural network deployment method has the problem of weak generality. Summary of the Invention
[0004] The purpose of the embodiments of this application is to propose a convolutional neural network deployment method, device, computer device, and storage medium applied to a mobile terminal to solve the problem of weak generality existing in the traditional neural network deployment method.
[0005] To solve the above technical problems, an embodiment of the present application provides a method for deploying a convolutional neural network applied to a mobile terminal, adopting the following technical solutions:
[0006] Receive a convolutional neural network deployment request carrying the model parameters to be deployed and instructions;
[0007] Transmit the model parameters to be deployed and the instructions to the DDR according to the PCIe interface;
[0008] Read the first instruction from the instruction area of the DDR;
[0009] Read the first weight data from the weight area of the DDR according to the first instruction and the counter, and store it in the Wt_RAM;
[0010] According to the first instruction and the counter, read the input feature map data from the feature Figure 1 area of the DDR and store it in the Fin_RAM;
[0011] Read the bias data from the bias area of the DDR according to the first instruction and the counter, and store it in the Bias_FIFO;
[0012] Perform a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and write the data in the Fout_RAM back to the feature Figure 2 area of the DDR;
[0013] When the convolution calculation operations for all instructions in the instruction area are completed, an inference result is obtained;
[0014] Output the inference result of the DDR according to the PCIe interface.
[0015] To solve the above technical problems, an embodiment of the present application further provides a convolutional neural network deployment device applied to a mobile terminal, adopting the following technical solutions:
[0016] A request receiving module, configured to receive a convolutional neural network deployment request carrying the model parameters to be deployed and instructions;
[0017] A data transmission module, configured to transmit the model parameters to be deployed and the instructions to the DDR according to the PCIe interface;
[0018] An instruction reading module, configured to read the first instruction from the instruction area of the DDR;
[0019] A weight data reading module, configured to read first weight data from a weight area of the DDR according to the first instruction and a counter, and store the data into the Wt_RAM;
[0020] An input feature map reading module, configured to read input feature map data from a feature Figure 1 area of the DDR according to the first instruction and the counter, and store the data into the Fin_RAM;
[0021] A bias data reading module, configured to read bias data from a bias area of the DDR according to the first instruction and the counter, and store the data into the Bias_FIFO;
[0022] A convolution calculation module, configured to perform a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and write the data in the Fout_RAM back to a feature Figure 2 area of the DDR;
[0023] An inference result obtaining module, configured to obtain an inference result after completing the convolution calculation operations for all instructions in the instruction area;
[0024] An output module, configured to output the inference result of the DDR according to the PCIe interface.
[0025] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solutions:
[0026] It includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the above-mentioned convolutional neural network deployment method applied to a mobile terminal are implemented.
[0027] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solutions:
[0028] Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the above-mentioned convolutional neural network deployment method applied to a mobile terminal are implemented.
[0029] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0030] The present application provides a method for deploying a convolutional neural network in a mobile terminal, including: receiving a convolutional neural network deployment request carrying model parameters to be deployed and instructions; transmitting the model parameters to be deployed and the instructions to the DDR according to the PCIe interface; reading the first instruction from the instruction area of the DDR; reading first weight data from the weight area of the DDR according to the first instruction and a counter, and storing it in the Wt_RAM; reading input feature map data from the feature Figure 1 area of the DDR according to the first instruction and the counter, and storing it in the Fin_RAM; reading bias data from the bias area of the DDR according to the first instruction and the counter, and storing it in the Bias_FIFO; performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and writing the data in the Fout_RAM back to the feature Figure 2 area of the DDR; when the convolution calculation operations for all instructions in the instruction area are completed, an inference result is obtained; outputting the inference result of the DDR according to the PCIe interface. The present application can, without any hardware knowledge, quickly deploy convolutional neural networks with different structures at different times by modifying the instruction set and data structure, or modify the structure of a deployed network. Standard convolution and depthwise separable convolution use almost the same hardware resources to complete, which improves generality while greatly reducing the consumption of hardware resources. Through the block and group of convolutions, it is possible to calculate convolutions with almost any feature map size, number of layers, kernel size, and convolution stride on limited hardware resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;
[0033] Figure 2 is a flowchart of the implementation of the method for deploying a convolutional neural network in a mobile terminal provided in Embodiment 1 of the present application;
[0034] Figure 3 is a block diagram of a general convolutional neural network acceleration system provided in Embodiment 1 of the present application;
[0035] Figure 4 is a schematic structural diagram of the Conv sub-module in the general convolutional calculation module provided in Embodiment 1 of the present application;
[0036] Figure 5 It is a schematic structural diagram of the Pooling sub-module in the general convolution calculation module provided in the first embodiment of the present application;
[0037] Figure 6 It is a schematic structural diagram of the convolutional neural network deployment device applied to a mobile terminal provided in the second embodiment of the present application:
[0038] Figure 7 It is a schematic structural diagram of an embodiment of a computer device according to the present application. Detailed implementation manners
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the description of the present application in the specification are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0040] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0041] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0042] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0043] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0044] Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.
[0045] Server 105 can be a server that provides various services, such as a background server that supports the pages displayed on terminal devices 101, 102, and 103.
[0046] It should be noted that the convolutional neural network deployment method applied to a mobile terminal provided in the embodiments of the present application is generally executed by a server / terminal device. Correspondingly, the convolutional neural network deployment device applied to a mobile terminal is generally set in a server / terminal device.
[0047] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0048] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to
[0049] shows the implementation flowchart of the convolutional neural network deployment method applied to a mobile terminal provided in Embodiment 1 of the present application. For the sake of convenience of description, only the parts related to the present application are shown.
[0050]
[0051] In step S101, receive a convolutional neural network deployment request carrying model parameters to be deployed and instructions; In step S102, transmit the model parameters to be deployed and instructions to the DDR according to the PCIe interface;
[0052] In step S103, read the first instruction from the instruction area of the DDR;
[0053] In step S104, according to the first instruction and the counter, read the first weight data from the weight area of the DDR and store it in the Wt_RAM;
[0054] In step S105, according to the first instruction and the counter, read the input feature map data from the Figure 1 feature area of the DDR and store it in the Fin_RAM;
[0055] In step S106, according to the first instruction and the counter, read the bias data from the bias area of the DDR and store it in the Bias_FIFO;
[0056] In step S107, perform a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and write the data in the Fout_RAM back to the Figure 2 feature area of the DDR;
[0057] In step S108, when the convolution calculation operations for all instructions in the instruction area are completed, an inference result is obtained;
[0058] In step S109, output the inference result of the DDR according to the PCIe interface.
[0059] In the embodiment of the present application, the step of performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data specifically includes:
[0060] (1) Determine the address of the output feature map data in the Fout_RAM according to the output feature map size, the counter, and whether pooling is performed;
[0061] (2) Determine the read data address of the Fin_RAM according to the input feature map size, the convolution kernel size, the padding value, and the counter;
[0062] (3) Determine the read data address of the Wt_RAM according to the counter;
[0063] (4) Determine the read data address of the Bias_FIFO according to the counter;
[0064] (5) Determine whether the current convolution calculation operation is completed according to the counter.
[0065] In the embodiment of the present application, refer to Figure 3The block diagram of the general convolutional neural network acceleration system shown stores the input feature map and output feature map of a certain layer, the weights and biases of all layers, and the instruction set for controlling the entire inference process in the DDR. After decoding the instructions, the Top_Ctrl module controls the data interaction between the DDR, RAM, and FIFO according to the decoded instructions. The general convolutional calculation module is the core of the design, and through instructions, it can implement standard convolution or depthwise separable convolution with variable feature map size, number of layers, kernel size, and convolution stride, and the option to enable or disable activation and pooling. Wt_RAM, Fin_RAM, and Fout_RAM use on-chip RAM resources, and Bias_FIFO is implemented using LUT due to its large data bit width. The Top_Ctrl module realizes specific functions through a state machine and determines when to start the entire system through the off-chip signal Start.
[0066] In the embodiment of this application, since the RAM resources on the FPGA are very limited, if the number of parameters of a certain layer is very large and the RAM cannot store them simultaneously, it is difficult to perform calculations. Therefore, the convolution is divided into blocks and groups. The convolution calculation in the convolutional layer of a neural network can be understood as: a three-dimensional array A (input feature map) passes through a four-dimensional array B (weights) to obtain a new three-dimensional array C (output feature map). The decomposition of the convolution calculation can be understood as: taking a small three-dimensional array a from a large three-dimensional array A, taking a small four-dimensional array b from a large four-dimensional array B, a passes through b to obtain c, and multiple c are combined to obtain C. Such a process is the data interaction process between the DDR, RAM, and FIFO, and the specific data decomposition method is reflected in the instructions.
[0067] In the embodiment of this application, refer to Figure 4 The structural schematic diagram of the Conv sub-module in the general convolutional calculation module shown can implement standard convolution calculation and depthwise separable convolution calculation. In addition to the data flow shown in the figure, the data in the register bank comes from Wt_RAM, the data in the multiplier array comes from Fin_RAM, and the data in the accumulation array comes from Bias_FIFO. There are two groups of registers in the register bank. Through ping-pong operation, one register group receives data from the RAM and the other register group transmits data downward. The multiplier array contains 8-bit multipliers and is not a systolic array, that is, the outputs of all multipliers are valid simultaneously. In the addition tree array, the number of levels is determined by the number of multipliers in the multiplier array, and it also includes data truncation and saturation operations. In the accumulation array, the number of accumulations is provided by the instruction. Specifically, it is determined by the block and group situation of the convolution and the kernel size, and it also includes data truncation and saturation operations.
[0068] In the embodiments of the present application, when performing standard convolution, the output result of the multiplier array is obtained as output data of the same type as the input data after passing through the adder tree array and the accumulation array; when performing depth convolution, a part of the multipliers in the multiplier array are disabled, the output result skips the adder tree array, and after passing through the accumulation array, output data of the same type as the input data is obtained; when performing point convolution, it is the same as standard convolution.
[0069] In the embodiments of the present application, Figure 5 The structure diagram of the Pooling sub-module in the general convolution calculation module is shown, and the pooling method is maximum pooling with 2X2 and a stride of 2. The pooling process is divided into row pooling and column pooling, and the result of row pooling needs to be row-cached. Due to the requirement of design generality, the row size of the output feature map is unpredictable during design. Generally, a sufficiently large row cache is required to meet the design generality requirement, but this method will cause a large waste of storage space when the row size is very small. In the present invention, by changing the calculation order in the convolution process, the size of the row cache is fixed, and convolution calculation can still be completed when the row size is several times the row cache.
[0070] In the embodiments of the present application, the output result of the Pooling module is the output feature map data of this layer, and its data structure is the same as that of the input feature map data, and can be directly used as the input of the next layer to participate in the calculation.
[0071] In the embodiments of the present application, the generality of the design is realized by variable instructions controlling invariant hardware. The instructions are mainly used to control the read and write addresses of the input and output data of each level of storage unit and the calculation mode of the convolution calculation unit. The calculation of the input and output data addresses of each level of storage unit depends on the counter inside the FPGA, the address calculation unit, and the counter counting information included in the instructions, such as the parameter information of the input feature map, output feature map, weight, and bias, the block and group information, the convolution, activation, pooling, and fully connected information, etc.
[0072] In the embodiments of the present application, the instruction set generation process can be regarded as the pre-blocking and grouping process of convolution, and this process is implemented by the host computer. The host computer generates the instruction set according to the structure of the convolutional neural network to be deployed, the capacity of each level of storage unit on the FPGA, and the designed hardware structure, while ensuring that each data will not overflow its respective storage area. At the same time, this makes there be an infinite number of combinations of the blocking and grouping of convolution.
[0073] In the embodiments of the present application, due to the chunking and grouping of convolutions, if the sorting methods of general weight data and original image data are not changed, the DDR read / write data addresses will be discontinuous, reducing the DDR read / write efficiency. Due to the structural characteristics of the multiplier array in the Conv sub-module, if the structures of the weight data and input feature map data are not transformed accordingly, the computational difficulty and the number of computational times of the RAM read / write addresses will increase. In the present invention, corresponding optimizations are made to the above problems.
[0074] In summary, the present application provides a method for deploying a convolutional neural network applied to a mobile terminal, including: receiving a convolutional neural network deployment request carrying to-be-deployed model parameters and instructions; transmitting the to-be-deployed model parameters and instructions to the DDR according to the PCIe interface; reading the first instruction from the instruction area of the DDR; reading the first weight data from the weight area of the DDR according to the first instruction and a counter, and storing it in the Wt_RAM; reading the input feature map data from the feature Figure 1 area of the DDR according to the first instruction and the counter, and storing it in the Fin_RAM; reading the bias data from the bias area of the DDR according to the first instruction and the counter, and storing it in the Bias_FIFO; performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and writing the data in the Fout_RAM back to the feature Figure 2 area of the DDR; when the convolution calculation operations for all the instructions in the instruction area are completed, obtaining an inference result; outputting the inference result of the DDR according to the PCIe interface. The present application can, without any hardware knowledge, quickly deploy convolutional neural networks with different structures at different times by modifying the instruction set and data structure, or modify the structure of a deployed network. Standard convolution and depthwise separable convolution use almost the same hardware resources to complete, which improves the generality while greatly reducing the consumption of hardware resources. Through the chunking and grouping of convolutions, it is possible to compute convolutions with almost any feature map size, number of layers, kernel size, and convolution stride on limited hardware resources.
[0075] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), etc., or a random access memory (RAM), etc.
[0076] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0077] Embodiment Two
[0078] Further referring to Figure 6 , as an implementation of the method shown above Figure 2 , this application provides an embodiment of a convolutional neural network deployment device applied to a mobile terminal. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0079] As Figure 7 shown, the convolutional neural network deployment device 100 applied to a mobile terminal in this embodiment includes: a request receiving module 310, a data transmission module 320, an instruction reading module 330, a weight data reading module 340, an input feature map reading module 350, a bias data reading module 360, a convolution calculation module 370, an inference result obtaining module 380, and an output module 390. Among them:
[0080] The request receiving module 310 is configured to receive a convolutional neural network deployment request carrying the model parameters to be deployed and instructions;
[0081] The data transmission module 320 is configured to transmit the model parameters to be deployed and instructions to the DDR according to the PCIe interface;
[0082] The instruction reading module 330 is configured to read the first instruction from the instruction area of the DDR;
[0083] The weight data reading module 340 is configured to read the first weight data from the weight area of the DDR according to the first instruction and the counter, and store it in the Wt_RAM;
[0084] The input feature map reading module 350 is configured to read the input feature map data from the feature Figure 1 area of the DDR according to the first instruction and the counter, and store it in the Fin_RAM;
[0085] The bias data reading module 360 is used to read bias data from the bias area of the DDR according to the first instruction and the counter, and store it into the Bias_FIFO;
[0086] The convolution calculation module 370 is used to perform a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and write the data of the Fout_RAM back to the feature Figure 2 area of the DDR;
[0087] The inference result obtaining module 380 is used to obtain the inference result when the convolution calculation operations of all instructions in the instruction area are completed;
[0088] The output module 390 is used to output the inference result of the DDR according to the PCIe interface.
[0089] In the embodiment of the present application, the convolution calculation module 370 includes: a first convolution calculation sub-module, a second convolution calculation sub-module, a third convolution calculation sub-module, a fourth convolution calculation sub-module, and a fifth convolution calculation sub-module, where:
[0090] The first convolution calculation sub-module is used to determine the address of the output feature map data in the Fout_RAM according to the output feature map size, the counter, and whether pooling is performed;
[0091] The second convolution calculation sub-module is used to determine the read data address of the Fin_RAM according to the input feature map size, the convolution kernel size, the padding value, and the counter;
[0092] The third convolution calculation sub-module is used to determine the read data address of the Wt_RAM according to the counter;
[0093] The fourth convolution calculation sub-module is used to determine the read data address of the Bias_FIFO according to the counter;
[0094] The fifth convolution calculation sub-module is used to determine whether the current convolution calculation operation is completed according to the counter.
[0095] In the embodiment of the present application, refer to Figure 3The block diagram of the general convolutional neural network acceleration system shown stores the input feature map and output feature map of a certain layer, the weights and biases of all layers, and the instruction set for controlling the entire inference process in the DDR memory. After decoding the instructions, the Top_Ctrl module controls the data interaction between the DDR, RAM, and FIFO according to the decoded instructions. The general convolutional calculation module is the core of the design, and through instructions, it can implement standard convolution or depthwise separable convolution with variable feature map size, number of layers, kernel size, and convolution stride, and the option to enable or disable activation and pooling. Wt_RAM, Fin_RAM, and Fout_RAM use on-chip RAM resources, and Bias_FIFO is implemented using LUT due to its large data bit width. The Top_Ctrl module realizes specific functions through a state machine and determines when to start the entire system through an off-chip signal Start.
[0096] In the embodiment of the present application, since the RAM resources on the FPGA are very limited, if the number of parameters of a certain layer is very large and the RAM cannot store it simultaneously, it is difficult to perform calculations. Therefore, the convolution is divided into blocks and groups. The convolution calculation in the convolutional layer of a neural network can be understood as: a three-dimensional array A (input feature map) passes through a four-dimensional array B (weights) to obtain a new three-dimensional array C (output feature map). The calculation decomposition of the convolution can be understood as: taking out a small three-dimensional array a from a large three-dimensional array A, taking out a small four-dimensional array b from a large four-dimensional array B, a passes through b to obtain c, and multiple c are combined to obtain C. Such a process is the process of data interaction between the DDR, RAM, and FIFO, and the specific data decomposition method is reflected in the instructions.
[0097] In the embodiment of the present application, refer to Figure 4 Refer to the structural schematic diagram of the Conv sub-module in the general convolutional calculation module shown. This module can implement standard convolution calculation and depthwise separable convolution calculation. In addition to the data flow shown in the figure, the data in the register bank comes from Wt_RAM, the data in the multiplier array comes from Fin_RAM, and the data in the accumulation array comes from Bias_FIFO. There are two groups of registers in the register bank. Through ping-pong operation, one register group receives data from the RAM, and one register group transfers data downward. The multiplier array contains 8-bit multipliers and is not a systolic array, that is, the outputs of all multipliers are valid simultaneously. In the adder tree array, the number of levels is determined by the number of multipliers in the multiplier array, and it also includes data truncation and saturation operations. In the accumulation array, the number of accumulations is provided by the instructions. Specifically, it is determined by the block and group situation of the convolution and the kernel size, and it also includes data truncation and saturation operations.
[0098] In the embodiments of the present application, when performing standard convolution, the output result of the multiplier array is obtained as output data of the same type as the input data after passing through the adder tree array and the accumulation array; when performing depth convolution, a part of the multipliers in the multiplier array are disabled, the output result skips the adder tree array, and the output data of the same type as the input data is obtained after passing through the accumulation array; when performing point convolution, it is the same as standard convolution.
[0099] In the embodiments of the present application, Figure 5 Fig. shows a schematic structural diagram of the Pooling sub-module in the general convolution calculation module. The pooling method is maximum pooling with a size of 2X2 and a stride of 2. The pooling process is divided into row pooling and column pooling, and the result of row pooling needs to be row-cached. Due to the requirement of design generality, the row size of the output feature map is unpredictable during design. Generally, a sufficiently large row cache needs to be ensured to meet the design generality requirement, but this method will cause a large waste of storage space when the row size is very small. In the present invention, by changing the calculation order during convolution, the size of the row cache is fixed, and convolution calculation can still be completed when the row size is several times the row cache.
[0100] In the embodiments of the present application, the output result of the Pooling module is the output feature map data of this layer, and its data structure is the same as that of the input feature map data, and can be directly used as the input for the next layer to participate in the calculation.
[0101] In the embodiments of the present application, the generality of the design is realized by variable instructions controlling invariant hardware. The instructions are mainly used to control the read and write addresses of the input and output data of each level of storage unit and the calculation mode of the convolution calculation unit. The calculation of the input and output data addresses of each level of storage unit depends on the counter, address calculation unit inside the FPGA and the counter counting information included in the instructions, such as the parameter information of the input feature map, output feature map, weight, and bias, the block and group information, the convolution, activation, pooling, and fully connected information, etc.
[0102] In the embodiments of the present application, the instruction set generation process can be regarded as the pre-blocking and grouping process of convolution, and this process is implemented by the host computer. The host computer generates the instruction set according to the convolutional neural network structure to be deployed, the capacity of each level of storage unit on the FPGA, and the designed hardware structure, while ensuring that each data will not overflow its respective storage area. At the same time, this makes the block and group of convolution have infinite combinations.
[0103] In the embodiments of the present application, due to the chunking and grouping of convolutions, if the sorting methods of general weight data and original image data are not changed, the DDR read / write data addresses will be discontinuous, reducing the DDR read / write efficiency; due to the structural characteristics of the multiplier array in the Conv sub-module, if the structures of the weight data and input feature map data are not transformed accordingly, the calculation difficulty and the number of calculation times of the RAM read / write addresses will be increased. In the present invention, corresponding optimizations are made to the above problems.
[0104] In summary, the present application provides a convolutional neural network deployment device for a mobile terminal, specifically including: a request receiving module, configured to receive a convolutional neural network deployment request carrying to-be-deployed model parameters and instructions; a data transmission module, configured to transmit the to-be-deployed model parameters and instructions to the DDR according to a PCIe interface; an instruction reading module, configured to read the first instruction from the instruction area of the DDR; a weight data reading module, configured to read first weight data from the weight area of the DDR according to the first instruction and a counter, and store it in the Wt_RAM; an input feature map reading module, configured to read input feature map data from the feature Figure 1 area of the DDR according to the first instruction and the counter, and store it in the Fin_RAM; a bias data reading module, configured to read bias data from the bias area of the DDR according to the first instruction and the counter, and store it in the Bias_FIFO; a convolution calculation module, configured to perform a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and write the data in the Fout_RAM back to the feature Figure 2 area of the DDR; an inference result obtaining module, configured to obtain an inference result when the convolution calculation operations of all instructions in the instruction area are completed; an output module, configured to output the inference result of the DDR according to the PCIe interface. The present application can quickly deploy convolutional neural networks with different structures at different times, or modify the structure of a deployed network, by modifying the instruction set and data structure without any hardware knowledge. Standard convolution and depthwise separable convolution use almost the same hardware resources to complete, while improving generality, greatly reducing the consumption of hardware resources. Through the chunking and grouping of convolutions, it is possible to calculate convolutions with almost any feature map size, number of layers, kernel size, and convolution stride on limited hardware resources.
[0105] To solve the above technical problems, embodiments of the present application also provide a computer device. Specifically, please refer to Figure 7 , Figure 7 which is the basic structure block diagram of the computer device in this embodiment.
[0106] The computer device 200 includes a memory 210, a processor 220, and a network interface 230 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 200 with components 210-230 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0107] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, etc.
[0108] The memory 210 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 210 can be an internal storage unit of the computer device 200, such as the hard disk or memory of the computer device 200. In other embodiments, the memory 210 can also be an external storage device of the computer device 200, such as a plug-in hard disk equipped on the computer device 200, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 210 can also include both the internal storage unit and the external storage device of the computer device 200. In this embodiment, the memory 210 is generally used to store the operating system and various application software installed on the computer device 200, such as computer-readable instructions for the convolutional neural network deployment method applied to a mobile terminal. In addition, the memory 210 can also be used to temporarily store various data that have been output or will be output.
[0109] In some embodiments, the processor 220 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 220 is generally used to control the overall operation of the computer device 200. In this embodiment, the processor 220 is used to run the computer-readable instructions stored in the memory 210 or process data, such as running the computer-readable instructions of the convolutional neural network deployment method applied to the mobile terminal.
[0110] The network interface 230 may include a wireless network interface or a wired network interface, and this network interface 230 is generally used to establish a communication connection between the computer device 200 and other electronic devices.
[0111] For the computer device provided in this application, without any hardware knowledge, this application can quickly deploy convolutional neural networks with different structures at different times by modifying the instruction set and data structure, or make structural modifications to a deployed network. Standard convolution and depthwise separable convolution are completed using almost the same hardware resources, which greatly reduces the consumption of hardware resources while improving generality. Through the block and group of convolutions, convolutions with almost any feature map size, number of layers, kernel size, and convolution stride can be calculated on limited hardware resources.
[0112] This application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to execute the steps of the convolutional neural network deployment method applied to the mobile terminal as described above.
[0113] For the computer-readable storage medium provided in this application, without any hardware knowledge, this application can quickly deploy convolutional neural networks with different structures at different times by modifying the instruction set and data structure, or make structural modifications to a deployed network. Standard convolution and depthwise separable convolution are completed using almost the same hardware resources, which greatly reduces the consumption of hardware resources while improving generality. Through the block and group of convolutions, convolutions with almost any feature map size, number of layers, kernel size, and convolution stride can be calculated on limited hardware resources.
[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0115] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is equally within the scope of the patent protection of the present application.
Claims
1. A method for deploying a convolutional neural network applied to a mobile terminal, characterized in that, it includes the following steps: Receiving a convolutional neural network deployment request carrying model parameters to be deployed and instructions; Transmitting the model parameters to be deployed and the instructions to the DDR according to the PCIe interface; Reading the first instruction from the instruction area of the DDR; Reading first weight data from the weight area of the DDR according to the first instruction and a counter, and storing it in the Wt_RAM; Reading input feature map data from the first feature map area of the DDR according to the first instruction and the counter, and storing it in the Fin_RAM; Reading bias data from the bias area of the DDR according to the first instruction and the counter, and storing it in the Bias_FIFO; Performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and writing the data in the Fout_RAM back to the second feature map area of the DDR; When the convolution calculation operations for all instructions in the instruction area are completed, obtaining an inference result; Outputting the inference result of the DDR according to the PCIe interface.
2. The method for deploying a convolutional neural network applied to a mobile terminal according to claim 1, characterized in that, The step of performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data specifically includes: Determining the address of the output feature map data in the Fout_RAM according to the output feature map size, the counter, and whether pooling is performed.
3. The method for deploying a convolutional neural network applied to a mobile terminal according to claim 1, characterized in that, The step of performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data specifically further includes: Determining the read data address of the Fin_RAM according to the input feature map size, the convolutional kernel size, the padding value, and the counter.
4. The method for deploying a convolutional neural network applied to a mobile terminal according to claim 1, characterized in that, The step of performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data specifically further includes: Determining the read data address of the Wt_RAM according to the counter.
5. The method for deploying a convolutional neural network applied to a mobile terminal according to claim 1, characterized in that, The step of performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data specifically further includes: Determining the read data address of the Bias_FIFO according to the counter.
6. The method for deploying a convolutional neural network applied to a mobile terminal according to claim 1, characterized in that, The step of performing a convolution calculation operation according to the weight data, the input feature map data, and the bias data specifically further includes: Determining whether the convolution calculation operation ends this time according to the counter.
7. A convolutional neural network deployment device applied to a mobile terminal, characterized in that, it includes: A request receiving module, configured to receive a convolutional neural network deployment request carrying model parameters to be deployed and instructions; A data transmission module, configured to transmit the model parameters to be deployed and the instructions to a DDR according to a PCIe interface; An instruction reading module, configured to read a first instruction from an instruction area of the DDR; A weight data reading module, configured to read first weight data from a weight area of the DDR according to the first instruction and a counter, and store the first weight data into a Wt_RAM; An input feature map reading module, configured to read input feature map data from a first feature map area of the DDR according to the first instruction and the counter, and store the input feature map data into a Fin_RAM; A bias data reading module, configured to read bias data from a bias area of the DDR according to the first instruction and the counter, and store the bias data into a Bias_FIFO; A convolution calculation module, configured to perform a convolution calculation operation according to the weight data, the input feature map data, and the bias data, and write data in the Fout_RAM back to a second feature map area of the DDR; An inference result obtaining module, configured to obtain an inference result after completing the convolution calculation operations for all instructions in the instruction area; An output module, configured to output the inference result of the DDR according to the PCIe interface.
8. The convolutional neural network deployment device applied to a mobile terminal according to claim 7, wherein, the convolution calculation module includes: A first convolution calculation sub-module, configured to determine an address of output feature map data in the Fout_RAM according to an output feature map size, the counter, and whether pooling is performed.
9. A computer device, wherein, it includes a memory and a processor, and computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the convolutional neural network deployment method applied to a mobile terminal according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, wherein, computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the convolutional neural network deployment method applied to a mobile terminal according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Neural network model processing method, device and terminal
CN109409518A
Scale-extensible convolutional neural network acceleration system and method
CN111242289A