Computing Method, Device, Equipment and Computer Readable Medium of Multi-Layer Neural Network
By saving the intermediate data of multi-layer neural networks in the edge buffer and iterating the edge-compensation strategy, the problem of low computing efficiency of multi-layer neural networks is solved, the memory bandwidth pressure is reduced, and the computing efficiency and density are improved.
Patent Information
- Application Number
- CN202011348888.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-11-26
AI Technical Summary
In the prior art, neural network computing efficiency is low, resulting in high memory bandwidth pressure, and the overlapping areas between adjacent segmented areas in the multi-layer convolutional neural network computing method is getting larger and larger, resulting in serious problems in repeated computing.
The intermediate data of each network layer of the target neural network is stored in the edge buffer. The intermediate data is the overlapping part between adjacent input segmentation areas of each network layer. It is loaded into the neural network processor through the edge buffer for calculation, and the iterative edge-compensation strategy is used to reduce repeated calculations.
It greatly reduces the memory bandwidth pressure, improves the computing efficiency of multi-layer neural networks, reduces repeated computing, and enhances the actual computing efficiency and density of NPUs.
Smart Images

Figure CN114548354B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of neural networks, and particularly to a calculation method, device, equipment, and computer-readable medium for a multi-layer neural network. Background Art
[0002] A neural network acceleration operation processor (NPU) generally includes a large number of MAC operation units and a relatively large SRAM buffer. In terminal applications that only contain the neural network inference part, the typical size of the NPU buffer varies from 1 to 4 MByte, and the buffer area is larger than the area of the MAC operation units.
[0003] Currently, in related technologies, in order to improve the computing power per unit area, the size of the NPU SRAM buffer can be reduced to add more MAC operation units. However, when the buffer area is reduced, it will cause the intermediate layer data of each layer to be imported and exported between the NPU buffer and the memory buffer, and will also cause the weight parameters of each layer to be reloaded multiple times, which not only has low computing efficiency, but also brings great pressure to the memory bandwidth, making the "memory wall" problem of the neural network chip more serious.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] This application provides a calculation method, device, equipment, and computer-readable medium for a multi-layer neural network to solve the technical problem of low computing efficiency of the multi-layer neural network.
[0006] According to one aspect of the embodiments of this application, this application provides a calculation method for a multi-layer neural network, including:
[0007] Storing the intermediate data of each network layer of the target neural network in an edge buffer, where the intermediate data is the data obtained by calculating the overlapping part between adjacent input segmentation regions of each network layer, and the edge buffer is a region for caching the intermediate data;
[0008] When loading the input data of the target network layer into the neural network processor, loading the intermediate data from the edge buffer into the neural network processor to calculate the target network layer by using the neural network processor.
[0009] Optionally, when loading the input data of the target network layer into the neural network processor, loading the intermediate data from the edge buffer into the neural network processor to calculate the target network layer by using the neural network processor includes:
[0010] Load the input segmentation region of the first network layer of the target neural network from memory into the input buffer of the neural network processor, so as to use the neural network processor to calculate the first network layer, obtain the output segmentation region of the first network layer, and save the edge region of the output segmentation region as intermediate data in the edge buffer. There is an overlapping region between two adjacent input segmentation regions of the first network layer, and an input segmentation region is the data input for all data input channels of the target row and column;
[0011] Starting from the second network layer, when loading the output segmentation region of the previous network layer into the input buffer of the neural network processor, load the edge region of the previous network layer from the edge buffer as the input padding region of the current network layer into the input buffer of the neural network processor, so as to use the neural network processor to calculate the current network layer, obtain the output segmentation region of the current network layer output by the output buffer of the neural network processor, and save the edge region in the output segmentation region of the current network layer as the input padding region of the next network layer until the output segmentation region of the last network layer is obtained; where,
[0012] The target network layer includes any network layer from the second network layer to the last network layer, and the input data includes the output segmentation region of the previous layer.
[0013] Optionally, before loading the input segmentation region of the first network layer of the target neural network into the input buffer of the neural network processor to use the neural network processor to calculate the first network layer, the method further includes:
[0014] Determine the overall input segmentation region and output segmentation region of each network layer in the target neural network. The output segmentation region includes an edge region, and the edge region is used to represent the overlapping region between adjacent overall input segmentation regions in the next layer network. The overall input segmentation region includes the output segmentation region of the previous layer of the target network layer and the input padding region. When the target network layer is the first network layer, the overall input segmentation region is the input segmentation region of the first network layer.
[0015] Optionally, determining the overall input segmentation region and output segmentation region of each network layer in the target neural network includes:
[0016] Determine that the output segmentation region of the last network layer of the target neural network is the maximum segmentation region;
[0017] Starting from the last network layer, determine the overall input segmentation region of the current network layer according to the size of the receptive field using the output segmentation region of the current network layer, and determine the input padding region of the current network layer using the size of the overlapping part between adjacent overall input segmentation regions of the current network layer, the output segmentation region of the previous network layer, and the edge region in the output segmentation region, until the input segmentation region of the first network layer is obtained; among them,
[0018] When the size of the overall input segmentation region of any network layer is greater than the size of the input buffer of the neural network processor, reduce the size of the output segmentation region of the last network layer, and re-determine the overall input segmentation region, input padding region, output segmentation region of the previous layer, and edge region in the output segmentation region layer by layer.
[0019] Optionally, after saving the edge region in the output segmentation region of the current network layer as the input padding region of the next network layer, the method further includes:
[0020] When the calculation result of the current network layer does not fill the output buffer of the neural network processor, load the next input segmentation region adjacent to the input segmentation region in the first network layer from the memory into the input buffer of the neural network processor, so as to obtain the next output segmentation region corresponding to the next input segmentation region using the neural network processor, and start from the second network layer again, layer by layer determine the next output segmentation region corresponding to the next input segmentation region of each network layer and the next edge region in the next output segmentation region, until the calculation result of the current network layer fills the output buffer of the neural network processor, and then continue to calculate the next network layer.
[0021] Optionally, when the output segmentation region of the last network layer is calculated, the method further includes:
[0022] Determine multiple output channels in the last network layer that can accommodate the output segmentation region according to the size of the output buffer of the neural network processor, and the multiple output channels divide the output segmentation region into multiple sub-output segmentation regions;
[0023] Extract the channels to be processed in the multiple output channels one by one, and the channels to be processed are the channels in the multiple output channels that have not undergone convolution operations;
[0024] Load the weight parameters of the channels to be processed and perform convolution operations on the sub-output buffer regions corresponding to the channels to be processed;
[0025] When convolution operations are completed for each output channel, obtain the convolution operation result of the last network layer.
[0026] Optionally, the edge region includes at least one of the following regions:
[0027] An overlapping row area between vertically adjacent overall input segmentation regions;
[0028] An overlapping column area between horizontally adjacent overall input segmentation regions;
[0029] An overlapping square area between adjacent overall input segmentation regions.
[0030] According to another aspect of the embodiments of the present application, the present application provides a computing device for a multi-layer neural network, including:
[0031] An intermediate data caching module, configured to save the intermediate data of each network layer of the target neural network in an edge buffer, where the intermediate data is data obtained by calculating the overlapping part between adjacent input segmentation regions of each network layer, and the edge buffer is an area for caching intermediate data;
[0032] A border supplement calculation module, configured to load the intermediate data from the edge buffer to the neural network processor when loading the input data of the target network layer to the neural network processor, so as to calculate the target network layer by using the neural network processor.
[0033] According to another aspect of the embodiments of the present application, the present application provides an electronic device, including a memory, a processor, a communication interface, and a communication bus. A computer program that can run on the processor is stored in the memory. The memory and the processor communicate through the communication bus and the communication interface. When the processor executes the computer program, the steps of the above method are implemented.
[0034] According to another aspect of the embodiments of the present application, the present application further provides a computer-readable medium having non-volatile program code executable by a processor, and the program code causes the processor to execute the above method.
[0035] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the related technologies:
[0036] The technical solution of the present application is to save the intermediate data of each network layer of the target neural network in the edge buffer. The intermediate data is data obtained by calculating the overlapping part between adjacent input segmentation regions of each network layer, and the edge buffer is an area for caching intermediate data. When loading the input data of the target network layer to the neural network processor, the intermediate data is loaded from the edge buffer to the neural network processor to calculate the target network layer by using the neural network processor. The present application adopts adding an input border supplement area to the input area of each network layer, which can make the operations of the multi-layer convolutional neural network have no repeated calculations, and a small amount of additional intermediate data is saved in the edge buffer, greatly reducing the memory bandwidth pressure, thereby solving the technical problem of low computing efficiency of the multi-layer neural network. Description of the Drawings
[0037] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or in related technologies, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or related technologies. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic diagram of the calculation of a multi-layer neural network;
[0040] Figure 2 It is a schematic diagram of the hardware environment of an optional calculation method for a multi-layer neural network provided according to an embodiment of this application;
[0041] Figure 3 It is a flowchart of an optional calculation method for a multi-layer neural network provided according to an embodiment of this application;
[0042] Figure 4 It is a schematic diagram of the calculation of an optional multi-layer neural network provided according to an embodiment of this application;
[0043] Figure 5 It is a block diagram of an optional calculation device for a multi-layer neural network provided according to an embodiment of this application;
[0044] Figure 6 It is a schematic diagram of the structure of an optional electronic device provided according to an embodiment of this application. Detailed implementation manners
[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.
[0046] In subsequent descriptions, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of description of this application, and they have no specific meaning in themselves. Therefore, "module" and "component" can be used interchangeably.
[0047] In the related art, for the most widely used convolutional neural network, due to the locality of its receptive field, in order to reduce the number of memory imports and exports, partial input row and column data can be loaded, and multiple layers of neural networks can be calculated at once and the partial row and column data of the last layer can be output. By iteratively loading and calculating all input segmentation regions, all output row and column data of the last layer of the multiple layers of neural networks can finally be obtained. For example:
[0048] As Figure 1 shown, due to the limitation of the NPU input buffer, the input data of the first layer of the multiple layers of neural networks is divided into 4 regions, namely I1, I2, I3, and I4. Each region contains partial adjacent input rows, or partial adjacent input columns, or all channel data of partial adjacent input rows and columns. Then, each region calculates 6 layers of neural networks at once and outputs partial output rows, or partial output columns, or partial output rows and columns of the last layer. Figure 1 The output segmentation regions corresponding to the input segmentation regions I1, I2, I3, and I4 are O1, O2, O3, and O4 respectively. Since the receptive field of the convolutional neural network increases layer by layer from back to front and the output of each layer decreases layer by layer, taking I2 and O2 as an example, O2 is much smaller than I2.
[0049] The inventor found during the research process that even if there are no overlapping regions in the outputs O1, O2, O3, and O4, there are overlapping calculation regions S1 in the calculation processes of the inputs I1 and I2, overlapping calculation regions S2 in the calculation processes of the inputs I2 and I3, and overlapping calculation regions S3 in the calculation processes of the inputs I3 and I4. At this time, for each calculation of each part of each layer, all the weight parameters of this layer need to be reloaded once. Therefore, the more times each layer is cut, the more times the weights are reloaded. Thus, the related art has the following problems:
[0050] 1. Although the operation of the convolutional neural network has locality, due to the receptive field increasing layer by layer from back to front, the overlapping regions between adjacent segmentation regions of the above-mentioned calculation method of the multiple layers of convolutional neural networks are getting larger and larger, making the problem of repeated calculation more and more serious;
[0051] 2. The more layers the calculation method of the multiple layers of convolutional neural networks spans, the smaller the effective output row and column are, the cutting times increase rapidly, and the number of times the weight parameters are reloaded also increases rapidly, which instead increases the memory bandwidth pressure and reduces the calculation efficiency.
[0052] In order to solve the problems mentioned in the background art, according to one aspect of the embodiments of the present application, an embodiment of a calculation method for a multiple layers of neural networks is provided.
[0053] Optionally, in the embodiments of the present application, the above-mentioned calculation method for the multiple layers of neural networks can be applied to the hardware environment composed of a terminal 201 and a server 203 as Figure 2 shown. AsFigure 2 As shown, the server 203 is connected to the terminal 201 through a network, and can be used to provide services for the terminal or the client installed on the terminal. A database 205 can be set up on the server or independently of the server to provide data storage services for the server 203. The above network includes, but is not limited to: wide area network, metropolitan area network or local area network. The terminal 201 includes, but is not limited to: PC, mobile phone, tablet computer, etc.
[0054] In the embodiments of the present application, a calculation method of a multi-layer neural network can be executed by the server 203, or can also be jointly executed by the server 203 and the terminal 201, as Figure 3 shown, the method may include the following steps:
[0055] Step S302, save the intermediate data of each network layer of the target neural network in the edge buffer. The intermediate data is the data obtained by calculating the overlapping part between adjacent input segmentation regions of each network layer. The edge buffer is a region for caching intermediate data.
[0056] In the embodiments of the present application, the above target neural network can be a convolutional neural network (Convolutional Neural Networks, CNN). A convolutional neural network can include an input layer, a hidden layer, a fully connected layer and an output layer. The input layer can process multi-dimensional data. For example, the input layer of a one-dimensional convolutional neural network receives a one-dimensional or two-dimensional array, where the one-dimensional array is usually a time or spectrum sample; the two-dimensional array may contain multiple channels. The hidden layer contains three common architectures: convolutional layer, pooling layer and fully connected layer. The convolutional kernel in the convolutional layer contains weight coefficients. In a convolutional neural network, the receptive field refers to the size of the area on the input picture mapped by the pixel points on the feature map output by each layer of the convolutional neural network.
[0057] Since the receptive field of the convolutional neural network increases layer by layer from back to front, the output of each layer decreases layer by layer, and the overlapping area between adjacent segmentation regions becomes larger and larger, resulting in a more and more serious problem of repeated calculation. Therefore, the technical solution of the present application can only save the part of the overlapping area (i.e., the above-mentioned edge area) caused by the overlap of the receptive fields between adjacent segmentation regions as intermediate data in the edge buffer to reduce the storage pressure of the memory.
[0058] Optionally, the above-mentioned edge area includes at least one of the following areas:
[0059] The row area overlapping between the upper and lower adjacent overall input segmentation regions;
[0060] The column area overlapping between the left and right adjacent overall input segmentation regions;
[0061] Overlapping square regions between adjacent overall input segmentation regions.
[0062] In step S304, when loading the input data of the target network layer into the neural network processor, load the intermediate data from the edge buffer into the neural network processor to calculate the target network layer using the neural network processor.
[0063] In the embodiments of the present application, a neural-network processing unit (NPU) adopts an architecture of "data-driven parallel computing" and is particularly good at processing massive amounts of multimedia data such as videos and images. It is specifically designed for Internet of Things artificial intelligence and is used to accelerate the operation of neural networks, solving the problem of low efficiency of traditional chips in neural network operations.
[0064] Optionally, when loading the input data of the target network layer into the neural network processor in step S304, loading the intermediate data from the edge buffer into the neural network processor to calculate the target network layer using the neural network processor may include the following steps:
[0065] Step 1, load the input segmentation region of the first network layer of the target neural network from the memory into the input buffer of the neural network processor to calculate the first network layer using the neural network processor, obtain the output segmentation region of the first network layer, save the edge region of the output segmentation region as intermediate data in the edge buffer, and there is an overlapping region between two adjacent input segmentation regions of the first network layer. One input segmentation region is the data input for all data input channels of the target row and column.
[0066] In the embodiments of the present application, the input of the first network layer of the convolutional neural network can be divided into multiple input segmentation regions, and each input segmentation region is the data to be input for all channels of a partial row and column of the feature map (i.e., the above-mentioned target row and column, which can be specifically selected according to the actual situation). Due to the overlap of the receptive fields of adjacent segmentation regions between the first network layer and the second network layer, the above-mentioned edge region is generated in the output segmentation region of the first network layer, specifically edge row data and edge column data.
[0067] Step 2: Starting from the second network layer, when loading the output segmentation region of the previous network layer into the input buffer of the neural network processor, load the edge region of the previous network layer from the edge buffer as the input padding region of the current network layer into the input buffer of the neural network processor, so as to use the neural network processor to calculate the current network layer, obtain the output segmentation region of the current network layer output by the output buffer of the neural network processor, and save the edge region in the output segmentation region of the current network layer as the input padding region of the next network layer until the output segmentation region of the last network layer is obtained; where
[0068] The target network layer includes any network layer from the second network layer to the last network layer, and the input data includes the output segmentation region of the previous layer.
[0069] In the embodiment of the present application, an iterative padding strategy is adopted. For the input segmentation region of the first layer, directly load the input segmentation region from the memory. For each input segmentation region of each layer of the convolutional neural network other than the row-first or column-first segmentation region (i.e., the above-mentioned original input segmentation), before the calculation starts, insert the edge rows and edge columns of the previous layer at the row-first and column-first, that is, the row data and column data of the overlapping region of the upper-left segmentation region and the receptive field of the current segmentation region.
[0070] To illustrate the present solution specifically and in detail, the following combines Figure 4 to describe the technical solution of the present application.
[0071] As Figure 4 shown, the input of the first layer of the multi-layer neural network can be divided into 4 regions, namely I1, I2, I3, and I4. Each input segmentation region can be all channel data of adjacent input rows, or all channel data of adjacent input columns, or all channel data of adjacent input rows and columns. For example, if the input row has 64 rows and 64 columns and is evenly divided, then I1 is rows 1-16, I2 is rows 16-32, I3 is rows 32-48, I4 is rows 48-64, or I1 is the square region [rows 1-32, columns 1-32], I2 is the square region [rows 1-32, columns 32-64], I3 is the square region [rows 32-64, columns 1-32], I4 is the square region [rows 32-64, columns 32-64]. Correspondingly, the outputs O1, O2, O3, and O4 are the output segmentation regions of the last layer of the multi-layer neural network, O1 corresponds to I1, O2 corresponds to I2, O3 corresponds to I3, and O4 corresponds to I4. It should be noted that Figure 3 only one-dimensional segmentation segments are shown, but generally, two-dimensional square segmentation regions are preferably used, which is beneficial to increasing the number of layers spanned.
[0072] After the calculation of each layer of the multi-layer neural network is completed, the edge regions of the output of each layer are saved in the memory buffer, so that the left edge column can be supplemented for the right adjacent input segmentation region of the next layer, or the upper edge row can be supplemented for the lower adjacent input segmentation region.
[0073] Taking I2 as an example, the inputs of each layer of the multi-layer neural network are I21, I22, I23, …, I2n, where I21 is equal to I2. The edge regions B11, B12, B13, …, B1n of the output of each layer obtained by dividing I1 are included in I21, I22, I23, …, I2n. That is, the edge region B12 of the output of the first layer of I1 is used as the input supplementary edge region of the second layer of I2 to supplement the left edge. Similarly, the edge region B22 of the output of the first layer of the I2 channel is used as the input supplementary edge region of the second layer of the I3 channel to supplement the left edge of the input of the second layer of the I3 channel. Typically, for a 3x3 convolutional kernel with a stride of 1, the edge region of each layer's output segmentation region contains all channel data of the last 2 rows and 2 columns of each layer's output segmentation region. For this convolutional kernel, all channel data of the last 2 rows and 2 columns of each layer's output segmentation region are saved as the edge region in the edge buffer.
[0074] For the calculation process of the multi-layer convolutional neural network that divides I2, although the output of each layer is smaller than the input of each layer, since the edge regions B12, B13, …, B1n of the output of the previous layer of I1 are inserted at the beginning of the rows and columns of each layer's input, therefore, I22, I23, …, I2n, which are used as the input of the next layer (regardless of the downsampling factor), are of equal size. Therefore, Figure 4 it will not be as Figure 1 shown a situation of decreasing layer by layer. It should be noted that Figure 4 only shows the one-dimensional edge region. For the two-dimensional edge region, the right edge column of the left adjacent segmentation region can be inserted at the beginning of the column, and the lower edge row of the upper adjacent segmentation region can be inserted at the beginning of the row.
[0075] Adopting the technical solution of the present application to supplement the input supplementary edge region in the input region of each network layer can make the operation of the multi-layer convolutional neural network have no repeated calculation, and a small amount of supplemented intermediate data is saved in the edge buffer, greatly reducing the memory bandwidth pressure, thereby solving the technical problem of low calculation efficiency of the multi-layer neural network. It makes up for the shortcomings of the calculation method of the multi-layer convolutional neural network, greatly increasing the number of layers that can be spanned. On the basis of reducing the SRAM buffer, it can still greatly reduce the import and export quantity between the NPU buffer and the memory buffer, greatly improving the actual calculation efficiency and density of the NPU. For example, using a 512K NPU buffer can approach the calculation efficiency of a 2M buffer, and reduce the NPU area by several times, and the pressure on the memory bandwidth is also small.
[0076] Optionally, before loading the input segmentation region of the first network layer of the target neural network into the input buffer of the neural network processor to perform calculations on the first network layer using the neural network processor, the method further includes:
[0077] Determine the overall input segmentation region and the output segmentation region of each network layer in the target neural network. The output segmentation region includes an edge region, and the edge region is used to represent the overlapping region between adjacent overall input segmentation regions in the next network layer. The overall input segmentation region includes the output segmentation region of the previous layer of the target network layer and the input padding region. When the target network layer is the first network layer, the overall input segmentation region is the input segmentation region of the first network layer.
[0078] This step may specifically include:
[0079] Determine that the output segmentation region of the last network layer of the target neural network is the maximum segmentation region;
[0080] Starting from the last network layer, use the output segmentation region of the current network layer to determine the overall input segmentation region of the current network layer according to the size of the receptive field, and use the size of the overlapping part between adjacent overall input segmentation regions of the current network layer to determine the input padding region of the current network layer, the output segmentation region of the previous network layer, and the edge region in the output segmentation region until the input segmentation region of the first network layer is obtained; wherein,
[0081] When the size of the overall input segmentation region of any network layer is greater than the size of the input buffer of the neural network processor, reduce the size of the output segmentation region of the last network layer, and re-determine the overall input segmentation region, the input padding region, the output segmentation region of the previous layer, and the edge region in the output segmentation region layer by layer.
[0082] In the embodiments of the present application, an iterative backtracking method may be adopted to start from the last network layer with the maximum output segmentation region and backtrack the size of the overall input segmentation region and the output segmentation region of each layer layer by layer. Among them, the overall input segmentation region includes the input padding region, such as Figure 4As shown, the size of the input padding area where I2 divides the second layer is the size of B12, and the size of the overall input segmentation area is the size of I22. The output segmentation area where I2 divides the first layer is the size of I22 minus B12. B22 is the size of the edge area and is also the size of the input padding area for the input where I3 divides the second layer. Check whether the NPU input buffer can accommodate all the channel data of the overall input segmentation area. If not, reduce the size of the output segmentation area of the last layer and reverse-infer layer by layer until all the channel data of the input segmentation areas of all layers can be accommodated by the NPU input buffer. Reducing the size of the output segmentation area of the last layer can be reducing one row and one column.
[0083] Optionally, after saving it as the edge area in the output segmentation area of the current layer network layer to be used as the input padding area for the next layer network layer, the method further includes:
[0084] When the calculation result of the current layer network layer does not fill the output buffer of the neural network processor, load the next input segmentation area adjacent to the input segmentation area in the first layer network layer from the memory into the input buffer of the neural network processor, so as to obtain the next output segmentation area corresponding to the next input segmentation area by using the neural network processor, and start from the second layer network layer again, layer by layer determine the next output segmentation area corresponding to the next input segmentation area of each layer network layer, and the next edge area in the next output segmentation area until the calculation result of the current layer network layer fills the output buffer of the neural network processor, and then continue to calculate the next layer network layer.
[0085] In the embodiments of the present application, when the number of rows and columns of a certain layer of the convolutional neural network is reduced by a multiple, it will cause insufficient data loading in the NPU output buffer, resulting in parameter overloading. At this time, a nested multi-layer neural network calculation strategy can be adopted. Stop calculating at this layer, then load the next input segmentation area of the first layer, and then obtain the next segmentation output area of the current layer through multi-layer operations until the output buffer of this layer is full.
[0086] Optionally, when the output segmentation area of the last layer network layer is calculated, the method may further include:
[0087] Determine multiple output channels that the last layer network layer can accommodate the output segmentation area according to the size of the output buffer of the neural network processor. The multiple output channels divide the output segmentation area into multiple sub-output segmentation areas;
[0088] Extract the channels to be processed one by one from the multiple output channels. The channels to be processed are the channels in the multiple output channels that have not undergone convolution operations;
[0089] Load the weight parameters of the channels to be processed and perform convolution operations on the sub-output buffer areas corresponding to the channels to be processed;
[0090] When the convolution operation is completed for each output channel, the convolution operation result of the last network layer is obtained.
[0091] In the embodiment of the present application, the output buffer of the last layer does not need to accommodate all output channels. First, calculate the data of some output channels and save it to the memory buffer, then without changing the input buffer, load the weight parameters of other output channels to calculate the data of other output channels until all output channels in the divided area are calculated.
[0092] According to another aspect of the embodiment of the present application, as Figure 5 shown, a computing device for a multi-layer neural network is provided, including:
[0093] An intermediate data caching module 501, configured to save the intermediate data of each network layer of the target neural network in the edge buffer, where the intermediate data is the data obtained by calculating the overlapping part between adjacent input divided areas of each network layer, and the edge buffer is an area for caching the intermediate data;
[0094] A border supplement calculation module 503, configured to load the intermediate data from the edge buffer to the neural network processor when loading the input data of the target network layer to the neural network processor, so as to calculate the target network layer by using the neural network processor.
[0095] It should be noted that the intermediate data caching module 501 in this embodiment can be used to execute step S302 in the embodiment of the present application, and the border supplement calculation module 503 in this embodiment can be used to execute step S304 in the embodiment of the present application.
[0096] It should be noted here that the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in the hardware environment as Figure 1 shown, and can be implemented by software or by hardware.
[0097] Optionally, the border supplement calculation module is specifically configured to:
[0098] Load the input divided area of the first network layer of the target neural network from the memory to the input buffer of the neural network processor, so as to calculate the first network layer by using the neural network processor to obtain the output divided area of the first network layer, and save the edge area of the output divided area as intermediate data in the edge buffer. There is an overlapping area between two adjacent input divided areas of the first network layer, and one input divided area is the data input for all data input channels of the target row and column;
[0099] Starting from the second network layer, when loading the output segmentation region of the previous network layer into the input buffer of the neural network processor, the edge region of the previous network layer is loaded from the edge buffer into the input buffer of the neural network processor as the input padding region of the current network layer, so as to use the neural network processor to calculate the current network layer, obtain the output segmentation region of the current network layer output from the output buffer of the neural network processor, and save the edge region in the output segmentation region of the current network layer as the input padding region of the next network layer until the output segmentation region of the last network layer is obtained; where
[0100] The target network layer includes any network layer from the second network layer to the last network layer, and the input data includes the output segmentation region of the previous layer.
[0101] Optionally, the computing device of the multi-layer neural network further includes an input segmentation region and an output segmentation region determination module for the multi-layer neural network, which is used for:
[0102] Determine the overall input segmentation region and output segmentation region of each network layer in the target neural network. The output segmentation region includes an edge region, and the edge region is used to represent the overlapping region between adjacent overall input segmentation regions in the next network layer. The overall input segmentation region includes the output segmentation region of the previous layer of the target network layer and the input padding region. When the target network layer is the first network layer, the overall input segmentation region is the input segmentation region of the first network layer.
[0103] Optionally, the input segmentation region and output segmentation region determination module of the multi-layer neural network is specifically used for:
[0104] Determine that the output segmentation region of the last network layer of the target neural network is the maximum segmentation region;
[0105] Starting from the last network layer, use the output segmentation region of the current network layer to determine the overall input segmentation region of the current network layer according to the size of the receptive field, and use the size of the overlapping part between adjacent overall input segmentation regions of the current network layer to determine the input padding region of the current network layer, the output segmentation region of the previous network layer, and the edge region in the output segmentation region until the input segmentation region of the first network layer is obtained; where
[0106] When the size of the overall input segmentation region of any network layer is greater than the size of the input buffer of the neural network processor, reduce the size of the output segmentation region of the last network layer, and re-determine the overall input segmentation region, input padding region, output segmentation region of the previous layer, and the edge region in the output segmentation region layer by layer.
[0107] Optionally, the computing device of the multi-layer neural network further includes an output buffer loading judgment module for:
[0108] When the calculation result of the current layer network layer does not fully load the output buffer of the neural network processor, load the next input segmentation area adjacent to the input segmentation area in the first layer network layer from the memory into the input buffer of the neural network processor, so as to obtain the next output segmentation area corresponding to the next input segmentation area by using the neural network processor, and start from the second layer network layer again, and layer by layer determine the next output segmentation area corresponding to the next input segmentation area in each layer network layer and the next edge area in the next output segmentation area, until the calculation result of the current layer network layer fully loads the output buffer of the neural network processor, and then continue to calculate the next layer network layer.
[0109] Optionally, the computing device of the multi-layer neural network further includes a convolution calculation module for:
[0110] Determine the multiple output channels that the last layer network layer can accommodate for the output segmentation area according to the size of the output buffer of the neural network processor, and the multiple output channels divide the output segmentation area into multiple sub-output segmentation areas;
[0111] Extract the channels to be processed in the multiple output channels one by one, and the channels to be processed are the channels in the multiple output channels that have not undergone convolution operations;
[0112] Load the weight parameters of the channel to be processed and perform convolution operations on the sub-output buffer area corresponding to the channel to be processed;
[0113] When convolution operations are completed for each output channel, obtain the convolution operation result of the last layer network layer.
[0114] Optionally, the edge area includes at least one of the following areas:
[0115] The overlapping row area between the vertically adjacent overall input segmentation areas;
[0116] The overlapping column area between the horizontally adjacent overall input segmentation areas;
[0117] The overlapping square area between the adjacent overall input segmentation areas.
[0118] According to another aspect of the embodiments of the present application, the present application provides an electronic device, as Figure 6 shown, including a memory 601, a processor 603, a communication interface 605 and a communication bus 607. A computer program that can run on the processor 603 is stored in the memory 601. The memory 601 and the processor 603 communicate through the communication interface 605 and the communication bus 607. When the processor 603 executes the computer program, the steps of the above method are implemented.
[0119] In the above electronic device, the memory and the processor communicate through a communication bus and a communication interface. The communication bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0120] The memory can include a Random Access Memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0121] The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0122] According to another aspect of the embodiments of the present application, there is also provided a computer-readable medium having non-volatile program code executable by a processor.
[0123] Optionally, in the embodiments of the present application, the computer-readable medium is configured to store program code for the processor to execute the following steps:
[0124] Save the intermediate data of each network layer of the target neural network in the edge buffer, where the intermediate data is the data obtained by calculating the overlapping part between adjacent input segmentation regions of each network layer, and the edge buffer is a region for caching the intermediate data;
[0125] When loading the input data of the target network layer into the neural network processor, load the intermediate data from the edge buffer into the neural network processor to calculate the target network layer using the neural network processor.
[0126] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and will not be elaborated herein.
[0127] In the specific implementation of the embodiments of the present application, reference may be made to the above embodiments, and corresponding technical effects are achieved.
[0128] It can be understood that these embodiments described herein can be implemented by hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.
[0129] For software implementation, the technologies described herein can be implemented by units that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented within the processor or outside the processor.
[0130] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0131] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0132] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0133] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0134] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0135] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application essentially, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0136] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A calculation method for a multi-layer neural network, characterized in that, Including: Storing the intermediate data of each network layer of the target neural network in an edge buffer, where the intermediate data is data obtained by calculating the overlapping part between adjacent input segmentation regions of each network layer, the edge buffer is a region for caching the intermediate data, and the target neural network includes a convolutional neural network; When loading the input data of the target network layer into the neural network processor, loading the intermediate data from the edge buffer into the neural network processor to calculate the target network layer by using the neural network processor, where the target network layer includes any network layer from the second network layer to the last network layer.
2. The method according to claim 1, wherein When loading the input data of the target network layer into the neural network processor, loading the intermediate data from the edge buffer into the neural network processor to calculate the target network layer by using the neural network processor includes: Loading the input segmentation region of the first network layer of the target neural network from the memory into the input buffer of the neural network processor to calculate the first network layer by using the neural network processor, obtaining the output segmentation region of the first network layer, where the edge region of the output segmentation region is stored as the intermediate data in the edge buffer, there is an overlapping region between two adjacent input segmentation regions of the first network layer, and one input segmentation region is data input for all data input channels of the target row and column; Starting from the second network layer, when loading the output segmentation region of the previous network layer into the input buffer of the neural network processor, loading the edge region of the previous network layer from the edge buffer into the input buffer of the neural network processor as the input padding region of the current network layer to calculate the current network layer by using the neural network processor, obtaining the output segmentation region of the current network layer output by the output buffer of the neural network processor, and storing the edge region in the output segmentation region of the current network layer as the input padding region of the next network layer until the output segmentation region of the last network layer is obtained; where the input data includes the output segmentation region of the previous layer.
3. The method according to claim 2, characterized in that, Before loading the input segmentation region of the first network layer of the target neural network into the input buffer of the neural network processor to calculate the first network layer by using the neural network processor, the method further includes: Determining the overall input segmentation region and the output segmentation region of each network layer in the target neural network, where the output segmentation region includes the edge region, the edge region is used to represent the overlapping region between adjacent overall input segmentation regions in the next network layer, the overall input segmentation region includes the output segmentation region of the previous layer of the target network layer and the input padding region, and when the target network layer is the first network layer, the overall input segmentation region is the input segmentation region of the first network layer.
4. The method according to claim 3, characterized in that, Determining the overall input segmentation region and the output segmentation region of each network layer in the target neural network includes: Determining the output segmentation region of the last network layer of the target neural network as the maximum segmentation region; Starting from the last network layer, using the output segmentation region of the current network layer to determine the overall input segmentation region of the current network layer according to the size of the receptive field, and using the size of the overlapping part between adjacent overall input segmentation regions of the current network layer to determine the input padding region of the current network layer, the output segmentation region of the previous network layer, and the edge region in the output segmentation region, until the input segmentation region of the first network layer is obtained; where When the size of the overall input segmentation region of any network layer is greater than the size of the input buffer of the neural network processor, reducing the size of the output segmentation region of the last network layer and re-determining the overall input segmentation region, the input padding region, the output segmentation region of the previous layer, and the edge region in the output segmentation region layer by layer.
5. The method according to claim 2, wherein After saving the edge region in the output segmentation region of the current network layer as the input padding region of the next network layer, the method further includes: When the calculation result of the current network layer does not fill the output buffer of the neural network processor, loading the next input segmentation region adjacent to the input segmentation region in the first network layer from the memory to the input buffer of the neural network processor, so as to obtain the next output segmentation region corresponding to the next input segmentation region by using the neural network processor, and starting from the second network layer again, determining the next output segmentation region corresponding to the next input segmentation region of each network layer and the next edge region in the next output segmentation region layer by layer until the calculation result of the current network layer fills the output buffer of the neural network processor, and then continuing to calculate the next network layer.
6. The method according to claim 5, wherein When the output segmentation region of the last network layer is calculated, the method further includes: Determining multiple output channels in which the last network layer can accommodate the output segmentation region according to the size of the output buffer of the neural network processor, where the multiple output channels divide the output segmentation region into multiple sub-output segmentation regions; Extracting the channels to be processed in the multiple output channels one by one, where the channels to be processed are the channels in the multiple output channels that have not undergone convolution operations; Loading the weight parameters of the channels to be processed and performing convolution operations on the sub-output buffer regions corresponding to the channels to be processed; When convolution operations are completed for each output channel, obtaining the convolution operation result of the last network layer.
7. The method according to any one of claims 3 to 4, characterized in that The edge region includes at least one of the following regions: The row region overlapping between the vertically adjacent overall input segmentation regions; The column region overlapping between the horizontally adjacent overall input segmentation regions; Square regions that overlap between adjacent overall input segmentation regions.
8. A computing device for a multi-layer neural network, characterized in that, Including: An intermediate data caching module, configured to save the intermediate data of each network layer of a target neural network in an edge buffer. Wherein, the intermediate data is data obtained by calculating the overlapping parts between adjacent input segmentation regions of each network layer, and the edge buffer is a region for caching the intermediate data. The target neural network includes a convolutional neural network; A border filling calculation module, configured to, when loading the input data of a target network layer into a neural network processor, load the intermediate data from the edge buffer into the neural network processor, so as to use the neural network processor to calculate the target network layer. The target network layer includes any network layer from the second network layer to the last network layer.
9. An electronic device, comprising a memory, a processor, a communication interface, and a communication bus. A computer program that can run on the processor is stored in the memory. The memory and the processor communicate through the communication bus and the communication interface, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7 above.
10. A computer-readable medium having non-volatile program code executable by a processor, characterized in that, The program code causes the processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Collaborative deep learning reasoning method for decentralized equipment
CN111522657A
Multi-domain cascade convolutional neural network
US20190042867A1