Data processing method, device, storage medium, and electronic device
By not performing convolution calculation when the weight of the feature map data or output channel is zero, the problem of low calculation efficiency of convolution part in the prior art is solved, and the effect of efficient acceleration and power consumption saving is achieved.
Patent Information
- Application Number
- CN201910569119.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2039-06-27
AI Technical Summary
The lack of efficient methods in the prior art to accelerate the calculation of convolutional parts, resulting in the disproportionate relationship between the computing complexity and storage complexity of deep learning algorithms, resulting in performance and power consumption bottlenecks.
By not performing convolution calculations when the weight of the feature map data or output channel is zero, and selecting one for convolution calculations when the values of multiple feature map data are the same, the multiplication and accumulation operations are reduced.
It realizes efficient acceleration of the convolutional part, reduces power consumption and improves computing efficiency.
Smart Images

Figure CN112149047B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular, to a method and apparatus for processing data, a storage medium, and an electronic device. Background Art
[0002] Artificial intelligence has been booming, but the basic architectures of existing chips such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), and FPGA (Field Programmable Gate Array) have existed long before this breakthrough in artificial intelligence and are not specifically designed for artificial intelligence. Therefore, they cannot perfectly undertake the task of implementing artificial intelligence. AI (Artificial Intelligence) algorithms are still in constant change, and a structure that can adapt to all algorithms needs to be found to make AI chips become energy-efficient general-purpose deep learning engines.
[0003] Currently, deep learning algorithms are based on multi-layer large-scale neural networks, which are essentially large computational function containing matrix products and convolution operations. Usually, a cost function including the variance of the regression problem and the cross-entropy during classification needs to be defined first, and then the data is passed into the network in batches, and the cost function value is derived according to the parameters to update the entire network model. This usually means at least millions of multiplication operations, with a huge amount of calculation. Generally speaking, it includes millions of calculations of A*B+C, consuming a huge amount of computing power. Therefore, deep learning algorithms mainly need to accelerate the convolution part, and the computing power is improved by accumulating in the convolution part. Compared with most past algorithms, the computational complexity of past algorithms was relatively high, while the relationship between the computational complexity and storage complexity of deep learning is inverted, and the performance bottleneck and power consumption bottleneck brought by the storage part are much greater than the computational part. Simply designing a convolution accelerator cannot improve the computational performance of deep learning.
[0004] It can be seen that there is currently no effective solution for how to efficiently accelerate the convolution part. Summary of the Invention
[0005] Embodiments of the present invention provide a method and apparatus for processing data, a storage medium, and an electronic device, so as to at least solve the problem in the related art that there is no effective solution for how to efficiently accelerate the convolution part in artificial intelligence.
[0006] According to an embodiment of the present invention, a method for processing data is provided, including: reading M*N feature map data of all input channels and weights of a preset number of output channels, wherein the value of M*N and the value of the preset number are respectively determined by a preset Y*Y weight; inputting the read feature map data and the weights of the output channels into a multiply-accumulate array of the preset number of output channels for convolution calculation; wherein, the convolution calculation method includes: not performing the convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple identical values of the feature map data, selecting one from the multiple identical values for the convolution calculation; outputting the result of the convolution calculation.
[0007] According to another embodiment of the present invention, a data processing device is provided, including: a reading module, configured to read M*N feature map data of all input channels and weights of a preset number of output channels, wherein the value of M*N and the value of the preset number are respectively determined by a preset Y*Y weight; a convolution module, configured to input the read feature map data and the weights of the output channels into a multiply-accumulate array of the preset number of output channels for convolution calculation; wherein, the convolution calculation method includes: not performing the convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple identical values of the feature map data, selecting one from the multiple identical values for the convolution calculation; an output module, configured to output the result of the convolution calculation.
[0008] According to still another embodiment of the present invention, a storage medium is further provided, in which a computer program is stored, and wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0009] According to still another embodiment of the present invention, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0010] Through the present invention, after reading the M×N feature map data of all input channels and the weights of a preset number of output channels, the convolution is performed in such a way that no convolution calculation is carried out when the feature map data or the weights of the output channels are zero; when there are multiple feature map data with the same value, one of the multiple identical values is selected for convolution calculation; that is to say, since if there are zero values in the feature map data and the weights, the multiplication result of these values must be 0, then this multiplication calculation and accumulation calculation can be saved to reduce power consumption. And when there are many identical values in the feature map data, there is no need to perform multiplication calculation for the subsequent identical values of the feature map data, and the result of the previous calculation can be directly used, which also reduces power consumption. Thus, the problem of how to efficiently accelerate the convolution part in artificial intelligence is solved, and the effect of efficiently accelerating the convolution part and saving power consumption is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0012] Figure 1 is a hardware structure block diagram of a terminal for a data processing method according to an embodiment of the present invention;
[0013] Figure 2 is a flowchart of a data processing method according to an embodiment of the present invention;
[0014] Figure 3 is a schematic overall design diagram according to an embodiment of the present invention;
[0015] Figure 4 is a schematic diagram of an AI processing architecture according to an embodiment of the present invention;
[0016] Figure 5 is a schematic diagram of the data flow of step S402 according to an alternative embodiment of the present invention;
[0017] Figure 6 is a schematic diagram of the data flow of step S403 according to an alternative embodiment of the present invention;
[0018] Figure 7 is a schematic diagram of the data flow of step S405 according to an alternative embodiment of the present invention;
[0019] Figure 8 is a schematic diagram of the CNN acceleration part according to an alternative embodiment of the present invention;
[0020] Figure 9 is a schematic diagram of power consumption reduction according to an embodiment of the present invention Figure 1 .
[0021] Figure 10 is a schematic diagram of power consumption reduction according to an embodiment of the present invention Figure 2
[0022] Figure 11 is a schematic structural diagram of a data processing device according to an embodiment of the present invention. Detailed implementation manners
[0023] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence.
[0025] Embodiment 1
[0026] The method embodiment provided in the first embodiment of the present application can be executed on a terminal, a computer terminal or a similar computing device. Taking running on a terminal as an example, Figure 1 is a hardware structure block diagram of a terminal of a data processing method according to an embodiment of the present invention. As Figure 1 shown, the terminal 10 may include one or more ( Figure 1 only one is shown in ) a processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal 10 may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 shown.
[0027] The memory 104 can be used to store computer programs, for example, software programs of application software and modules, such as the computer program corresponding to the data processing method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal 10 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0028] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0029] In this embodiment, a data processing method running on the above terminal is provided. Figure 2 It is a flowchart of the data processing method according to the embodiments of the present invention, as Figure 2 shown, and the process includes the following steps:
[0030] Step S202, read the feature map data of M*N of all input channels and the weights of a preset number of output channels, where the values of M*N and the preset number are respectively determined by a preset Y*Y weight; M, N, and Y are all positive integers;
[0031] Among them, if the preset Y*Y weight (weights) is 3*3 / 1*1, M*N = (15 + 2)*(9 + 2); if weights is 5*5, M*N = (15 + 4)*(25 + 4); if weights is 7*7, M*N = (15 + 6)*(49 + 6); if weights is 11*11, M*N = (15 + 10)*(121 + 10).
[0032] If the preset weights are 3*3 / 1*1, oc_num (preset quantity) = 16; if weights are 5*5, oc_num = 5; if weights are 7*7, oc_num = 3; if weights are 11*11, oc_num = 1.
[0033] Step S204, input the read feature map data and the weights of the output channels into the multiply-accumulate array of the preset number of output channels for convolution calculation; wherein, the convolution calculation method includes: not performing convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple feature map data with the same value, select one from the multiple identical values for convolution calculation;
[0034] Step S206, output the result of the convolution calculation.
[0035] Through the above steps S202 to S206, after reading the M*N feature map data of all input channels and the weights of the preset number of output channels, the convolution method is not to perform convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple feature map data with the same value, select one from the multiple identical values for convolution calculation; that is to say, since if there are zero values in the feature map data and the weights, the multiplication result of these values must be 0, then this multiplication calculation and accumulation calculation can be saved to reduce power consumption. And when there are many identical values in the feature map data, there is no need to perform multiplication calculation for the subsequent identical feature map data values, and the result of the previous calculation can be directly used, which also reduces power consumption. Thus, the problem of how to efficiently accelerate the convolution part in artificial intelligence is solved, and the effect of efficiently accelerating the convolution part and saving power consumption is achieved.
[0036] In an alternative implementation manner of this embodiment, the method of reading the M*N feature map data of all input channels and the weights of the preset number of output channels involved in step S202 of this embodiment can be:
[0037] Step S202-11, read the M*N feature map data of all input channels and save them to the memory;
[0038] Step S202-12, read the weights of the preset number of output channels and save them to the memory.
[0039] In a specific application scenario, step S202 can be: reading the M*N feature map data of all input channels and storing it in the internal SRAM (Static Random Access Memory), and reading the weights of the oc_num output channels and storing them in the internal SRAM.
[0040] In an alternative embodiment of this embodiment, for the method of inputting the read feature map data and the weights of the output channels into the multiplication and addition array of a preset number of output channels for convolution calculation involved in step S204 of this application, it can be implemented by the following steps:
[0041] Step S1: Input the M*1 feature map data of the first input channel into the calculation array of a preset number of output channels, and use the Z*1 multiplication and addition units in the first group to perform multiplication and addition calculations to obtain Z calculation results, where Z is determined by the preset Y*Y weights;
[0042] Step S2: In the next cycle, sequentially input the M*1 feature map data of the next row into the calculation array of a preset number of output channels until the Yth cycle of the read operation, and then perform an overall replacement of all the feature map data, where the read operation is: reading the M*N feature map data of all input channels and the weights of a preset number of output channels;
[0043] Among them, step S2 further includes:
[0044] Step S21: In the next cycle, send the M*1 feature map data of the next row of the first input channel to the calculation array of a preset number of output channels, use the Z*1 multiplication and addition units in the second group to perform multiplication and addition calculations to obtain the intermediate results of the next row of Z points, and shift the feature map data of the first row to the left so that all multiplications and additions of the same output point are implemented in the same multiplication and addition unit;
[0045] Step S22: Continue to input the M*1 feature map data of the next row and perform the same processing as in step S21;
[0046] Step S23: After the Yth cycle of the read operation, continue to input the M*1 feature map data of the next row and perform the same processing as in step S21, and perform an overall replacement of all the feature map data.
[0047] Step S3, continue to input the M*1 feature map data of the next line into the computing array of the preset number of output channels, and sequentially use the next group of Z*1 multiplier-accumulator units for multiplier-accumulation calculations to obtain Z calculation results. Until after the Y*Y-th cycle of the read operation, all the multiplier-accumulation calculations of the Z data in the first row on the first input channel are completed;
[0048] Step S4, input the feature map data of the next input channel of the first input channel into the computing array, and repeat the above steps S1 to S4;
[0049] Step S5, after the Y*Y*preset number of cycles of the read operation, all the multiplier-accumulation calculations of the Z data in the first row are completed, and the calculation results are output;
[0050] Step S6, read the next M*N feature map data of all input channels, and repeat the above steps S1 to S5 until all the feature map data of all input channels are calculated.
[0051] For the above steps S1 to S6, in a specific application scenario, it can be:
[0052] Step S301, send the M*1 featuremap data of input channel0 to the computing array of oc_num output channels, and use the first group of 15*1 multiplier-accumulator units to calculate the multiplier-accumulation calculations of the first row to obtain the intermediate results of 15 points;
[0053] Among them, if weights is 3*3 / 1*1, the computing array contains 15*9 multiplier-accumulator units; if weights is 5*5, the computing array contains 15*25 multiplier-accumulator units; if weights is 7*7, the computing array contains 15*49 multiplier-accumulator units; if weights is 11*11, the computing array contains 15*121 multiplier-accumulator units.
[0054] Step S302, in the next cycle, send the M*1 featuremap data of the next row of input channel0 to the computing array of oc_num output channels, and use the second group of 15*1 multiplier-accumulator units to calculate the multiplier-accumulation calculations of the second row to obtain the intermediate results of the next row of 15 points; at the same time, shift the data register0 0~25 of the first row to the left so that all the multiplier-accumulations of the same output point are realized in the same multiplier-accumulator unit;
[0055] Step S303, continue to input the M*1 featuremap data of the next row and perform the same processing;
[0056] Step S304: After K cycles in the first step, continue to input the next line of M*1 feature map data and perform the same processing. Then, all data registers are replaced as a whole. The value of data register1 is assigned to data register0, the value of data register2 is assigned to data register1, and so on, to achieve the reuse of row data;
[0057] Step S305: Continue to input the next line of M*1 feature map data and perform the same processing as in the fourth step;
[0058] Step S306: After K*K cycles in Step S202 (it should be noted that this K*K is the same as the above Y*Y, that is, the meanings of K and Y are the same, and the same applies to K and K*K that appear below), all multiplication and addition calculations of the 15 data in the first row on input channel0 are completed. Send the M*1 feature map data of input channel1 into the computing array, and repeat Steps S301 to S306;
[0059] Step S307: After K*K*ic_num (the number of input channels) cycles in Step S202, all multiplication and addition calculations of the 15 data in the first row are completed, and output them to the DDR;
[0060] Step S308: Read the next M*N of all input channels, and repeat Steps S301 to S307 until all input channel data is processed.
[0061] The present application will be described in detail below in conjunction with the optional embodiments of the present application;
[0062] This optional embodiment provides an efficient AI processing method. This processing method analyzes the convolution algorithm, such as Figure 3As shown, the feature maps of F input channels are convolved (corresponding to F K*K weights) and accumulated to output a feature map of one output channel. When multiple output channels of feature maps are required, they are obtained by accumulating the feature maps of the same F input channels (corresponding to another F K*K weights). Then the number of times the feature map data is reused is the number of output channels. Therefore, the feature map data should be read only once as much as possible to reduce the bandwidth and power consumption requirements for DDR reading.
[0063] Since the number of multiplications and additions (i.e., computing power) is fixed, the number of output channels that can be calculated within one cycle is determined. When we increase / decrease the computing power, we can achieve the expansion and reduction of computing power by adjusting the number of output channels calculated at one time. That is, there are some 0 values in the feature map and weights, and the multiplication results of these values must be 0, so this multiplication calculation and accumulation calculation can be saved to reduce power consumption. Due to the relationship of fixed-point quantization, there are many identical values in the feature map. When encountering the same feature map values later, there is no need to perform multiplication calculations again, and the results of the previous calculation can be directly used.
[0064] Adopting this embodiment, the data stored in DDR only needs to be read once, reducing bandwidth consumption; during the calculation process, all data realizes data reuse through the shift method, reducing the power consumption of multiple SRAM reads.
[0065] Figure 4 It is a schematic diagram of the AI processing architecture according to an embodiment of the present invention, based on Figure 4 , the efficient AI processing method of this alternative embodiment includes the following steps:
[0066] Step S401, read the M*N feature map data of all input channels (if the weights are 3*3 / 1*1, M*N = (15 + 2)*(9 + 2); if the weights are 5*5, M*N = (15 + 4)*(25 + 4); if the weights are 7*7, M*N = (15 + 6)*(49 + 6); if the weights are 11*11, M*N = (15 + 10)*(121 + 10)) and store them in the internal SRAM. Read the weights of oc_num (if the weights are 3*3 / 1*1, oc_num = 16; if the weights are 5*5, oc_num = 5; if the weights are 7*7, oc_num = 3; if the weights are 11*11, oc_num = 1) output channels and store them in the internal SRAM.
[0067] Step S402, send the M*1 feature map data of input channel 0 to the computing array of oc_num output channels (if the weights are 3*3 / 1*1, the computing array contains 15*9 multiply-accumulate units; if the weights are 5*5, the computing array contains 15*25 multiply-accumulate units; if the weights are 7*7, the computing array contains 15*49 multiply-accumulate units; if the weights are 11*11, the computing array contains 15*121 multiply-accumulate units), use the first group of 15*1 multiply-accumulate units to calculate the multiply-accumulate calculation of the first row, and obtain the intermediate results of 15 points;
[0068] Among them, the data flow of step S402 is as Figure 5 shown.
[0069] Step S403, in the next cycle, send the next row of M*1 feature map data of input channel 0 to the computing array of oc_num output channels, use the second group of 15*1 multiply-accumulate units to calculate the multiply-accumulate calculation of the second row, and obtain the intermediate results of the next row of 15 points; at the same time, shift the data register 0 0~25 of the first row to the left so that all multiply-accumulations of the same output point are implemented in the same multiply-accumulate unit;
[0070] Among them, the data flow of step S403 is as Figure 6 shown.
[0071] Step S404, continue to input the next row of M*1 feature map data and perform the same processing;
[0072] Step S405, after K cycles of step S401, continue to input the next line of M*1 feature map data and perform the same processing. Then, all data registers are replaced as a whole. The value of data register1 is assigned to data register0, the value of data register2 is assigned to data register1, and so on, to achieve the reuse of row data;
[0073] Among them, the data flow of step S405 is as Figure 7 shown.
[0074] Step S406, continue to input the next line of M*1 feature map data and perform the same processing as step S404;
[0075] Step S407, after K*K cycles of step S401, all the multiply-accumulate calculations of the 15 data in the first row on input channel0 are completed. Send the M*1 feature map data of input channel1 into the computing array and repeat steps S402 to S406;
[0076] Step S408, after K*K*ic_num (the number of input channels) cycles of step S401, all the multiply-accumulate calculations of the 15 data in the first row are completed, and output them to the DDR;
[0077] Step S409, read the next M*N of all input channels, and repeat steps S401 to S406 until all the input channel data are processed.
[0078] If the above steps S401 to S409 are divided into three parts and executed by three modules respectively, the three modules include: INPUT_CTRL, convolution acceleration, and OUTPUT_CTRL. Their function descriptions and corresponding steps are as follows:
[0079] A. INPUT_CTRL
[0080] Corresponding to the above step S401, this module mainly reads the featuremap and weights from the DDR through the AXI bus and stores them in the SRAM for subsequent convolutional acceleration reading. Since the SRAM space is limited, according to the different sizes of the weights, a small piece of data corresponding to all input channel featuremaps is read and stored in the SRAM, and after all the output channel data of this piece of data range is calculated, it is released, and then the next small piece of data of all input channel featuremaps is used.
[0081] B. Convolutional Acceleration
[0082] Corresponding to the above steps S402 to S407, this module mainly performs hardware acceleration on the CNN convolutional network. As Figure 8 shown, the data sent by the INPUT_CTRL is scattered into the multiply-accumulate array for convolutional calculation, and then the calculation result is returned to the OUTPUT_CTRL.
[0083] Among them, during the calculation process, the power consumption during the operation is reduced in the following two ways:
[0084] 1) Method 1, as Figure 9 shown, when the featuremap or weights are 0, no multiplication and accumulation calculations are performed;
[0085] 2) Method 2, as Figure 10 shown, when multiple data values of the featuremap are the same, only one data is multiplied, and the other data do not perform multiplication calculations and directly use the result of the multiplication of the first data;
[0086] C. OUTPUT_CTRL
[0087] Corresponding to the above steps S408 and S409, this module mainly writes all the output channel featuremap data after convolutional acceleration to the DDR through the AXI bus after arbitration and address management control for use in the next layer of convolutional acceleration.
[0088] The following takes 2160 multiply-accumulate resources and a kernel of 3*3 as an example to illustrate the efficient AI processing process of this optional implementation manner. The steps of this processing process are:
[0089] Step S501: Read the 17*11 feature map data of all input channels and store it in the internal SRAM. Read the weights of 16 output channels and store them in the internal SRAM.
[0090] Step S502: Send the 17*1 feature map data of input channel 0 to the computing array of 16 output channels, and use the first group of 15*1 multiply-accumulate units to calculate the multiply-accumulate operations of the first row, obtaining the intermediate results of 15 points.
[0091] Step S503: In the next cycle, send the next row of 17*1 feature map data of input channel 0 to the computing array of 16 output channels, and use the second group of 15*1 multiply-accumulate units to calculate the multiply-accumulate operations of the second row, obtaining the intermediate results of the next row of 15 points. At the same time, shift data register 0 0~25 of the first row to the left so that all multiply-accumulate operations for the same output point are implemented in the same multiply-accumulate unit.
[0092] Step S504: Continue to input the next row of 17*1 feature map data and perform the same processing.
[0093] Step S505: After 3 cycles of step S501, continue to input the next row of 17*1 feature map data and perform the same processing. Then, perform an overall replacement of all data registers, assign the value of data register 1 to data register 0, assign the value of data register 2 to data register 1, and so on, to achieve the reuse of row data.
[0094] Step S506: Continue to input the next row of 17*1 feature map data and perform the same processing as in step S504.
[0095] Step S507: After 9 cycles of step S501, all multiply-accumulate operations for the 15 data points in the first row on input channel 0 are completed. Send the 17*1 feature map data of input channel 1 to the computing array and repeat steps 2 to 6.
[0096] Step S508: After 2304 (when the number of input channels is 256) cycles of step S501, all multiply-accumulate operations for the 15 data points in the first row are completed, and output them to the DDR.
[0097] Step S509: Read the next 17*11 of all input channels, and repeat Steps S501 to S507 until all the input channel data are processed.
[0098] Through this optional embodiment, data storage in the DDR only requires one read, reducing bandwidth consumption; during the calculation process, all data realizes data multiplexing through the shift method, reducing the power consumption of multiple reads from the SRAM.
[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0100] Embodiment 2
[0101] In this embodiment, a data processing device is further provided. This device is used to implement the above embodiments and preferred embodiments, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the device described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0102] Figure 11 is a structural block diagram of the data processing device according to an embodiment of the present invention, as Figure 11As shown, the device includes: a reading module 92, configured to read the M*N feature map data of all input channels and the weights of a preset number of output channels, where the values of M*N and the preset number are respectively determined by a preset Y*Y weight; M, N, and Y are all positive integers; a convolution module 94, coupled to the reading module 92, configured to input the read feature map data and the weights of the output channels into a multiply-accumulate array of a preset number of output channels for convolution calculation; where the convolution calculation method includes: not performing convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple identical values in the feature map data, selecting one of the multiple identical values for convolution calculation; an output module 96, coupled to the convolution module 94, configured to output the result of the convolution calculation.
[0103] Optionally, the reading module 92 in the present application may further include: a first reading unit, configured to read the M*N feature map data of all input channels and store them in a memory; a second reading unit, configured to read the weights of a preset number of output channels and store them in the memory.
[0104] Optionally, the convolution module 94 in the present application is configured to perform the following steps:
[0105] Step S1, input the M*1 feature map data of the first input channel into the calculation array of a preset number of output channels, and use the first set of Z*1 multiply-accumulate units to perform multiply-accumulate calculation to obtain Z calculation results, where Z is determined by a preset Y*Y weight;
[0106] Step S2, in the next cycle, sequentially input the M*1 feature map data of the next row into the calculation array of a preset number of output channels until after the Yth cycle of the read operation, and perform an overall replacement of all the feature map data, where the read operation is: reading the M*N feature map data of all input channels and the weights of a preset number of output channels;
[0107] Step S3, continue to input the M*1 feature map data of the next row into the calculation array of a preset number of output channels, and sequentially use the next set of Z*1 multiply-accumulate units to perform multiply-accumulate calculation to obtain Z calculation results until after the Y*Yth cycle of the read operation, and all the multiply-accumulate calculations of the Z data in the first row on the first input channel are completed;
[0108] Step S4, input the feature map data of the next input channel of the first input channel into the calculation array, and repeat steps S1 to S4 above;
[0109] Step S5, after the Y*Y*preset number of cycles of the read operation, all the multiply-accumulate calculations of the Z data in the first row are completed, and the calculation results are output;
[0110] Step S6, read the next M*N feature map data of all input channels, and repeatedly execute the above steps S1 to S5 until the feature map data of all input channels are all calculated.
[0111] Among them, step S2 can further include the following steps:
[0112] Step S21, in the next cycle, send the next row of M*1 feature map data of the first input channel to the calculation array of a preset number of output channels, perform multiplication and addition calculations using the Z*1 multiplication and addition units in the second group, obtain the intermediate results of Z points in the next row, and shift the feature map data of the first row to the left so that all multiplications and additions of the same output point are realized in the same multiplication and addition unit;
[0113] Step S22, continue to input the next row of M*1 feature map data, and perform the same processing as step S21;
[0114] Step S23, after the Yth cycle of the read operation, continue to input the next row of M*1 feature map data, perform the same processing as step S21, and perform an overall replacement on all feature map data.
[0115] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0116] Embodiment 3
[0117] The embodiment of the present invention also provides a storage medium, in which a computer program is stored, and wherein the computer program is set to execute the steps in any one of the above method embodiments when running.
[0118] Optionally, in this embodiment, the above storage medium can be set to store a computer program for executing the following steps:
[0119] S1, read the M*N feature map data of all input channels and the weights of a preset number of output channels, where the values of M*N and the preset number are respectively determined by the preset Y*Y weights;
[0120] S2, input the read feature map data and the weights of the output channels into the multiplication and addition array of a preset number of output channels for convolution calculation; among them, the convolution calculation method includes: not performing convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple identical values of the feature map data, select one from the multiple identical values for convolution calculation;
[0121] S3. Output the result of the convolution calculation.
[0122] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0123] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0124] Optionally, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0125] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0126] S1. Read the feature map data of M*N for all input channels and the weights of a preset number of output channels, where the values of M*N and the preset number are respectively determined by a preset Y*Y weight;
[0127] S2. Input the read feature map data and the weights of the output channels into the multiply-accumulate array of a preset number of output channels for convolution calculation; among them, the convolution calculation method includes: not performing convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple feature map data with the same value, selecting one from the multiple identical values for convolution calculation;
[0128] S3. Output the result of the convolution calculation.
[0129] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0130] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be centralized on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0131] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for processing data, characterized in that, Including: Reading the feature map data of M*N for all input channels and the weights of a preset number of output channels, where the values of M*N and the preset number are respectively determined by a preset Y*Y weight, and M, N, and Y are all positive integers; Inputting the read feature map data and the weights of the output channels into a multiply-accumulate array of the preset number of output channels for convolution calculation; wherein, the convolution calculation method includes: not performing the convolution calculation when the feature map data or the weights of the output channels are zero; and when there are multiple identical values of the feature map data, selecting one from the multiple identical values for the convolution calculation; Outputting the result of the convolution calculation; Wherein, the inputting the read feature map data and the weights of the output channels into a multiply-accumulate array of the preset number of output channels for convolution calculation includes: Step S1, inputting the feature map data of M*1 for the first input channel into a calculation array of a preset number of output channels, and performing multiply-accumulate calculation using Z*1 multiply-accumulate units in the first group to obtain Z calculation results, where Z is determined by the preset Y*Y weight; Step S2, in the next cycle, sequentially input the feature map data of M*1 for the next row into the calculation array of the preset number of output channels until after the Yth cycle of the read operation, and perform an overall replacement of all the feature map data, where the read operation is: reading the feature map data of M*N for all input channels and the weights of the preset number of output channels, including: Step S21, in the next cycle, sending the feature map data of M*1 for the next row of the first input channel to the calculation array of the preset number of output channels, performing multiply-accumulate calculation using Z*1 multiply-accumulate units in the second group to obtain the intermediate results of Z points for the next row, and shifting the feature map data of the first row to the left so that all multiply-accumulations for the same output point are implemented in the same multiply-accumulate unit; Step S22, continue to input the feature map data of M*1 for the next row and perform the same processing as in Step S21; Step S23, after the Yth cycle of the read operation, continue to input the feature map data of M*1 for the next row and perform the same processing as in Step S21, and perform an overall replacement of all the feature map data; Step S3, continue to input the feature map data of M*1 for the next row into the calculation array of the preset number of output channels and sequentially perform multiply-accumulate calculation using the next group of Z*1 multiply-accumulate units to obtain Z calculation results until after the Y*Yth cycle of the read operation, and all the multiply-accumulate calculations for the Z data in the first row on the first input channel are completed; Step S4, input the feature map data of the next input channel of the first input channel into the calculation array and repeat Steps S1 to S4 above; Step S5, after performing the read operation for Y*Y*preset number of cycles, all the multiply-accumulate calculations for the Z data in the first row are completed, and the calculation results are output; Step S6, read the next M*N feature map data of all input channels, and repeat the above steps S1 to S5 until the feature map data of all input channels are calculated.
2. The method according to claim 1, wherein The reading of the M*N feature map data of all input channels and the weights of a preset number of output channels includes: Read the M*N feature map data of all input channels and save them in the memory; Read the weights of a preset number of output channels and save them in the memory.
3. A data processing device, characterized in that, It includes: A reading module for reading the M*N feature map data of all input channels and the weights of a preset number of output channels, where the values of M*N and the preset number are respectively determined by the preset Y*Y weights, and M, N, and Y are all positive integers; A convolution module for inputting the read feature map data and the weights of the output channels into the multiply-accumulate array of the preset number of output channels for convolution calculation; where the convolution calculation method includes: not performing the convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple identical values of the feature map data, selecting one from the multiple identical values for the convolution calculation; An output module for outputting the result of the convolution calculation; Among them, the convolution module is used to perform the following steps: Step S1, input the M*1 feature map data of the first input channel into the calculation array of the preset number of output channels, and use the first group of Z*1 multiply-accumulate units for multiply-accumulate calculation to obtain Z calculation results, where Z is determined by the preset Y*Y weights; Step S2, in the next cycle, sequentially input the M*1 feature map data of the next row into the calculation array of the preset number of output channels until after the Yth cycle of the read operation, and perform an overall replacement of all feature map data, where the read operation is: read the M*N feature map data of all input channels and the weights of a preset number of output channels, including: Step S21, in the next cycle, send the M*1 feature map data of the next row of the first input channel to the calculation array of the preset number of output channels, use the second group of Z*1 multiply-accumulate units for multiply-accumulate calculation to obtain the intermediate results of Z points in the next row, and shift the feature map data of the first row to the left so that all multiply-accumulates of the same output point are realized in the same multiply-accumulate unit; Step S22, continue to input the M*1 feature map data of the next row and perform the same processing as in Step S21; Step S23, after the Yth cycle of the read operation, continue to input the M*1 feature map data of the next row and perform the same processing as in Step S21, and perform an overall replacement of all feature map data; Step S3, continue to input the M*1 feature map data of the next row into the calculation array of the preset number of output channels, and sequentially use the next group of Z*1 multiply-accumulate units for multiply-accumulate calculation to obtain Z calculation results until after the Y*Yth cycle of the read operation, and all multiply-accumulate calculations of the Z data in the first row on the first input channel are completed; Step S4: Input the feature map data of the next input channel of the first input channel into the computing array, and repeatedly execute the above steps S1 to S4; Step S5: After executing the read operation for Y*Y* a preset number of cycles, all the multiply-accumulate calculations of the Z data in the first row are completed, and the calculation results are output; Step S6: Read the next M*N feature map data of all input channels, and repeatedly execute the above steps S1 to S5 until the feature map data of all input channels are calculated.
4. The device according to claim 3, characterized in that, The reading module includes: A first reading unit for reading the M*N feature map data of all input channels and storing them in the memory; A second reading unit for reading the weights of a preset number of output channels and storing them in the memory.
5. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is set to execute the method described in any one of claims 1 to 2 when running.
6. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is set to run the computer program to execute the method described in any one of claims 1 to 2.
Citation Information
Patent Citations
Method and system for deep learning algorithm acceleration on field-programmable gate array platform
CN106228238A
Method for achieving and executing neural network and computer readable medium
CN107392305A
Cited By
Data processing method and device, storage medium and electronic device
WO2020259031A1