Method, apparatus, device and storage medium for implementing convolutional neural network operations

By using multiple buffers to store image and weight data in convolutional neural networks and using high parallel expansion method to perform convolution operations, the problem of computation delay and power consumption of convolutional neural networks is solved, and more efficient calculations are achieved.

CN114611683BActive Publication Date: 2025-06-24HANGZHOU WEIMING XINKE TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210141616.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-06-24
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

The existing convolutional neural network models have problems with large computational delays, and heterogeneous processors such as GPUs and ASICs have insufficient power consumption and development cycles.

Method used

By storing the image data to be processed and the weight data in multiple buffers respectively, and convolution operations are performed on the data in the buffer based on a high parallelism expansion method, the calculation time is reduced and the power consumption of the processor is reduced.

Benefits of technology

It effectively solves the problem of large computing delay of convolutional neural networks, reduces the power consumption and cost of the processor, and improves the computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611683B_ABST
    Figure CN114611683B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and storage medium for implementing convolutional neural network operations. The method includes: respectively reading the image data to be processed into a first preset number of first buffers; respectively reading the weight data of the convolutional neural network into a second preset number of second buffers; the second preset number is less than the first preset number; performing convolutional neural network operations with a third preset degree of parallelism on the image data to be processed based on the weight data; the third preset number is an integer multiple of the first preset number. The present application essentially solves the problem of large computational latency of convolutional neural networks, and greatly reduces the power consumption and cost of the processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of convolutional neural networks, and particularly relates to a method, device, equipment and storage medium for implementing convolutional neural network operations. Background Art

[0002] A convolutional neural network model (Convolutional Neural Networks, CNN) is a type of feedforward neural network that contains convolutional calculations and has a deep structure, and is one of the representative algorithms of deep learning. Due to the relatively complex calculation process and large amount of data processed by the convolutional neural network model, currently, most convolutional neural network models have the problem of calculation delay.

[0003] In the prior art, to improve the processing speed of the convolutional neural network, the convolutional neural network is usually unloaded from the CPU and accelerated by means of heterogeneous processing. The main heterogeneous processors currently used include GPUs, FPGAs, ASICs, etc. However, GPUs can only achieve instruction pipelining, cannot achieve data pipelining, and have too high power consumption. And ASICs often only support specific convolutional neural network operations, and have a long development cycle. In the current era of rapid iteration of neural network algorithms, the algorithm is often outdated when the chip is fabricated. Summary of the Invention

[0004] This application proposes a method, device, equipment and storage medium for implementing convolutional neural network operations, which essentially solves the problem of large calculation delay of convolutional neural networks, and greatly reduces the power consumption and cost of the processor.

[0005] The first aspect embodiment of this application proposes a method for implementing convolutional neural network operations, and the method includes:

[0006] Read the to-be-processed image data into a first preset number of first buffers respectively;

[0007] Read the weight data of the convolutional neural network into a second preset number of second buffers respectively; the second preset number is less than the first preset number;

[0008] Perform convolutional neural network operations with a third preset degree of parallelism on the to-be-processed image data based on the weight data; the third preset number is an integer multiple of the first preset number.

[0009] In some embodiments of this application, performing convolutional neural network operations with a third preset degree of parallelism on the to-be-processed image data based on the weight data includes:

[0010] Read out corresponding numbers of to-be-processed image data from each first buffer based on the size of the preset convolutional kernel;

[0011] Perform convolution operations on the read image data to be processed in parallel with a third preset number of degrees according to the weight data in the second buffer and the preset convolution kernel.

[0012] In some embodiments of the present application, performing convolution operations on the read image data to be processed in parallel with a third preset number of degrees according to the weight data in the second buffer and the preset convolution kernel includes:

[0013] According to the preset convolution kernel, respectively accumulate and calculate the image data to be processed read from each first buffer with the corresponding number of weight data to obtain a third preset number of first intermediate data;

[0014] Perform an accumulation operation on the first intermediate data according to the preset convolution kernel and the preset dependency relationship;

[0015] Obtain the convolution operation result of the image data to be processed based on the results of all accumulation operations.

[0016] In some embodiments of the present application, each buffer stores a fourth preset number of rows, the fourth preset number of rows is equal to the second preset number, and the product of the two is equal to the first preset number;

[0017] According to the preset convolution kernel, respectively accumulate and calculate the image data to be processed read from each first buffer with the corresponding number of weight data to obtain a third preset number of first intermediate data, including:

[0018] Perform multiplication calculations on each row of the image data to be processed read from each first buffer with the corresponding fifth preset number of rows of weight data; the fifth preset number is equal to the height of the preset convolution kernel.

[0019] In some embodiments of the present application, the performing an accumulation operation on the first intermediate data according to the preset convolution kernel and the preset dependency relationship includes:

[0020] Overlap the bottom row of the preset convolution kernel with the target number row of the first intermediate data, and perform an accumulation operation on each row of the target number row of the image data to be processed and the preset convolution kernel respectively;

[0021] Translate the preset convolution kernel downward by one step in sequence, and perform an accumulation operation on each row of the preset convolution kernel and the previous row of the image data to be processed in the previous operation respectively until all pixel points are traversed.

[0022] In some embodiments of the present application, the obtaining the convolution operation result of the image data to be processed based on the results of all accumulation operations includes:

[0023] For each row of pixel points, respectively determine a set of intermediate operation results corresponding to the image data to be processed in this row and the image data to be processed in multiple adjacent rows;

[0024] According to a preset rule, select one intermediate operation result from each of the intermediate operation results corresponding to the image data to be processed in each row, and form the convolution operation result of the pixel points in this row.

[0025] In some embodiments of the present application, after reading out the corresponding number of image data to be processed from each first buffer, it further includes:

[0026] Adopt the buffer pipelining method to cache the image data to be processed read out from each first buffer in sequence.

[0027] An embodiment of the second aspect of the present application provides a convolutional neural network operation implementation device, and the device includes:

[0028] A first reading module, configured to read the image data to be processed into a first preset number of first buffers respectively;

[0029] A second reading module, configured to read the weight data of the convolutional neural network into a second preset number of second buffers respectively; the second preset number is less than the first preset number;

[0030] A convolution operation module, configured to perform a convolutional neural network operation with a third preset degree of parallelism on the image data to be processed based on the weight data; the third preset number is an integer multiple of the first preset number.

[0031] An embodiment of the third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor runs the computer program to implement the method as described in the first aspect.

[0032] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method as described in the first aspect.

[0033] The technical solutions provided in the embodiments of the present application at least have the following technical effects or advantages:

[0034] The method for implementing convolutional neural network operations provided by the embodiments of the present application first stores the image data to be processed and the weight data in the first buffer and the second buffer respectively. When convolutional operations need to be performed, data can be read from the buffers, and based on a highly parallel unfolding method, highly parallel convolutional neural network operations are performed on the image data to be processed in the buffer, reducing the overall calculation time for performing convolutional neural network operations on the image data, thus essentially solving the problem of large computational latency in convolutional neural networks, and greatly reducing the power consumption and cost of the processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.

[0036] In the drawings:

[0037] Figure 1 A flowchart showing the process of a traditional convolutional neural network operation is shown;

[0038] Figure 2 A flowchart showing the method for implementing convolutional neural network operations proposed by the embodiments of the present application is shown;

[0039] Figure 3 A schematic diagram showing a storage method of image data to be processed is shown;

[0040] Figure 4 A schematic diagram showing a storage method of weight data is shown;

[0041] Figure 5 A schematic diagram showing the reading of image data to be processed and weight data is shown;

[0042] Figure 6 A schematic diagram showing the overall unfolding of the convolutional neural network operation proposed by the embodiments of the present application is shown;

[0043] Figure 7 A schematic diagram showing the local unfolding and magnification of the convolutional neural network operation proposed by the embodiments of the present application is shown;

[0044] Figure 8 A schematic diagram showing the specific unfolding of the convolutional neural network operation proposed by the embodiments of the present application is shown;

[0045] Figure 9 A schematic diagram showing the structure of a device for implementing convolutional neural network operations proposed by the embodiments of the present application is shown;

[0046] Figure 10The figure shows a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0047] Figure 11 The figure shows a schematic diagram of a storage medium provided by an embodiment of the present application. Detailed implementation manners

[0048] Hereinafter, exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0049] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meanings understood by those skilled in the art to which the present application belongs.

[0050] The basic calculation process of the convolutional neural network is as Figure 1 shown (taking one-dimensional convolution as an example, the convolution kernel is 3*3, and the stride is 1). For the image data to be processed, the convolution kernel usually starts from the upper left corner of the image data to perform convolution calculation. First, it is translated one stride to the right in the horizontal direction. After traversing all the pixel points in the horizontal direction, it starts from the leftmost side again and is translated one stride downward, repeating the above horizontal movement until all the pixel points are traversed. For the convolution operation with a 3*3 convolution kernel, 9 multiplication operations need to be performed in a single convolution operation, and all the multiplication results are accumulated to obtain the result of one pixel point. It can be seen that the calculation amount of the convolution operation is relatively large, and this method for implementing the convolutional neural network operation will consume a large amount of power and processing time of the processor.

[0051] In view of the above problems, this embodiment proposes a method, device, equipment and storage medium for implementing convolutional neural network operations. The device for implementing convolutional neural network operations can be a server or a processor of the server, and the device can be implemented based on FPGA (Field Programmable Gate Array). By utilizing the advantages of infinite reprogramming and fast processing speed of FPGA, it can better execute the method for implementing convolutional neural network operations provided in this embodiment. The method for implementing convolutional neural network operations is based on a high-parallelism expansion method, and can perform high-parallelism convolution operations, reducing the overall calculation time for performing convolutional neural network operations on image data, thereby essentially solving the problem of large calculation delay of convolutional neural networks, and greatly reducing the power consumption and cost of the processor.

[0052] As Figure 2As shown, the method for implementing convolutional neural network operations provided in this embodiment can implement convolutional neural network operations based on a highly parallel unfolding method. This method may include the following steps:

[0053] Step S1: Read the image data to be processed into a first preset number of first buffers respectively.

[0054] Step S2: Read the weight data of the convolutional neural network into a second preset number of second buffers respectively.

[0055] To support this highly parallel unfolding method, this embodiment designs the storage methods for the image data to be processed and the weight data (such as the weight parameters in the feature extraction process) used in the process of performing convolutional neural network operations.

[0056] For the convenience of calculation in this embodiment, the image data to be processed and the weight data can be stored in buffers (BRAMs) respectively based on the highest parallelism. Specifically, the image data to be processed can be stored in a first preset number of first buffers (BRAMs), and the weight data can be stored in a second preset number of second buffers. Usually, this highest parallelism is often an integer multiple of the first preset number, and the weight data is relatively less than the image data to be processed. The second preset number of the second buffers for storing the weight data can be set to be less than the first preset number. For example, the first preset number can be 521, 1024, 2048, etc., and the second preset number can be 16, 32, 64, etc.

[0057] In addition, when storing the image data to be processed and the weight data, each buffer (including the first buffer and the second buffer) can store a fourth preset number of rows. The fourth preset number of rows is equal to the second preset number, and the product of the two is equal to the first preset number. This fourth preset number is usually a multiple of 8, such as 32, 64, etc.

[0058] The following takes the unfolding method with 32 input channels, 32 rows per channel, and a convolutional kernel size of 3*3 as an example. As Figure 3 shown (the data in the box in the figure is an example of image data), the image data to be processed is distributed in 1024 first BRAMs, which are the 32 rows of data for channel 0, the 32 rows of data for channel 1, until the 32 rows of data for channel 31, totaling 1024 BRAMs. The data of each channel exceeding 32 rows is stored in the subsequent space of the same BRAM.

[0059] It should be noted that the above storage methods for the image data to be processed and the weight data and the specific values of the first preset number and the second preset number are only the preferred implementation manners of this embodiment, and this embodiment is not limited thereto.

[0060] Step S3: Perform a convolutional neural network operation on the image data to be processed with a third preset number of parallelism based on the weight data.

[0061] Among them, the third preset number is usually an integer multiple of the first preset number, and this integer multiple is usually the product of the convolutional kernel sizes. For example, in this embodiment, the third preset number can be 9 (3*3) times the first preset number.

[0062] When performing a convolutional neural network operation with a third preset number of parallelism, corresponding numbers of image data to be processed can be read out from each first buffer based on the size of the preset convolutional kernel.

[0063] Similar to the storage method of the image data to be processed, the storage method of the weight data is as Figure 4 shown (the data in the box in the figure is an example of weight data). The weight data is distributed in 32 BRAMs. The weight data of channels 0 to 31 are stored in each BRAM in sequence, and the weights exceeding 32 channels are stored again from the beginning. For a preset convolutional kernel with a size of 3*3, the above storage method ensures that 32*32 data can be read out in one cycle, and 32*32*3 data can be read out after 3 cycles for parallel calculation. The read data can be cached using registers for calculation.

[0064] When performing a convolutional neural network operation, data is first read out from each buffer, as Figure 5 shown. Each time, 32 rows * 32 channels of data can be read out. After three clock cycles, 32 rows * 32 channels * 3 data can be obtained. The read data can be cached using registers for calculation. Similarly, the weight data also uses the same reading method, and parallel calculation can be performed only after the data reading is completed. Among them, the register pipelining method can be used to cache the image data to be processed read out from each first buffer in sequence.

[0065] After reading out the data, convolutional operations with a third preset number of parallelism can be performed on the read image data to be processed respectively according to the weight data in the second buffer and the preset convolutional kernel.

[0066] Specifically, in this embodiment, when performing a convolutional operation with a third preset number of parallelism, the image data to be processed read out from each first buffer can be respectively accumulated with the corresponding numbers of weight data according to the preset convolutional kernel to obtain third preset number of first intermediate data.

[0067] In this embodiment, one line of data can be read each time, and each line of image data to be processed read from each first buffer is respectively subjected to multiplication calculation with the weight data of the corresponding fifth preset number of lines. Among them, the corresponding relationship between the image data to be processed and the weight data is usually related to the size of the preset convolution kernel, and can be equal to the height size of the preset convolution kernel, that is, the fifth preset number is equal to the height of the preset convolution kernel. For example, each line of image data to be processed read out can be accumulated with each line of the three lines of weight data, so that 32*32*3*3 multiplication calculations can be completed in a single cycle. Compared with the method of performing one multiplication calculation in one cycle in the traditional convolutional neural network operation implementation method, the calculation efficiency can be greatly improved, the calculation time can be reduced, and the problem of large operation delay in the convolutional neural network can be essentially solved. Compared with the implementation method of the traditional convolutional neural network operation, the parallelism of this embodiment is 32*32*3*3 times that of it.

[0068] After obtaining the first intermediate data obtained by accumulating the above-mentioned image data to be processed and the weight data, the first intermediate data can be accumulated based on the preset convolution kernel and the preset dependency relationship, and then the convolution operation result of the image data to be processed, that is, the value of the corresponding pixel point, can be obtained based on the results of all the accumulation operations.

[0069] To accumulate the first intermediate data according to the preset convolution kernel and the preset dependency relationship, the bottom row of the preset convolution kernel can be first overlapped with the target number row of the first intermediate data, and the target number row of the image data to be processed and each row of the preset convolution kernel are respectively subjected to accumulation operations. Then, the preset convolution kernel is translated downward by one step length in turn, and the image data of the previous row of the previous operation and each row of the preset convolution kernel are respectively subjected to accumulation operations until all pixel points are traversed.

[0070] Based on the results of all the accumulation operations to obtain the convolution operation result of the image data to be processed, for each row of pixel points, a set of intermediate operation results corresponding to the image data to be processed in that row and the image data to be processed in multiple adjacent rows can be first determined respectively, and then according to the preset rule, one intermediate operation result is selected from each of the intermediate operation results corresponding to the image data to be processed in each row to form the convolution operation result of the pixel points in that row.

[0071] The preset dependency relationship can be as Figure 6 and Figure 7 shown. Taking the translation calculation of the preset convolution kernel from top to bottom as an example (due to the high parallelism of this embodiment, it can be realized that one row of pixel points can be convolved within the same convolution time, so the convolution kernel does not need to be horizontally moved, only vertically moved), the third row of the preset convolution kernel can be first overlapped with the (n-1)th row of the image data to be processed, and the three rows of this row and the preset convolution kernel are respectively subjected to accumulation operations to obtain M n-1 0, M n-1 1, Mn-1 2 three intermediate operation results. Then, shift the preset convolution kernel downward by one step. The second and third rows of the preset convolution kernel overlap with the (n - 1)-th row and the n-th row of the image. Perform cumulative addition operations on the n-th row and the three rows of the preset convolution kernel respectively to obtain M n 0, M n 1, M n 2 intermediate operation results. Continue to shift the preset convolution kernel downward by one step. The first, second, and third rows of the convolution kernel overlap with the (n - 1)-th row, the n-th row, and the (n + 1)-th row of the image. Perform cumulative addition operations on the (n + 1)-th row and the three rows of the preset convolution kernel respectively to obtain M n+1 0, M n+1 1, M n+1 2 intermediate operation results.

[0072] Among them, M n-1 0, M n 1, M n+1 The sum of 2, M

[0073] This embodiment uses the register pipelining method to cache data to ensure that there are 32 * 32 * 9 multiplication calculations in each cycle. The specific calculation process is as Figure 8 shown. The image data to be processed is cached in BRAM (buffer_img). After the data is read out, it is cached in the register pipelining method (pipelined for two beats). Then, the first three pixel data of the first row are respectively multiplied and accumulated with each row of the preset convolution kernel to obtain three intermediate operation results m00, m01, and m02. The first three pixel data of the second row are respectively multiplied and accumulated with each row of the preset convolution kernel to obtain three intermediate operation results m10, m11, and m12, and so on. The three intermediate operation results of the third row are m20, m21, and m22. According to the requirements of convolution calculation, the sum of m00, m11, and m22 is used to obtain the result of one pixel point, and the data is cached in BRAM (buffer_img_tmp). Other intermediate operation results are the components of other pixel points, and corresponding accumulations are performed according to the requirements of convolution calculation.

[0074] Complete the calculation of 32 input channels in one go. When the number of input data channels is greater than 32, it is necessary to add the data in buffer_img_tmp to the result of the current calculation. After all channels are calculated, the result is output to buffer_img for the next convolution calculation.

[0075] In addition, when calculating 32 lines of image data to be processed in one go, it is necessary to distinguish whether the current 32 lines are the first 32 lines of the entire image, the middle 32 lines of data (which may include multiple sets of middle 32 lines of data), or the last 32 lines of data. If the current 32 lines are the first 32 lines of the entire image, the convolution operation of the 32nd line depends on the intermediate operation result of the 33rd line. Therefore, the intermediate operation result needs to be cached. When calculating in the next cycle, it is accumulated to obtain the convolution operation result of the 32nd line. Similarly, the convolution operation of the 33rd line depends on the intermediate operation result of the 32nd line. Therefore, the intermediate operation result of the 32nd line of image data to be processed needs to be cached. When calculating in the next cycle, it is accumulated with other intermediate operation results to obtain the convolution operation result of the 33rd line. If the current 32 lines are the middle 32 lines of image data to be processed in the entire image, the convolution operation of the first line of the current 32 lines depends on the intermediate operation data of the 32nd line in the previous round, and the convolution operation of the 32nd line of the current 32 lines depends on the intermediate operation data of the first line in the next cycle. Therefore, as Figure 8 shown, corresponding judgments need to be made and data caching needs to be performed. In the calculation of the next cycle, the final convolution operation result is obtained. When the current 32 lines are the last 32 lines of data in the entire image, the convolution operation of the 32nd line has no other dependencies and can be directly output.

[0076] The method for implementing the convolutional neural network operation provided in this embodiment first stores the image data to be processed and the weight data in the first buffer and the second buffer respectively. When a convolution operation needs to be performed, data can be read from the buffers, and based on a highly parallel expansion method, a highly parallel convolutional neural network operation is performed on the image data to be processed in the buffer, reducing the overall calculation time for performing the convolutional neural network operation on the image data, thus essentially solving the problem of large calculation latency in the convolutional neural network, and greatly reducing the power consumption and cost of the processor.

[0077] Based on the same concept as the above method for implementing the convolutional neural network operation, this embodiment also provides a device for implementing the convolutional neural network operation, which is used to implement the method for implementing the convolutional neural network operation in any of the above embodiments, as Figure 9 shown. The device includes:

[0078] A first reading module, configured to read the image data to be processed into a first preset number of first buffers respectively;

[0079] A second reading module, configured to read the weight data of the convolutional neural network into a second preset number of second buffers respectively; the second preset number is less than the first preset number;

[0080] A convolution operation module, configured to perform a convolutional neural network operation with a third preset degree of parallelism on the image data to be processed based on the weight data; the third preset number is an integer multiple of the first preset number.

[0081] The convolutional neural network operation implementation device provided in this embodiment, based on the same concept as the above encoding method, can at least achieve the beneficial effects of the above encoding method, which will not be elaborated here.

[0082] The implementation manner of this application also provides an electronic device to execute the above convolutional neural network operation implementation method. Please refer to Figure 10 , which shows a schematic diagram of an electronic device provided in some implementation manners of this application. As Figure 10 shown, the electronic device 8 includes: a processor 800, a memory 801, a bus 802, and a communication interface 803. The processor 800, the communication interface 803, and the memory 801 are connected through the bus 802; a computer program that can run on the processor 800 is stored in the memory 801, and when the processor 800 runs the computer program, it executes the convolutional neural network operation implementation method provided in any of the foregoing implementation manners of this application.

[0083] Among them, the memory 801 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 803 (which can be wired or wireless), a communication connection is realized between this device network element and at least one other network element, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.

[0084] The bus 802 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 801 is used to store programs. After receiving an execution instruction, the processor 800 executes the program, and the convolutional neural network operation implementation method disclosed in any of the foregoing implementation manners of this application can be applied to the processor 800 or implemented by the processor 800.

[0085] The processor 800 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 800 or the instructions in the form of software. The above-mentioned processor 800 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 801, and the processor 800 reads the information in the memory 801 and combines its hardware to complete the steps of the above method.

[0086] The electronic device provided in the embodiments of the present application and the method for implementing convolutional neural network operations provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by them.

[0087] The embodiments of the present application also provide a computer-readable storage medium corresponding to the method for implementing convolutional neural network operations provided in the foregoing embodiments. Please refer to Figure 11 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the method for implementing convolutional neural network operations provided in any of the foregoing embodiments.

[0088] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0089] The computer-readable storage medium provided by the above embodiments of the present application and the method for implementing convolutional neural network operations provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0090] It should be noted that:

[0091] In the specification provided herein, a large number of specific details are set forth. However, it is understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0092] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the following schematic: that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim stands on its own as a separate embodiment of the present application.

[0093] In addition, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not others, the combination of features of different embodiments is within the scope of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0094] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for implementing convolutional neural network operations, characterized in that, The method includes: Read the image data to be processed into a first preset number of first buffers respectively; Read the weight data of the convolutional neural network into a second preset number of second buffers respectively; the second preset number is less than the first preset number; Perform a convolutional neural network operation with a third preset degree of parallelism on the image data to be processed based on the weight data, and obtain a third preset number of first intermediate data; the third preset number is an integer multiple of the first preset number; Perform a multiply-accumulate operation on the first intermediate data according to a preset convolution kernel and a preset dependency relationship; obtain the convolution operation result of the image data to be processed based on the results of all multiply-accumulate operations, including: Translate the preset convolution kernel from top to bottom, convolve the pixel points of one row within the same convolution time, overlap the third row of the preset convolution kernel with the (n - 1)-th row of the first intermediate data, and perform multiplication and accumulation operations on this row with the three rows of the preset convolution kernel respectively to obtain M n-1 0, M n-1 1, M n-1 2 three intermediate operation results; translate the preset convolution kernel downward by one step size so that the second and third rows of the preset convolution kernel overlap with the (n - 1)-th row and the n-th row of the first intermediate data respectively, and perform multiplication and accumulation operations on the n-th row with the three rows of the preset convolution kernel respectively to obtain M n 0, M n 1, M n 2 intermediate operation results; continue to translate the preset convolution kernel downward by one step size so that the first, second, and third rows of the preset convolution kernel overlap with the (n - 1)-th row, the n-th row, and the (n + 1)-th row of the first intermediate data respectively, and perform multiplication and accumulation operations on the (n + 1)-th row with the three rows of the preset convolution kernel respectively to obtain M n+1 0, M n+1 1, M n+1 2 intermediate operation results; add M n-1 0, M n 1, M n+1 2 together to obtain the result of the pixel point; where the preset convolution kernel is a 3*3 convolution kernel.

2. The method according to claim 1, characterized in that, Perform a convolutional neural network operation with a third preset degree of parallelism on the image data to be processed based on the weight data, including: Read out a corresponding number of image data to be processed from each first buffer based on the size of the preset convolution kernel; Perform a convolution operation with a third preset degree of parallelism on the read image data to be processed according to the weight data in the second buffer and the preset convolution kernel.

3. The method according to claim 2, wherein The obtaining of the third preset number of first intermediate data includes: According to the preset convolution kernel, perform a multiply-accumulate calculation on the image data to be processed read out from each first buffer respectively with a corresponding number of weight data, and obtain a third preset number of first intermediate data.

4. The method according to claim 3, characterized in that, Each buffer stores a fourth preset number of rows, the fourth preset number of rows is equal to the second preset number, and the product of the two is equal to the first preset number; Perform a multiply-accumulate operation on the image data to be processed read out from each first buffer respectively with a corresponding number of weight data according to the preset convolution kernel, and obtain a third preset number of first intermediate data, including: Perform a multiplication calculation on each row of the image data to be processed read out from each first buffer respectively with a corresponding fifth preset number of rows of weight data; The fifth preset number is equal to the height of the preset convolution kernel.

5. The method according to any one of claims 2-4, characterized in that, After reading out a corresponding number of image data to be processed from each first buffer, it further includes: Adopt a register pipelining method to cache the image data to be processed read out from each first buffer in sequence.

6. An apparatus for implementing convolutional neural network operations, characterized in that It includes: A first reading module, configured to read the image data to be processed into a first preset number of first buffers respectively; A second reading module, configured to read the weight data of the convolutional neural network into a second preset number of second buffers respectively; The second preset number is less than the first preset number; A convolution operation module, configured to perform a convolutional neural network operation with a third preset degree of parallelism on the image data to be processed based on the weight data, and obtain a third preset number of first intermediate data; The third preset number is an integer multiple of the first preset number; Perform a multiply-accumulate operation on the first intermediate data according to a preset convolution kernel and a preset dependency relationship; Obtain the convolution operation result of the image data to be processed based on the results of all multiply-accumulate operations, including: Translate the preset convolution kernel from top to bottom, convolve the pixel points of one row within the same convolution time, overlap the third row of the preset convolution kernel with the (n - 1)-th row of the first intermediate data, and perform multiplication and accumulation operations on this row with the three rows of the preset convolution kernel respectively to obtain M n-1 0, M n-1 1, M n-1 2 three intermediate operation results; translate the preset convolution kernel downward by one step size so that the second and third rows of the preset convolution kernel overlap with the (n - 1)-th row and the n-th row of the first intermediate data respectively, and perform multiplication and accumulation operations on the n-th row with the three rows of the preset convolution kernel respectively to obtain M n 0, M n 1, M n 2 intermediate operation results; continue to translate the preset convolution kernel downward by one step size so that the first, second, and third rows of the preset convolution kernel overlap with the (n - 1)-th row, the n-th row, and the (n + 1)-th row of the first intermediate data respectively, and perform multiplication and accumulation operations on the (n + 1)-th row with the three rows of the preset convolution kernel respectively to obtain M n+1 0, M n+1 1, M n+1 2 intermediate operation results; add M n-1 0, M n 1, M n+1 2 together to obtain the result of the pixel point; wherein, the preset convolution kernel is a 3*3 convolution kernel.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Neural network processor, current neural network data multiplexing method and related apparatus

    CN109740732A