Convolution computing method and apparatus, and electronic device and storage medium
By dividing the data into sub-data and storing and computing it in segments between registers and computing units, the problem of low efficiency in traditional convolution operations is solved, and more efficient convolution operations are achieved.
Patent Information
- Application Number
- PCT/CN2024/130154
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2024-11-06
- Publication Date
- 2026-02-12
AI Technical Summary
Traditional convolutional neural networks are inefficient when performing convolution operations, resulting in a waste of computing resources.
The data to be processed is divided into multiple sub-data and stored in registers respectively. The sub-data is input into the computation unit through a selector and a buffer queue. The computation unit to be operated on is determined according to the convolution kernel, and elements and parameters that do not participate in the operation are filtered out.
It improves the efficiency and accuracy of convolution operations and avoids unnecessary waste of computing resources.
Smart Images

Figure CN2024130154_12022026_PF_FP_ABST
Abstract
Description
Convolution calculation method and device, electronic equipment and storage medium TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to a convolution calculation method and device, electronic equipment and storage medium.
[0002] The present application claims priority to the Chinese patent application No. 202411100109.2, filed on August 9, 2024, and entitled "Convolution calculation method and device, electronic equipment and storage medium", the content of which is incorporated herein by reference in its entirety. BACKGROUND
[0003] At present, with the development of data science, deep learning technology has received more and more extensive attention. Among them, the convolutional neural network has attracted widespread attention in the field of processing image data and signal data, and the convolutional neural network has a relatively ideal performance for large image processing. And compared with the traditional image analysis algorithm, the convolutional neural network model avoids the complex pre-processing of the image, and can directly receive and process the original image, so it is more widely used.
[0004] However, in the process of convolution operation of the traditional convolutional neural network, a plurality of parameters in the convolution kernel are usually used to perform convolution operation with a plurality of elements in the data to be processed, and then the corresponding convolution operation result is obtained. In the process of convolution operation, each parameter in the convolution kernel and each element need to participate in the convolution operation, so this convolution operation method may have the problem of low efficiency. TECHNICAL PROBLEM
[0005] In view of the above, it is necessary to provide a convolution calculation method, device, electronic equipment and storage medium to solve the technical problem of low efficiency of convolution calculation.
[0006] The present application provides a convolution calculation method applied to an electronic device, the method comprising: dividing the data to be processed into a plurality of sub-data according to a preset convolution kernel; inputting a plurality of corresponding sub-data to a plurality of preset registers, each sub-data corresponding to a register; inputting a plurality of elements included in each sub-data from the register to a plurality of calculation units, wherein each calculation unit corresponds to an element; determining a calculation unit to be operated from the plurality of calculation units according to the convolution kernel; using the calculation unit to be operated to calculate the element corresponding to the calculation unit to be operated and the convolution kernel, and determining the convolution operation result of each sub-data; and determining the convolution operation result of the data to be processed based on the convolution operation result corresponding to each sub-data in the plurality of sub-data.
[0007] The embodiment of the present application further provides a convolution calculation device, which comprises: a division module configured to divide to-be-processed data into a plurality of sub-data according to a preset convolution kernel; a cache module configured to input the plurality of sub-data into a plurality of preset registers, each of the sub-data corresponding to one register; the cache module is further configured to input a plurality of elements included in each of the sub-data from the register into a plurality of calculation units, each of the calculation units corresponding to one of the elements; an operation module configured to determine a to-be-operated calculation unit from the plurality of calculation units according to the convolution kernel; the operation module is further configured to perform calculation on the element corresponding to the to-be-operated calculation unit and the convolution kernel by using the to-be-operated calculation unit, to determine a convolution operation result of each of the sub-data; and the operation module is further configured to determine a convolution operation result of the to-be-processed data based on the convolution operation result corresponding to each of the plurality of sub-data.
[0008] The embodiment of the present application further provides an electronic device, which comprises: a memory configured to store at least one instruction; and a processor configured to execute the instruction stored in the memory to implement the convolution calculation method.
[0009] The embodiment of the present application further provides a computer storage medium, which stores a computer program, and the computer program is executed by a processor to implement the convolution calculation method.
[0010] As can be seen from the above technical solutions, the embodiment of the present application divides the to-be-processed data into a plurality of sub-data, and stores each of the sub-data in a corresponding preset register, so that different parts of the to-be-processed data can be stored in fragments, thereby improving the efficiency of the electronic device in reading or modifying the to-be-processed data, and the convolution operation can be performed on a certain part of the to-be-processed data each time, thereby improving the efficiency of the convolution operation. Furthermore, a plurality of elements included in each of the sub-data are input from the register into a plurality of calculation units, a to-be-operated calculation unit is determined from the plurality of calculation units according to the convolution kernel, and the element corresponding to the to-be-operated calculation unit and the convolution kernel are calculated by using the to-be-operated calculation unit, to obtain the convolution operation result of each of the sub-data. In this way, elements in the to-be-processed data that do not participate in the convolution operation and parameters in the convolution kernel that do not participate in the convolution operation can be screened out in the process of the convolution operation, so that all parameters and elements do not participate in the convolution operation, thereby improving the efficiency of the convolution operation while ensuring the accuracy of the convolution operation. BRIEF DESCRIPTION OF DRAWINGS
[0011] FIG. 1 is an application scenario diagram of a convolution calculation method according to an embodiment of the present application.
[0012] FIG. 2 is a flowchart of a convolution calculation method according to an embodiment of the present application.
[0013] FIG. 3 is a schematic diagram of determining sub-data according to an embodiment of the present application.
[0014] FIG. 4 is a schematic diagram of convolution operation according to an embodiment of the present application.
[0015] FIG. 5 is a flowchart of a method of storing sub-data into a register according to an embodiment of the present application.
[0016] FIG. 6 is a schematic diagram of a preset register according to an embodiment of the present application.
[0017] FIG. 7 is a flowchart of a method of inputting sub-data into a plurality of calculation units according to an embodiment of the present application.
[0018] FIG. 8 is a schematic diagram of a hole convolution operation according to an embodiment of the present application.
[0019] FIG. 9 is a schematic diagram of a plurality of calculation units according to an embodiment of the present application.
[0020] FIG. 10 is a functional block diagram of a convolution calculation device according to an embodiment of the present application.
[0021] FIG. 11 is a schematic diagram of an electronic device according to an embodiment of the present application. Embodiments of the present application
[0022] In order to more clearly understand the purpose, features and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict. In the following description, a large number of specific details are set forth in order to facilitate a full understanding of the present application, and the described embodiments are only a part of the embodiments of the present application, but not all the embodiments.
[0023] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0025] The embodiment of the present application provides a convolution calculation method, which can be applied to one or more electronic devices. The electronic device is a device capable of automatically performing numerical calculation and / or information processing according to a pre-set or stored instruction. The hardware of the electronic device includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0026] The electronic device can be any electronic product capable of human-computer interaction with a user, for example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive internet protocol television (IPTV), a smart wearable device, etc.
[0027] The electronic device can further include a network device and / or a client device. The network device includes but is not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0028] The network in which the electronic device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0029] The convolution operation is a linear operation, and the convolution operation also has translational invariance, that is, if the input signal is shifted, the result of the convolution operation will also be shifted accordingly. The convolution operation can be applied to multiple data analysis fields, including image processing, signal processing and machine learning, etc. Specifically, in the field of image processing, the convolution operation can be applied to image filtering, edge detection, feature extraction, etc.; in the field of signal processing, the convolution operation can be used to analyze the spectral characteristics of a signal, to realize filtering of the signal, etc.; in the field of machine learning, the convolution operation can be applied to a convolutional neural network (CNN), and a convolution layer in the convolutional neural network extracts features of an image through the convolution operation.
[0030] For example, when convolution operation is used for feature extraction, the convolution operation can extract the edge, texture and other features in the image. By designing different convolution kernels, different features of the image can be extracted. For example, a Sobel operator can be used as a convolution kernel to extract edge information of the image. In a convolutional neural network, a convolution layer performs convolution operation with an input image through multiple convolution kernels to generate multiple feature maps, each of which contains a specific feature of the input image. Convolution operation can also be used for image filtering, such as smoothing filtering, sharpening filtering, etc. By selecting appropriate convolution kernels, smoothing, denoising, sharpening and other effects of the image can be achieved. Convolution operation can also be used for image recognition and classification. In image recognition and classification tasks, a convolutional neural network extracts high-level features of an image layer by layer through the stacking of multiple convolution layers and pooling layers, and finally realizes classification through a fully connected layer. Convolution operation plays a crucial role in this process. Convolution operation can also be applied to image segmentation tasks, and plays an important role in image segmentation tasks. Through the image features extracted by the convolutional neural network, different regions in the image can be segmented and recognized.
[0031] As shown in FIG. 1, the convolution calculation method provided by the present application can be applied to an electronic device 100, which is in communication connection with a server 200. The electronic device 100 is configured to acquire to-be-processed data from the server 200, and perform convolution operation on the to-be-processed data to obtain a corresponding convolution operation result. The electronic device 100 includes at least one register (for example, the first register 111, the second register 112 and the third register 113 shown in FIG. 1), a selector 120, a cache queue 130 and at least one calculation unit (for example, the first calculation unit 140, the second calculation unit 150 and the third calculation unit 160 shown in FIG. 1). Specifically, the at least one register divides the to-be-processed data into a plurality of sub-data according to a preset convolution kernel; stores the plurality of sub-data in the register (for example, the first register 111, the second register 112 and the third register 113 shown in FIG. 1), wherein each sub-data corresponds to one register; the selector 120 periodically determines the corresponding sub-data from the at least one register; and inputs the corresponding sub-data to the cache queue 130. The electronic device 100 further inputs a plurality of elements in the sub-data from the cache queue 130 to the at least one calculation unit (for example, the first calculation unit 140, the second calculation unit 150 and the third calculation unit 160 shown in FIG. 1), wherein each calculation unit corresponds to one element; the electronic device 100 further determines a to-be-operated calculation unit from the plurality of calculation units (for example, the first calculation unit 140, the second calculation unit 150 and the third calculation unit 160 shown in FIG. 1) according to a parameter of the convolution kernel; and determines a convolution operation result according to the element corresponding to the to-be-operated calculation unit and the parameter of the convolution kernel. The electronic device can be a DMA (Direct Memory Access) device, which is a hardware device configured to exchange data with a system memory directly through a CPU, and can reduce the burden of the CPU and improve the overall performance of the system.
[0032] As shown in FIG. 2, it is a flowchart of a convolution calculation method provided by an embodiment of the present application. The order of the steps in the flowchart can be changed according to different requirements, and some steps can be omitted. The convolution calculation method provided by the embodiment of the present application includes the following steps.
[0033] S20, dividing the to-be-processed data into a plurality of sub-data according to a preset convolution kernel.
[0034] In an embodiment of the present application, the data to be processed can be any data that needs to be processed by convolution operation. For example, the data to be processed can be a grayscale image, each pixel point of the grayscale image includes a luminance value (an integer between 0 and 255), and the convolution operation can be used for edge detection, blur processing or sharpening processing of the grayscale image; the data to be processed can also be a digital image with color (for example, an RGB image), each pixel point of which includes three pixel values representing the values of red, green and blue color channels. Specifically, the convolution operation can be applied to each color channel respectively, and the data of each color channel can be processed by the convolution operation respectively; the data to be processed can also be audio data, which can be one-dimensional time series data, and the convolution operation can be used for audio filtering (for example, low-pass filtering, high-pass filtering, band-pass filtering, etc.), echo cancellation, noise reduction, etc.; the data to be processed can also be video data, which can be a sequence of digital images. Specifically, the convolution operation can be applied to each frame of digital image in the video data, and the convolution operation can also be used to process multiple digital images in the video data in the time dimension, thereby realizing video filtering, motion detection, video stabilization, etc.; the data to be processed can also be one-dimensional or multi-dimensional time series data collected by various sensors (such as accelerometers, gyroscopes, temperature sensors, etc.), and the convolution operation can be used for feature extraction, pattern recognition or filtering processing of the one-dimensional or multi-dimensional time series data; the data to be processed can also be medical image data (such as MRI images or CT scan images, etc.), which is usually two-dimensional or three-dimensional, and the convolution operation is very important in medical image processing, which is used for image enhancement, lesion detection, tissue segmentation, etc.; the data to be processed can also be text data, and the convolution operation can be used for natural language processing (NLP) of the text data, such as text classification, sentiment analysis, etc. Specifically, the text data is converted into a word embedding matrix or other forms of vector representation, and then the convolution operation is applied to extract features.
[0035] In an embodiment of the present application, in order to ensure that the processing of the to-be-processed data is more flexible, thereby improving the efficiency of the convolution operation on the to-be-processed data, the to-be-processed data can be divided into a plurality of sub-data according to a preset convolution kernel, wherein each sub-data corresponds to a clock cycle. Subsequently, one of the sub-data can be selected for convolution operation in each clock cycle, which can ensure the flexibility of the convolution operation. The preset convolution kernel can be a convolution kernel used for convolution operation, which is a matrix smaller than the dimension of the to-be-processed data, and its size is usually an odd number x an odd number (for example, 3 x 3, 5 x 5, etc.), which is used to divide the to-be-processed data into sub-data with the same size as the convolution kernel, and then to realize the convolution operation according to the parameters in the convolution kernel and the elements in the sub-data. The clock cycle is the inverse of the operating frequency of the electronic device, which is used to represent the time required for the electronic device to perform one convolution operation. For example, when the operating frequency of the electronic device is 2.4 GHz, it means that the electronic device performs 2.4*10 9 -10 -10
[0036] In an embodiment of the present application, in order to determine in advance the to-be-processed data to be processed by the electronic device in each clock cycle, thereby adjusting the data participating in the convolution operation in real time to improve the efficiency of the convolution operation, the to-be-processed data can be divided into a plurality of sub-data, specifically, the to-be-processed data is divided into a plurality of sub-data according to a preset convolution kernel, including: determining the size of the convolution kernel; dividing the to-be-processed data based on the size and a preset sliding direction to obtain a plurality of sub-data; wherein the size of each sub-data is the same as the size of the convolution kernel, and any two sub-data include different to-be-processed data.
[0037] In an embodiment of the present application, the to-be-processed data is divided based on the size and a preset sliding direction to obtain a plurality of sub-data, including: determining the corresponding position of the convolution kernel in the to-be-processed data according to the preset sliding direction and the size; determining the data at the position of the to-be-processed data as a sub-data. For example, when the to-be-processed data is two-dimensional data and the preset sliding direction is from the leftmost side to the right side of the to-be-processed data, the sliding mode is to determine the to-be-processed data with the same dimension as the convolution kernel as the first sub-data from the first dimension of the to-be-processed data; determine the to-be-processed data adjacent to the first sub-data on the right side of the first sub-data and with the same dimension as the convolution kernel as the second sub-data, and so on to determine the remaining sub-data.
[0038] As shown in FIG. 3 is a schematic diagram of determining sub-data. For example, when the size of the convolution kernel 310 is 3*3, when the data to be processed is a two-dimensional matrix 320, and the size of the data to be processed is 9*9, the first sub-data is determined as 330, and the second sub-data is determined as 340.
[0039] In an embodiment of the present application, when the data to be processed is two-dimensional image data, the process of the convolution operation includes: starting from the top left corner of the image, aligning the convolution kernel with the corresponding region of the image with the same size; multiplying each parameter in the convolution kernel with the elements of the corresponding region of the image, and adding all the products to obtain the new pixel value of the center pixel of the region; after completing one convolution operation, determining the corresponding region of the convolution kernel in the image for the next convolution operation according to the preset step size and sliding direction. The preset step size can be the number of columns of the convolution kernel, or other numerical values. For example, when the preset step size is 1 and the sliding direction is to the right, the region of the image participating in the last convolution operation with the same size as the convolution kernel can be moved one column to the right to determine the region of the image participating in the current convolution operation, and the above calculation process is repeated until all regions of the image are traversed; after the above operation, a new image matrix is obtained, which is the result of the convolution operation. Specifically, the parameters of the convolution kernel will affect the result of the convolution operation, and different convolution kernels can be used to achieve different image processing effects. For example, when performing convolution operation, the boundary problem of the image needs to be handled. Common boundary handling methods include padding (such as zero padding, edge padding, etc.) and ignoring the boundary region. Step size setting, the step size is the distance that the convolution kernel moves on the image. The larger the step size, the smaller the dimension of the data obtained by the convolution operation, and the smaller the step size, the larger the dimension of the data obtained by the convolution operation.
[0040] As shown in FIG. 4 is a schematic diagram of convolution operation. 410 is a convolution kernel, wherein K11 to K33 are parameters of the convolution kernel; 420 is data to be processed, wherein A11 to A33 are data participating in the convolution operation for the first time in the data to be processed; 430 is feature data, wherein B11 is feature data 430 obtained by performing convolution operation on the data to be processed 420 based on the convolution kernel 410.
[0041] S21, inputting a plurality of sub-data to a plurality of preset registers, each sub-data corresponding to a register.
[0042] In an embodiment of the present application, since the electronic device needs to frequently access the to-be-processed data in the process of performing convolution operation, in order to improve the efficiency of the electronic device in accessing the to-be-processed data, after the to-be-processed data is divided into a plurality of sub-data according to the preset convolution kernel, the plurality of sub-data can be stored in the register. The register is a high-speed storage unit in the electronic device, which is used to temporarily store data, instructions or address information. In the process of executing instructions by the central computing unit, the data needs to be frequently accessed and modified, and the register can be used to quickly access these data, and the response speed of the central computing unit accessing the register is faster than that of accessing the memory (RAM), thereby improving the efficiency of the electronic device in performing data processing or instruction processing.
[0043] In an embodiment of the present application, each register corresponds to a number, which can be used to represent the order of the register being accessed by the central processor, and the smaller the number is, the earlier the access time is. In order to ensure that the order of accessing the register to obtain the sub-data is consistent with the order of determining the sub-data in the to-be-processed data according to the convolution kernel, the registers are arranged in order from small to large, and the sub-data in the to-be-processed data is extracted in the direction from left to right, and the extracted sub-data is input into the register. For specific method of storing sub-data into the register, please refer to the corresponding flowchart of FIG. 5.
[0044] As shown in FIG. 6 is a schematic diagram of the preset register. When the sub-data in the to-be-processed data are sequentially sub-data 1, sub-data 2 and sub-data 3 according to the position in the to-be-processed data, the number of the register 610 is Reg_Group0, and the corresponding sub-data is sub-data 1; the number of the register 620 is Reg_Group1, and the corresponding sub-data is sub-data 2; the number of the register 630 is Reg_Group2, and the corresponding sub-data is sub-data 3.
[0045] In an embodiment of the present application, in the case of continuously performing convolution operation on a plurality of sub-data, in order to ensure that all elements in each sub-data can be processed by a plurality of computing units at the same time when performing convolution operation on each sub-data subsequently, the elements in the next sub-data can be input into the computing unit after obtaining the convolution operation result corresponding to one sub-data, so as to ensure that only the elements in one sub-data are operated at the same time, and the accuracy of the convolution operation can be improved.
[0046] S22, a plurality of elements included in each of the sub-data are input from the register to a plurality of computing units, wherein each of the computing units corresponds to one of the elements.
[0047] In a possible implementation, before the plurality of elements included in each sub-data is input to the plurality of computing units, the sub-data in the registers are stored in the cache queue in a time-first manner according to the time sequence in which the sub-data is stored in the registers, until all the sub-data in the registers are stored in the cache queue, and the sub-data is extracted from the cache queue and input to the plurality of computing units based on the first-in first-out rule, and after all the elements in the sub-data currently operated are processed, the plurality of elements in the next sub-data is input to the plurality of computing units, so as to ensure that only the elements in one sub-data are operated at the same time, and the accuracy of the convolution operation is improved.
[0048] For example, when the first determined sub-data is data 1, the second determined sub-data is data 2, and the third determined sub-data is data 3, data 1 is stored in the first register, data 2 is stored in the second register, and data 3 is stored in the third register. In the process of the convolution operation, in order to ensure that only the elements in one sub-data are operated at the same time, the sub-data in the registers can also be input to the cache queue in sequence, for example, the sub-data input to the cache queue can be determined as data 1, data 2, and data 3 in sequence. Since the cache queue is a first-in first-out queue, the order of the sub-data output from the cache queue is data 1, data 2, and data 3 in sequence.
[0049] In an embodiment of the present application, in the process of the convolution operation on the to-be-processed data, in order to determine whether each sub-data participates in the convolution operation to improve the efficiency of the convolution operation, the plurality of elements corresponding to the sub-data in the cache queue can be input to the plurality of computing units according to a preset extraction order. The preset extraction order can be a left-to-right order or a clockwise order, which is not limited herein. For details of the method of inputting the plurality of elements in the sub-data to the plurality of computing units, please refer to the corresponding detailed description of FIG. 7.
[0050] S23, determining a computing unit to be operated from the plurality of computing units according to the convolution kernel.
[0051] In an embodiment of the present application, when the convolution operation is a dilated convolution, the calculation unit to be operated can be determined from the plurality of calculation units according to the clock period and the parameters of the convolution kernel, so as to avoid the waste of computing power caused by the parameters with a value of 0 in the convolution kernel of the dilated convolution participating in the convolution operation. The dilated convolution is used to increase the receptive field in the convolutional neural network, and by injecting "holes" (i.e. zero values) into the standard convolution kernel, the actual convolution effect area (i.e. receptive field) of the convolution kernel can be increased while keeping the size of the convolution kernel unchanged. This feature enables the dilated convolution to effectively capture larger image information without increasing the amount of calculation and the number of parameters. Specifically, the implementation process of the dilated convolution includes: an operation process, setting a dilation rate: the dilation rate is a key parameter in the dilated convolution, which determines how many zero values are inserted between the elements of the convolution kernel. For example, for a 3x3 convolution kernel, if the dilation rate is 2, zero values can be inserted between the elements of the convolution kernel participating in the convolution operation, forming a larger size convolution kernel. When performing the convolution operation, the dilated convolution determines the corresponding region with the same size as the convolution kernel on the input feature map according to the preset step size and sliding direction, and sums the multiplication of the convolution kernel and the input feature map elements of the corresponding region. However, due to the existence of the dilation rate, only part of the input feature map elements actually participate in the calculation. The dilated convolution can increase the size of the receptive field, which enables the network to capture more extensive context information. The calculation of the receptive field of the convolution kernel is related to multiple parameters such as the size of the convolution kernel, the stride, the padding, and the dilation rate.
[0052] As shown in FIG. 8, which is a schematic diagram of a dilated convolution operation, 810 is a dilated convolution kernel used to perform a dilated convolution operation, wherein K11 to K55 are parameters of the dilated convolution kernel. Since the operation strategy of the dilated convolution is to insert a 0 value every other element in the dilated convolution kernel to increase the size of the receptive field of the convolution kernel, the values of a plurality of parameters in K11 to K55 are 0. Specifically, the elements with a value of 0 in the dilated convolution kernel shown in FIG. 8 include K12, K14, K21 to K25, K32, K34, K41 to K45, K52, and K54. 820 is the data to be processed, wherein A11 to A56 are data in the data to be processed that participate in the dilated convolution operation for the first time and the second time. 830 is feature data, wherein B11 is feature data obtained by performing a dilated convolution operation on the data to be processed 820 based on the dilated convolution kernel 810. Specifically, when the step of the convolution operation is the number of columns of the convolution kernel, that is, the step is 5, the first dilated convolution operation uses the dilated convolution kernel 810 to process elements A11 to A55 in the data to be processed 820 to obtain element B11 in the feature data 830, and then the dilated convolution kernel 810 is horizontally translated by a distance of 5 pixel points along the data to be processed 820, and the dilated convolution kernel 810 is used to process elements A16 to A60 (not shown in the figure) in the data to be processed 820 to obtain element B12 in the feature data 830. Or, when the step is 1, the first dilated convolution operation uses the dilated convolution kernel 810 to process elements A11 to A55 in the data to be processed 820 to obtain element B11 in the feature data 830, determines that the dilated convolution kernel 810 is horizontally translated by a distance of 1 pixel point along the data to be processed 820, and uses the dilated convolution kernel 810 to process elements A12 to A56 in the data to be processed 820 to obtain element B12 in the feature data 830. In this way, the effect of sliding window operation of the dilated convolution kernel 810 on the data to be processed 820 can be achieved.
[0053] Specifically, each time a dilated convolution operation is performed, since a 0 value is inserted every other element in the dilated convolution kernel, when the dilated convolution kernel is used to process the data to be processed, the data in the data to be processed that is opposite to the position of the 0 value in the dilated convolution kernel does not participate in the convolution calculation. As shown in FIG. 8, when the first dilated convolution operation is performed, since the elements with a value of 0 in the dilated convolution kernel 810 include K12, K14, K21 to K25, K32, K34, K41 to K45, K52, and K54, the elements in the data to be processed 820 that do not participate in the convolution operation when the first dilated convolution operation is performed include A12, A14, A21 to A25, A32, A34, A41 to A45, A52, and A54.
[0054] In an embodiment of the present application, in order to determine the correspondence between the elements in the sub-data and the calculation units in the subsequent step, and thus determine the order in which the elements in the sub-data participate in the convolution operation in the process of performing convolution operation on the data to be processed, the arrangement order of the plurality of calculation units can be determined first. As shown in FIG. 9, which is a schematic diagram of the plurality of calculation units, the first calculation unit 901, the second calculation unit 902, the third calculation unit (not shown in the figure), the fourth calculation unit (not shown in the figure), the fifth calculation unit 905 and the sixth calculation unit 906 are connected in series.
[0055] In an embodiment of the present application, the convolution kernel includes at least one parameter, each of the calculation units corresponds to a parameter of the convolution kernel, and the determining of the calculation unit to be operated from the plurality of calculation units according to the convolution kernel includes: traversing the parameters in the convolution kernel according to a preset traversal direction, and traversing the plurality of calculation units according to the preset extraction order, the position of each of the calculation units in the plurality of calculation units being the same as the position of the element corresponding to each of the calculation units in each of the sub-data; in the traversal process, in the case that the current parameter being traversed is not the first value, the calculation unit being currently traversed is determined as the calculation unit to be operated; in the case that the current parameter being traversed is the first value, it is determined that the calculation unit being currently traversed does not participate in the convolution operation. The preset traversal direction and the preset extraction order direction are the same, which can be from left to right or from top to bottom, and are not limited herein. Thus, the corresponding position of the parameter being traversed in the convolution kernel is the same as the position of the element corresponding to the calculation unit being traversed in the sub-data. For example, in the case that the size of the convolution kernel and the size of the sub-data are both 3*3, when the position of the parameter being traversed in the convolution kernel is the second row and the first column, the position of the calculation unit corresponding to the parameter being traversed in the sub-data is also the second row and the first column.
[0056] The arrangement order of the calculation units is the same as the extraction order of the plurality of elements in the extracted sub-data. For example, when the plurality of elements extracted from a certain sub-data according to the preset extraction order are 1, 2, 3, 4, 5, and 6 in sequence from front to back, if the corresponding calculation units in FIG. 9 are used to process the sub-data, the arrangement order of the plurality of calculation units is the first calculation unit 901 corresponding to the element 1, the second calculation unit 902 corresponding to the element 2, the third calculation unit (not shown in the figure) corresponding to the element 3, the fourth calculation unit (not shown in the figure) corresponding to the element 4, the fifth calculation unit 905 corresponding to the element 5, and the sixth calculation unit 906 corresponding to the element 6 in sequence from front to back. For another example, when the plurality of elements extracted from a certain sub-data according to the preset extraction order are 6, 5, 4, 3, 2, and 1 in sequence from front to back, if the corresponding calculation units in FIG. 9 are used to process the sub-data, the arrangement order of the plurality of calculation units is the sixth calculation unit 906 corresponding to the element 6, the fifth calculation unit 905 corresponding to the element 5, the fourth calculation unit (not shown in the figure) corresponding to the element 4, the third calculation unit (not shown in the figure) corresponding to the element 3, the second calculation unit 902 corresponding to the element 2, and the first calculation unit 901 corresponding to the element 1 in sequence from front to back.
[0057] For example, when there are 25 calculation units, and each calculation unit corresponds to an element in the hole convolution kernel 810 in FIG. 8, the parameter corresponding to a certain calculation unit in the convolution kernel can be 0. When the parameter corresponding to a certain calculation unit in the convolution kernel is not 0, it indicates that the sub-data corresponding to the calculation unit participates in the convolution operation, and then the calculation unit can be determined as a calculation unit to be budgeted. When the arrangement order of the parameters in the hole convolution kernel 810 is determined as a clockwise arrangement, that is, the arrangement order of the parameters is K11, K12, K13, K14, K15, K25, …, K55, K54, …, K51, …, K33 in sequence, the arrangement order of the calculation units corresponding to each parameter is the same as the arrangement order of the parameters.
[0058] In an embodiment of the present application, when the current parameter corresponding to the calculation unit being traversed is the first value, it is determined that neither the sub-data corresponding to the calculation unit nor the parameter in the convolution kernel participates in the convolution operation. For example, when the parameter corresponding to a certain calculation unit is 0, it indicates that the element corresponding to the calculation unit does not participate in the convolution operation, and then the calculation unit can be skipped in the process of performing the convolution operation on the data to be processed. In this way, the accuracy of the hole convolution operation on the data to be processed can be ensured, and the efficiency of the convolution operation can be improved because some elements do not participate in the convolution operation.
[0059] S24, using the to-be-operated calculation unit, calculating the element corresponding to the to-be-operated calculation unit and the convolution kernel to determine the convolution operation result of each sub-data.
[0060] In an embodiment of the present application, the to-be-operated calculation unit is used to calculate the element corresponding to the to-be-operated calculation unit and the parameter corresponding to the element in the convolution kernel, so that the product between the convolution kernel parameter corresponding to each calculation unit and the element can be determined, and the mean of all the products is determined as the feature data (for example, B11 in FIG. 8) obtained after one convolution operation of one sub-data of the to-be-processed data in the current clock cycle using the convolution kernel.
[0061] S25, determining the convolution operation result of the to-be-processed data based on the convolution operation result corresponding to each sub-data in the plurality of sub-data.
[0062] In an embodiment of the present application, after the convolution operation result corresponding to each sub-data is determined, the operation result corresponding to each sub-data is arranged according to the position of the sub-data in the to-be-processed data to obtain the operation result corresponding to the to-be-processed data. For example, as shown in FIG. 4, the convolution operation result corresponding to the to-be-processed data 420 is the feature data 430; as shown in FIG. 8, the convolution operation result corresponding to the to-be-processed data 820 is the feature data 830.
[0063] As can be seen from the above technical solutions, in the embodiments of the present application, the to-be-processed data is divided into a plurality of sub-data, and each sub-data is stored in a corresponding preset register, so that different parts of the to-be-processed data can be stored in fragments, thereby improving the efficiency of the electronic device in reading or modifying the to-be-processed data, and the convolution operation can be performed on a certain part of the to-be-processed data each time, thereby improving the efficiency of the convolution operation. Then, a plurality of elements included in each sub-data are input from the register to a plurality of calculation units, a to-be-operated calculation unit is determined from the plurality of calculation units according to the convolution kernel, and the element corresponding to the to-be-operated calculation unit and the convolution kernel are calculated using the to-be-operated calculation unit to obtain the convolution operation result of each sub-data. In this way, elements in the to-be-processed data that do not participate in the convolution operation and parameters in the convolution kernel that do not participate in the convolution operation can be screened out in the process of the convolution operation, so that all the parameters and elements do not participate in the convolution operation, thereby improving the efficiency of the convolution operation while ensuring the accuracy of the convolution operation.
[0064] As shown in FIG. 5, it is a flowchart of a method for storing sub-data into a register according to an embodiment of the present application. The order of the steps in the flowchart can be changed according to different needs, and some steps can be omitted. The method for storing sub-data into a register provided by the embodiments of the present application includes the following steps.
[0065] S50, traversing a plurality of preset registers to determine a number of the traversed register.
[0066] In an embodiment of the present application, each register corresponds to a number, which is used to represent the identity of the corresponding register, and the number can also be used to represent the degree to which the data in the register is preferentially calculated in the convolution operation. As shown in FIG. 6, the numbers of the registers can be Reg_Group0, Reg_Group1 and Reg_Group2, wherein the data stored in the register numbered Reg_Group0 is preferentially calculated, followed by the data stored in the register numbered Reg_Group1, and then the data stored in the register numbered Reg_Group2.
[0067] S51, determining the number corresponding to each sub-data according to the position of the sub-data in the to-be-processed data.
[0068] In an embodiment of the present application, the position of the sub-data in the to-be-processed data is used to represent the order in which the sub-data is generated. For example, as shown in FIG. 3, the first sub-data 330 is generated earlier than the second sub-data 340, which indicates that the first sub-data 330 is located in front of the second sub-data 340. The position of the sub-data in the to-be-processed data is also used to represent the degree to which the sub-data is preferentially processed when participating in the convolution operation. Therefore, the number corresponding to the sub-data can be determined according to the position of the sub-data in the to-be-processed data. The number can also be used to represent the degree to which the sub-data is processed.
[0069] For example, the closer the position of the sub-data in the to-be-processed data to the starting position of the convolution operation on the to-be-processed data, the higher the degree to which the sub-data is preferentially processed when participating in the convolution operation. The farther the position of the sub-data in the to-be-processed data from the starting position of the convolution operation on the to-be-processed data, the lower the degree to which the sub-data is preferentially processed when participating in the convolution operation.
[0070] As shown in FIG. 3, because the first sub-data 330 is generated earlier than the second sub-data 340, the degree to which the first sub-data 330 is preferentially processed when participating in the convolution operation is higher than the degree to which the second sub-data 340 is preferentially processed when participating in the convolution operation.
[0071] S52, storing the sub-data in the register corresponding to the number.
[0072] In an embodiment of the present application, after the correspondence between the sub-data and the number of the register is determined, the sub-data can be stored in the register corresponding to the number, so as to ensure that the to-be-processed data is stored in the register, and the efficiency of reading or modifying the to-be-processed data by the electronic device is improved, thereby improving the efficiency of performing convolution operation on the to-be-processed data.
[0073] As shown in FIG. 7, it is a flow chart of a method for inputting sub-data to a calculation unit according to an embodiment of the present application. The order of steps in the flow chart can be changed according to different requirements, and some steps can be omitted. The method for inputting sub-data to a calculation unit according to an embodiment of the present application includes the following steps.
[0074] S70, extracting an element from each of the sub-data according to a preset extraction order.
[0075] In an embodiment of the present application, after the element contained in each sub-data is determined, each element in the sub-data can be traversed according to a preset extraction order, and the traversed element is extracted. The preset extraction order can be any order, for example, when the preset extraction order is clockwise, the order of extracting elements from the sub-data shown in FIG. 4 is A11, A12, A13, A23, A33, A32, A31, A21 and A22 in turn; when the preset extraction order is from left to right in each row, the order of extracting elements from the sub-data shown in FIG. 4 is A11, A12, A13, A21, A22, A23, A31, A32 and A33 in turn.
[0076] For example, as shown in FIG. 4, the element corresponding to the parameter K11 of the convolution kernel 410 is A11 in the to-be-processed data 420, the element corresponding to the parameter K12 of the convolution kernel 410 is A12 in the to-be-processed data 420, the element corresponding to the parameter K13 of the convolution kernel 410 is A13 in the to-be-processed data 420, the element corresponding to the parameter K21 of the convolution kernel 410 is A21 in the to-be-processed data 420, the element corresponding to the parameter K22 of the convolution kernel 410 is A22 in the to-be-processed data 420, the element corresponding to the parameter K23 of the convolution kernel 410 is A23 in the to-be-processed data 420, the element corresponding to the parameter K31 of the convolution kernel 410 is A31 in the to-be-processed data 420, the element corresponding to the parameter K32 of the convolution kernel 410 is A32 in the to-be-processed data 420, and the element corresponding to the parameter K33 of the convolution kernel 410 is A33 in the to-be-processed data 420.
[0077] S71, inputting the extracted element to the calculation unit until the elements in each of the sub-data are all input to the calculation unit.
[0078] In an embodiment of the present application, the plurality of elements can be input into the calculation units according to the arrangement order of the calculation units, and the calculation units can be arranged according to the positions of the elements corresponding to the calculation units in the sub-data. Since each calculation unit corresponds to an element, and each calculation unit corresponds to a parameter of the convolution kernel, the corresponding position of the parameter corresponding to the calculation unit in the convolution kernel can be ensured to be the same as the position of the element corresponding to the calculation unit in the sub-data, so that whether the element in the calculation unit participates in the convolution operation can be determined according to the parameter of the convolution kernel, and the efficiency of the convolution operation can be improved.
[0079] Referring to FIG. 10, FIG. 10 is a functional module diagram of a convolution calculation device according to an embodiment of the present application. The convolution calculation device 101 includes a division module 1010, a cache module 1011, and an operation module 1012. The module / unit referred to in the present application refers to a series of computer readable instructions capable of being executed by the processor 13 and capable of completing a fixed function, which is stored in the memory 12. In the present embodiment, the functions of the modules / units will be described in detail in subsequent embodiments.
[0080] The division module 1010 is configured to divide the to-be-processed data into a plurality of sub-data according to a preset convolution kernel.
[0081] The cache module 1011 is configured to input the plurality of sub-data into a plurality of registers.
[0082] The cache module 1011 is further configured to input a plurality of elements included in each of the sub-data from the registers to a plurality of calculation units, wherein each of the calculation units corresponds to one of the elements.
[0083] The operation module 1012 is configured to determine a to-be-operated calculation unit from the plurality of calculation units according to the convolution kernel.
[0084] The operation module 1012 is further configured to calculate the element corresponding to the to-be-operated calculation unit and the convolution kernel by using the to-be-operated calculation unit, and determine a convolution operation result of each of the sub-data.
[0085] The operation module 1012 is further configured to determine a convolution operation result of the to-be-processed data based on the convolution operation result corresponding to each of the plurality of sub-data.
[0086] The dividing module 1010 divides the to-be-processed data into a plurality of sub-data according to a preset convolution kernel, including: determining a size of the convolution kernel; dividing the to-be-processed data based on the size and a preset sliding direction to obtain a plurality of the sub-data; wherein a size of each sub-data is the same as the size of the convolution kernel, and any two sub-data include different to-be-processed data.
[0087] The caching module 1011 inputs a plurality of elements included in each of the sub-data from the cache to a plurality of computing units, including: extracting elements from each of the sub-data according to a preset extraction order, and inputting the extracted elements to the computing units until the elements in each of the sub-data are all input to the computing units.
[0088] Please refer to FIG. 11, which is a structural schematic diagram of an electronic device provided in an embodiment of the present application. The electronic device 100 includes a memory 12 and a processor 13. The memory 12 is configured to store computer readable instructions, and the processor 13 is configured to execute the computer readable instructions stored in the memory to implement the convolution calculation method described in any of the above embodiments.
[0089] In an embodiment of the present application, the electronic device 100 further includes a bus, a computer program stored in the memory 12 and executable on the processor 13, such as a convolution calculation program.
[0090] FIG. 11 only shows the electronic device 100 with the memory 12 and the processor 13, and those skilled in the art can understand that the structure shown in FIG. 11 does not constitute a limitation on the electronic device 100, which can include fewer or more components than shown, or combine certain components, or different component arrangements.
[0091] In combination with FIG. 2, the memory 12 in the electronic device 100 stores a plurality of computer readable instructions to implement a convolution calculation method, and the processor 13 can execute the plurality of instructions to implement: dividing to-be-processed data into a plurality of sub-data according to a preset convolution kernel; inputting a plurality of the sub-data to a plurality of preset registers, each of the sub-data corresponding to a register; inputting a plurality of elements included in each of the sub-data from the registers to a plurality of computing units, wherein each of the computing units corresponds to an element; determining to-be-operated computing units from the plurality of computing units according to the convolution kernel; using the to-be-operated computing units to calculate elements corresponding to the to-be-operated computing units and the convolution kernel to determine convolution operation results of each of the sub-data; and determining a convolution operation result of the to-be-processed data based on the convolution operation result corresponding to each of the sub-data in a plurality of the sub-data.
[0092] Specifically, the processor 13 can refer to the description of the relevant steps in the corresponding embodiment of FIG. 2 for the specific implementation method of the above instructions, which will not be described here.
[0093] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 100 and does not constitute a limitation on the electronic device 100. The electronic device 100 can be a bus type structure or a star type structure. The electronic device 100 can also include more or less other hardware or software or different component arrangements, for example, the electronic device 100 can also include an input / output device, a network access device, etc.
[0094] It should be noted that the electronic device 100 is only an example. Other existing or future electronic products, such as those adaptable to the present application, should also be included in the protection scope of the present application and are hereby incorporated by reference.
[0095] The memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 12 can be an internal storage unit of the electronic device 100 in some embodiments, for example, a mobile hard disk of the electronic device 100. The memory 12 can also be an external storage device of the electronic device 100 in other embodiments, for example, a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 100. The memory 12 can be used to store application software and various data installed on the electronic device 100, for example, the code of a convolution calculation program, etc., and can also be used to temporarily store data that has been output or will be output.
[0096] The processor 13 can be composed of an integrated circuit in some embodiments, for example, can be composed of a single packaged integrated circuit or can be composed of multiple packaged integrated circuits with the same function or different functions, including one or more combinations of a central processing unit (CPU), a microprocessor, a digital processing chip, a graphics processor, and various control chips, etc. The processor 13 is the control core of the electronic device 100, which connects all components of the electronic device 100 through various interfaces and lines, executes programs or modules stored in the memory 12 (for example, executes a convolution calculation program, etc.), and calls data stored in the memory 12 to perform various functions of the electronic device 100 and process data.
[0097] The processor 13 executes an operating system of the electronic device 100 and various application programs installed. The processor 13 executes the application programs to implement the steps in each of the above-described convolution calculation method embodiments, such as the steps shown in FIG. 2.
[0098] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units can be a series of computer-readable instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the electronic device 100. For example, the computer program can be divided into a division module 1010, a cache module 1011, and an operation module 1012.
[0099] The integrated units implemented in the form of software function modules described above can be stored in a computer-readable storage medium. The software function modules described above are stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor (Processor) to execute part of the convolution calculation method described in each embodiment of the present application.
[0100] The modules / units integrated in the electronic device 100, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiment methods can also be instructed by a computer program to complete related hardware devices, and the computer program can be stored in a computer-readable storage medium. The computer program is executed by the processor, and the steps of each method embodiment described above can be implemented.
[0101] The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, and other memories.
[0102] Further, the computer-readable storage medium can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; and the data storage area can store data created according to the use of the blockchain node, etc.
[0103] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, only one arrow is used in FIG. 11, but this does not mean that there is only one bus or only one type of bus. The bus is configured to enable connection and communication between the memory 12, the at least one processor 13, and the like.
[0104] The embodiment of the present application further provides a computer readable storage medium (not shown in the figure), which stores computer readable instructions. The computer readable instructions are executed by a processor in an electronic device to implement the convolution calculation method according to any one of the above embodiments.
[0105] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other manners. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner.
[0106] The modules described as separate components can or can not be physically separated, and the components shown as modules can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. According to actual needs, some or all of the modules can be selected to achieve the purpose of the embodiment.
[0107] In addition, each functional module in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional modules.
[0108] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The plurality of units or devices stated in the specification can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not mean any particular order.
[0109] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A method of convolution computation applied to an electronic device, comprising: The method comprises: dividing the to-be-processed data into a plurality of sub-data according to a preset convolution kernel; inputting the plurality of sub-data into a plurality of preset registers, each of the sub-data corresponding to one register; inputting a plurality of elements included in each of the sub-data from the register into a plurality of calculation units, each of the calculation units corresponding to one of the elements; determining a calculation unit to be operated from the plurality of calculation units according to the convolution kernel; performing calculation on the element corresponding to the calculation unit to be operated and the convolution kernel by using the calculation unit to be operated, to determine a convolution operation result of each of the sub-data; determining a convolution operation result of the to-be-processed data based on the convolution operation result corresponding to each of the sub-data.
2. The convolution calculation method of claim 1, wherein, The method comprises: determining a size of the convolution kernel; dividing the to-be-processed data based on the size and a preset sliding direction to obtain the plurality of sub-data, wherein a size of each of the sub-data is the same as the size of the convolution kernel.
3. The method of convolution calculation of claim 1, wherein, The method comprises: extracting an element from each of the sub-data according to a preset extraction order, and inputting the extracted element into the calculation unit, until the element in each of the sub-data is input into the calculation unit.
4. The convolution calculation method according to any one of claims 1 to 3, characterized in that, The convolution kernel comprises at least one parameter; the method comprises: traversing the parameters in the convolution kernel according to a preset traversal direction, and traversing the plurality of calculation units according to the preset extraction order, wherein a position of each of the calculation units in the plurality of calculation units is the same as a position of the element corresponding to each of the calculation units in each of the sub-data; in a case where a current traversed parameter is not a first value, determining the calculation unit currently traversed as the calculation unit to be operated in the traversal process; in a case where the current traversed parameter is the first value, determining that the calculation unit currently traversed does not participate in the convolution operation.
5. The convolution calculation method according to any one of claims 1 to 3, characterized in that, The method further comprises: after dividing the to-be-processed data into the plurality of sub-data according to the preset convolution kernel, traversing the plurality of preset registers to determine a number of the register traversed; determining the number corresponding to each of the sub-data according to a position of the sub-data in the to-be-processed data; storing the sub-data into the register corresponding to the number.
6. A convolution computing device, characterized by, The device comprises a module for implementing the convolution calculation method according to any one of claims 1 to 5, and the device comprises: a division module configured to divide to-be-processed data into a plurality of sub-data according to a preset convolution kernel; a cache module configured to input the plurality of sub-data into a plurality of preset registers, each of the sub-data corresponding to one register; the cache module is further configured to input a plurality of elements included in each of the sub-data from the register into a plurality of calculation units, each of the calculation units corresponding to one of the elements; The operation module is configured to determine, according to the convolution kernel, a calculation unit to be operated from the plurality of calculation units; The operation module is further configured to perform calculation on the elements corresponding to the calculation unit to be operated and the convolution kernel by using the calculation unit to be operated, and determine a convolution operation result of each of the sub-data. The operation module is further configured to determine a convolution operation result of the to-be-processed data based on the convolution operation result corresponding to each of the plurality of sub-data.
7. The convolution calculation apparatus of claim 6, wherein The division module divides the to-be-processed data into a plurality of sub-data according to a preset convolution kernel, including: determining a size of the convolution kernel; dividing the to-be-processed data based on the size and a preset sliding direction to obtain the plurality of sub-data; wherein a size of each of the sub-data is the same as the size of the convolution kernel.
8. The convolution calculation apparatus of claim 6, wherein The cache module inputs a plurality of elements included in each of the sub-data from the register to the plurality of calculation units, including: extracting elements from each of the sub-data according to a preset extraction order, and inputting the extracted elements to the calculation units until the elements in each of the sub-data are all input to the calculation units.
9. An electronic device, comprising: The electronic device includes a processor and a memory, and the processor is configured to implement the convolution calculation method of any one of claims 1-5 when executing the computer program stored in the memory.
10. A computer storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the convolution calculation method of any one of claims 1-5.
Citation Information
Patent Citations
Data processing system and method, and medium
CN110516799A
Convolution calculation method, device and apparatus and storage medium
CN111199273A
Convolution calculation method, system and equipment and storage medium
CN113870091A
Data processing method and device, chip, electronic equipment and medium
CN115019054A