Data processing method, electronic device, medium and computer program product

By blocking in parallel according to batches and output channels in the reinforcement learning network model, and selecting appropriate storage modes for data loading and computing, the problems of hardware accelerator resource occupation and computational efficiency are solved, and efficient convolutional calculation and model operation are achieved.

CN119719595BActive Publication Date: 2025-05-13LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510245690.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-13
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

In the reinforcement learning network model, the network model with convolutional computing added to the front-end increases the resource usage of the hardware accelerator and reduces the computing efficiency.

Method used

By chunking in parallel according to batch and output channels, the parallel computing architecture of block matrix multiplication is fully utilized, and appropriate storage modes (on-chip storage or off-chip memory access) are selected for loading and computing feature data and weight data.

Benefits of technology

It improves the efficiency of convolutional calculation, optimizes the calculation process, reduces the time overhead of data addressing and reading, ensures that the model runs efficiently at different scales, and reduces the complexity and resource consumption of hardware implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719595B_ABST
    Figure CN119719595B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method, electronic device, medium and computer program product, and relates to the field of computer technology, including: firstly arranging feature data blocks with batches as elements, arranging weight data blocks with output channels as elements, and being able to access feature data and weight data in a more orderly manner in convolution calculation, and then selecting a suitable storage mode by judging the network data scale and the set threshold value; when the data scale is less than the set threshold, the on-chip storage mode is used to load the data and then perform convolution calculation, which can reduce data reading delay and improve the calculation response speed; when the data scale is large, the off-chip access memory mode is used to load the data while performing convolution calculation, so as to avoid idle computing resources caused by waiting for all data to be loaded. In this way, the convolution calculation part of the reinforcement learning front end for feature extraction can be collaboratively deployed on the block matrix computing architecture, so as to realize resource reuse and effectively reduce the complexity of hardware implementation and the consumption of hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data processing method, electronic equipment, medium and computer program product. Background Art

[0002] Reinforcement learning network models often need to process massive amounts of data, and involve a large number of complex calculations in the training and inference stages. Hardware accelerators can read, write, and calculate this data. Reinforcement learning network models are usually composed of convolution, generative pre-trained transformer (GPT), and feed forward neural network (FFN), where convolution is used in the front-end state encoding part for feature extraction. For GPT networks, the main calculation is matrix multiplication, such as the implementation of block matrix multiplication. However, this network model with convolution calculation added to the front end usually uses additional convolution components, which increases the resource usage of the hardware accelerator and reduces the computing efficiency. Summary of the invention

[0003] The purpose of the present invention is to provide a data processing method, electronic device, medium and computer program product to at least solve the problem of increasing resource occupation of hardware accelerator and reducing computing efficiency in the related art.

[0004] In order to solve the above technical problems, the present invention provides a data processing method, which comprises:

[0005] Arrange the feature data blocks of the convolutional layer according to a set feature data format with batches as elements;

[0006] Arrange the weight data blocks of the convolutional layer according to the set weight data format with output channels as elements;

[0007] After the feature data blocks and weight data blocks are arranged, it is determined whether the size of the network data is less than the set threshold;

[0008] If yes, the on-chip storage mode is used to load the feature data and weight data, and the convolution calculation is performed after the loading is completed;

[0009] If not, the off-chip memory access mode is used to load the feature data and weight data, and perform convolution calculations at the same time.

[0010] In order to solve the above technical problem, the present invention further provides an electronic device, the electronic device comprising:

[0011] Memory for storing computer programs;

[0012] A processor is used to implement the steps of the above-mentioned data processing method when executing the computer program.

[0013] In order to solve the above technical problem, the present invention also provides a non-volatile storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above data processing method are implemented.

[0014] In order to solve the above technical problem, the present invention also provides a computer program product, including a computer program / instruction, which implements the steps of the above data processing method when executed by a processor.

[0015] It can be seen from the above technical scheme that the present invention provides a data processing method, which includes: arranging the feature data blocks of the convolution layer according to the set feature data format with batches as elements; arranging the weight data blocks of the convolution layer according to the set weight data format with output channels as elements; after the feature data blocks and the weight data blocks are arranged, judging whether the size of the network data is less than the set threshold; if so, using the on-chip storage mode to load the feature data and weight data, and performing convolution calculations after the loading is completed; if not, using the off-chip memory access mode to load the feature data and weight data, and performing convolution calculations at the same time.

[0016] The beneficial effect of the present invention is that the above-mentioned data processing method provided by the present invention first arranges the feature data blocks with batches as elements, which is in line with the characteristics of the neural network processing data in batches, and is conducive to continuously accessing the feature data belonging to the same batch when processing batch feature data, and improving the efficiency and cache hit rate of subsequent feature data reading; arranging the weight data blocks with output channels as elements can allow the weights to be accessed more orderly during convolution calculation, making full use of the parallel computing architecture of block matrix multiplication, speeding up the data reading and processing speed during convolution calculation, optimizing the calculation process, and reducing the time overhead of data addressing and reading. Then, by judging the network data scale and the set threshold size, a suitable storage mode is selected. When the data scale is less than the set threshold, the on-chip storage mode is used to load the data and then perform convolution calculation, which has a fast access speed, can reduce data reading delay, improve the calculation response speed, and speed up the model training or reasoning process; when the data scale is large, the off-chip access mode is used to load the data while performing convolution calculation, which can break through the on-chip storage capacity limit, ensure that the system can operate normally under large data volume, avoid idle computing resources caused by waiting for all data to be loaded, and improve the utilization rate of overall computing resources. In this way, the convolution calculation part of the reinforcement learning front end for feature extraction can be collaboratively deployed on the block matrix computing architecture to achieve resource reuse, effectively solve the adaptation problem between storage resources and computing needs, ensure that the model can run efficiently at different scales, and effectively reduce the complexity of hardware implementation and the consumption of hardware resources.

[0017] In addition, the present invention also provides corresponding electronic devices, non-volatile storage media and computer program products for the data processing method, which have the same or corresponding technical features as the above-mentioned data processing method and have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 A flowchart of a data processing method provided by an embodiment of the present invention;

[0020] Figure 2 A schematic diagram of a characteristic data block arrangement format provided by an embodiment of the present invention;

[0021] Figure 3 A schematic diagram of the arrangement format of weight data blocks provided by an embodiment of the present invention;

[0022] Figure 4 A schematic diagram of accessing characteristic data blocks of different groups of on-chip storage modes provided by an embodiment of the present invention;

[0023] Figure 5 A schematic diagram of accessing weight data blocks in different groups of on-chip storage modes provided by an embodiment of the present invention;

[0024] Figure 6 A schematic diagram of accessing output results of different groups of on-chip storage modes provided by an embodiment of the present invention;

[0025] Figure 7 A schematic diagram of a feature cache structure for an off-chip memory access mode provided by an embodiment of the present invention;

[0026] Figure 8 A schematic diagram of a characteristic cache loading timing of an off-chip memory access mode provided by an embodiment of the present invention;

[0027] Fig. 9 A schematic diagram of the structure of a convolution acceleration device provided by an embodiment of the present invention;

[0028] Fig.10 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Figure 1 A flowchart of a data processing method provided by an embodiment of the present invention, such as Figure 1 As shown, the method includes:

[0031] S101. Arrange feature data blocks of the convolutional layer according to a set feature data format with batches as elements.

[0032] It should be noted that a batch refers to a collection of data samples that are fed into a neural network for processing at one time. Feature data is a set of values ​​used to describe data features, including feature width, height, input channel, batch, etc. Feature data blocks are local data units divided from feature data. Feature data blocks can contain local feature information such as specific areas and specific channels in feature data.

[0033] Since convolution calculation requires data to be accumulated on the input channel and parallel on the output channel, but the data block arrangement format used by general matrix multiplication is M rows × M columns, in order to fully utilize the parallel computing capability of the matrix multiplication calculation array, combined with the fact that the BEV (Bird's Eye View) features in the current neural network are usually multiple batches, when executing step S101, the feature data blocks of the convolution layer are arranged in accordance with the set feature data format with batches as elements. The set feature data format with batches as elements is to organize the storage order, dimensional arrangement, etc. of the feature data in a predetermined form around the batch dimension. This method is consistent with the characteristics of neural networks processing data in batches, which is conducive to continuously accessing feature data belonging to the same batch when processing batch feature data, and improves the efficiency of subsequent feature data reading and cache hit rate.

[0034] S102, arranging the weight data blocks of the convolutional layer according to a set weight data format with output channels as elements.

[0035] It should be noted that weight data is a numerical value used to measure the strength or importance of the connection between neurons in the network, including convolution kernels, input channels, output channels, etc. The weight data block is a data set formed by grouping or dividing the weight data according to the set rules and structure. For weight data, since the input channels are accumulated and the output channels are independent, the present invention adopts the output channel parallel mode. When executing step S102, the weight data block of the convolution layer is arranged in accordance with the set weight data format with the output channel as the element. The set weight data format with the output channel as the element is to organize the arrangement of weight data in a predetermined form around the dimension of the output channel. This method can better utilize the parallel computing capabilities of the hardware. For example, when performing convolution operations, multiple output channels can be calculated simultaneously, which greatly improves the calculation speed, and each output channel can focus on extracting specific features, and the features of different channels can be more flexibly integrated in subsequent network layers.

[0036] The present invention divides blocks into batches and output channels in parallel, fully utilizes the parallel computing architecture of block matrix multiplication, accumulates input channels, and performs parallel computing between output channels.

[0037] After the feature data blocks and the weight data blocks are arranged, step S103 is executed.

[0038] S103: Determine whether the size of the network data is smaller than a set threshold.

[0039] In practical applications, the network here can be a reinforcement learning network, which has many application scenarios, such as autonomous driving. The network data scale can be the size of the amount of data processed by the network.

[0040] If yes, execute step S104; if no, execute step S105.

[0041] S104, using the on-chip storage mode to load feature data and weight data, and performing convolution calculation after the loading is completed.

[0042] It should be pointed out that the on-chip storage mode refers to the way of storing and managing data on the chip. When the present invention executes step S104, the feature data and weight data are loaded through the on-chip storage mode. After the feature data and weight data are successfully loaded into the on-chip storage, the convolution calculation can be carried out. During the calculation process, the data can be accessed quickly, the calculation response speed is improved, and the model training or reasoning process is accelerated.

[0043] S105 , using an off-chip memory access mode to load feature data and weight data, and performing convolution calculations at the same time.

[0044] It should be pointed out that the off-chip memory access mode refers to the way the chip obtains data from an off-chip storage device. When the present invention executes step S105, the feature data and weight data are gradually loaded into the cache through the off-chip memory access mode, and the data is loaded while being calculated. This utilizes the time overlap characteristics of data loading and calculation, which can improve the overall efficiency of the system and improve the utilization rate of the overall computing resources.

[0045] In the above data processing method provided by the embodiment of the present invention, firstly, the feature data blocks are arranged with batches as elements, which is in line with the characteristics of the neural network processing data in batches, and is conducive to continuously accessing the feature data belonging to the same batch when processing batch feature data, thereby improving the efficiency and cache hit rate of subsequent feature data reading; arranging the weight data blocks with output channels as elements can allow the weights to be accessed more orderly during convolution calculation, speed up the data reading and processing speed during convolution calculation, optimize the calculation process, and reduce the time overhead of data addressing and reading. Then, by judging the network data scale and the set threshold size, a suitable storage mode is selected. When the data scale is less than the set threshold, the on-chip storage mode is used to load the data and then perform convolution calculation, which has a fast access speed, can reduce data reading delay, improve the calculation response speed, and speed up the model training or reasoning process; when the data scale is large, the off-chip access mode is used to load the data while performing convolution calculation, which can break through the on-chip storage capacity limit, ensure that the system can operate normally under large data volume, avoid idle computing resources caused by waiting for all data to be loaded, and improve the utilization rate of overall computing resources. In this way, the convolution calculation part of the reinforcement learning front end for feature extraction can be collaboratively deployed on the block matrix computing architecture to achieve resource reuse, effectively solve the adaptation problem between storage resources and computing needs, ensure that the model can run efficiently at different scales, and effectively reduce the complexity of hardware implementation and the consumption of hardware resources.

[0046] Furthermore, in a specific implementation, in the above-mentioned data processing method provided in an embodiment of the present invention, step S101 arranges the feature data blocks of the convolutional layer according to a set feature data format with batches as elements, which may specifically include: organizing the feature data blocks of the convolutional layer according to a set feature data format of M batches multiplied by M input channels multiplied by the bit width of each element; wherein M is a positive integer.

[0047] In implementation, since convolution calculation requires data to be accumulated on the input channel and parallel on the output channel, the data blocks used in general matrix multiplication are generally arranged in matrix form, and the arrangement format is M rows × M columns. In the convolution operation of the convolutional neural network, the elements at the corresponding positions are accumulated for different channels of the input data, in order to combine the feature information of different channels. In the output channel dimension, the convolution calculation is processed in parallel, that is, the calculation of each output channel can be performed simultaneously. This matrix arrangement method has specific application methods and advantages in matrix multiplication calculations, but it is different from the requirements of convolution calculations for data processing in the form of data organization. In order to make full use of the parallel computing capabilities of the computing array, combined with the BEV features in the current neural network, which usually exist in the form of multiple batches, the present invention organizes the feature data blocks of the convolution layer into the form of M batches × M input channels × each element bit width. This organization method helps to adapt the parallel computing requirements of the computing array in subsequent calculations, making the convolution calculation more efficient and improving the overall computing efficiency.

[0048] Figure 2 A schematic diagram of the characteristic data block arrangement format provided by an embodiment of the present invention. Figure 2 The arrow in the middle indicates the memory depth direction. The memory depth direction refers to a dimension of the memory cell in the memory array. This direction is related to the number of memory cells and the address space. When the memory cells of the memory are arranged in a two-dimensional array, the depth direction can be understood as the number of memory cells in the row or column direction. The memory here can be a random access memory (RAM).

[0049] like Figure 2As shown, each feature data block is composed of {the 1st to Mth input channels in batch 1, the 1st to Mth input channels in batch 2, ..., the 1st to Mth input channels in batch M}, that is, each feature data block can specifically include: the data of the 1st to Mth input channels in batch 1, the data of the 1st to Mth input channels in batch 2, and so on, until the data of the 1st to Mth input channels in batch M are jointly composed. Such a feature data block organization is convenient for subsequent convolution calculations according to specific rules, and can match the structure of the calculation array. The 1st to Mth input channel data in each batch can be regarded as a row or column of elements in the matrix, so that when performing calculations such as matrix multiplication, the parallel computing capability of the calculation array can be fully utilized, so that M data can be processed in parallel, greatly improving the calculation speed and reducing the calculation time. In addition, this data block composition method uses batches as units, selects a fixed range of input channels in each batch, and provides convenience for subsequent parallel computing on the output channel. When performing convolution operations, input channel data at the same position in different batches can participate in the calculation of the convolution kernel at the same time, and the calculations corresponding to different output channels can be carried out in parallel, thereby improving the parallelism of the overall convolution calculation.

[0050] Furthermore, in the specific implementation, in the above data processing method provided by the embodiment of the present invention, the convolution calculation is performed in step S104 and / or step S105, which may specifically include: for the feature data, according to The convolution calculation is performed on the feature data block in a loop traversal mode; where K is the convolution kernel size, Cin is the number of input channels, S is the stride, C is the feature map row width, and R is the number of feature map rows; and Respectively represent the number of times the convolution kernel slides in the column direction and row direction of the feature map, Indicates rounding down; Batch is the total number of batches, Cout is the number of output channels after the convolution operation; the width of the feature buffer is set to the data bit width of a data block, and the depth of the feature buffer is set to .

[0051] In the implementation, it is assumed that the convolution kernel size is K, which represents the size of the convolution kernel in the spatial dimension. The step size is S, which determines the step size of the convolution kernel sliding on the feature map. The feature map row width (number of columns) is C, and the number of rows is R, which describes the spatial size of the feature map. The number of input channels is Cin, that is, the number of channels of the input data. The number of convolution kernel output channels is Cout, that is, the number of channels of the output features after the convolution operation. The process of traversing the feature data block of the present invention can be as follows: The circulation mode. Among them, is the spatial size of the convolution kernel. When the convolution kernel slides on the feature map, it processes one Size of the area. This is because the input channels are divided into groups of M previously, and here represents the number of input channels in each group. This expression calculates the number of times the convolution kernel can slide completely in the column direction of the feature map. It is rounded down because the convolution kernel must be guaranteed to operate within the range of the feature map. This expression calculates the number of times the convolution kernel can slide completely in the row direction, also rounded down. Indicates that the data blocks are grouped by batch, and the number of batches in each group is M. This is based on the previous setting that the data blocks are grouped by batch. The batch dimension needs to be considered during traversal calculation. This loop traversal method can make data access and calculation more orderly and efficient. The data organization method in each data block is unified. During the traversal process, it is more convenient to perform multiplication and addition operations between the convolution kernel and the data according to the established rules, reducing confusion and errors in the data processing process.

[0052] Figure 2 The A in the table represents batches 1 to M. data blocks, B represents batches M+1-2M The present invention sets the width of the feature cache to the data bit width of a data block. This design is to store the data of a data block in the cache at the same time, so as to meet the requirements of one data block. In the convolution calculation process, this width setting can ensure that enough data is obtained in the cache at one time for parallel operation, thus improving the calculation efficiency. The feature cache depth is determined by The depth is set to cache enough data to meet the data storage requirements of different stages of the entire convolution calculation process, ensuring that enough data is cached and processed when the convolution calculation is performed according to the previous loop traversal method.

[0053] Furthermore, in a specific implementation, in the above-mentioned data processing method provided in an embodiment of the present invention, step S102 arranges the weight data blocks of the convolutional layer according to a set weight data format with the output channel as an element, which may specifically include: organizing the weight data blocks of the convolutional layer according to a set weight data format of M output channels multiplied by M input channels multiplied by the bit width of each element; wherein M is a positive integer.

[0054] In implementation, in the convolution calculation, the weight data is used to perform a convolution operation on the input feature map to obtain an output feature map. For weight data, the input channels are accumulated and the output channels are independent. Input channel accumulation refers to adding the convolution results of the corresponding positions on each channel of the input feature map and the weights when calculating the value of a certain position of the output feature map. For example, in a convolution layer, the input feature map has multiple channels (assuming Cin channels), and for each pixel of the output feature map, the convolution results of the corresponding positions on these Cin channels need to be accumulated. Output channel independence means that the calculation of each output channel is independent of each other, and different output channels have their own independent set of weights. Each output channel corresponds to a convolution kernel group, and these convolution kernel groups are convolved with the input feature map to obtain their own output feature map channels. Based on this, in order to improve the calculation efficiency, the present invention utilizes the independent characteristics of the output channels to calculate multiple output channels at the same time, that is, the output channels are parallelized to organize the weight data blocks into the form of M output channels × M input channels × each element bit width. This parallel computing can make full use of hardware resources and speed up convolution operations.

[0055] Figure 3 The following is a schematic diagram of the arrangement format of weight data blocks provided by an embodiment of the present invention. Figure 3 As shown, each weight data block is composed of {the 1st-Mth input channels in output channel 1, the 1st-Mth input channels in output channel 2, ..., the 1st-Mth input channels in output channel M}, that is, the 1st-Mth input channel data corresponding to output channels 1 to M are integrated together to form a weight data block. This organization of weight data blocks, combined with the parallel computing requirements of the computing array, helps to process and calculate data more efficiently. It breaks the conventional organization idea of ​​taking the input data as a whole or a single output channel as a unit, and constructs data blocks from the perspective of a specific combination of output channels and input channels to adapt to specific computing processes and hardware acceleration requirements.

[0056] Figure 3 The D in the figure indicates the output channels 1 to M. data blocks, E represents the output channel M+1-2M data blocks, and so on.

[0057] Further, in the specific implementation, in the above data processing method provided by the embodiment of the present invention, the convolution calculation is performed in step S104 and / or step S105, which may specifically include: for the weight data, according to The weight data block is convolved in a loop traversal mode; where K is the convolution kernel size, Cin is the number of input channels, and Cout is the number of output channels; the width of the weight cache is set to the data bit width of a data block, and the depth of the weight cache is set to .

[0058] In implementation, the process of traversing the weight data block of the present invention can be as follows: Here is the cycle method. The loop means traversing each position of the convolution kernel to complete the convolution operation. This loop is used to iterate over all input channel groups. This loop is used to traverse all output channel groups. The loop traverses, that is, traverses each position of the convolution kernel, the input channel group, and the output channel group in turn to complete the entire convolution operation.

[0059] The weight cache of the present invention is used to temporarily store weight data for convolution calculations. The weight data is organized into the form of M output channels × M input channels × the bit width of each element. The bit width of a data block is the number of binary bits corresponding to this organization. The width of the weight cache is set to the data bit width of a data block in order to be able to store a complete data block at one time, thereby supporting subsequent parallel calculations. Since a data block contains weight data for M output channels and M input channels, when performing convolution calculations, the weights of these M output channels and M input channels can be multiplied and added to the input feature map at the same time. This parallel calculation can significantly improve computing efficiency. The depth of the weight cache represents the number of data blocks that can be stored in the cache, and is set to , is to be able to store data blocks of all input channel groups and output channel groups combined to meet the computational requirements of the entire convolution operation.

[0060] The above-mentioned loop traversal method for feature data and weight data realizes the result accumulation between input channels and the data arrangement format between output channels and batches so that the input and output arrangement formats of each layer are consistent.

[0061] Furthermore, in a specific implementation, in the above-mentioned data processing method provided by an embodiment of the present invention, step S104 adopts an on-chip storage mode to load feature data, which may specifically include: organizing the memory access operation of the feature data into a first multi-stage loop structure; the first multi-stage loop structure includes a corresponding convolution kernel row block count, a corresponding convolution kernel row count, an input channel group count, each group of feature row block counts, a feature row group count, an output channel group count, and a batch group count; and the first multi-stage loop structure is used to load the feature data.

[0062] In implementation, the control method for feature data loading in the on-chip storage mode can organize the memory access operation of feature data into a multi-level loop structure, and accurately control the access and loading of data through counting in multiple dimensions. The multi-level loop structure can be a seven-level loop structure, that is, when performing feature data access, seven layers of nested loops are used to control the data access process, including the corresponding convolution kernel row block count, the corresponding convolution kernel row count, the input channel group count, each group of feature row block count, feature row group count, output channel group count, and batch group count. Each level of loop is responsible for controlling data access of a specific dimension. Through this layered approach, multi-dimensional feature data can be processed in an orderly manner.

[0063] Furthermore, in specific implementation, the first multi-level loop structure is used in the above steps to load feature data of different dimensions, which may specifically include: traversing the matrix blocks in the row grouping of the corresponding convolution kernel in turn; if all the matrix blocks in the row grouping are traversed, then continue to traverse the row groups of the corresponding convolution kernel in turn; if all the row groups in the convolution kernel are traversed, then continue to traverse the input channel groups of the corresponding convolution kernel in turn; if all the input channel groups are traversed, then continue to traverse the row groups of the corresponding feature data in turn; if the traversal in each row group of the feature data is completed, then continue to traverse the row groups of the corresponding feature data in turn; if the traversal of all row groups of the feature data is completed, then continue to traverse the output channel groups of the corresponding weights in turn; if the traversal of all output channel groups of the weights is completed, then continue to traverse the batch groups of the corresponding weights in turn; if the traversal of all batch groups of the weights is completed, then continue to traverse the batch groups of the corresponding feature data in turn; if the traversal of all batch groups of the feature data is completed, it means that the loading of the feature data is completed.

[0064] In the implementation, first enter the matrix blocks in a row group of the corresponding convolution kernel and traverse them one by one, which corresponds to the "count of blocks in the row of the corresponding convolution kernel", that is, visit the small blocks in each row of the convolution kernel one by one. Then, determine whether all blocks in a row group have been traversed. If not, continue to traverse the matrix blocks in the group; if completed, enter the next level of traversal, corresponding to the "count between rows of the corresponding convolution kernel", and traverse the row groups of the corresponding convolution kernel one by one, that is, start to visit different row groups of the convolution kernel. Determine whether all row groups in the corresponding convolution kernel have been traversed. If not, continue to traverse the row groups; if so, enter the next level of traversal, corresponding to the "input channel group count", and traverse the input channel groups of the corresponding convolution kernel one by one, that is, process different input channel groups in sequence. Determine whether all input channel groups have been traversed. If not completed, continue to traverse the input channel grouping; if completed, enter the traversal of the feature (feature), corresponding to the "block count within a group of feature rows", traverse the row grouping of the corresponding feature in turn, that is, visit the small blocks in a row grouping of the feature data one by one. Determine whether the traversal of the row grouping of the corresponding feature is completed. If not, continue to traverse in the row grouping; if yes, enter the next level of traversal, corresponding to the "feature row group count", traverse all row groups of the corresponding feature in turn, that is, start to access different row groups of the feature data. Determine whether the traversal of all row groups of the corresponding feature is completed. If not completed, continue to traverse the row grouping; if completed, enter the weight-related traversal, corresponding to the "output channel group count", traverse all output channel groups of the corresponding weight in turn, that is, process different output channel groups in sequence. Determine whether the traversal of all output channel groups of the corresponding weight is completed. If not, continue to traverse the output channel grouping; if yes, enter the batch-related traversal, corresponding to the "batch group count", traverse all batch groups of the corresponding feature in turn, that is, start to process different batches of data. Determine whether the traversal of all batch groups of the corresponding feature is completed. If not completed, continue to traverse the batch grouping; if completed, the process reaches the end node, which means that the seven-level loop structure of the entire feature data access memory has been executed. In this way, in the on-chip storage mode, through counting and traversal of different dimensions, the feature data is accessed and loaded in an orderly and organized manner to meet the data processing requirements of deep learning and other tasks.

[0065] Furthermore, in a specific implementation, in the above-mentioned data processing method provided by an embodiment of the present invention, step S104 adopts an on-chip storage mode to load weight data, which may specifically include: organizing the memory access operation of the weight data into a second multi-stage loop structure; the second multi-stage loop structure includes a corresponding convolution kernel intra-row block count, a corresponding convolution kernel inter-row count, an input channel group count, and an output channel group count; and the weight data is loaded using the second multi-stage loop structure.

[0066] In implementation, the present invention can perform read or write operations on stored weight data. When processing access to weight data, a multi-layer nested loop is used to control the entire process. Each level of loop is responsible for access control of a specific dimension. This hierarchical approach realizes orderly processing of weight data, and can systematically and orderly load weight data, meeting the requirements for efficient access to weights during neural network calculations.

[0067] Furthermore, in the specific implementation, in the above steps, a second multi-level loop structure is used to load the weight data, which may specifically include: traversing the matrix blocks in the row grouping corresponding to the convolution kernel in turn; if all the matrix blocks in the row grouping are traversed, then continuing to traverse the row groups of the corresponding convolution kernel in turn; if all the row groups of the convolution kernel are traversed, then continuing to traverse the input channel groups of the corresponding convolution kernel in turn; if all the input channel groups are traversed, then judging whether the number of loops equal to the total number of feature map elements is completed; if the number of loops equal to the total number of feature map elements is completed, then traversing all the output channel groups of the corresponding weight in turn; if all the output channel groups of the weight are traversed, then judging whether the number of loops equal to the batch grouping value is completed; if the number of loops equal to the batch grouping value is completed, it indicates that the weight data loading is completed.

[0068] In the implementation, first enter the matrix blocks in a row group of the corresponding convolution kernel and traverse them one by one, which corresponds to the "count of blocks within the row of the corresponding convolution kernel", that is, visit the small blocks in each row of the convolution kernel one by one. Then, determine whether all blocks in a row group have been traversed. If not, continue to traverse the matrix blocks in the group; if completed, enter the next level of traversal, corresponding to the "count between rows of the corresponding convolution kernel", and traverse the row groups of the corresponding convolution kernel one by one, that is, start to visit different row groups of the convolution kernel. Determine whether all row groups in the corresponding convolution kernel have been traversed. If not, continue to traverse the row groups; if so, enter the next level of traversal, corresponding to the "input channel group count", and traverse the input channel groups of the corresponding convolution kernel one by one, that is, process different input channel groups in sequence. Determine whether all input channel groups have been traversed. If not, continue to traverse the input channel groups; if completed, determine Is the cycle completed? If the loop is not completed, the step of traversing the matrix blocks in a row group of the corresponding convolution kernel will be restarted. After the second cycle is completed, the output channel group traversal will begin, corresponding to the "output channel group count", and all output channel groups of the corresponding weight will be traversed in sequence, that is, different output channel groups will be processed in sequence. Determine whether the traversal of all output channel groups of the corresponding weight has been completed. If not, continue to traverse the output channel group; if so, determine Is the cycle completed? If the loop is not completed, the step of traversing the matrix blocks in a row group of the corresponding convolution kernel will be restarted. The completion of the loop and the process reaching the end node means that the four-level loop structure of the entire weight data access is completed. In this way, through traversal and judgment at different levels, the weight data can be loaded and accessed in an orderly and conditional manner to meet the needs of weight data processing during neural network calculations.

[0069] Further, in the specific implementation, in the above data processing method provided by the embodiment of the present invention, step S104 and / or step S105 performs convolution calculation, which may specifically include: setting the accumulation period of each group of result data blocks to ; The weight data block of the current group output channel keeps repeating, and the feature data block is traversed according to the accumulation cycle After completing the calculation of the current group of output channels, switch to the weight data block of the next group of output channels and repeat the traversal of the corresponding group of batch feature data blocks; if all output channels are traversed, take the next group of batches and re-traverse the weights of all output channels until all batches of feature data are traversed.

[0070] In implementation, Figure 4 A schematic diagram of memory access to characteristic data blocks of different groups of on-chip storage modes provided by an embodiment of the present invention. Figure 4 The feature data blocks of different groups of batches and input channels are shown in Figure 1. Taking the expression "Group 1 batch Group 2 input channel" as an example, the situation after the input data is grouped by batch and input channel is presented. The data blocks of different groups are displayed in the form of a matrix, reflecting the organization of the data in the on-chip storage. The feature data blocks are traversed according to the accumulation cycle. These input data blocks are the basis for subsequent calculations and will be used in conjunction with the weight data blocks for operations such as convolution. Figure 4 The Column, No. The information of the convolution operation is shown in Figure 1, which corresponds to the C (number of columns), R (number of rows) of the feature map, as well as the parameters such as the convolution kernel size K and the step size S. These annotations reflect the dimensional information of the feature map during the convolution operation, and explain the range and number of times the convolution kernel moves on the feature map, and the traversal of the feature data block according to the accumulation cycle. The calculation logic of this is consistent with the previous one. Figure 4 middle Indicates the traversal order of feature data blocks: For a single kernel, For all input channel groups, For the output result line, To output result rows, Between batch groups.

[0071] Figure 5 Schematic diagram of accessing weight data blocks in different groups of on-chip storage modes provided by an embodiment of the present invention. Figure 5 As shown in the figure, the weight data blocks of the same group of output channels are repeatedly traversed. After the current group of batch feature data blocks are traversed, the weight data blocks of the next group of output channels are replaced and the group of batch feature data blocks are repeatedly traversed. If all output channels are traversed, the next group of batches are taken and the weights of all output channels are traversed again, indicating that in the calculation process, the data blocks of different output channel groups will be processed in turn to complete the conversion process from input to output. Finally, the feature data of all batches are traversed.

[0072] Figure 6 Schematic diagram of accessing the output results of different groups of on-chip storage modes provided by an embodiment of the present invention. Figure 6 As shown, the result data block is obtained by inputting the data block (corresponding to Figure 4 ) and weight data block (corresponding to Figure 5 ) is calculated, which reflects the final output of the entire data access and calculation process. The cumulative cycle of each group of result data blocks is .

[0073] Furthermore, in the specific implementation, in the above data processing method provided in the embodiment of the present invention, before the off-chip memory access mode is adopted to load the feature data and the weight data, it may specifically include: setting Group cache, each group cache size is ; Each group of cache is preloaded before the convolution calculation begins, where As a ping cache, As pong cache.

[0074] In practice, when processing a large feature matrix, the entire feature matrix cannot be stored at one time due to the limited on-chip storage capacity. Therefore, the present invention can use an off-chip storage device to gradually load the data of the feature matrix into the on-chip cache. The off-chip storage device can use a double data rate synchronous dynamic random access memory. During the calculation process, instead of waiting for all data to be loaded before starting the calculation, data is loaded from the off-chip storage device to the on-chip cache while the existing data in the on-chip cache is used for calculation. This method effectively solves the problem of insufficient on-chip storage capacity.

[0075] The computing array usually uses pipeline to perform calculations. In order to make the pipeline work continuously and efficiently without data waiting, the resource constraint condition that the on-chip cache capacity is limited must be considered. According to the characteristics of data block matrix data acquisition during convolution operation (for example, the size of the convolution kernel, step length and other factors affect the range and order of data acquisition), the present invention provides a method for determining the minimum cache configuration to balance the computing efficiency and the use of cache resources, that is, pre-setting The purpose of the group buffer is to store and manage data grouped by different input channels. The size of each group buffer is set to , such a size setting is to meet the demand for data storage during the convolution calculation process, and can store a certain amount of data related to the convolution operation. Before the calculation starts, each group of cache is pre-filled with corresponding data so that the data in the cache can be used immediately after the calculation starts, reducing the calculation waiting time. In addition, the present invention further divides each group of cache into two parts, As a ping cache, The part is used as pong cache. In this way, the data to be calculated is preloaded by ping-pong cache. At the same time, according to the characteristics of convolution data reuse, the memory access overhead can be reduced to ensure that the computing array can continuously load and calculate data.

[0076] Furthermore, in the specific implementation, in the above data processing method provided by the embodiment of the present invention, step S105 adopts the off-chip memory access mode to load the feature data and weight data, and performs convolution calculation at the same time, which may specifically include: after each calculation of a set of results, the data block used for the next set of calculations in the ping cache plus S rows of data read from the off-chip storage device are output as new data to the calculation array; when the calculation array is completed After the calculation of the next row result, choose to switch to the next set of output channel weights for calculation or read the feature data of the next batch to reload the first K rows of data blocks in the current or next batch to the starting ping cache position, and the memory access address loop returns to the initial state; After cycles, the convolution calculation is completed.

[0077] In practice, after the present invention starts computing, when a group of results is calculated, the remaining data blocks in the cache that can be used for the next group of calculations are combined with S rows of data newly read from the off-chip storage device to form new data and provided to the computing array for the next group of calculations. In this way, continuous data supply is achieved to ensure the continuous operation of the computing array.

[0078] In the convolution calculation process, the data of different input channels need to be accumulated in time to obtain the calculation result. In order to ensure that the accumulation operation can be performed continuously without interruption due to data loss, the present invention can cache the data of each input channel group, so that when accumulation is required, the corresponding data can be obtained in the cache, thereby ensuring the continuity of the calculation.

[0079] When completed After the result of the first row is calculated (this number represents the effective number of convolution operations in the row direction of the feature map), you can choose to switch to the next set of output channel weights for calculation, or read the feature data of the next batch. No matter which operation is chosen, the first K rows of data blocks in the current or next batch need to be preloaded to the starting ping cache position, and the memory address is returned to the initial state for the next round of calculation and data loading operations to realize the data processing cycle. After a complete cycle of operations (including data loading, calculation, cache update, etc.), the entire calculation task is completed. This indicates the number of cycles of the entire calculation process and the conditions for the end of the calculation.

[0080] Figure 7 A schematic diagram of the off-chip memory access mode feature cache structure provided by an embodiment of the present invention. Figure 7 The structure and data flow of the feature cache in the off-chip emulation mode are shown in the case of large amounts of data, with the examples of “1st batch, 1st input channel” and “1st batch, 1st input channel”. Taking "Group Input Channel" as an example, it shows that the feature data is stored and managed in groups by batches and input channels. This grouping method helps to access and calculate in an orderly manner when processing large-scale data. The cache structure adopts the ping-pong cache mechanism. Figure 7 The cache area of ​​each group in the RAID is divided into two parts: ping and pong. Data is stored and provided alternately in the form of ping cache and pong cache. When the data in the ping cache is sent to the computing array for calculation, new data can be read from the off-chip storage device into the pong cache, and vice versa. This mechanism ensures the continuous supply of data, enables the computing array to run continuously, and improves the efficiency of data processing. Feature data is read from the off-chip storage device, enters the cache for temporary storage and processing, and is sent to the computing array for calculation operations such as convolution. Figure 7 In , C is used to identify the number of columns of the feature map, indicating that the width of the cache area is related to the number of columns of the feature map, reflecting the relationship between the cache structure and the dimension of the feature map data.

[0081] Table 1 shows that for K=3, S=2, When the first set of caches is written and read out, the lines are written and read out sequentially.

[0082] Table 1 The order of rows written and read from the first cache group

[0083]

[0084] Table 2 shows that for K=3, S=1, The second set of caches writes and reads lines sequentially.

[0085] Table 2 The row order of writing and reading the second cache

[0086]

[0087] Table 3 shows that for K=5, S=2, When , the third set of caches writes and reads lines sequentially.

[0088] Table 3 The row order of writing and reading the second cache

[0089]

[0090] Figure 8 A schematic diagram of the characteristic cache loading timing of the off-chip memory access mode provided by an embodiment of the present invention. Figure 8 Different stages are divided from left to right in the figure, including data preparation, calculation of different rows, calculation of different groups of output channels, and batch calculation, which reflects the time advancement of the entire calculation process and the tasks of different stages. Figure 8 The first preload, the second preload, and so on are marked in Preload. The preload operation is to preload the feature data into the cache before the calculation starts to ensure the timely supply of data during the calculation. For example, the first preload is performed in the data preparation stage to prepare for the subsequent calculation of row 1. Figure 8 Different groups of input channel buffers are shown in Figure 1, such as the first group of Cin ping buffers, the second group of Cin ping buffers, and the Group Cin ping cache, etc. In different preloading stages, data will be loaded into the cache of the corresponding group to support subsequent calculation operations. This shows that during the calculation process, the data of different input channel groups will be cached and processed in a certain order. In the stages of calculating the first row and the second row, the data in the cache will be combined with other data such as weights and provided to the calculation array for convolution calculation. As the calculation proceeds, preloading and cache data are continuously updated to ensure the continuity of the calculation. For example, after each row calculation is completed, the next preloading may be performed to load new data into the cache to prepare for the calculation of the next row. Figure 8 The right side shows the first set of output channels calculated to Group output channel calculation and This means that after the calculation of the feature map row direction is completed, the calculation of different output channel groups and the calculation of different batch groups are performed. Figure 8 It can be seen that the present invention can ensure efficient calculation through reasonable cache loading timing when calculating large amounts of data.

[0091] In the above embodiments, the data processing method is described in detail, and the present invention also provides corresponding embodiments of the convolution acceleration device and the electronic device. It should be noted that the present invention describes the embodiments of the device part from two perspectives, one is based on the functional module perspective, and the other is based on the hardware perspective.

[0092] Fig. 9 A schematic diagram of the structure of a convolution acceleration device provided in an embodiment of the present invention. Based on the perspective of functional modules, the device includes:

[0093] A first arrangement module 10, for arranging feature data blocks of the convolutional layer according to a set feature data format with batches as elements;

[0094] A second arrangement module 11 is used to arrange the weight data blocks of the convolution layer according to a set weight data format with output channels as elements;

[0095] The scale judgment module 12 is used to judge whether the scale of the network data is less than a set threshold after the feature data block and the weight data block are arranged; if so, the first loading module is started; if not, the second loading module is started;

[0096] A first loading module 13 is used to load feature data and weight data using an on-chip storage mode, and perform convolution calculation after loading is completed;

[0097] The second loading module 14 is used to load feature data and weight data using an off-chip memory access mode and perform convolution calculations at the same time.

[0098] In the above-mentioned convolution acceleration device provided in the embodiment of the present invention, through the interaction of the above-mentioned five modules, the convolution calculation part of the reinforcement learning front end for feature extraction can be collaboratively deployed on the block matrix computing architecture to achieve resource reuse, effectively solve the adaptation problem between storage resources and computing requirements, ensure that the model can run efficiently at different scales, and effectively reduce the complexity of hardware implementation and the consumption of hardware resources.

[0099] Since the embodiments of the device part correspond to the embodiments of the method part, the embodiments of the device part refer to the description of the embodiments of the method part, which will not be described here. And it has the same beneficial effects as the above-mentioned data processing method.

[0100] Furthermore, in a specific implementation, in the above-mentioned convolution acceleration device provided by an embodiment of the present invention, the first arrangement module 10 can be specifically used to organize the feature data blocks of the convolution layer according to a set feature data format of M batches multiplied by M input channels multiplied by the bit width of each element; wherein M is a positive integer.

[0101] The second arrangement module 11 can be specifically used to organize the weight data block of the convolution layer according to a set weight data format of M output channels multiplied by M input channels multiplied by the bit width of each element; wherein M is a positive integer.

[0102] Further, in the specific implementation, in the above-mentioned convolution acceleration device provided by the embodiment of the present invention, the first loading module 13 and / or the second loading module 14 can be specifically used for characteristic data, according to The convolution calculation is performed on the feature data block in a loop traversal mode; where K is the convolution kernel size, Cin is the number of input channels, S is the stride, C is the feature map row width, and R is the number of feature map rows; and Respectively represent the number of times the convolution kernel slides in the column direction and row direction of the feature map, Indicates rounding down; Batch is the total number of batches, Cout is the number of output channels after the convolution operation; the width of the feature buffer is set to the data bit width of a data block, and the depth of the feature buffer is set to For weight data, according to The weight data block is convolved in a loop traversal mode; where K is the convolution kernel size, Cin is the number of input channels, and Cout is the number of output channels; the width of the weight cache is set to the data bit width of a data block, and the depth of the weight cache is set to .

[0103] Furthermore, in a specific implementation, in the above-mentioned convolution acceleration device provided by an embodiment of the present invention, the first loading module 13 can be specifically used to organize the memory access operation of the feature data into a first multi-stage loop structure; the first multi-stage loop structure includes a corresponding convolution kernel row block count, a corresponding convolution kernel row count, an input channel group count, each group of feature row block counts, a feature row group count, an output channel group count, and a batch group count; the first multi-stage loop structure is used to load the feature data.

[0104] Furthermore, in a specific implementation, in the above-mentioned convolution acceleration device provided by an embodiment of the present invention, the second loading module 14 can be specifically used to organize the memory access operation of the weight data into a second multi-stage loop structure; the second multi-stage loop structure includes a corresponding convolution kernel intra-row block count, a corresponding convolution kernel inter-row count, an input channel group count, and an output channel group count; and the second multi-stage loop structure is used to load the weight data.

[0105] Furthermore, in a specific implementation, the convolution acceleration device provided in the embodiment of the present invention may further include: a cache module for setting Group cache, each group cache size is ; Each group of cache is preloaded before the convolution calculation begins, where As a ping cache, As pong cache.

[0106] Furthermore, in a specific implementation, in the convolution acceleration device provided in the embodiment of the present invention, the second loading module 14 can be used to output to the calculation array as new data each time a set of results is calculated, using the data blocks used for the next set of calculations in the ping cache plus S rows of data read from the off-chip storage device; when the calculation array is completed After the calculation of the next row result, choose to switch to the next set of output channel weights for calculation or read the feature data of the next batch to reload the first K rows of data blocks in the current or next batch to the starting ping cache position, and the memory access address loop returns to the initial state; After cycles, the convolution calculation is completed.

[0107] Fig.10 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. This embodiment is based on the hardware perspective, such as Fig.10 As shown, the electronic equipment includes:

[0108] A memory 20, used for storing computer programs;

[0109] The processor 21 is used to implement the steps of the data processing method mentioned in the above embodiment when executing the computer program.

[0110] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0111] The memory 20 may include one or more non-volatile storage media, which may be non-transitory. The memory 20 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201, wherein, after the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the data processing method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 20 may also include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the data involved in the above-mentioned data processing method, etc.

[0112] In some embodiments, the electronic device may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26. Those skilled in the art will appreciate that Fig.10 The structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure. The electronic device provided by the embodiment of the present invention includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the following method: data processing method, the effect is the same as above.

[0113] Finally, the present invention also provides an embodiment corresponding to a non-volatile storage medium. The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.

[0114] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and executes all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code. The non-volatile storage medium provided by the present invention can implement the above-mentioned data processing method, and the effect is the same as above.

[0115] Finally, the present invention also provides an embodiment corresponding to a computer program product. The computer program product includes a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps described in the above data processing method embodiment are implemented. The computer program product provided by the present invention can implement the above-mentioned data processing method, and the effect is the same as above.

[0116] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0117] The data processing method, electronic device, medium and computer program product provided by the present invention are introduced in detail above. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A data processing method, characterized in that: The method comprises: Arranging the feature data blocks of the convolutional layer according to a set feature data format with batches as elements, including: organizing the feature data blocks of the convolutional layer according to a set feature data format of M batches multiplied by M input channels multiplied by a bit width of each element; wherein M is a positive integer; Arrange the weight data blocks of the convolutional layer according to the set weight data format with output channels as elements; After the feature data blocks and weight data blocks are arranged, it is determined whether the size of the network data is less than the set threshold; If yes, the on-chip storage mode is used to load the feature data and weight data, and the convolution calculation is performed after the loading is completed; If not, the off-chip memory access mode is used to load the feature data and weight data, and the convolution calculation is performed at the same time; Perform convolution calculations, including: For feature data, according to The convolution calculation is performed on the feature data block in a loop traversal manner; Among them, K is the convolution kernel size, Cin is the number of input channels, S is the step size, C is the feature map row width, and R is the number of feature map rows; and Respectively represent the number of times the convolution kernel slides in the column direction and row direction of the feature map, Indicates rounding down; Batch is the total number of batches, and Cout is the number of output channels after the convolution operation; Set the width of the feature buffer to the data bit width of a data block, and the depth of the feature buffer to .

2. The data processing method according to claim 1, characterized in that: The weight data blocks of the convolutional layer are arranged according to the set weight data format with the output channel as the element, including: The weight data block of the convolutional layer is organized according to a set weight data format of M output channels multiplied by M input channels multiplied by the bit width of each element; where M is a positive integer.

3. The data processing method according to claim 2, characterized in that: Perform convolution calculations, including: For weight data, according to The weight data block is convolved in a loop traversal manner; Among them, K is the convolution kernel size, Cin is the number of input channels, and Cout is the number of output channels; The width of the weight cache is set to the data bit width of a data block, and the depth of the weight cache is set to .

4. The data processing method according to claim 1, characterized in that: Use on-chip storage mode to load feature data, including: Organizing the memory access operation of feature data into a first multi-stage loop structure; the first multi-stage loop structure includes a corresponding convolution kernel row block count, a corresponding convolution kernel row count, an input channel group count, a feature row block count of each group, a feature row group count, an output channel group count, and a batch group count; The first multi-level loop structure is used to load the feature data.

5. The data processing method according to claim 4, characterized in that: The first multi-level loop structure is used to load the characteristic data, including: Traverse the matrix blocks in the row group corresponding to the convolution kernel in turn; If all matrix blocks in a row group are traversed, continue to traverse the row groups of the corresponding convolution kernels in sequence; If all row groups in the convolution kernel are traversed, continue to traverse the input channel groups of the corresponding convolution kernel in sequence; If all input channel groups are traversed, continue to traverse the corresponding feature data rows in turn; If the traversal within each row group of the feature data is completed, continue to traverse the row groups of the corresponding feature data in sequence; If the traversal between all row groups of feature data is completed, continue to traverse the output channel groups of corresponding weights in sequence; If the traversal between all output channel groups of the weight is completed, continue to traverse the batch groups of the corresponding weight in sequence; If all batch groups of weights are traversed, continue to traverse the batch groups of corresponding feature data in sequence; If all batch groups of feature data are traversed, it means that the feature data loading is complete.

6. The data processing method according to claim 1, characterized in that: The on-chip storage mode is used to load weight data, including: Organizing the memory access operation of the weight data into a second multi-stage loop structure; the second multi-stage loop structure includes a corresponding convolution kernel row block count, a corresponding convolution kernel row count, an input channel group count, and an output channel group count; The weight data is loaded using the second multi-level loop structure.

7. The data processing method according to claim 6, characterized in that: The weight data is loaded using the second multi-level loop structure, including: Traverse the matrix blocks in the row group corresponding to the convolution kernel in turn; If all matrix blocks in a row group are traversed, continue to traverse the row groups of the corresponding convolution kernels in sequence; If all row groups of the convolution kernel are traversed, continue to traverse the input channel groups of the corresponding convolution kernel in sequence; If all input channel group traversals are completed, determine whether the number of loops equal to the total number of feature map elements is completed; If the number of cycles equal to the total number of feature map elements is completed, all output channel groups with corresponding weights are traversed in turn; If the traversal between all output channel groups of the weight is completed, then determine whether the number of cycles with the same batch group value is completed; If the number of cycles equal to the batch grouping value is completed, it means that the weight data loading is complete.

8. The data processing method according to claim 1, characterized in that: Perform convolution calculations, including: Set the accumulation period for each group of result data blocks to ; The weight data block of the current group output channel keeps repeating, and the feature data block is traversed according to the accumulation cycle times, traverse the feature data blocks of the current batch; After completing the calculation of the current group of output channels, switch to the weight data block of the next group of output channels and repeatedly traverse the corresponding group of batch feature data blocks; If all output channels are traversed, take the next batch and traverse the weights of all output channels again until all batches of feature data are traversed.

9. The data processing method according to claim 1, characterized in that: Before using the off-chip memory access mode to load feature data and weight data, it includes: set up Group cache, each group cache size is ; Each group of cache is preloaded before the convolution calculation begins. As a ping cache, As pong cache.

10. The data processing method according to claim 9, characterized in that: The off-chip memory access mode is used to load feature data and weight data, and convolution calculations are performed at the same time, including: After each set of results is calculated, the data blocks used for the next set of calculations in the ping cache plus S rows of data read from the off-chip storage device are used as new data and output to the calculation array; When the calculation array is completed After the result of the next row is calculated, the next set of output channel weights is switched to be calculated or the feature data of the next batch is read to reload the first K rows of data blocks in the current or next batch to the starting ping cache position, and the memory access address loop returns to the initial state; In progress After cycles, the convolution calculation is completed.

11. An electronic device, characterized in that: The electronic device comprises: Memory for storing computer programs; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 10 when executing the computer program.

12. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 10 are implemented.

13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the data processing method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Parallel convolution method, acceleration framework and computer readable storage medium

    CN118627554A