Data processing method, accelerator, device and readable storage medium

By determining the number of channels of each packet in the convolution neural network and distributing data and coefficients based on the number of convolution kernels, the problem of low hardware resource utilization in packet convolution is solved, and a faster convolution operation speed is achieved.

CN114065911BActive Publication Date: 2025-07-29SHENZHEN CORERAIN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111251072.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-07-29
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

In convolutional neural networks, the hardware resource utilization rate is low during group convolution, resulting in slower computing speed.

Method used

By obtaining the number of channels and groups of input data, the number of channels of each packet is determined, and when the number of channels of each packet is less than or equal to half of the maximum parallelism degree of the convolution hardware, the data and convolution coefficients are distributed to the calculation unit according to the number of convolution kernels of the packets to perform convolution operations.

Benefits of technology

It improves the resource utilization rate of convolution hardware and improves the speed of convolutional operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065911B_ABST
    Figure CN114065911B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method, an accelerator, a device and a readable storage medium. The data processing method includes: obtaining the number of channels of input data and the number of groups when performing grouped convolution on the input data; determining the number of channels of each group according to the number of channels of the input data and the number of groups; when the number of channels of each group is less than or equal to half of the maximum parallelism of the convolution hardware, distributing the input data and convolution coefficients to a computing unit according to the number of convolution kernels of each group, so that the computing unit performs convolution operations on the input data. The purpose of the present invention is to improve the operation speed of convolution operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data processing method, an accelerator, a device and a readable storage medium. Background Art

[0002] In a convolutional neural network, each convolutional kernel includes a channel dimension and two spatial dimensions, and each output point is a weighted sum of all points in these three dimensions. When using a convolutional neural network model for data processing, grouped convolution is usually adopted to reduce the number of model parameters and thus reduce the amount of computation. However, in grouped convolution, when the number of channels in each group and the number of convolutional kernels are relatively small, the utilization rate of hardware resources during the convolutional operation process is relatively low, resulting in a slow operation speed. Summary of the Invention

[0003] The main purpose of the present invention is to provide a data processing method, an accelerator, a device and a readable storage medium, aiming to improve the operation speed of convolutional operations.

[0004] To achieve the above purpose, the present invention provides a data processing method, which includes:

[0005] Obtain the number of channels of the input data and the number of groups during grouped convolution of the input data;

[0006] Determine the number of channels in each group according to the number of channels of the input data and the number of groups;

[0007] When the number of channels in each group is less than or equal to half of the maximum parallelism of the convolutional hardware, distribute the input data and convolutional coefficients to the computing units according to the number of convolutional kernels in each group, so that the computing units perform convolutional operations on the input data.

[0008] Optionally, the step of distributing the input data and convolutional coefficients to the computing units according to the number of convolutional kernels in each group includes:

[0009] Determine the grouped parallelism of the convolutional hardware;

[0010] Distribute the input data and the convolutional coefficients to the computing units according to the grouped parallelism and the number of convolutional kernels in each group.

[0011] Optionally, the step of determining the grouped parallelism of the convolutional hardware includes:

[0012] Calculate the ratio of the maximum parallelism of the convolutional hardware to the number of channels in each group;

[0013] Determine the grouped parallelism of the convolutional hardware according to the ratio.

[0014] Optionally, after the step of determining the number of channels of each group according to the number of channels of the input data and the number of groups, the method further includes:

[0015] When the number of channels of each group is greater than half of the maximum parallelism of the convolution hardware, obtaining the number of convolution kernels of the convolution hardware;

[0016] Distributing the input data and convolution coefficients to the computing units according to the number of convolution kernels of the convolution hardware, so that the computing units perform convolution operations on the input data.

[0017] Optionally, before the step of obtaining the number of convolution kernels of the convolution hardware, the method further includes:

[0018] Determining whether the number of channels of each group is greater than the maximum parallelism of the convolution hardware;

[0019] When the number of channels of each group is greater than the maximum parallelism of the convolution hardware, performing the step of obtaining the number of convolution kernels of the convolution hardware.

[0020] Optionally, after the step of determining whether the number of channels of each group is greater than the maximum parallelism of the convolution hardware, the method further includes:

[0021] When the number of channels of each group is less than or equal to the maximum parallelism of the convolution hardware, obtaining the number of convolution kernels of each group;

[0022] Distributing the input data and convolution coefficients to the computing units according to the number of convolution kernels of each group, so that the computing units perform convolution operations on the input data.

[0023] In addition, to achieve the above object, the present invention further provides an accelerator, which includes a memory, a processor, and a data processing program stored on the memory and executable on the processor. When the data processing program is executed by the processor, the steps of the data processing method described in any one of the above are implemented.

[0024] In addition, to achieve the above object, the present invention further provides a data processing device, which includes the above accelerator.

[0025] In addition, to achieve the above object, the present invention further provides a data processing device, which includes:

[0026] An obtaining module, configured to obtain the number of channels of the input data and the number of groups during grouped convolution of the input data;

[0027] A determining module, configured to determine the number of channels of each group according to the number of channels of the input data and the number of groups;

[0028] A distribution module, configured to distribute the input data and convolution coefficients to a computing unit according to the number of convolution kernels in each group when the number of channels in each group is less than or equal to half of the maximum parallelism of the convolution hardware, so that the computing unit can perform a convolution operation on the input data.

[0029] In addition, to achieve the above object, the present invention further provides a readable storage medium, on which a data processing program is stored. When the data processing program is executed by a processor, the steps of the data processing method described in any one of the above are implemented.

[0030] The present invention provides a data processing method, an accelerator, a device, and a readable storage medium. By obtaining the number of channels of input data and the number of groups during grouped convolution of the input data, determining the number of channels in each group according to the number of channels and the number of groups of the input data, and when the number of channels in each group is less than or equal to half of the maximum parallelism of the convolution hardware, distributing the input data and convolution coefficients to a computing unit according to the number of convolution kernels in each group, so that the computing unit can perform a convolution operation on the input data. In this solution, when the number of channels in each group is less than or equal to half of the maximum parallelism of the convolution hardware, distributing data and coefficients according to the number of convolution kernels in each group can make full use of the hardware resources of the convolution hardware, improve the utilization rate of hardware resources, and improve the operation speed of the convolution operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic diagram of the hardware architecture of the accelerator related to the solution of the embodiment of the present invention;

[0032] Figure 2 is a schematic flowchart of an embodiment of the data processing method of the present invention;

[0033] Figure 3 is a schematic flowchart of an embodiment of the data processing method of the present invention;

[0034] Figure 4 is a schematic flowchart of an embodiment of the data processing method of the present invention;

[0035] Figure 5 is a schematic diagram of the module structure of the data processing device related to the solution of the embodiment of the present invention.

[0036] The implementation, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] As an implementation solution, please refer to Figure 1 , Figure 1 which is a schematic diagram of the hardware architecture of the accelerator involved in the embodiment of the present invention. As shown in Figure 1 , the accelerator may include a processor 101, such as a CPU, a memory 102, and a communication bus 103. Among them, the communication bus 103 is used to realize the connection and communication between these modules.

[0039] The memory 102 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. As shown in Figure 1 , the memory 102, as a readable storage medium, may include a data processing program; and the processor 101 may be used to call the data processing program stored in the memory 102 and perform the following operations:

[0040] Obtain the number of channels of the input data and the number of groups when performing grouped convolution on the input data;

[0041] Determine the number of channels of each group according to the number of channels of the input data and the number of groups;

[0042] When the number of channels of each group is less than or equal to half of the maximum parallelism of the convolution hardware, distribute the input data and convolution coefficients to the calculation unit according to the number of convolution kernels of each group, so that the calculation unit performs convolution operations on the input data.

[0043] Further, the processor 101 may be used to call the data processing program stored in the memory 102 and perform the following operations:

[0044] Determine the grouped parallelism of the convolution hardware;

[0045] Distribute the input data and the convolution coefficients to the calculation unit according to the grouped parallelism and the number of convolution kernels of each group.

[0046] Further, the processor 101 may be used to call the data processing program stored in the memory 102 and perform the following operations:

[0047] Calculate the ratio of the maximum parallelism of the convolution hardware to the number of channels of each group;

[0048] Determine the grouped parallelism of the convolution hardware according to the ratio.

[0049] Further, the processor 101 may be used to call the data processing program stored in the memory 102 and perform the following operations:

[0050] When the number of channels in each of the groups is greater than half of the maximum parallelism of the convolution hardware, obtain the number of convolution kernels of the convolution hardware;

[0051] Distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels of the convolution hardware for the computing units to perform convolution operations on the input data.

[0052] Further, the processor 101 may be used to call a data processing program stored in the memory 102 and perform the following operations:

[0053] Determine whether the number of channels in each of the groups is greater than the maximum parallelism of the convolution hardware;

[0054] When the number of channels in each of the groups is greater than the maximum parallelism of the convolution hardware, perform the step of obtaining the number of convolution kernels of the convolution hardware.

[0055] Further, the processor 101 may be used to call a data processing program stored in the memory 102 and perform the following operations:

[0056] When the number of channels in each of the groups is less than or equal to the maximum parallelism of the convolution hardware, obtain the number of convolution kernels of each of the groups;

[0057] Distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels of each of the groups for the computing units to perform convolution operations on the input data.

[0058] In the prior art, when using a convolutional neural network model for data processing, group convolution is usually used to reduce the number of model parameters to reduce the amount of computation. However, in group convolution, when the number of channels in each group and the number of convolution kernels are relatively small, the utilization rate of hardware resources during the convolution operation process is low, resulting in a slow operation speed.

[0059] Based on the above problems existing in the prior art, the present application proposes a data processing method. During group convolution, by comparing the size relationship between the number of channels in each group and half of the maximum parallelism of the convolution hardware, the input data is distributed to different computing units in a shifted manner according to the comparison result for the computing units to perform convolution operations, improving the utilization rate of the convolution hardware and the operation speed. The data processing method proposed by the present application will be further explained below through specific embodiments.

[0060] Please refer to Figure 2 , in an embodiment of the present invention, the data processing method includes the following steps:

[0061] Step S10, obtain the number of channels of the input data and the number of groups during group convolution of the input data;

[0062] In this embodiment, the execution subject of the data processing method is an accelerator or a data processing device including the accelerator, such as an image processing device. In other embodiments, the data processing device may also be other devices that can execute the data processing method of this application. This application will be described with the execution subject being a data processing device.

[0063] In this embodiment, the data processing device obtains the number of channels of the input data and the number of groups for grouped convolution of the input data. The input data may be image data. In other embodiments, the input data may also be other data that can perform convolution operations. The number of groups for grouped convolution of the input data refers to the number of groups into which the input data is divided during grouped convolution, and the number of groups for grouped convolution of the input data can be set in advance. It should be noted that the number of channels of the input data is usually a multiple of the number of groups.

[0064] After the data processing device obtains the input data, when grouped convolution operation needs to be performed on the input data, the data processing device can automatically read the number of channels of the input data and the number of groups for grouped convolution of the input data set in advance.

[0065] For example, the input data is image data, the image data includes 768 channels / 768 convolution kernels, and the number of groups for grouped convolution preset by the convolution hardware is 16. After the data processing device obtains the image data, it automatically reads the number of channels 768 of the image data and the number of groups 16.

[0066] Step S20: Determine the number of channels of each group according to the number of channels of the input data and the number of groups;

[0067] In this embodiment, after the data processing device obtains the number of channels of the input data and the number of groups for grouped convolution of the input data, it determines the number of channels of each group according to the number of channels of the input data and the number of groups. The number of channels of each group is calculated as follows: the number of channels of each group = the number of channels of the input data / the number of groups. Correspondingly, the number of convolution kernels of each group is the same as the number of channels of each group.

[0068] For example, after the data processing device obtains the number of channels 768 of the image data and the number of groups 16, the number of channels of each group = 768 / 16 = 48. Correspondingly, the number of convolution kernels of each group is 48. That is, each group includes 48 channels / 48 convolution kernels.

[0069] Step S30: When the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels in each of the groups, so that the computing units perform convolution operations on the input data.

[0070] In this embodiment, after determining the number of channels in each of the groups, the data processing device determines whether the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware. When it is determined that the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, obtain the number of convolution kernels in each of the groups, and distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels in each of the groups. After receiving the input data and convolution coefficients, each computing unit performs a convolution operation on the input data.

[0071] For example, the number of channels in each of the groups is 48, and the maximum parallelism of the convolution hardware is 128 channels in parallel / 128 convolution kernels in parallel. Since the number of channels 48 in each of the groups is less than half of the maximum parallelism 128 of the convolution hardware, which is 64, distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels 48 in each of the groups. That is, in the 128-convolution-kernel parallel convolution hardware, the first 48 convolution kernels perform the operations of the first group, and the last 48 convolution kernels perform the operations of the second group. At this time, the remaining 32 convolution kernels are idle. However, compared with the existing technical solution in which only 48 convolution kernels are used to run the operations of a group in the 128-convolution-kernel parallel convolution hardware and the remaining 80 convolution kernels are idle, the utilization rate of hardware resources is greatly improved, and the operation rate of the convolution operation is increased.

[0072] In the technical solution provided in this embodiment, by obtaining the number of channels of the input data and the number of groups during the grouped convolution of the input data, determine the number of channels in each of the groups according to the number of channels of the input data and the number of groups. When the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels in each of the groups, so that the computing units perform convolution operations on the input data. In this solution, when the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, distributing the data and coefficients according to the number of convolution kernels in each of the groups can make full use of the hardware resources of the convolution hardware, improve the utilization rate of hardware resources, and increase the operation speed of the convolution operation.

[0073] Please refer to Figure 3, in an embodiment of the present invention, when the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, the step of distributing the input data and the convolution coefficients to the computing units according to the number of convolution kernels in each of the groups for the computing units to perform convolution operations on the input data includes:

[0074] Step S31, when the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, determine the group parallelism of the convolution hardware;

[0075] In this embodiment, when the data processing device determines that the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, it determines the group parallelism of the convolution hardware. The group parallelism of the convolution hardware refers to the number of groups that perform convolution operations simultaneously in the same round of operations. It should be noted that the product of the number of convolution kernels in each of the groups and the group parallelism cannot be greater than the maximum parallelism of the convolution hardware.

[0076] Step S32, distribute the input data and the convolution coefficients to the computing units according to the group parallelism and the number of convolution kernels in each of the groups for the computing units to perform convolution operations on the input data.

[0077] In this embodiment, after the data processing device determines the group parallelism of the convolution hardware, it distributes the input data and the convolution coefficients to the computing units according to the group parallelism and the number of convolution kernels in each of the groups. After each computing unit receives the input data and the convolution coefficients, it performs convolution operations on the input data.

[0078] For example, the convolution hardware is 128 channels in parallel / 128 convolution kernels in parallel, that is, the maximum parallelism of the convolution hardware is 128, and each group includes 48 channels / 48 convolution kernels, that is, the number of convolution kernels in each group is 48, then the group parallelism is 2.

[0079] In the technical solution provided in this embodiment, when the data processing device determines that the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, by determining the group parallelism of the convolution hardware, and distributing the input data and the convolution coefficients to the computing units according to the group parallelism and the number of convolution kernels in each of the groups. This solution can make full use of the hardware resources of the convolution hardware according to the number of convolution kernels in each group and the group parallelism, improve the utilization rate of the hardware resources, and improve the operation speed of the convolution operation.

[0080] In an embodiment of the present invention, the step of determining the group parallelism of the convolution hardware includes:

[0081] Calculate the ratio of the maximum parallelism of the convolution hardware to the number of channels in each of the groups;

[0082] Determine the group parallelism of the convolution hardware according to the ratio.

[0083] In this embodiment, when the data processing device determines that the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, it calculates the ratio of the maximum parallelism of the convolution hardware to the number of channels in each of the groups, and determines the group parallelism of the convolution hardware according to the ratio. The group parallelism is calculated according to the following formula: group parallelism = maximum parallelism of the convolution hardware / number of channels in each group, and the group parallelism takes an integer.

[0084] For example, the convolution hardware is 128-channel parallel / 128-convolution-kernel parallel, that is, the maximum parallelism of the convolution hardware is 128, and each of the groups includes 48 channels / 48 convolution kernels, that is, the number of convolution kernels in each of the groups is 48, then the group parallelism = 128 / 48, and the group parallelism is rounded to 2.

[0085] In the technical solution provided in this embodiment, when the data processing device determines that the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware, it calculates the ratio of the maximum parallelism of the convolution hardware to the number of channels in each of the groups, determines the group parallelism of the convolution hardware according to the ratio, and distributes the input data and the convolution coefficients to the computing unit according to the group parallelism and the number of convolution kernels in each of the groups, so that the computing unit can perform convolution operations on the input data. This solution distributes data and coefficients according to the number of convolution kernels in each of the groups and the group parallelism, which can make full use of the hardware resources of the convolution hardware, improve the utilization rate of the hardware resources, and improve the operation speed of the convolution operation.

[0086] In an embodiment of the present invention, after the step of determining the number of channels in each group according to the number of channels of the input data and the number of groups, it further includes:

[0087] Step S40, when the number of channels in each of the groups is greater than half of the maximum parallelism of the convolution hardware, obtain the number of convolution kernels of the convolution hardware;

[0088] In this embodiment, after the data processing device determines the number of channels in each of the groups according to the number of channels of the input data and the number of groups, it determines whether the number of channels in each of the groups is less than or equal to half of the maximum parallelism of the convolution hardware. When the number of channels in each of the groups is greater than half of the maximum parallelism of the convolution hardware, it obtains the number of convolution kernels of the convolution hardware.

[0089] For example, the input data is image data, the image data includes 768 channels / 768 convolutional kernels, the number of groups of grouped convolutions preset in the convolutional hardware is 2, and the convolutional hardware is parallel with 128 channels / 128 convolutional kernels in parallel. Then, the number of channels of each group determined according to the number of channels of the image data and the number of groups is 384, and the number of channels of each group, 384, is greater than half of the maximum parallelism of the convolutional hardware, which is 64. At this time, the data processing device automatically reads the number of convolutional kernels of the convolutional hardware, which is 128.

[0090] Step S50: Distribute the input data and convolutional coefficients to the computing unit according to the number of convolutional kernels of the convolutional hardware for the computing unit to perform a convolutional operation on the input data.

[0091] In this embodiment, after the data processing device obtains the number of convolutional kernels of the convolutional hardware, it distributes the input data and convolutional coefficients to the computing unit according to the number of convolutional kernels of the convolutional hardware for the computing unit to perform a convolutional operation on the input data.

[0092] For example, when the number of channels of each group, 384, is greater than half of the maximum parallelism of the convolutional hardware, which is 64, the data processing device distributes the input data and convolutional coefficients to the computing unit according to the number of convolutional kernels of the convolutional hardware, which is 128, for the computing unit to perform a convolutional operation on the input data.

[0093] In the technical solution provided in this embodiment, after the data processing device determines the number of channels of each group according to the number of channels of the input data and the number of groups, when the number of channels of each group is greater than half of the maximum parallelism of the convolutional hardware, it obtains the number of convolutional kernels of the convolutional hardware, and distributes the input data and convolutional coefficients to the computing unit according to the number of convolutional kernels of the convolutional hardware for the computing unit to perform a convolutional operation on the input data. This solution can simplify the complexity of hardware design when the number of channels in each group is relatively large in grouped convolutions.

[0094] In an embodiment of the present invention, before the step of obtaining the number of convolutional kernels of the convolutional hardware, it further includes:

[0095] When the number of channels of each group is greater than half of the maximum parallelism of the convolutional hardware, determine whether the number of channels of each group is greater than the maximum parallelism of the convolutional hardware;

[0096] When the number of channels of each group is greater than the maximum parallelism of the convolutional hardware, obtain the number of convolutional kernels of the convolutional hardware.

[0097] In this embodiment, when the data processing device determines that the number of channels of each group is greater than half of the maximum parallelism of the convolution hardware, it further determines whether the number of channels of each group is greater than the maximum parallelism of the convolution hardware. When the number of channels of each group is greater than the maximum parallelism of the convolution hardware, it obtains the number of convolution kernels of the convolution hardware, and then distributes the input data and convolution coefficients to the computing units according to the number of convolution kernels of the convolution hardware, so that the computing units can perform convolution operations on the input data.

[0098] For example, when the data processing device determines that the number of channels of each group, which is 384, is greater than half of the maximum parallelism of the convolution hardware, which is 64, it further determines whether the number of channels of each group, which is 384, is greater than the maximum parallelism of the convolution hardware, which is 128. When it determines that the number of channels of each group, which is 384, is greater than the maximum parallelism of the convolution hardware, which is 128, it obtains the number of convolution kernels of the convolution hardware, which is 128, and distributes the input data and convolution coefficients to the computing units according to the number of convolution kernels of the convolution hardware, which is 128, so that the computing units can perform convolution operations on the input data.

[0099] In the technical solution provided in this embodiment, when the data processing device determines that the number of channels of each group is greater than half of the maximum parallelism of the convolution hardware, it further determines whether the number of channels of each group is greater than the maximum parallelism of the convolution hardware. When the number of channels of each group is greater than the maximum parallelism of the convolution hardware, it obtains the number of convolution kernels of the convolution hardware, and distributes the input data and convolution coefficients to the computing units according to the number of convolution kernels of the convolution hardware, so that the computing units can perform convolution operations on the input data. This solution can simplify the complexity of hardware design when the number of channels in each group is large in grouped convolution.

[0100] In an embodiment of the present invention, after the step of determining whether the number of channels of each group is greater than the maximum parallelism of the convolution hardware, it further includes:

[0101] When the number of channels of each group is less than or equal to the maximum parallelism of the convolution hardware, obtain the number of convolution kernels of each group;

[0102] Distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels of each group, so that the computing units can perform convolution operations on the input data.

[0103] In this embodiment, when the data processing device determines that the number of channels in each group is greater than half of the maximum parallelism of the convolution hardware, it further determines whether the number of channels in each group is greater than the maximum parallelism of the convolution hardware. When the number of channels in each group is less than or equal to the maximum parallelism of the convolution hardware, it obtains the number of convolution kernels in each group, and distributes the input data and convolution coefficients to the computing units according to the number of convolution kernels in each group for the computing units to perform convolution operations on the input data. At this time, only the data within one group is processed in each round of operation, and the data is padded to the convolution hardware for operation.

[0104] For example, the input data is image data, the image data includes 768 channels / 768 convolution kernels, the number of groups for grouped convolution preset in the convolution hardware is 8, and the convolution hardware is 128-channel parallel / 128-convolution-kernel parallel. Then, the number of channels in each group determined according to the number of channels in the image data and the number of groups is 96. The number of channels 96 in each group is greater than half of the maximum parallelism of the convolution hardware, which is 64, and less than the maximum parallelism of the convolution hardware, which is 128. At this time, the input data and convolution coefficients are distributed to the computing units according to the number of convolution kernels 96 in each group for the computing units to perform convolution operations on the input data. Only the data within one group is processed in each round of operation, and the data is padded to the convolution hardware with 128 channels / 128 convolution kernels for operation.

[0105] In the technical solution provided in this embodiment, when the data processing device determines that the number of channels in each group is greater than half of the maximum parallelism of the convolution hardware, it further determines whether the number of channels in each group is greater than the maximum parallelism of the convolution hardware. When the number of channels in each group is less than or equal to the maximum parallelism of the convolution hardware, it obtains the number of convolution kernels in each group; and distributes the input data and convolution coefficients to the computing units according to the number of convolution kernels in each group for the computing units to perform convolution operations on the input data. This solution can simplify the complexity of hardware design when the number of channels in each group is relatively large in grouped convolution.

[0106] To describe the inventive concept of the present application more clearly, the following is an illustration through a specific example:

[0107] Assume that the convolution hardware is 128-channel parallel / 128-convolution-kernel parallel, and the input data is image data, and the image data includes 768 channels / 768 convolution kernels.

[0108] When performing ordinary convolution, 6 * 6 = 36 rounds of operations are required. Each round of operation calculates 128 channels, and every 6 operations complete the convolution operations of 768 channels of 128 convolution kernels.

[0109] When performing grouped convolution:

[0110] When the number of groups of grouped convolution is 2, each group includes 384 channels / 384 convolution kernels. Since the number of channels in each group, 384, is greater than the maximum parallelism of the convolution hardware, 128, the image data and convolution parameters are distributed to the computing units according to the number of convolution kernels of the convolution hardware, 128, for the computing units to perform convolution operations on the image data. Each group requires 3 * 3 rounds of operations, and a total of 3 * 3 * 2 = 18 rounds of operations are required.

[0111] When the number of groups of grouped convolution is 8, each group has 96 channels / 96 convolution kernels. The number of channels in each group, 96, is less than the maximum parallelism of the convolution hardware, 128, and greater than half of the maximum parallelism of the convolution hardware, 64. At this time, the image data and convolution parameters are distributed to the computing units according to the number of convolution kernels in each group, 96, for the computing units to perform convolution operations on the image data. Each round of operation still only processes the data within one group, and at the same time, the data is padded to 128 channels / 128 convolution kernels of the convolution hardware. Each group requires one round of operation, and a total of 8 rounds of operations are required.

[0112] When the number of groups of grouped convolution is 16, each group has 48 channels / 48 convolution kernels. The number of channels in each group, 48, is less than half of the maximum parallelism of the convolution hardware, 64. First, calculate the grouped parallelism. The sum of the number of channels in each group and the grouped parallelism cannot be greater than the maximum parallelism of the convolution hardware. In this example, the grouped parallelism is 128 / 48, rounded down to 2. In this mode, at this time, the image data and convolution parameters are distributed to the computing units according to the number of convolution kernels in each group, 48, and the grouped parallelism, 2, for the computing units to perform convolution operations on the image data. The amount of data calculated in each round becomes twice the original, and 2 is the grouped parallelism. Specifically, among the 128 parallel convolution kernels in the convolution hardware, the first 48 convolution kernels perform the operations of the first group, and the subsequent 48 convolution kernels perform the operations of the second group, with the remaining 32 idle (there will be a certain waste of hardware resources at this time, but without using this design, 80 convolution kernels will be idle. Therefore, this design greatly reduces waste). Two groups of calculations are completed in each round of operation, and a total of 16 / 2 = 8 rounds of operations are required.

[0113] Based on the above embodiments, the present invention further provides an accelerator, which may include a memory, a processor, and a data processing program stored in the memory and executable on the processor. When the processor executes the data processing program, the steps of the data processing method described in any of the above embodiments are implemented.

[0114] Based on the above embodiments, the present invention further provides a data processing device, and the data processing device may include the accelerator as described above.

[0115] Based on the above embodiments, the present invention further provides a data processing device, which includes:

[0116] An acquisition module 100, configured to acquire the number of channels of input data and the number of groups when performing grouped convolution on the input data;

[0117] A determination module 200, configured to determine the number of channels of each group according to the number of channels of the input data and the number of groups;

[0118] A distribution module 300, configured to, when the number of channels of each group is less than or equal to half of the maximum parallelism of the convolution hardware, distribute the input data and convolution coefficients to a computing unit according to the number of convolution kernels of each group, so that the computing unit performs a convolution operation on the input data.

[0119] Based on the above embodiments, the present invention further provides a readable storage medium, on which a data processing program is stored. When the data processing program is executed by a processor, the steps of the data processing method described in any of the above embodiments are implemented.

[0120] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.

[0121] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) as described above and includes several instructions for causing a terminal device (which can be a smart TV, mobile phone, computer, etc.) to execute the methods described in various embodiments of the present invention.

[0123] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A data processing method, characterized in that, The data processing method includes: Obtaining the number of channels of the input data and the number of groups when performing grouped convolution on the input data. The number of groups when performing grouped convolution on the input data refers to the number of groups into which the input data is divided when performing grouped convolution, and is preset; Determining the number of channels of each group according to the number of channels of the input data and the number of groups; Comparing the size relationship between the number of channels of each group and half of the maximum parallelism of the convolution hardware, and distributing the input data to different computing units in a shifted manner according to the comparison result; When the number of channels of each group is less than or equal to half of the maximum parallelism of the convolution hardware, distributing the input data and convolution coefficients to the computing units according to the number of convolution kernels of each group. The number of convolution kernels of each group is equal to the number of channels of each group, so that the computing units perform convolution operations on the input data. The number of channels of each group is a first preset value, and the maximum parallelism of the convolution hardware is a second preset value. In a convolution hardware with the second preset value of convolution kernels in parallel, the amount of data calculated in each round becomes twice the original. The first first preset value of convolution kernels perform the operations of the first group, and the last first preset value of convolution kernels perform the operations of the second group.

2. The data processing method according to claim 1, characterized in that The step of distributing the input data and convolution coefficients to the computing units according to the number of convolution kernels of each group includes: Determining the grouped parallelism of the convolution hardware; Distributing the input data and the convolution coefficients to the computing units according to the grouped parallelism and the number of convolution kernels of each group.

3. The data processing method according to claim 2, wherein The step of determining the grouped parallelism of the convolution hardware includes: Calculating the ratio of the maximum parallelism of the convolution hardware to the number of channels of each group; Determining the grouped parallelism of the convolution hardware according to the ratio.

4. The data processing method according to claim 1, characterized in that After the step of determining the number of channels of each group according to the number of channels of the input data and the number of groups, it further includes: When the number of channels of each group is greater than half of the maximum parallelism of the convolution hardware, obtaining the number of convolution kernels of the convolution hardware; Distributing the input data and convolution coefficients to the computing units according to the number of convolution kernels of the convolution hardware, so that the computing units perform convolution operations on the input data.

5. The data processing method according to claim 4, characterized in that Before the step of obtaining the number of convolution kernels of the convolution hardware, it further includes: Judging whether the number of channels of each group is greater than the maximum parallelism of the convolution hardware; When the number of channels of each group is greater than the maximum parallelism of the convolution hardware, performing the step of obtaining the number of convolution kernels of the convolution hardware.

6. The data processing method according to claim 5, wherein After the step of judging whether the number of channels of each group is greater than the maximum parallelism of the convolution hardware, it further includes: When the number of channels of each group is less than or equal to the maximum parallelism of the convolution hardware, obtaining the number of convolution kernels of each group; Distributing the input data and convolution coefficients to the computing units according to the number of convolution kernels of each group, so that the computing units perform convolution operations on the input data.

7. An accelerator, characterized in that, The accelerator includes a memory, a processor, and a data processing program stored on the memory and executable on the processor. When the data processing program is executed by the processor, it implements the steps of the data processing method according to any one of claims 1-6.

8. A data processing device, characterized in that, The data processing device includes the accelerator according to claim 7.

9. A data processing device, characterized in that, The data processing device includes: an acquisition module, configured to acquire the number of channels of input data and the number of groups when performing grouped convolution on the input data. The number of groups when performing grouped convolution on the input data refers to the number of groups into which the input data is divided when performing grouped convolution, and is preset; a determination module, configured to determine the number of channels of each group according to the number of channels of the input data and the number of groups; a distribution module, configured to compare the relationship between the number of channels of each group and half of the maximum parallelism of the convolution hardware, and distribute the input data to different computing units in a shifted manner according to the comparison result; when the number of channels of each group is less than or equal to half of the maximum parallelism of the convolution hardware, distribute the input data and convolution coefficients to the computing units according to the number of convolution kernels of each group. The number of convolution kernels of each group is equal to the number of channels of each group, so that the computing units perform convolution operations on the input data. The number of channels of each group is a first preset value, and the maximum parallelism of the convolution hardware is a second preset value. In a convolution hardware with the second preset value of convolution kernels in parallel, the amount of data calculated in each round becomes twice the original. The first first preset value of convolution kernels perform the operations of the first group, and the latter first preset value of convolution kernels perform the operations of the second group.

10. A readable storage medium, characterized in that, A data processing program is stored on the readable storage medium. When the data processing program is executed by a processor, it implements the steps of the data processing method according to any one of claims 1-6.

Citation Information

Patent Citations

  • An OPU instruction set definition method for CNN acceleration

    CN110058882A

  • Grouping convolution process optimization method for embedded platform

    CN110516796A

  • Grouping convolution hardware accelerator based on FPGA and method thereof

    CN111445012A

  • Convolutional neural network accelerator and working method thereof

    CN113312285A