Pixel parallel depth operation implementation method and device, medium, equipment and product

Through pixel parallel strategy, deep computing is performed on the artificial intelligence processor, the problem of mismatch between thread scheduling granularity and task allocation is solved, the maximum utilization of hardware resources is achieved, and the computing efficiency is improved.

CN120298196AActive Publication Date: 2025-07-11SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510796726.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-11
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

When the prior art performs deep operations on artificial intelligence processors, the thread scheduling granularity does not match the task allocation, resulting in a decrease in computing power utilization, especially in deep-deep convolution operations.

Method used

The pixel parallel strategy is adopted to divide the input tensor into multiple channel groups, and each thread is assigned the calculation task of the local sliding window field. Through the pixel parallel strategy between threads, independent computing task allocation is realized and hardware resource utilization is maximized.

Benefits of technology

It improves computing efficiency, makes full use of hardware resources, reduces thread idleness, and improves the computing power utilization rate of artificial intelligence processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298196A_ABST
    Figure CN120298196A_ABST
Patent Text Reader

Abstract

The invention discloses a pixel parallel depth operation implementation method and device, a medium, equipment and a product, and the method comprises the steps: dividing an input tensor into m channel groups, and distributing the channel groups to a thread bundle containing NT threads for independent processing; obtaining the number of elements, participating in coverage operation, of a sliding window operator in a two-dimensional plane, a sliding step length and row and column starting point coordinates of initial coverage of each element; allocating register groups with the same number as the elements for each thread; each register group is used for storing a pixel group of a corresponding element; in the channel group, NT row and column positions are scanned by rows from each row and column starting point coordinate according to the sliding step length, and corresponding pixel groups are loaded to register groups of corresponding threads in sequence, so that each thread can process a calculation task of a local sliding window view. According to the method, thread-level independent calculation task allocation is realized through a pixel parallel strategy among threads, and allocated hardware resources can be utilized to the maximum extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, computer-readable storage medium, electronic device, and computer program product for implementing depth operations in parallel with pixels. Background Art

[0002] When performing depth operations on an artificial intelligence processor, such as depthwise convolution operations, pooling operations, and grouped convolution operations, etc., the existing methods adopt a channel parallel strategy among threads, which will lead to a reduction in the computing power utilization rate of the artificial intelligence processor. The reason lies in the mismatch between the thread scheduling granularity and the task allocation.

[0003] Taking the depthwise convolution operation as an example, the existing method assigns the convolution calculation task of a single channel to each thread. Since the artificial intelligence processor uses a warp as the scheduling unit, and each warp contains N T threads, when the number of channels C of the input tensor cannot be divided evenly by N T , it will cause some threads in a warp to be idle, that is, the enabled threads cannot be fully utilized, and the computing power utilization rate of the artificial intelligence processor drops to . Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method, apparatus, computer-readable storage medium, electronic device, and computer program product for implementing depth operations in parallel with pixels. Through the pixel parallel strategy among threads, independent computing task allocation at the thread level is realized, and the computing task of the corresponding local sliding window view is assigned to each thread, which can maximize the utilization of the allocated hardware resources and thus improve the computing efficiency.

[0005] The first aspect embodiment of the present invention provides a method for implementing depth operations in parallel with pixels, including: Dividing an input tensor into m channel groups and allocating them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of channels; N T ≥1; N C ≥1; m≥1; Obtaining the number of elements, the sliding step, and the starting row and column coordinates initially covered by each element when a sliding window operator participates in the covering operation in a two-dimensional plane; Allocating a register group equal to the number of elements to each of the threads; where each register group is used to store a pixel group corresponding to the elements; each pixel group is composed of all the pixels at the same row and column position in the corresponding channel group; In the channel group, starting from each of the row-column starting coordinates, with the sliding step as the span, scan N row-column positions row by row, and sequentially load the pixel groups at the scanned N row-column positions into the register groups of each thread within the corresponding warp; T The N T pixel groups at the row-column positions are loaded into the register groups of each thread within the corresponding warp; Each thread calculates the loaded pixel group to obtain the output result within the corresponding local sliding window field of view.

[0006] Optionally, the input tensor is composed of one feature data block, or is obtained by splicing multiple feature data blocks in the vertical direction of the original space, or is composed of any sub-data block split from one feature data block.

[0007] Optionally, the sliding window operator is a pooling operator; The step of each thread calculating the loaded pixel group to obtain the output result within the corresponding local sliding window field of view includes: In any warp, each thread performs a channel-isolated pooling operation on the loaded pixel group to obtain the pooling results of N channels within the corresponding local sliding window field of view. C The pooling results of the N channels are obtained.

[0008] Optionally, the sliding window operator is a single-channel convolution kernel; The step of each thread calculating the loaded pixel group to obtain the output result within the corresponding local sliding window field of view includes: Broadcast the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any warp, each thread performs a channel-isolated multiply-accumulation operation on the loaded pixel group and the convolution weights to obtain the convolution results of N channels within the corresponding local sliding window field of view. C The convolution results of the N channels are obtained.

[0009] Optionally, the sliding window operator is a multi-channel convolution kernel; The step of each thread calculating the loaded pixel group to obtain the output result within the corresponding local sliding window field of view includes: Broadcast the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any warp, each thread performs a channel-fused multiply-accumulation operation on the loaded pixel group and the convolution weights to obtain an output feature value within the corresponding local sliding window field of view; wherein, each channel group is used to represent an input group in the grouped convolution operation.

[0010] Optionally, the method further includes: Before the thread performs the multiply-add operation, perform an out-of-bounds check on the coordinates of the pixel values in the pixel group, and mask the pixel values corresponding to the out-of-bounds coordinates with a specified value to obtain new pixel values.

[0011] An embodiment of the second aspect of the present invention provides a device for implementing pixel-parallel depth operations, including: A warp configuration module, configured to divide the input tensor into m channel groups and allocate them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of a channel; N T ≥1; N C ≥1; m≥1; A sliding window parameter acquisition module, configured to acquire the number of elements involved in the covering operation of the sliding window operator in the two-dimensional plane, the sliding step, and the starting row and column coordinates initially covered by each element; A thread configuration module, configured to allocate a register group equal to the number of elements to each of the threads; where each register group is used to store a pixel group corresponding to the element; each pixel group is composed of all pixels at the same row and column position in the corresponding channel group; A pixel loading module, configured to, in the channel group, starting from each of the row and column starting coordinates, scan N T row and column positions with the sliding step as the span, and sequentially load the pixel groups at the scanned N T row and column positions into the register groups of the respective threads within the corresponding warp; A thread calculation module, configured to calculate the pixel groups loaded by each of the threads to obtain the output results within the corresponding local sliding window view.

[0012] An embodiment of the third aspect of the present invention provides a computer-readable storage medium, where the computer-readable storage medium includes a stored computer program; where the computer program, when running, controls the device where the computer-readable storage medium is located to execute the pixel-parallel depth operation implementation method according to any one of the above first aspects.

[0013] An embodiment of the fourth aspect of the present invention provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements the pixel-parallel depth operation implementation method according to any one of the above first aspects.

[0014] An embodiment of the fifth aspect of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the pixel-parallel depth operation implementation method according to any one of the above first aspects.

[0015] Compared with the prior art, the embodiments of the present invention provide a method, device, computer-readable storage medium, electronic device and computer program product for implementing depth operation with pixel parallelism, having the following beneficial effects: In the embodiments of the present invention, an input tensor is first divided into multiple channel groups and assigned to different warps; then, the number of elements involved in the covering operation of the sliding window operator in the two-dimensional plane, the sliding step, and the starting row and column coordinates initially covered by each element are obtained; subsequently, a corresponding number of register groups are assigned to each thread, and each register group stores a pixel group corresponding to an element; then, starting from each row and column starting coordinate of each channel group, with the sliding step as the span, N T row and column positions are scanned row by row, and the scanned pixel groups are successively loaded into the register groups of the corresponding threads, so that each thread can process the calculation tasks of the corresponding local sliding window view. Therefore, through the pixel parallelism strategy among threads, the embodiments of the present invention realize the allocation of independent calculation tasks at the thread level, assign the calculation tasks of the corresponding local sliding window view to each thread, can maximize the use of the allocated hardware resources, and thus improve the calculation efficiency. Description of the Drawings

[0016] Figure 1 is a schematic flowchart of an embodiment of the implementation of depth operation with pixel parallelism provided by the present invention.

[0017] Figure 2 is a schematic diagram of an embodiment of the sliding of the sliding window operator in the input tensor provided by the present invention; Figure 3 is a schematic diagram of an embodiment of the projection of the input tensor on the two-dimensional plane provided by the present invention; Figure 4 is a schematic diagram of an embodiment of the mapping relationship between the warp and the pixel group provided by the present invention; Figure 5 is a schematic diagram of an embodiment of loading the pixel group into the warp provided by the present invention; Figure 6 is a schematic diagram of another embodiment of loading the pixel group into the warp provided by the present invention; Figure 7 is an example diagram of channel parallelism when a thread executes a depthwise convolution operation provided by the prior art; Figure 8 is a schematic structural diagram of an embodiment of the device for implementing depth operation with pixel parallelism provided by the present invention; Figure 9 is a schematic structural diagram of an embodiment of the electronic device provided by the present invention. Detailed Embodiments

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art in the technical field of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] The artificial intelligence processor involved in the present invention can be any one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural network Processing Unit), a DPU (Deeplearning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit), which is determined when the embodiments of the present invention are applied to specific products or technologies.

[0020] In addition, in the embodiments of the present invention, the so-called "pixel" is not the traditional image pixel value, but an abstract unit used to represent any type of data. Specifically, a "pixel" corresponds to a data unit at a certain position in the input tensor.

[0021] Taking the GPU as an example, the following describes a method, device, computer-readable storage medium, electronic device, and computer program product for implementing pixel-parallel depth operations provided by the embodiments of the present invention.

[0022] See Figure 1 , which is a schematic flowchart of an embodiment for implementing pixel-parallel depth operations provided by the present invention.

[0023] The first aspect of the embodiments of the present invention provides a method for implementing pixel-parallel depth operations, including steps S1 to S5, which are specifically as follows: Step S1: Divide the input tensor into m channel groups and allocate them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of channels; N T ≥1; N C ≥1; m≥1; Step S2: Obtain the number of elements that the sliding window operator participates in the covering operation in the two-dimensional plane, the sliding step, and the starting row and column coordinates of the initial coverage of each element; Step S3: Allocate a register group equal to the number of elements to each of the threads; wherein, each register group is used to store a pixel group of the corresponding element; each pixel group is composed of all the pixels at the same row and column position in the corresponding channel group; Step S4: In the channel group, starting from each of the row and column starting coordinates, with the sliding step as the span, scan N T row and column positions, and sequentially load the pixel groups at the scanned N T row and column positions into the register groups of the respective threads within the corresponding warp; Step S5: Calculate the pixel groups loaded by each of the threads to obtain the output result within the corresponding local sliding window view.

[0024] It should be noted that the sliding window operator in the embodiments of the present invention includes: a pooling operator, a single-channel convolution kernel (which can be used for depthwise convolution operations), or a multi-channel convolution kernel (which can be used for grouped convolution operations).

[0025] In step S1, the input tensor with dimensions H×W×C is grouped along the channel dimension to generate m channel groups, and each channel group contains N C pixels of consecutive channels; wherein, H is the height of the input tensor, W is the width of the input tensor, and C is the total number of channels of the input tensor; the dimension of each channel group is H×W×N C . Further, allocate an independent warp to each channel group; wherein, each warp contains N T threads (such as 32 threads).

[0026] In step S2, obtain the attribute parameters of the sliding window operator in the two-dimensional plane, including: the number of elements that the sliding window operator participates in the covering operation in the two-dimensional plane, the starting row and column coordinates of the initial coverage of each element, and the sliding step; wherein, the two-dimensional plane is a plane determined by the vertical and horizontal directions of the sliding window operator or the input tensor, and has nothing to do with the depth / channel dimension. As Figure 2 shown, it is a schematic diagram of an embodiment of the sliding window operator provided by the present invention sliding in the input tensor. Figure 2 The coordinates of the pixel data are , that is, the th row and the th column in the data block, which also has nothing to do with the channel dimension; wherein, , , , and and co-locate pixel data at an actual row position on a two-dimensional plane (i.e., row position = + -1). As Figure 3 shown, it is a schematic diagram of an embodiment of the projection of the input tensor provided by the present invention on a two-dimensional plane. In Figure 3 , are the row and column position coordinates of the pixel data; the sliding window operator is a 3×3 single-channel convolution kernel, and the number of elements it participates in the covering operation in the two-dimensional plane is 9, and the elements are convolution weights ; among them, is the position offset of the convolution weight from the center of the convolution kernel in the vertical direction, is the position offset of the convolution weight from the center of the convolution kernel in the horizontal direction. Therefore, the 9 elements are respectively , , , , , , , and .

[0027] As Figure 2 and Figure 3 shown, the gray area ① with the center point is used as the starting local sliding window view (initial covering area) of the 3×3 convolution kernel in the current calculation batch. The row and column starting coordinates of the weight are , the row and column starting coordinates of the weight are , the row and column starting coordinates of the weight are , ……, the row and column starting coordinates of the weight are , that is, the set of row and column starting coordinates of the sliding window operator on the two-dimensional plane is { , , , , , , , , }; the sliding step is 1, that is, starting from the gray area ① with the center point , sliding 1 step by row, to obtain the dotted-line enclosed area ② with the center . Since the above coordinates are independent of the depth (channel) dimension, the 3×3 single-channel convolution kernel is extended to a multi-channel convolution kernel (such as 3×3×N CAfter that, its attribute parameters in the two-dimensional plane are still the same as the above content.

[0028] In step S3, K dedicated register groups are allocated to each thread in the warp, forming a sliding window data buffer; where K is the height of the sliding window operator and K is the width of the sliding window operator. For example, for a 3×3 convolution kernel, each thread will be allocated 9 dedicated register groups. Each register group is used to store the pixel values of N channels at the same row and column position. When the register can store 32-bit floating-point numbers and each pixel value is a 16-bit floating-point number, if N = 2, each register group (labeled Rx) consists of only 1 register (labeled rx); if N > 2, each register group consists of H ×K W registers. As shown in H K W , it is a schematic diagram of an embodiment of the mapping relationship between the warp and the pixel group provided by the present invention. In C , the input tensor contains 8 channels (i.e., C0, C1,..., C7). According to the partitioning rule of N = 2, the input tensor is divided into 4 channel groups. Therefore, 4 warps are enabled, and any pixel group processed by each thread in the warp is 2-channel pixels at the same row and column position in the corresponding channel group. C =2, then each register group (labeled Rx) consists of only 1 register (labeled rx); if N C >2, then each register group consists of registers. As Figure 4 shown, it is a schematic diagram of an embodiment of the mapping relationship between the warp and the pixel group provided by the present invention. In Figure 4 , the input tensor contains 8 channels (i.e., C0, C1,..., C7). According to the partitioning rule of N C =2, the input tensor is divided into 4 channel groups. Therefore, 4 warps are enabled, and any pixel group processed by each thread in the warp is 2-channel pixels at the same row and column position in the corresponding channel group.

[0029] In addition, there is a mapping relationship between each register group in the thread and the corresponding element in the sliding window operator. As Figure 5 shown, it is a schematic diagram of an embodiment of loading the pixel group into the warp provided by the present invention. In Figure 5 , the sliding window operator is a 3×3 convolution kernel. Warp 0 processes the channel group containing channels C0 and C1, and each thread has 9 dedicated register groups (i.e., R0, R1,..., R8) respectively; each register group corresponds to a set of ( , ) convolution weights. For example, register group R0 (which can be represented by register r1 when N C =2) is used to store the pixel group , } corresponding to a set of convolution weights with an offset of (-1,-1) (composed of two channel pixels corresponding to the coordinates ); similarly, register group R1 is used to store the pixel group corresponding to a set of convolution weights with an offset of { , } It can be seen therefrom that one thread in any warp can hold N channel pixels corresponding to a local window view of the sliding window operator. C Pixels of N channels.

[0030] In step S4, in the i-th channel group, the j-th row-column starting coordinate in the set of row-column starting coordinates of the sliding window operator is used as the starting point of the j-th row scan route, and with a sliding step as the span, N row-column starting coordinates are scanned, and the pixel groups at these N row-column positions are sequentially loaded into the j-th register group of each thread in the i-th warp; where 1 ≤ i ≤ m; 1 ≤ j ≤ K T Row-column starting coordinates, and the pixel groups at these N T Row-column positions are loaded into the j-th register group of each thread in the i-th warp; where 1 ≤ i ≤ m; 1 ≤ j ≤ K H ×K W .

[0031] To describe the technical solution provided by the embodiments of the present invention more clearly, the following provides some specific embodiments for reference: Example 1: As shown in Figure 2 , 3 And 5, the pixel group of the 1st channel group (corresponding to channels C0 and C1) is loaded into 32 threads of warp 0. For the 1st register group R0 (or register r0) exclusive to each of the 32 threads, starting from the 1st row-column starting coordinate In the set of row-column starting coordinates of the sliding window operator, with a sliding step of 1 as the span, 32 row-column coordinates are continuously read (the last row-column coordinate read is ), and the pixel groups corresponding to these 32 row-column coordinates To Are sequentially filled into the register group R0 exclusive to each of the 32 threads; for the 2nd register group R1 (or register r1) exclusive to each of the 32 threads, starting from the 2nd row-column starting coordinate In the set of row-column starting coordinates of the sliding window operator, with a sliding step of 1 as the span, 32 row-column coordinates are continuously read (the last row-column coordinate read is ), and the pixel groups corresponding to these 32 row-column coordinates To Are sequentially filled into the register group R1 exclusive to each of the 32 threads; following the above similar operations, continue to complete the data loading of the register groups R2~R8 exclusive to each of the 32 threads of warp 0. Therefore, the local window view corresponding to thread T0 is the gray area ① centered on , the local window view corresponding to thread T1 is the dotted-line enclosed area ② centered on , ……, the local window view corresponding to thread T31 is the dotted-line enclosed area ③ centered on .

[0032] Example 2: As shown inFigure 6 As shown, it is a schematic diagram of another embodiment in which the pixel groups provided by the present invention are loaded into a warp. In Figure 6 , the input tensor has been stored in the shared memory, and each memory cell stores two-channel pixels corresponding to a row-column position, and the pixels are marked with the sequential number of the row-by-row loading order and the channel number; for example, (p -1 C0, p -1 C1) is equivalent to the above-mentioned pixel group , and p -1 C0 is the pixel value of channel C0 in this pixel group. In the current calculation batch, the input tensor contains 8 channels (i.e., C0, C1,..., C7), and the channel groups are divided according to N C = 2, then 4 warps are enabled as execution units (ExecutiveUnit, EU). Each thread is allocated K H × K W dedicated register groups. If the sliding window operator is a 3×3 convolution kernel, then each thread is allocated 9 dedicated register groups (i.e., R0, R1,..., R8). In the channel group corresponding to each execution unit, starting from the 9 row-column starting coordinates of the sliding window operator, 9-way data loading operations are performed to continuously load 32 groups of pixel data in a row-by-row scanning manner, and the data is filled into the register groups with corresponding numbers.

[0033] In step S5, each thread independently processes K H × K W × N C pixel values and supports multiple operation modes, such as convolution operation and pooling operation.

[0034] The existing method adopts a channel parallel strategy among threads; for example, if the input tensor has 90 channels, then 2 warps (128 threads) need to be enabled, but when these two warps perform a convolution calculation, they can only process 90 local sliding window fields of view (corresponding to the calculation tasks of 90 channels), and the computing power of the hardware resources cannot be fully utilized. In contrast, the embodiment of the present invention adopts a pixel parallel strategy among threads, that is, a local sliding window field of view parallel strategy; when the above 2 warps perform a convolution calculation, they can process N C channel pixels corresponding to 128 local sliding window fields of view. Therefore, the embodiment of the present invention realizes the allocation of independent calculation tasks at the thread level through the pixel parallel strategy among threads, allocates the calculation tasks of the corresponding local sliding window field of view to each thread, can maximize the utilization of the allocated hardware resources, and thus improves the calculation efficiency.

[0035] In an optional embodiment, the input tensor is composed of one feature data block, or is obtained by splicing multiple feature data blocks in the vertical direction of the original space, or is composed of any sub-data block split from one feature data block.

[0036] It should be noted that the input tensor in the embodiments of the present invention can be obtained through the following three main data organization methods, as long as the condition of warp parallel processing of data and reasonable resource allocation is met: (1) Single feature data block method: If the dimension of the feature data block is [H, W, C], and on the two-dimensional plane formed by the vertical and horizontal directions, the number N of positions that can be covered by the center point of the local sliding window field of view center ≤N T , and at the same time the warp resources are sufficient to cover the total number C of channels of the entire feature data block, then the input tensor is directly composed of the complete feature data block; this method is applicable to the scenario where the feature data block size is small and the computing resources are sufficient.

[0037] (2) Vertical splicing method: If q feature data blocks are vertically spliced along the height dimension to form a composite tensor with the dimension [qH, W, C], and the corresponding N of this composite tensor center ≤N T , and at the same time the warp resources are sufficient to cover the total number C of channels of the entire feature data block, then this composite tensor is used as the input tensor; this method is suitable for the scenario where multiple data blocks need to be integrated for unified processing.

[0038] (3) Dynamic splitting method: ① Row and column position splitting, when the corresponding N of the feature data block center >N T , then it can be split into multiple sub-data blocks perpendicular to the two-dimensional plane direction according to the condition that N center =N T ; ② Channel dimension splitting: If the warp resources cannot cover the total number C of channels of the feature data block, then according to the limitation of the warp resources, the feature data block is split into multiple sub-data blocks in the channel dimension. These sub-data blocks will be processed in batches, and when calculating each batch, one of the sub-data blocks is used as the input tensor. This method is applicable to the scenario where the feature data block size is large or the computing resources are limited, and the reasonable utilization of computing resources is achieved through splitting.

[0039] It is worth noting that the method for obtaining the input tensor is not limited to the individual use of the above three methods, and can also be any combination of multiple methods, and the embodiments of the present invention do not limit this. For example, multiple feature data blocks can be first vertically spliced along the vertical direction of the original space into a new feature data block (such as Figure 2 ), and then, when the condition N center =N T is met, the new feature data block is split perpendicular to the two-dimensional plane direction to obtain multiple input tensors.

[0040] The embodiments of the present invention have a high degree of flexibility in constructing the input tensor and can adapt to different characteristic data scales and computing resource conditions.

[0041] In an optional embodiment, the sliding window operator is a pooling operator; The step of calculating, by each of the threads, the loaded pixel groups to obtain the output result within the corresponding local sliding window view includes: In any one of the thread warps, each thread performs a channel-isolated pooling operation on the loaded pixel groups to obtain the pooling results of N C channels within the corresponding local sliding window view.

[0042] It should be noted that, for each thread, the total data dimension formed by all the pixel groups it loads is K H ×K W ×N C . During the calculation process, the thread will perform independent pooling operations for each channel. Common pooling operations include max pooling, average pooling, etc. Taking max pooling as an example, any one thread will operate on each channel separately within its responsible local sliding window view, select the maximum pixel value in the channel as the pooling result of the channel. In this way, each thread finally obtains the pooling results of N C channels within the corresponding local sliding window view.

[0043] In an optional embodiment, the sliding window operator is a single-channel convolution kernel; The step of calculating, by each of the threads, the loaded pixel groups to obtain the output result within the corresponding local sliding window view includes: Broadcasting the convolution weights corresponding to each register group to all threads within the corresponding thread warp in scalar form; In any one of the thread warps, each thread performs a channel-isolated multiply-accumulate operation on the loaded pixel groups and the convolution weights to obtain the convolution results of N C channels within the corresponding local sliding window view.

[0044] It should be noted that, as Figure 7 shown, it is an example diagram of channel parallelism when a thread executes a depthwise convolution operation provided by the prior art. In Figure 7Among them, the total number of channels of the input tensor is 12. Therefore, only 12 threads (T0 to T11) participate in the convolution operation task, and the remaining threads (T12 to T31) are idle. While each thread loads the pixel data within the local sliding window view of the corresponding channel, it also needs to load the convolution kernel corresponding to that channel. For example, thread T0 loads the 3×3 convolution kernel W0 corresponding to channel C0, and also loads the covered pixels of the current convolution kernel W0 in channel C0; similarly, thread T10 loads the 3×3 convolution kernel W10 corresponding to channel C 10 and also loads the covered pixels of the current convolution kernel W10 in channel C 10 . Obviously, the pixel data and convolution kernels loaded by different threads are different. Therefore, when performing the convolution operation, each thread needs to process three-way input vector operands: the first is the vector of input feature data, the second is the vector of convolution weights, and the third is the vector of accumulated inputs. However, the current GPU architecture is more suitable for processing instructions with no more than two-way vector operands. Therefore, the third operand usually needs to wait, resulting in a decrease in the instruction execution efficiency. Based on this, the existing technology has a high requirement for the read bandwidth of vector registers, further restricting the improvement of computing performance.

[0045] To solve the above problems, the embodiment of the present invention changes the input convolution weights (the second-way vector operand) to scalar operands, thereby reducing the input operands of the multiply-accumulate instruction to two-way vector operands. This improvement not only better adapts to the characteristics of the GPU architecture, enabling the two-way instructions to be executed in parallel, but also significantly reduces the read bandwidth requirement of the vector registers. At the same time, the scalar operand and the two-way vector operands can be synchronized and parallel, thereby reducing the instruction waiting time and further improving the computing efficiency.

[0046] Specifically, in the embodiment of the present invention, each thread within any warp holds the data within the local sliding window view of its corresponding channel group. Since these sliding window data are in the same channel group, all threads within the warp need to access the same convolution weights. Based on this characteristic, the present invention broadcasts the convolution weights in scalar form to all threads within the warp, enabling all threads within the same warp to share the same convolution weights. In other words, the convolution weights are shared in scalar form within the warp. As Figure 5 shown, T0 to T31 of warp 0 share K H ×K W ×N C convolution weights. Finally, each thread independently processes the pixels on each channel and the shared convolution weights, performs element-wise multiplication operations and then accumulates the results to finally obtain the convolution result of the local sliding window view within the corresponding channel.

[0047] In an optional embodiment, the sliding window operator is a multi-channel convolution kernel; Calculating, by each of the threads, the loaded pixel groups to obtain output results within a corresponding local sliding window field of view, including: Broadcasting the convolution weights corresponding to each of the register groups to all threads within the corresponding warp in scalar form; In any one of the warps, performing a multiply-accumulation operation for channel fusion on the loaded pixel groups and the convolution weights through each thread to obtain an output feature value within a corresponding local sliding window field of view; wherein each of the channel groups is used to represent an input group in the grouped convolution operation.

[0048] It should be noted that in the grouped convolution operation, each channel group is used to represent an input group in the grouped convolution operation. For each input group, the corresponding warp will independently complete the calculation of one output channel in the grouped calculation through 1 shared multi-channel convolution kernel. For example, assume that the first channel group is used as the first input group (with dimensions H×W×N C ), and the number of output channels per group is 5. Then the first warp needs to complete the calculations of these 5 output channels respectively and share 5 multi-channel convolution kernels (with dimensions K H ×K W ×N C ), that is, when performing the calculation of one output channel, the 32 threads of the first warp jointly access one multi-channel convolution kernel.

[0049] During the calculation process, each thread is responsible for processing one output feature value within the local sliding window field of view. For the K H ×K W ×N C pixel values obtained by any one thread, after multiplying each by the corresponding convolution weight element-wise and then accumulating all the product values, the thread outputs a convolution result value, which serves as an element in the corresponding output feature layer.

[0050] In an optional embodiment, the method further includes: Before the threads perform the multiply-accumulation operation, determining whether the coordinates of the pixel values in the pixel groups are out of bounds, and masking the pixel values corresponding to the out-of-bounds coordinates with a specified value to obtain new pixel values.

[0051] It should be noted that the operation of filling the specified value in the embodiments of the present invention (taking filling 0 as an example) includes but is not limited to the following two methods: (1) After the register groups of the threads obtain the pixel groups, based on the coordinates of the center points within the local sliding window field of view and the offsets of each convolution weight, determining whether the coordinates of the corresponding convolution weight at the current covered position point are out of bounds; if out of bounds, masking the pixel values corresponding to the out-of-bounds coordinates with a specified value (such as masking with 0).

[0052] (2) During the execution of the scanning operation, while loading the pixel groups, actively determine whether the pixel coordinates are out of bounds; if out of bounds, mask the pixel values corresponding to the out-of-bounds coordinates as a specified value, and write the specified value into the corresponding register group; otherwise, still write the original pixel values into the register group.

[0053] In addition to the above two methods, other suitable masking strategies can also be adopted according to actual requirements and hardware architectures, and the embodiments of the present invention do not limit this.

[0054] See Figure 8 , which is a schematic structural diagram of an embodiment of the apparatus for realizing pixel-parallel depth operation provided by the present invention.

[0055] An embodiment of the second aspect of the present invention provides an apparatus for realizing pixel-parallel depth operation, including: A warp configuration module 11, configured to divide an input tensor into m channel groups and allocate them to m warps for independent processing; wherein, each warp contains N T threads; each channel group contains N C pixels of channels; N T ≥1; N C ≥1; m≥1; A sliding window parameter acquisition module 12, configured to acquire the number of elements, the sliding step, and the initial row and column starting coordinates covered by each element of a sliding window operator in a two-dimensional plane; A thread configuration module 13, configured to allocate a register group equal to the number of elements to each of the threads respectively; wherein, each register group is used to store a pixel group corresponding to the element; each pixel group is composed of all pixels at the same row and column position in the corresponding channel group; A pixel loading module 14, configured to, in the channel group, starting from each of the row and column starting coordinates, with the sliding step as the span, scan N T row and column positions by rows, and sequentially load the pixel groups at the scanned N T row and column positions into the register groups of each thread within the corresponding warp; A thread calculation module 15, configured to calculate the pixel groups loaded by each thread to obtain the output results within the corresponding local sliding window view.

[0056] It should be noted that the pixel-parallel depth operation implementation device provided in the embodiments of the second aspect of the present invention can implement all the processes of the pixel-parallel depth operation implementation method described in any of the embodiments of the first aspect. The functions and the achieved technical effects of each module and unit in the device respectively correspond to and are the same as those of the pixel-parallel depth operation implementation method described in any of the embodiments of the first aspect, and will not be elaborated here.

[0057] An embodiment of the third aspect of the present invention provides a computer-readable storage medium, which includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the pixel-parallel depth operation implementation method described in any of the embodiments of the first aspect.

[0058] An embodiment of the fourth aspect of the present invention provides a computer program product, including a computer program, which when executed by a processor, implements the pixel-parallel depth operation implementation method described in any of the embodiments of the first aspect.

[0059] See Figure 9 , which is a schematic structural diagram of an embodiment of the electronic device provided by the present invention.

[0060] An embodiment of the fifth aspect of the present invention provides an electronic device, including a processor 21, a memory 22, and a computer program stored in the memory 22 and configured to be executed by the processor 21. When the processor executes the computer program, it implements the pixel-parallel depth operation implementation method described in any of the embodiments of the first aspect.

[0061] Preferably, the computer program can be divided into one or more modules / units (such as computer program 1, computer program 2,...). The one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.

[0062] The processor 21 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor 21 may also be any conventional processor. The processor 21 is the control center of the electronic device, and connects various parts of the electronic device through various interfaces and circuits.

[0063] The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc., and the data storage area can store relevant data, etc. In addition, the memory 22 may be a high-speed random access memory, or may also be a non-volatile memory, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., or the memory 22 may also be other volatile solid-state storage devices.

[0064] It should be noted that the above-mentioned electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 9 The structural block diagram shown is only an example of the structure of the above-mentioned electronic device, and does not constitute a limitation on the structure of the above-mentioned electronic device. The above-mentioned electronic device may include more or fewer components than shown in the figure, or combine certain components, or different components.

[0065] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can still be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A method for implementing pixel-parallel depth operations, characterized in that, including: Divide the input tensor into m channel groups and assign them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of channels; N T ≥ 1; N C ≥ 1; m ≥ 1; obtaining the number of elements involved in the covering operation of the sliding window operator in the two-dimensional plane, the sliding step, and the starting row and column coordinates initially covered by each element; allocating a register group equal to the number of elements to each of the threads; wherein each register group is used to store a pixel group corresponding to the corresponding element; each pixel group is composed of all pixels at the same row and column position in the corresponding channel group; In the channel group, starting from each of the row-column starting coordinates respectively, with the sliding step as the span, scan N row-column positions row by row, and sequentially load the pixel groups at the scanned N row-column positions into the register groups of the respective threads within the corresponding warp; T where N T is the number of row-column positions, and load the pixel groups at the scanned N row-column positions into the register groups of the respective threads within the corresponding warp in sequence; calculating, by each of the threads, the loaded pixel group to obtain an output result within the corresponding local sliding window view.

2. The method for implementing depth operation in pixel parallelism according to claim 1, wherein The input tensor is composed of one feature data block, or is obtained by splicing multiple feature data blocks in the vertical direction of the original space, or is composed of any sub-data block split from one feature data block.

3. The method for implementing depth operation in parallel with pixels according to claim 1, wherein The sliding window operator is a pooling operator; The calculating, by each of the threads, the loaded pixel group to obtain an output result within the corresponding local sliding window view includes: In any one of the thread bundles, each thread performs a channel isolation pooling operation on the loaded pixel group to obtain the pooling results of N C channels within the corresponding local sliding window view.

4. The method for implementing depth operation in pixel parallelism according to claim 1, wherein The sliding window operator is a single-channel convolution kernel; The calculating, by each of the threads, the loaded pixel group to obtain an output result within the corresponding local sliding window view includes: broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any one of the warp, each thread performs a channel-isolated multiply-add operation on the loaded pixel group and the convolution weights to obtain the convolution results of N C channels within the corresponding local sliding window view.

5. The method for implementing depth operation in pixel parallelism according to claim 1, wherein The sliding window operator is a multi-channel convolution kernel; The calculating, by each of the threads, the loaded pixel group to obtain an output result within the corresponding local sliding window view includes: broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; in any one of the warps, performing a multiply-accumulate operation of channel fusion on the loaded pixel group and the convolution weights by each thread to obtain an output feature value within the corresponding local sliding window view; wherein each of the channel groups is used to represent an input group in the grouped convolution operation.

6. The method for implementing pixel-parallel depth operation according to claim 4 or 5, characterized in that The method further includes: before the threads perform the multiply-accumulate operation, determining whether the coordinates of the pixel values in the pixel group are out of bounds, and masking the pixel values corresponding to the out-of-bounds coordinates with a specified value to obtain new pixel values.

7. An apparatus for implementing depth operations in parallel for pixels, characterized in that, including: The warp configuration module is used to divide the input tensor into m channel groups and allocate them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of channels; N T ≥ 1; N C ≥ 1; m ≥ 1; a sliding window parameter acquisition module, configured to obtain the number of elements involved in the covering operation of the sliding window operator in the two-dimensional plane, the sliding step, and the starting row and column coordinates initially covered by each element; a thread configuration module, configured to allocate a register group equal to the number of elements to each of the threads; wherein each register group is used to store a pixel group corresponding to the corresponding element; each pixel group is composed of all pixels at the same row and column position in the corresponding channel group; A pixel loading module, configured to, in the channel group, respectively start from each of the row-column starting coordinates, with the sliding step as the span, scan N row-column positions row by row, and sequentially load the pixel groups at the scanned N row-column positions into the register groups of the respective threads within the corresponding warp; T where N T is the number of row-column positions, and load the pixel groups at the scanned N row-column positions into the register groups of the respective threads within the corresponding warp in sequence; a thread calculation module, configured to calculate, by each of the threads, the loaded pixel group to obtain an output result within the corresponding local sliding window view.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein the computer program, when running, controls the device where the computer-readable storage medium is located to execute the pixel-parallel depth operation implementation method according to any one of claims 1 to 6.

9. A computer program product, characterized in that, including a computer program, which, when executed by a processor, implements the pixel-parallel depth operation implementation method according to any one of claims 1 to 6.

10. An electronic device, characterized in that, Comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for implementing pixel-parallel depth operation as described in any one of claims 1 to 6 is realized.

Citation Information

Patent Citations

  • Data reading method and device, computer equipment and storage medium

    CN117851080A

  • Model operation optimization method, product, equipment and medium

    CN118277133A

  • Data processing method and device, processor, electronic equipment and storage medium

    CN119312003A

  • Convolution operation filling value generation method, convolution operation filling value application method, convolution operation filling value generation device, convolution operation filling value application device, medium, equipment and product

    CN119939095A

  • Object supply system and control method thereof

    KR1020250064426A

Cited By

  • Convolution weight gradient calculation method and device, medium and equipment

    CN120952065A

  • Operator calculation method, electronic equipment and storage medium

    CN121052300A