Vertical direction pixel parallel depth operation implementation method and device, medium, equipment and product
By employing a vertical pixel-parallel computation strategy that divides tasks among threads, the method addresses inefficiencies in thread scheduling, resulting in improved computational efficiency and resource utilization on AI processors.
Patent Information
- Application Number
- CN202510796725.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-16
AI Technical Summary
When the prior art performs deep operations on artificial intelligence processors, the thread scheduling granularity does not match the task allocation, resulting in a decrease in computing power utilization, especially in deep-deep convolution operations.
The vertical pixel parallel strategy is adopted to divide the input tensor into multiple channel groups, and each thread is assigned the calculation task of the local sliding window field. Through the pixel parallel strategy between threads, independent computing task allocation and resource maximization utilization is achieved.
Improve computing efficiency, make full use of hardware resources, reduce memory access conflicts, and improve computing performance.
Smart Images

Figure CN120318057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, computer-readable storage medium, electronic device, and computer program product for implementing depth operations with pixel parallelism in the vertical direction. Background Art
[0002] When performing depth operations on an artificial intelligence processor, such as depthwise convolution operations, pooling operations, and grouped convolution operations, etc., the existing methods adopt a channel parallel strategy among threads, which will lead to a reduction in the computing power utilization rate of the artificial intelligence processor. The reason is the mismatch between the thread scheduling granularity and task allocation.
[0003] Taking the depthwise convolution operation as an example, the existing method assigns the convolution calculation task of a single channel to each thread. Since the artificial intelligence processor uses a warp as the scheduling unit, and each warp contains N T threads, when the number of channels C of the input tensor cannot be divided by N T evenly, it will cause some threads in a warp to be idle, that is, the enabled threads cannot be fully utilized, and the computing power utilization rate of the artificial intelligence processor is reduced to . Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method, apparatus, computer-readable storage medium, electronic device, and computer program product for implementing depth operations with pixel parallelism in the vertical direction. Through the pixel parallel strategy among threads, independent calculation task allocation at the thread level is realized, and the calculation task of the corresponding local sliding window view is assigned to each thread, which can maximize the utilization of the allocated hardware resources and thus improve the calculation efficiency.
[0005] The first aspect embodiment of the present invention provides a method for implementing depth operations with pixel parallelism in the vertical direction, including:[[]] Dividing the input tensor into m channel groups and allocating them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of channels; N T ≥1; N C ≥1; m≥1; Obtaining the jump step for fetching numbers in the vertical direction in the channel group, as well as the height, width, vertical step of the sliding window operator, and N H row base addresses; where N H is equal to the height; Allocating N H register clusters to each of the threads; where each register cluster contains y register groups; each register group is used to store NC Pixels of a channel form a pixel group; y is not less than the width and is equal to the number of column slices participating in the calculation in each channel group; Based on the row base address and the vertical step, determine N H row starting coordinates for the current scan round; When performing the current scan round, in each column slice participating in the calculation, starting from each of the row starting coordinates, with the jump step as the span, scan N T row-column positions, and sequentially load the pixel groups at the scanned N T row-column positions into the register groups of each thread within the corresponding warp; Each thread calculates the loaded pixel group to obtain the output result within the corresponding local sliding window view.
[0006] Optionally, the jump step is the product of the height and the vertical step.
[0007] Optionally, the method further includes: Through the execution of N H rounds of scans, each warp completes the coverage of all sliding window areas of the sliding window operator in the corresponding channel group.
[0008] Optionally, when performing the current scan round, in each column slice participating in the calculation, starting from each of the row starting coordinates, scan N T row-column positions according to the jump step, and sequentially load the pixel groups at the scanned N T row-column positions into the register groups of each thread within the corresponding warp, including: When performing the i-th round of vertical data scanning, in the w-th column slice participating in the calculation within the j-th channel group, use the h-th row starting coordinate as the starting point of the h-th vertical scanning route, with the jump step as the span, scan N T row-column positions, and sequentially load the pixel groups at the N T row-column positions into the w-th register group of the h-th register cluster of each thread within the j-th warp; where 1 ≤ i ≤ N H ; 1 ≤ j ≤ m; 1 ≤ w ≤ y; 1 ≤ h ≤ N H .
[0009] Optionally, the input tensor is composed of one feature data block, or is obtained by splicing multiple feature data blocks in the vertical direction of the original space, or is composed of any sub-data block split from one feature data block.
[0010] Optionally, the column slice participating in the calculation is the column slice traversed by the sliding window operator during horizontal sliding.
[0011] Optionally, determining N row starting coordinates of the current scan round according to the row base address and the vertical step length includes: H Determining the row starting coordinates through the following formula: ; where is the th row starting coordinate in the th scan round; is the th row base address; is the vertical step length.
[0012] Optionally, the sliding window operator is a pooling operator; The calculating, by each of the threads, the loaded pixel group to obtain an output result within a corresponding local sliding window view includes: In any one of the warps, performing a channel-isolated pooling operation on the loaded pixel group by each thread to obtain pooling results of N channels within a corresponding local sliding window view. C
[0013] Optionally, the sliding window operator is a single-channel convolution kernel; The calculating, by each of the threads, the loaded pixel group to obtain an output result within a corresponding local sliding window view includes: Broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any one of the warps, performing a channel-isolated multiply-accumulate operation on the loaded pixel group and the convolution weights by each thread to obtain convolution results of N channels within a corresponding local sliding window view. C
[0014] Optionally, the sliding window operator is a multi-channel convolution kernel; The calculating, by each of the threads, the loaded pixel group to obtain an output result within a corresponding local sliding window view includes: Broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any one of the warps, performing a channel-fused multiply-accumulate operation on the loaded pixel group and the convolution weights by each thread to obtain an output feature value within a corresponding local sliding window view; wherein each channel group is used to represent an input group in grouped convolution operations.
[0015] Optionally, the method further includes: Before the thread performs the multiply-accumulate operation, perform an out-of-bounds determination on the coordinates of the pixel values in the pixel group, and mask the pixel values corresponding to the out-of-bounds coordinates with a specified value to obtain new pixel values.
[0016] An embodiment of the second aspect of the present invention provides a vertical-direction pixel parallel depth operation implementation device, including: A warp configuration module, configured to divide an input tensor into m channel groups and allocate them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of N T channels; N C ≥ 1; N ≥ 1; m ≥ 1; A parameter acquisition module, configured to acquire a jump step for vertical-direction data fetching in the channel group, and the height, width, vertical step of the sliding window operator, and N H row base addresses; where N H is equal to the height; A thread configuration module, configured to allocate N H register clusters to each of the threads respectively; where each register cluster contains y register groups; each register group is used to store a pixel group composed of N C pixels of N channels at the same row and column position; y is not less than the width and is equal to the number of column slices participating in the calculation in each channel group; A starting point determination module, configured to determine N H row starting point coordinates of the current scan round according to the row base address and the vertical step; A pixel loading module, configured to, when performing the current scan round, in each column slice participating in the calculation, respectively start from each of the row starting point coordinates, scan N T row-column positions with the jump step as the span, and sequentially load the pixel groups at the scanned N T row-column positions into the register groups of each thread within the corresponding warp; A thread calculation module, configured to calculate the pixel groups loaded by each of the threads to obtain the output results within the corresponding local sliding window view.
[0017] An embodiment of the third aspect of the present invention provides a computer-readable storage medium, where the computer-readable storage medium includes a stored computer program; where the computer program, when running, controls the device where the computer-readable storage medium is located to execute the vertical-direction pixel parallel depth operation implementation method according to any one of the above first aspects.
[0018] An embodiment of the fourth aspect of the present invention provides a computer program product, including a computer program which, when executed by a processor, implements the method for implementing depth operation with vertical pixel parallelism according to any one of the above-mentioned first aspects.
[0019] An embodiment of the fifth aspect of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for implementing depth operation with vertical pixel parallelism according to any one of the above-mentioned first aspects.
[0020] Compared with the prior art, the embodiment of the present invention provides a method for implementing depth operation with vertical pixel parallelism, which has the following beneficial effects: In the embodiment of the present invention, the input tensor is first divided into multiple channel groups and assigned to different warps for independent processing; then, combined with the parameter configuration of the sliding window operator, register resources are allocated for each thread to form a sliding window data buffer; then, the N H starting coordinates of rows in the current scanning round are calculated, and in each column slice participating in the calculation, starting from each starting coordinate of the row, with the jump step as the span, N T row-column positions are scanned, and the pixel groups at the N T scanned row-column positions are sequentially loaded into the register groups of each thread within the corresponding warp, so that each thread can process the calculation tasks of the corresponding local sliding window view. Therefore, through the pixel parallel strategy between threads, the embodiment of the present invention realizes the allocation of independent calculation tasks at the thread level, assigns the calculation tasks of the corresponding local sliding window view to each thread, can maximize the utilization of the allocated hardware resources, and thus improves the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic flowchart of an embodiment of the implementation of depth operation with vertical pixel parallelism provided by the present invention; Figure 2 is a schematic diagram of an embodiment of the row base address provided by the present invention; Figure 3 is a schematic diagram of an embodiment of performing multi-channel vertical scanning in a channel group provided by the present invention; Figure 4 is a schematic diagram of an embodiment of the relationship between a pixel group and a register group provided by the present invention; Figure 5 is a schematic diagram of an embodiment of a thread holding local sliding window view data provided by the present invention; Figure 6 is a schematic diagram of another embodiment of the relationship between a pixel group and a register group provided by the present invention; Figure 7It is a schematic diagram of another embodiment of the thread holding local sliding window vision data provided by the present invention; Figure 8 It is an example diagram of channel parallelism when a thread performs depthwise convolution operations provided by the prior art; Figure 9 It is a schematic structural diagram of an embodiment of a depth operation implementation device with pixel parallelism in the vertical direction provided by the present invention; Figure 10 It is a schematic structural diagram of an embodiment of an electronic device provided by the present invention. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art in the technical field without creative efforts belong to the scope of protection of the present invention.
[0023] The artificial intelligence processor involved in the present invention can be any one of a CPU (Central Processing Unit, central processing unit), a GPU (Graphics Processing Unit, graphics processing unit), a TPU (Tensor Processing Unit, tensor processing unit), an NPU (Neural network Processing Unit, neural network processing unit), a DPU (Deeplearning Processing Unit, deep learning processing unit), an APU (Accelerated Processing Unit, accelerated processing unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit, general-purpose graphics processing unit), which is determined when the embodiments of the present invention are applied to specific products or technologies.
[0024] In addition, in the embodiments of the present invention, the so-called "pixel" is not the traditional image pixel value, but an abstract unit used to represent any type of data. Specifically, a "pixel" corresponds to a data unit at a certain position in the input tensor.
[0025] Next, taking the GPU as an example, the method, device, computer-readable storage medium, electronic device, and computer program product for implementing depth operation with pixel parallelism in the vertical direction provided by the embodiments of the present invention will be described.
[0026] See Figure 1, which is a schematic flowchart of an embodiment for implementing depth operation with vertical pixel parallelism provided by the present invention.
[0027] An embodiment of the first aspect of the present invention provides a method for implementing depth operation with vertical pixel parallelism, including steps S1 to S6, specifically as follows: Step S1: Divide the input tensor into m channel groups and assign them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of N T channels; N C ≥ 1; N ≥ 1; m ≥ 1; Step S2: Obtain the jump step for vertical data fetching in the channel group, as well as the height, width, vertical step of the sliding window operator, and N H row base addresses; where N H is equal to the height; Step S3: Allocate N H register clusters for each of the threads; where each register cluster contains y register groups; each register group is used to store a pixel group composed of N C pixels of N channels at the same row and column positions; y is not less than the width and is equal to the number of column slices participating in the calculation in each channel group; Step S4: Determine the N H row start coordinates of the current scan round according to the row base address and the vertical step; Step S5: When executing the current scan round, in each column slice participating in the calculation, starting from each of the row start coordinates, scan N T row and column positions with the jump step as the span, and sequentially load the pixel groups at the scanned N T row and column positions into the register groups of the respective threads within the corresponding warp; Step S6: Calculate the pixel groups loaded by each of the threads to obtain the output results within the corresponding local sliding window view.
[0028] It should be noted that the sliding window operator in the embodiments of the present invention includes: a pooling operator, a single-channel convolution kernel (which can be used for depthwise convolution operation), or a multi-channel convolution kernel (which can be used for grouped convolution operation).
[0029] In step S1, the input tensor with dimensions H×W×C is divided into m channel groups along the channel dimension, and each channel group forms a binding relationship with a warp containing N T threads (N TThe normal value is 32); where H is the height of the input tensor, W is the width of the input tensor, C is the total number of channels of the input tensor, and N C is the number of channels included in each channel group; the dimension of each channel group is H×W×N C .
[0030] In step S2, obtain the skip step (i.e., the distance between rows) for fetching data in the vertical direction, which is used to achieve cross-row access in the vertical direction in the channel group; for example, the skip step S jump is 3. If the currently read data is on the first row, the next read data will be on the fourth row.
[0031] Meanwhile, obtain the height K H and width K W of the sliding window operator, the vertical step , and the set of row base addresses ; where is used to represent the set of row coordinates corresponding to the starting positions covered by each element in the sliding window operator. Obviously the number N H of row base addresses in H is equal to the height K Figure 2 of the sliding window operator. As Figure 2 shown, it is a schematic diagram of an embodiment of the row base address provided by the present invention. In , are the row and column position coordinates of the pixel data; where is the row coordinate and is the column coordinate. The sliding window operator is a 3×3 single-channel convolutional kernel, and the number of elements participating in the covering operation in the two-dimensional plane is 9, and the elements are convolutional weights ; where is the position offset of the convolutional weight from the convolutional center in the vertical direction, and is the position offset of the convolutional weight from the convolutional center in the horizontal direction; the two-dimensional plane is determined by the vertical and horizontal directions of the input tensor and has nothing to do with the depth / channel dimension. When the 3×3 single-channel convolutional kernel performs the first covering operation, it covers a part of the pixel data in 3 rows in the two-dimensional plane, and the set of row coordinates of these 3 rows is represented as . In other words, the row base address refers to the row coordinate where each row of elements in the sliding window operator first covers in the two-dimensional plane; for example, if the row coordinate where the first row of convolutional weights in the 3×3 single-channel convolutional kernel first covers in the two-dimensional plane is -1, then the first row base address is -1; similarly, for the third row of convolutional weights The initially covered row coordinate in the two-dimensional plane is 1, so the 3rd row base address is 1. Since the above row base address is independent of the depth (channel) dimension, when expanding the 3×3 single-channel convolution kernel into a multi-channel convolution kernel (such as 3×3×N C ) the resulting remains equal to .
[0032] In step S3, N H register clusters are allocated to each thread within a warp, and each register cluster contains y register groups, that is, a single thread is allocated N H ×y dedicated register groups to form a sliding window data buffer; where y is the number of column slices participating in the calculation in each channel group, and the dimension of the column slice is H×1×N C . In addition, each register group is used to store a pixel group composed of N C channel pixels at the same row and column position, so the dimension of the pixel group is 1×1×N C . When the register can store 32-bit floating-point numbers and each pixel value is a 16-bit floating-point number, if N C = 2, each register group consists of only 1 register; if N C > 2, each register group consists of registers.
[0033] In the case where the number of column slices y is equal to the sliding window operator width K W , each thread has K H ×K W register groups. At this time, each thread can store N C channel pixels within a local sliding window operator (that is, the data dimension held by the thread is K H ×K W ×N C ). In the case where the number of column slices y is greater than the sliding window operator width K W , it indicates that the thread has a coverage data preloading mechanism for horizontal sliding of the sliding window operator to avoid repeated memory access; for example, when y = K W + 1 and the horizontal step = 1, the thread can process the data of two adjacent local sliding window fields in the horizontal direction. Through this preloading mechanism, when calculating the second local sliding window field, the overlapping part of the data with the first local sliding window field can be directly reused, avoiding rereading the memory, thereby effectively reducing unnecessary memory access overhead.
[0034] In step S4, a rolling mechanism for the row start coordinate is established according to the vertical step . In the first round of vertical data scanning, N HThe starting coordinates of each row are equal to N H units. In each subsequent round of vertical data scanning, increment the starting coordinates of the previous round by units to generate new starting coordinates for the rows. For example, when the set of starting coordinates for the first round of scanning is {-1, 0, 1}, and the vertical step size is 1, the set of starting coordinates for the second round of scanning will become {0, 1, 2}. This rolling mechanism ensures that after multiple rounds of scanning, each warp can cover the entire sliding window area within the corresponding channel group.
[0035] In step S5, when performing the current scanning round, each column slice initiates N H vertical scanning routes. Each route starts from the corresponding starting coordinate of the row, with a jump step size as the span, and continuously accesses N T row-column positions, and sequentially loads the pixel groups at these N T row-column positions into the register groups corresponding to each thread. For example, assume that the set of starting coordinates for the first round of scanning is {-1, 0, 1}, and N T = 32. The first vertical scanning route of the first column slice in the first channel group starts from the starting coordinate -1, with a jump step size as the span, continuously obtains 32 pixel groups, and sequentially loads these 32 pixel groups into the register group R0 in register cluster 0 of the 32 threads in warp 0; similarly, the second vertical scanning route of this column slice starts from the starting coordinate 0, with a jump step size as the span, continuously obtains 32 pixel groups, and sequentially loads these 32 pixel groups into the register group R0 in register cluster 1 of the 32 threads in warp 0.
[0036] In step S6, each thread independently processes K H ×y×N C pixel values (y ≥ K W ). As mentioned above, when y = K W , each thread holds K H ×K W ×N C pixel values; when y > K W , the thread can pre-load the coverage data for the horizontal sliding of the sliding window operator. Therefore, each thread holds at least N C channel pixels of a local sliding window view and supports multiple operation modes, such as convolution operation and pooling operation.
[0037] It should be noted that the existing method adopts a channel parallel strategy among threads. For example, for an input tensor with 90 channels, 2 warps (i.e., 128 threads) need to be enabled to complete the calculation. However, when these two warps perform a convolution calculation, they can only process 90 local sliding window views (corresponding to the calculation tasks of 90 channels), and the computing power of the hardware resources cannot be fully utilized.
[0038] In contrast, the embodiment of the present invention adopts a pixel parallel strategy among threads, that is, a local sliding window view parallel strategy. Based on this strategy, when the above-mentioned 2 warps perform a convolution calculation, they can process at least N C channel pixels corresponding to 128 local sliding window views. Therefore, the embodiment of the present invention realizes the allocation of independent calculation tasks at the thread level through the pixel parallel strategy among threads, assigns the calculation tasks of the corresponding local sliding window views to each thread, and can maximize the utilization of the allocated hardware resources, thereby improving the calculation efficiency.
[0039] In an optional embodiment, the skip step is the product of the height and the vertical step.
[0040] Further, through the execution of N H rounds of scans, each warp completes the coverage of all sliding window areas of the sliding window operator in the corresponding channel group.
[0041] It should be noted that the embodiment of the present invention determines the skip step through the following formula: S jump = × a; where S jump is the skip step; is the vertical step; a is the coverage ratio factor, and a is a positive integer, which is used to control the ratio of the local sliding window views covered by each warp after one round of scan in the corresponding channel group to the total local sliding window views (i.e., 1 / a).
[0042] In order to describe the technical solution provided by the embodiment of the present invention more clearly, the following provides some specific embodiments for reference: Example 1: When a = 1, S jump = , then when warp 0 performs one round of scan in channel group 0, the sliding window operator slides in the vertical direction according to the vertical step (column-first order). Therefore, such a scanning method is a 1:1 global coverage and only one round of scan is required.
[0043] Example 2: When a = K H , S jump = × K H, when warp 0 performs a round of scans in channel group 0, the sliding window operator slides vertically with a step size of ×K H . This is equivalent to sparsely reading windows by skipping (K - 1) local sliding window views from the originally dense local sliding window view (corresponding to the step size H ). Therefore, such a scanning method is a sparse coverage of 1:K H , and K H rounds of misaligned scans need to be performed to vertically cover all sliding window areas in channel group 0. In other words, each warp only covers 1 / K H = 1 / N H of the sliding window area (sparse coverage) after a round of scans. This phased sparse coverage design splits the originally dense sliding window task into K H (N H ) rounds, and each round processes different sparse subsets.
[0044] It should be noted that a = K H is a preferred solution in the embodiment of the present invention. In any round of scans, the warp jumps the complete height distance of the sliding window operator each time, and the data of the local sliding window views read are completely staggered vertically (no overlapping areas), so that when performing multiple vertical scan routes in the column slice, the row and column position points of each scan are also staggered and non-overlapping, effectively avoiding memory access conflicts.
[0045] In an alternative embodiment, determining the N H row starting coordinates for the current scan round according to the row base address and the vertical step size includes: Determining the row starting coordinates through the following formula: ; where is the -th row starting coordinate in the -th round of scans; is the -th row base address; is the vertical step size.
[0046] Combining the above embodiments, in order to achieve complete sliding window coverage in the channel group, multiple rounds of misaligned scans in the vertical direction need to be performed, and the original sliding rule of the sliding window operator in the channel (i.e., sliding according to the vertical step size ) needs to be followed. Therefore, the row starting coordinates of the previous round are shifted by the vertical step size Increment to achieve the rolling update of the row starting coordinates in the current scan round, ensuring that the remaining sliding window area is covered in the next round of scanning. Thus, after multiple rounds of scanning are completed, each warp can cover the entire sliding window area within the corresponding channel group.
[0047] As Figure 3 shown, it is a schematic diagram of an embodiment of multi-channel vertical scanning in a channel group provided by the present invention. In Figure 3 , channel group 0 contains pixels of 2 channels (N C = 2); represents a pixel group and is composed of all pixels at the same row and column positions in the channel group. The pixel group has coordinates of ; where is the row coordinate, and is the column coordinate. The sliding window operator is a 3×3 single-channel convolution kernel, and the horizontal step is 1. Therefore, the column slices 0, 1, and 2 participating in the calculation are adjacent and continuous. The height K H of the sliding window operator is 3, which means that 3 vertical scanning routes need to be started for each column slice. The vertical step is 1, and the coverage ratio factor a = K H . Therefore, in the current scan round of the warp, each time it jumps a complete height distance of the sliding window operator, the locally read sliding window view data is completely staggered in the vertical direction, as shown in Table 1.
[0048] Table 1. Example table of data coverage based on vertical scanning In Figure 3 , all column slices participating in the calculation have the same scanning paths indicated by the gray arrow, black arrow, and dashed arrow. To make the picture display clearer, the embodiments of the present invention only selectively show some scanning paths. As can be seen from Table 1, the first scan of the first round corresponds to the scanning path indicated by the gray arrow; the second scan of the first round corresponds to the scanning path indicated by the black arrow; the third scan of the first round corresponds to the scanning path indicated by the dashed arrow. In any round of scanning, the row positions scanned by the 3 vertical scanning routes of each column slice are completely staggered (no overlapping area), so there will be no memory access conflict. In addition, since each round of scanning only covers 1 / K H proportion of the sliding window area (sparse coverage), through the rolling update of the row starting coordinates, different sparse subsets are processed in each round of scanning. After K H = 3 rounds of scanning, the sparse coverage is accumulated into full coverage to achieve the complete access of the channel group.
[0049] In an optional embodiment, when performing the current scan round, in each column slice participating in the calculation, starting from each of the row starting coordinates, according to the jump step size, scan N T row-column positions, and sequentially load the pixel groups at the scanned N T row-column positions into the register groups of each thread within the corresponding warp, including: When performing the data scan in the vertical direction in the i-th round, in the w-th column slice participating in the calculation within the j-th channel group, use the h-th row starting coordinate as the starting point of the h-th vertical scan route, with the jump step size as the span, scan N T row-column positions, and sequentially load the pixel groups at the N T row-column positions into the w-th register group of the h-th register cluster of each thread within the j-th warp; where 1 ≤ i ≤ N H ; 1 ≤ j ≤ m; 1 ≤ w ≤ y; 1 ≤ h ≤ N H .
[0050] As Figure 4 shown, it is a schematic diagram of an embodiment of the relationship between the pixel group and the register group provided by the present invention. In Figure 4 , Rx represents the label of the register group, and Tt represents the label of the thread within a warp; where 0 ≤ x ≤ y - 1, 0 ≤ t ≤ N T -1. Combining Figure 3 and Figure 4 it can be known that the specific mapping relationship between the pixel group and the register group in the embodiment of the present invention is: in the w-th column slice participating in the calculation within the j-th channel group, the k-th pixel group on the h-th vertical scan route is stored in the w-th register group of the h-th register cluster of the k-th thread within the j-th warp; where 1 ≤ k ≤ N T .
[0051] In an optional embodiment, the input tensor is composed of one feature data block, or is obtained by splicing multiple feature data blocks in the vertical direction of the original space, or is composed of any sub-data block split from one feature data block.
[0052] It should be noted that the input tensor in the embodiment of the present invention can be obtained through the following 3 main data organization methods, as long as the condition of warp parallel data processing and reasonable resource allocation is satisfied: (1) Single feature data block method: If the dimension of the feature data block is [H, W, C], all the sliding window regions it contains can meet the requirement of the warp to perform multiple rounds of data scanning in the vertical direction, and the warp resources are sufficient to cover the total number of channels C of the entire feature data block, then the input tensor is directly composed of the complete feature data block. This method is applicable to the scenario where the size of the feature data block is small and the computing resources are sufficient.
[0053] (2) Vertical splicing method: When there are q feature data blocks, they can be vertically spliced along the height dimension to form a composite tensor with the dimension of [qH, W, C]. If the sliding window region of this composite tensor meets the multi-round scanning requirement of the warp in the vertical direction, and the warp resources can cover the total number of channels C of the composite tensor, then this composite tensor is used as the input tensor. This method is applicable to the scenario where multiple feature data blocks need to be integrated for unified processing.
[0054] (3) Dynamic splitting method: ① Horizontal splitting: If the number of column slices participating in the calculation in any channel group of the feature data block exceeds the number of register groups that a single register cluster can accommodate, the feature data block can be split into multiple sub-data blocks along the horizontal direction to ensure that the number of column slices in each sub-data block does not exceed the capacity of the register cluster; ② Vertical splitting: If the number of pixel groups contained in any column slice of the feature data block exceeds the data requirement that the vertical scanning route can handle, the data block can be split along the vertical direction to generate multiple sub-data blocks to ensure that the number of vertical pixel groups in each column slice adapts to the processing capacity of the scanning route; ③ Channel dimension splitting: If the warp resources cannot cover the total number of channels C of the feature data block, the feature data block can be split into multiple sub-data blocks according to the limitation of the warp resources in the channel dimension. These sub-data blocks will be processed in batches. In each batch, only one of the sub-data blocks is used as the input tensor for calculation. This method is applicable to the scenario where the size of the feature data block is large or the computing resources are limited, and reasonable utilization of computing resources is achieved through splitting.
[0055] It should be noted that the method for obtaining the input tensor is not limited to the single use of the above three methods, but can also be any combination of multiple methods. The embodiments of the present invention do not make any limitations in this regard. For example, multiple feature data blocks can be first vertically spliced along the vertical direction of the original space into a large-sized feature data block, and then the large-sized feature data block can be split along the vertical direction according to whether the number of pixel groups in any column slice exactly meets the data requirement of the vertical scanning route, and finally multiple input tensors are obtained for batch processing.
[0056] In summary, the embodiments of the present invention can flexibly select the optimal construction method of the input tensor according to the actual hardware resource limitations and specific computing requirements, thereby maximizing the computing efficiency while ensuring the rationality of resource allocation.
[0057] In an alternative embodiment, the column slices participating in the calculation are the column slices traversed by the sliding window operator when sliding in the horizontal direction.
[0058] It should be noted that in the embodiments of the present invention, the dimension of the column slice is H×1×N C . When selecting the column slices participating in the calculation, the original sliding rule of the sliding window operator in the channel also needs to be followed (i.e., sliding according to the horizontal step ). Therefore, the column slices participating in the calculation are the column slices traversed by the sliding window operator when sliding in the horizontal direction.
[0059] In addition, the number y of column slices participating in the calculation in any channel group is greater than or equal to the width K of the sliding window operator W .
[0060] For example Figure 3 in the shown channel group, the number y of column slices participating in the calculation = K W = 3, and the corresponding warp executes N H = K H = 3 rounds of scans. After that, each thread in the warp holds K H ×K W ×N C pixel values. As Figure 5 shown, it is a schematic diagram of an embodiment in which a thread holds local sliding window view data provided by the present invention. In Figure 5 , each thread holds N C channel pixels of a local sliding window view (N C = 2), and C0 and C1 respectively represent the first channel and the second channel.
[0061] Refer to Figure 6 and Figure 7 , Figure 6 is a schematic diagram of another embodiment of the relationship between the pixel group and the register group provided by the present invention, Figure 7 is a schematic diagram of another embodiment in which a thread holds local sliding window view data provided by the present invention. In Figure 6 , in the channel group (N C = 2), the number y of column slices participating in the calculation = 4. Therefore, each register cluster has 4 register groups, and the corresponding warp executes N H = K H = 3 vertical scan routes. After that, each thread in the warp holds K H ×y×N Cpixel values. In the embodiments of the present invention, when y > K W , it is possible to preload the covered data when the sliding window operator slides in the horizontal direction. In Figure 7 , since the horizontal step size = 1, each thread holds N C -channel pixels of 2 local sliding window views. In short, each thread holds at least N C -channel pixels of one local sliding window view to support various operation modes, such as convolution operation and pooling operation.
[0062] In an optional embodiment, the sliding window operator is a pooling operator; The calculation of the loaded pixel groups by each thread to obtain the output result within the corresponding local sliding window view includes: In any one of the warps, each thread performs a channel-isolated pooling operation on the loaded pixel group to obtain the pooling results of N C channels within the corresponding local sliding window view.
[0063] It should be noted that for each thread, the total data dimension formed by all the loaded pixel groups is K H × y × N C . During the calculation process, the thread will perform independent pooling operations for each channel. Common pooling operations include max pooling, average pooling, etc. Taking max pooling as an example, any one thread will operate on each channel separately within the local sliding window view (at least 1) it is responsible for, and select the maximum pixel value in the channel as the pooling result. In this way, each thread finally obtains the pooling results of N C channels within the corresponding local sliding window view.
[0064] In an optional embodiment, the sliding window operator is a single-channel convolution kernel; The calculation of the loaded pixel groups by each thread to obtain the output result within the corresponding local sliding window view includes: Broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any one of the warps, each thread performs a channel-isolated multiply-accumulate operation on the loaded pixel group and the convolution weights to obtain the convolution results of N C channels within the corresponding local sliding window view.
[0065] It should be noted that, as Figure 8 shows, it is an example diagram of channel parallelism when a thread executes a depthwise convolution operation provided by the prior art. In Figure 8 , the total number of channels of the input tensor is 12 (i.e., channels C0, C1, …, C11 ), so only 12 threads (T0 to T11) are involved in the convolution operation task, and the remaining threads (T12 to T31) are idle. While each thread loads the pixel data of the local sliding window view within the corresponding channel, it also needs to load the convolution kernel corresponding to that channel. For example, thread T0 loads the 3×3 convolution kernel W0 corresponding to channel C0, and also loads the covered pixels of the current convolution kernel W0 in channel C0; similarly, thread T10 loads the 3×3 convolution kernel W10 corresponding to channel C 10 and also loads the covered pixels of the current convolution kernel W10 in channel C 10 . Obviously, the pixel data and convolution kernels loaded by different threads are different. Therefore, when performing the convolution operation, each thread needs to process three-way input vector operands: the first way is the vector of input feature data, the second way is the vector of convolution weights, and the third way is the vector of accumulated inputs. However, the current GPU architecture is more suitable for processing instructions with no more than two-way vector operands. Therefore, the third-way operand usually needs to wait, resulting in a decrease in the instruction execution efficiency. Based on this, the prior art has a high requirement for the read bandwidth of vector registers, further restricting the improvement of computing performance.
[0066] To solve the above problems, the embodiments of the present invention change the input convolution weights (the second-way vector operand) to scalar operands, thereby reducing the input operands of the multiply-accumulate instruction to two-way vector operands. This improvement not only better adapts to the characteristics of the GPU architecture, enabling the two-way instructions to be executed in parallel, but also significantly reduces the read bandwidth requirement of the vector registers. At the same time, the scalar operand and the two-way vector operands can be synchronized and parallel, thereby reducing the instruction waiting time and further improving the computing efficiency.
[0067] Specifically, in the embodiments of the present invention, each thread within a warp holds the data within the local sliding window view of its corresponding channel group. Since these sliding window data are in the same channel group, all threads within the warp need to access the same convolution weights. Based on this characteristic, the present invention broadcasts the convolution weights in scalar form to all threads within the warp, enabling all threads within the same warp to share the same convolution weights. In other words, the convolution weights are shared in scalar form within the warp.
[0068] As Figure 4 shown, each register group corresponds to a set of ( , ) convolution weights, that is, ; where is the position offset of the convolution weight from the convolution center in the vertical direction, is the position offset of the convolution weight from the convolution center in the horizontal direction; is the single-channel convolution kernel corresponding to channel C0 with an offset position of ( , convolution weights of (). Specifically, in warp 0 corresponding to channel group 0, threads T0 to T31 share K H ×K W ×N C convolution weights; for example, when calculating the first local sliding window view, register group R0 in register cluster 0 of threads T0 to T31 all need to read convolution weights , register group R1 in register cluster 0 of threads T0 to T31 all need to read convolution weights , ……, register group R2 in register cluster 2 of threads T0 to T31 all need to read convolution weights . Finally, each thread independently processes the pixels on each channel and the shared convolution weights, performs element-wise multiplication operations and then accumulates the results to finally obtain the convolution result of the local sliding window view within the corresponding channel.
[0069] In an optional embodiment, the sliding window operator is a multi-channel convolution kernel; The step of each thread calculating the loaded pixel group to obtain the output result within the corresponding local sliding window view includes: Broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any warp, each thread performs a multiply-accumulate operation for channel fusion on the loaded pixel group and convolution weights to obtain an output feature value within the corresponding local sliding window view; wherein, each channel group is used to represent an input group in the grouped convolution operation.
[0070] It should be noted that in the grouped convolution operation, each channel group is used to represent an input group in the grouped convolution operation. For each input group, the corresponding warp will independently complete the calculation of one output channel in the grouped calculation through 1 shared multi-channel convolution kernel. For example, assuming that the first channel group is the first input group (dimension H×W×N C ), and the number of output channels per group is 5, then the first warp needs to complete the calculations of these 5 output channels respectively and share 5 multi-channel convolution kernels (dimension K H ×K W ×N C ), that is, when performing the calculation of one output channel, the 32 threads of the first warp jointly access one multi-channel convolution kernel.
[0071] During the calculation process, each thread is responsible for processing one output feature value within the local sliding window view. For any K obtained by a thread H ×y×N Cpixel values, the thread extracts a data block (dimension K H ×K W ×N C ) from the local sliding window view, and then the thread multiplies the data of the local sliding window view element by element with the corresponding convolution weights, and accumulates all the product values to obtain the convolution result of the local sliding window view, which is used as an element in the corresponding output feature layer.
[0072] In an optional embodiment, the method further includes: Before the thread performs the multiplication and addition operation, it determines whether the coordinates of the pixel values in the pixel group are out of bounds, and masks the pixel values corresponding to the out-of-bounds coordinates with a specified value to obtain new pixel values.
[0073] It should be noted that the operation of filling the specified value in the embodiment of the present invention (taking filling 0 as an example) includes but is not limited to the following two methods: (1) After the thread's register group obtains the pixel group, through the coordinates of the center point in the local sliding window view and the offsets of each convolution weight, it determines whether the coordinates of the corresponding convolution weight at the current coverage position point are out of bounds; if out of bounds, the pixel values corresponding to the out-of-bounds coordinates are masked with the specified value (such as masked as 0).
[0074] (2) During the process of performing the scanning operation, while loading the pixel group, it actively determines whether the pixel coordinates are out of bounds; if out of bounds, the pixel values corresponding to the out-of-bounds coordinates are masked with the specified value, and the specified value is written into the corresponding register group; otherwise, the original pixel values are still written into the register group.
[0075] In addition to the above two methods, other suitable masking strategies can also be adopted according to actual requirements and hardware architectures, and the embodiments of the present invention do not limit this.
[0076] See Figure 9 , which is a schematic structural diagram of an embodiment of the vertical pixel parallel depth operation implementation device provided by the present invention.
[0077] An embodiment of the second aspect of the present invention provides a vertical pixel parallel depth operation implementation device, including: A warp configuration module 11, configured to divide the input tensor into m channel groups and allocate them to m warps for independent processing; wherein, each warp contains N T threads; each channel group contains N C channels of pixels; N T ≥1; N C ≥1; m≥1; A parameter acquisition module 12, configured to acquire the skip step for vertical fetching in the channel group, and the height, width, vertical step of the sliding window operator, and NH row base addresses; where N H is equal to the height; The thread configuration module 13 is configured to allocate N H register clusters to each of the threads; where each register cluster includes y register groups; each register group is used to store a pixel group composed of N C channel pixels at the same row and column position; y is not less than the width and is equal to the number of column slices participating in the calculation in each channel group; The starting point determination module 14 is configured to determine N H row starting point coordinates of the current scanning round according to the row base address and the vertical step size; The pixel loading module 15 is configured to, when performing the current scanning round, in each column slice participating in the calculation, respectively start from each of the row starting point coordinates, with the jump step size as the span, scan N T row and column positions, and sequentially load the pixel groups at the scanned N T row and column positions into the register groups of each thread within the corresponding warp; The thread calculation module 16 is configured to calculate the pixel groups loaded by each thread to obtain the output results within the corresponding local sliding window view.
[0078] It should be noted that the vertical pixel parallel depth operation implementation device provided in the second aspect embodiment of the present invention can implement all the processes of the vertical pixel parallel depth operation implementation method described in any embodiment of the first aspect above. The functions and achieved technical effects of each module and unit in the device respectively correspond to the functions and achieved technical effects of the vertical pixel parallel depth operation implementation method described in any embodiment of the first aspect above, and will not be elaborated here.
[0079] The third aspect embodiment of the present invention provides a computer-readable storage medium, and the computer-readable storage medium includes a stored computer program; where the computer program controls the device where the computer-readable storage medium is located to execute the vertical pixel parallel depth operation implementation method described in any embodiment of the first aspect above when running.
[0080] The fourth aspect embodiment of the present invention provides a computer program product, including a computer program, and the computer program implements the vertical pixel parallel depth operation implementation method described in any embodiment of the first aspect above when being executed by a processor.
[0081] See Figure 10 , which is a schematic structural diagram of an embodiment of an electronic device provided by the present invention.
[0082] An embodiment of the fifth aspect of the present invention provides an electronic device, including a processor 21, a memory 22, and a computer program stored in the memory 22 and configured to be executed by the processor 21. When the processor executes the computer program, the method for implementing depth operation with vertical pixel parallelism according to any embodiment of the first aspect described above is implemented.
[0083] Preferably, the computer program can be divided into one or more modules / units (such as computer program 1, computer program 2,...). The one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.
[0084] The processor 21 can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 21 can also be any conventional processor. The processor 21 is the control center of the electronic device, and connects various parts of the electronic device through various interfaces and lines.
[0085] The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc., and the data storage area can store relevant data, etc. In addition, the memory 22 can be a high-speed random access memory, or can also be a non-volatile memory, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., or the memory 22 can also be other volatile solid-state storage devices.
[0086] It should be noted that the above-mentioned electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 10The structural block diagram shown is only an example of the structure of the above-mentioned electronic device and does not constitute a limitation on the structure of the above-mentioned electronic device. The above-mentioned electronic device may include more or fewer components than those shown, or combine certain components, or different components.
[0087] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A method for implementing depth operation with vertical pixel parallelism, characterized in that Including: Divide the input tensor into m channel groups and assign them to m warps for independent processing; where each warp contains N T threads; each channel group contains the pixels of N C channels; N T ≥ 1; N C ≥ 1; m ≥ 1; Obtain the jump step for vertical data extraction in the channel group, as well as the height, width, vertical step, and N H row base addresses of the sliding window operator; where N H is equal to the height; Allocate N for each of the said threads H register clusters; wherein each register cluster contains y register groups; each register group is used to store a pixel group composed of N C channel pixels at the same row and column position; y is not less than the width and is equal to the number of column slices participating in the calculation in each channel group; Determine the N H row starting point coordinates of the current scan round according to the row base address and the vertical step size; When executing the current scan round, in each column slice participating in the calculation, starting from each of the row starting coordinates, with the jump step as the span, scan N T row-column positions, and sequentially load the pixel groups at the scanned N T row-column positions into the register groups of the respective threads within the corresponding warp; Calculating the loaded pixel group through each of the threads to obtain an output result within the corresponding local sliding window field of view.
2. The method for implementing depth operation with vertical pixel parallelism according to claim 1, wherein The jump step size is the product of the height and the vertical step size.
3. The method for implementing depth operation with vertical pixel parallelism according to claim 2, wherein The method further includes: Through N H rounds of scans are performed, so that each warp completes the coverage of all sliding window regions of the sliding window operator in the corresponding channel group.
4. The method for implementing depth operation with vertical pixel parallelism according to claim 2, wherein When performing the current scan round, in each column slice participating in the calculation, starting from each of the row starting coordinates, scanning N row-column positions according to the jump step size, and sequentially loading the pixel groups at the scanned N row-column positions into the register groups of each thread in the corresponding warp, including: T N T pixel groups at the row-column positions into the register groups of each thread in the corresponding warp, including: When performing the vertical data scan in the i-th round, in the w-th column slice participating in the calculation within the j-th channel group, the h-th row starting coordinate is used as the starting point of the h-th vertical scan route, and with the jump step as the span, scan N T row-column positions, and successively load the pixel groups at the N T row-column positions into the w-th register group of the h-th register cluster of each thread within the j-th warp; where 1 ≤ i ≤ N H ; 1 ≤ j ≤ m; 1 ≤ w ≤ y; 1 ≤ h ≤ N H .
5. The method for implementing depth operation with vertical pixel parallelism as described in claim 1, characterized in that, The input tensor is composed of one feature data block, or is obtained by splicing multiple feature data blocks in the vertical direction of the original space, or is composed of any sub-data block split from one feature data block.
6. The method for implementing depth operation with vertical pixel parallelism as described in claim 1, wherein The column slices participating in the calculation are the column slices traversed when the sliding window operator slides in the horizontal direction.
7. The method for implementing depth operation with vertical pixel parallelism according to claim 2, characterized in that Determining N row starting point coordinates of the current scanning round according to the row base address and the vertical step size, including: H Determining the row starting coordinate through the following formula: ; Among them, is the starting coordinate of the th row in the th scan; is the base address of the th row; is the vertical step size.
8. The method for implementing depth operation with vertical pixel parallelism according to claim 1, characterized in that, The sliding window operator is a pooling operator; The calculating the loaded pixel group through each of the threads to obtain an output result within the corresponding local sliding window field of view includes: In any one of the warps, each thread performs a channel isolation pooling operation on the loaded pixel group to obtain the pooling results of N C channels within the corresponding local sliding window view.
9. The method for implementing depth operation with vertical pixel parallelism according to claim 1, characterized in that The sliding window operator is a single-channel convolution kernel; The calculating the loaded pixel group through each of the threads to obtain an output result within the corresponding local sliding window field of view includes: Broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any one of the warps, each thread performs a channel-isolated multiply-accumulation operation on the loaded pixel group and the convolution weights to obtain the convolution results of N C channels within the corresponding local sliding window field of view.
10. The method for implementing depth operation with vertical pixel parallelism according to claim 1, characterized in that The sliding window operator is a multi-channel convolution kernel; The calculating the loaded pixel group through each of the threads to obtain an output result within the corresponding local sliding window field of view includes: Broadcasting the convolution weights corresponding to each register group to all threads within the corresponding warp in scalar form; In any one of the warps, performing a multiply-accumulate operation of channel fusion on the loaded pixel group and the convolution weights through each thread to obtain an output feature value within the corresponding local sliding window field of view; wherein, each of the channel groups is used to represent an input group in the grouped convolution operation.
11. The method for implementing depth operation with vertical pixel parallelism according to claim 9 or 10, characterized in that, The method further includes: Before the threads perform the multiply-accumulate operation, determining whether the coordinates of the pixel values in the pixel group are out of bounds, and masking the pixel values corresponding to the out-of-bounds coordinates with a specified value to obtain new pixel values.
12. An apparatus for implementing depth operation with pixel parallelism in the vertical direction, characterized in that, Including: The warp configuration module is used to divide the input tensor into m channel groups and allocate them to m warps for independent processing; where each warp contains N T threads; each channel group contains N C pixels of channels; N T ≥ 1; N C ≥ 1; m ≥ 1; A parameter acquisition module, configured to acquire a jump step for vertical data acquisition in a channel group, as well as the height, width, vertical step, and N of a sliding window operator; where N H is equal to the height; H A thread configuration module for allocating N register clusters to each of the threads respectively H where each register cluster contains y register groups; each register group is used to store a pixel group composed of N channel pixels at the same row and column position C ; y is not less than the width and is equal to the number of column slices participating in the calculation in each channel group A starting point determination module, configured to determine N H row starting point coordinates of the current scanning round according to the row base address and the vertical step size; A pixel loading module, when executing the current scan round, in each column slice participating in the calculation, starting from each of the row starting coordinates, with the jump step as the span, scans N T row-column positions, and sequentially loads the pixel groups at the scanned N T row-column positions into the register groups of the respective threads within the corresponding warp; A thread calculation module, configured to calculate the loaded pixel group through each of the threads to obtain an output result within the corresponding local sliding window field of view.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for implementing vertical-direction pixel parallel depth operation according to any one of claims 1 to 11.
14. A computer program product, characterized in that, Including a computer program, which implements the method for implementing vertical-direction pixel parallel depth operation according to any one of claims 1 to 11 when executed by a processor.
15. An electronic device, characterized in that, Including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the method for implementing vertical-direction pixel parallel depth operation according to any one of claims 1 to 11 when executing the computer program.
Citation Information
Patent Citations
Air-to-air signal processing modularization design method based on GPU acceleration
CN115827215A
Data processing method and device, processor, electronic equipment and storage medium
CN119312003A
Method for accelerating sparse multi-scalar multiplication by using image processor
CN119540025A
Method and device for processing data
US20060248317A1
Multiple register allocation sizes for threads
US20220413916A1
Cited By
Data processing method and device, processor and electronic equipment
CN120632269A
Data processing method and device, processor and electronic equipment
CN120632269B
Operator execution method and device, equipment, storage medium and product
CN121636221A