Method and apparatus for computing pooling operations on a gpu architecture

By allocating computational tasks on the GPU based on batch dimensions and performing row pooling and column pooling operations, the pooling process is optimized, solving the problems of high computational complexity and frequent memory access in pooling operations on the GPU, thus achieving performance improvement and efficiency enhancement.

CN122152492APending Publication Date: 2026-06-05INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTEL CORP
Filing Date
2025-10-31
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies lack efficient pooling operation computation methods on graphics processing unit (GPU) architectures, resulting in high computational complexity, reduced concurrent processes, and frequent memory access, which affects the performance of deep learning models.

Method used

By allocating computation tasks on GPUs based on batch dimensions, performing row pooling and column pooling operations, computation time complexity and memory access are reduced. Intermediate outputs are stored by swapping dimensions to ensure continuous memory access. The pooling process is optimized using custom step size and fill operations.

Benefits of technology

It enables efficient computation through pooling operations on the GPU, improving performance by 13.7%, reducing computation and memory accesses, and improving the operating efficiency of the computer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122152492A_ABST
    Figure CN122152492A_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods and apparatus for computing pooling operations on a GPU architecture. An example apparatus comprises: interface circuitry; machine readable instructions; and at least one processor circuitry programmed by the machine readable instructions to: allocate a computation task among one or more cores based on a batch dimension; perform a row pooling operation based on the batch dimension and a first pooling size to generate a first intermediate data output; perform a column pooling operation based on the batch dimension and a second pooling size to produce a second intermediate data output; and generate a final output based on the first intermediate data output and the second intermediate data output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications This patent claims the benefit of U.S. Provisional Patent Application No. 63 / 727,543, filed December 3, 2024, entitled “Method and Apparatus for Computational Pooling Operations on a GPU Architecture”. The entire disclosure of U.S. Provisional Patent Application No. 63 / 727,543 is incorporated herein by reference. Technical Field

[0002] This disclosure relates to methods and apparatus for computing pooling operations on a graphics processing unit (GPU) architecture. Background Technology

[0003] Pooling, as part of image processing, is applied in deep learning to reduce feature dimensionality and computational complexity. This technique is commonly used in convolutional neural networks (CNNs) to reduce the resolution of feature maps while preserving the features needed for classification-based tasks. Therefore, pooling allows for the preservation of key features while reducing computational complexity and the risk of overfitting. Summary of the Invention

[0004] According to one aspect of this disclosure, an apparatus is provided, comprising: interface circuitry; machine-readable instructions; and at least one processor circuitry of a graphics processing unit, the at least one processor circuitry being programmed via the machine-readable instructions to: distribute computational tasks among one or more cores based on a batch dimension; perform row pooling operations based on the batch dimension and a first pooling size to generate a first intermediate data output; perform column pooling operations based on the batch dimension and a second pooling size to generate a second intermediate data output; and generate a final output based on the first intermediate data output and the second intermediate data output.

[0005] According to another aspect of this disclosure, at least one machine-readable medium is provided, including machine-readable instructions to cause at least one processor circuitry of a graphics processing unit to at least: allocate computational tasks among one or more cores based on a batch dimension; perform a row pooling operation based on the batch dimension and a first pooling size to generate a first intermediate data output; perform a column pooling operation based on the batch dimension and a second pooling size to generate a second intermediate data output; and generate a final output based on the first intermediate data output and the second intermediate data output.

[0006] According to another aspect of this disclosure, an apparatus is provided, comprising: means for distributing computational tasks among one or more cores based on a batch dimension; means for performing a row pooling operation based on the batch dimension and a first pooling size to generate a first intermediate data output; means for performing a column pooling operation based on the batch dimension and a second pooling size to generate a second intermediate data output; and means for generating a final output based on the first intermediate data output and the second intermediate data output.

[0007] According to another aspect of this disclosure, a method is provided, comprising: distributing computational tasks among one or more cores based on a batch dimension; performing a row pooling operation based on the batch dimension and a first pooling size to generate a first intermediate data output; performing a column pooling operation based on the batch dimension and a second pooling size to generate a second intermediate data output; and generating a final output based on the first intermediate data output and the second intermediate data output. Attached Figure Description

[0008] Figure 1 This is a block diagram of an example implementation of a pooling executor circuit for computing pooling operations on a graphics processing unit (GPU) architecture, constructed in accordance with the teachings of this disclosure.

[0009] Figure 2 It is a flowchart representing example machine-readable instructions and / or example operations that can be implemented by example programmable circuitry running, instantiating, and / or executing. Figure 1 Example pooled actuator circuit.

[0010] Figure 3 It is a flowchart representing example machine-readable instructions and / or example operations that can be implemented by example programmable circuitry running, instantiating, and / or executing. Figure 1 An example pooling actuator circuit is provided to perform m-dimensional average pooling.

[0011] Figure 4 The diagram illustrates an example two-dimensional matrix representing an input image of a given size (NxM) and the output of the pooling operation.

[0012] Figure 5A The illustration shows how to use it. Figure 1 The example computation shown is a portion of the pooling operation performed by the pooling actuator circuit, associated with one or more selected data rows (or rows thereof).

[0013] Figure 5B The illustration shows how to use it. Figure 1 The example computation shown is a portion of the pooling operation performed by the pooling actuator circuit, associated with one or more selected data columns.

[0014] Figure 6A The illustration shows how to use it. Figure 1 The pooling executor circuit shown performs a portion of the pooling operation, based on a computation associated with one or more selected data rows (or rows thereof), as an example dimension reordering.

[0015] Figure 6B The diagram illustrates the use of Figure 1 The example pooling operation performed by the pooling actuator circuit shown is an example of a pooling operation.

[0016] Figure 7A The illustration shows an example memory access using a known pooling technique.

[0017] Figure 7B The illustration shows when using Figure 1 The example shown illustrates the reduction in memory accesses when the pooling actuator circuit performs a pooling operation.

[0018] Figure 8 The illustration shows the results of example pooling operations with and without the Single Instruction Multiple Data (SIMD) optimization.

[0019] Figure 9 This is a block diagram of an example processing platform that includes programmable circuitry configured to run, instantiate, and / or execute. Figure 2-3 Example machine-readable instructions and / or execution Figure 2-3 Example operations to implement Figure 1 The pooled actuator circuit.

[0020] Figure 10 yes Figure 9 A block diagram illustrating an example implementation of a programmable circuit.

[0021] Figure 11 yes Figure 9 A block diagram of another example implementation of a programmable circuit.

[0022] Figure 12 This is a block diagram of an example software / firmware / instruction distribution platform (e.g., one or more servers) used to distribute software, instructions, and / or firmware (e.g., with...) Figure 2-3 (Corresponding to example machine-readable instructions) are distributed to client devices associated with: end users and / or consumers (e.g., for licensing, sales and / or use), retailers (e.g., for sales, resale, licensing and / or sublicensing), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products to be distributed to, for example, retailers and / or other end users such as direct purchase customers).

[0023] Generally, the same reference numerals will be used throughout the drawings and accompanying written descriptions to refer to the same or similar parts. The drawings are not necessarily to scale. Detailed Implementation

[0024] Pooling is a common operation in deep learning models, representing the second most computationally intensive operation associated with convolutional neural networks (CNNs), consuming over 20% of inference time. However, the application of pooling operations associated with graphics processing units (GPUs) has not provided any fundamental performance improvements. Current GPU-related techniques include a brute-force approach to pooling operations in mainstream known frameworks such as PyTorch, TensorFlow, and the Computational Unified Device Architecture (CUDA). For example, assuming an input size of O( N 2 The kernel size is O( ), K 2 Known pooling techniques define time complexity as O( ). N 2 K 2 While average pooling operations (e.g., AvgPool) compute the average for local patches of a feature map, they use O(...). N 2 While this approach achieves a time complexity of O(n), it is undesirable for GPU architectures as it leads to a reduction in concurrent processes. Therefore, compared to using O(n),... N 2 Compared to the time complexity of brute-force methods, they execute computations much faster on GPUs. Similarly, for pooling operations that compute maximum values ​​for local blocks of a feature map (e.g., MaxPool), there is a lack of computational techniques that can be implemented on GPUs. Therefore, pooling operations associated with GPU-based architectures are limited to brute-force pooling.

[0025] The methods and apparatus disclosed herein address the lack of computation on GPUs based on existing efficient pooling operations by: (1) reducing the time complexity of 2D pooling by a factor of K (kernel size) (meaning a reduction to 1 / K of the previous value in this paper, and similar expressions below have similar meanings); (2) reducing memory accesses in the GPU; and (3) reducing the number of operands and / or single instruction multiple data (SIMD) instructions for 2D pooling by a factor of K / 2 (meaning a reduction to 2 / K of the previous value in this paper, and similar expressions below have similar meanings). The examples disclosed herein determine the pooling on one dimension at a time and process the dimensions in a predetermined order (e.g., from right to left). For a specific element in a particular dimension, the subsequent K elements can be used to determine the 1D pooling, thereby updating the input(s) for the subsequent(one or more) dimensions(s). In some examples, intermediate outputs are stored by swapping dimensions to ensure contiguous memory access to the subsequent(one or more) dimensions(s). Therefore, the methods and apparatus disclosed herein introduce an efficient method for computing pooling operations using GPU-based architecture(s). The methods and apparatus disclosed herein deliver up to 13.7% performance improvement for end-to-end inference and training associated with CNN-based models. Furthermore, the examples disclosed herein improve the performance of GPU-based 2D pooling operations by K / 2 times and reduce associated computation and / or memory accesses by K / 2 times. Therefore, the examples disclosed herein improve the operation of computers and / or computing systems.

[0026] Figure 1 This is a block diagram 100 illustrating an example implementation of a pooling executor circuit 105 for performing pooling-based computations on a GPU, constructed in accordance with the teachings of this disclosure. Figure 1 The pooled actuator circuit 105 can be instantiated by a programmable circuit (e.g., creating instances of it, making it exist for any length of time, materializing it, implementing it, etc.). For example, the programmable circuit can be implemented by: a central processing unit (CPU) executing first instructions, a field-programmable gate array (FPGA), a programmable logic device (PLD), a general-purpose array logic (GAL) device, a programmable array logic (PAL) device, a complex programmable logic device (CPLD), a simple programmable logic device (SPLD), a microcontroller (MCU), a programmable system-on-a-chip (PSoC), etc. Additionally or alternatively, Figure 1 The pooled actuator circuit 105 can be instantiated (e.g., its instance is created, it is materialized, implemented, etc.) by (i) an application-specific integrated circuit (ASIC) and / or (ii) a field-programmable gate array (FPGA) (e.g., another form of programmable circuit), constructed and / or configured to perform an operation corresponding to the first instruction in response to the execution of the second instruction. It should be understood that... Figure 1Some or all of the circuits can thus be instantiated at the same or different times. Figure 1 Some or all of the circuitry can be instantiated, for example, in one or more threads, which execute concurrently and / or serially on the hardware. Furthermore, in some examples, Figure 1 Some or all of the circuits can be implemented by microprocessor circuits executing instructions and / or FPGA circuits performing operations to implement one or more virtual machines and / or containers.

[0027] exist Figure 1 In the example, the pooling actuator circuit 105 includes an example input identifier circuit 110, an example row pooler circuit 120, an example column pooler circuit 125, an example output generator circuit 130, and an example data storage device 140. Figure 1 In the example, the input identifier circuit 110, the row pooler circuit 120, the column pooler circuit 125, the output generator circuit 130, and the data storage device 140 communicate with the example bus 145.

[0028] Input identifier circuit 110 receives N-dimensional input, which includes: (1) an NM-dimensional input batch (B) (e.g., {B1, B2, …, B…}). N-M}), including M-dimensional input working data (I) (e.g., { I1, I2, …, I M}); and (2) the M-dimensional input pooling convolution kernel size (K) (e.g., { K1, K2, …, K}); and (2) the input pooling convolution kernel size (K) of dimension M (e.g., { K1, K2, …, K}). M The input data received by the input identifier circuit 110 can be used by the pooling actuator circuit 105 to traverse the received batches (B) of NM dimensions and process the input working data (I) of M dimensions. As described in more detail in the examples disclosed herein, the pooling actuator circuit 105 starts from the rightmost dimension (e.g., I). M Process each dimension of the input working data (I) up to the leftmost dimension (e.g., I1), and process the current dimension (D) for each dimension. i All elements in ) have subsequent elements of input pooling size (e.g., K). i Summing (the sum of elements). For example, for d1, d2, …, d… M All possible values ​​(where d) j The value corresponding to the j-th dimension of the working data, for example, satisfying 0 ≤ d. j < I i The summation can be expressed according to Equation 1: Once the summation of subsequent elements is complete, the pooling executor circuit 105 performs average pooling according to Equation 2 by dividing each element by the convolution kernel size (K), where the modified input represents the final output of the pooling operation: For example, in the case of two-dimensional pooling, the input identifier circuit 110 receives input in NCHW data format, where N and C are batch dimensions (e.g., N = number of images, C = number of channels), and H = height, while W = width. In the example disclosed herein, the input identifier circuit 110 identifies the size of the pooling convolution kernel as K1 x K2, and the output size is represented by the batch dimensions (N, C) and the following equation result: H - K1 + 1 and W - K2 + 1 (e.g., as combined with...). Figure 4 As shown below, the pooling executor circuit 105 further performs the following steps using the received input data: (1) distributing workload(s) (e.g., computational tasks) among different cores based on the batch dimension; (2) processing the first dimension (e.g., performing row-based processing using the row pooling circuit 120); and (3) processing the second dimension (e.g., performing column-based processing using the column pooling circuit 125). For example, the input identifier circuit 110 (e.g., using an Intel® Xe GPU architecture) distributes workloads and / or computational tasks among different cores based on the batch dimension (NC). In the examples disclosed herein, the pooling dimension is represented by HW, and each core receives a set of data units (HxW) for processing (e.g., represented as mHW, where m is a component from the batch dimension NC).

[0029] In some examples, the apparatus includes means for allocating computational tasks. For example, this means for allocating computational tasks may be implemented by input identifier circuitry 110. In some examples, input identifier circuitry 110 may be implemented by programmable circuitry (e.g., Figure 9 Example programmable circuit 912) can be instantiated. For example, input identifier circuit 110 can be instantiated using... Figure 10 The example microprocessor 1000 executes machine-executable instructions (e.g., at least...). Figure 2 The input identifier circuit 110 is instantiated by the machine-executable instructions implemented by block 210. In some examples, the input identifier circuit 110 can be instantiated by hardware logic circuitry, which can be an ASIC, XPU, or other hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 11The FPGA circuit 1100 is used for implementation. Alternatively or additionally, the input identifier circuit 110 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the input identifier circuit 110 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGA, ASIC, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0030] Row pooling circuit 120 performs processing on the first dimension (e.g., row-based pooling operation) based on input received from input identifier circuit 110. In the example disclosed herein, row pooling circuit 120 receives input represented as... input[H] [W] Input, create intermediate data input_v2[H][W] For example, considering a single data unit (HxW) and a single core, the methods and apparatus disclosed herein accommodate as many rows (e.g., of size W) as possible in shared local memory (SLM). Furthermore, the row pooling circuit 120 can perform a set of memory transfers to the L1 cache and / or SLM, such that a thread (e.g., a virtual sequence of instructions executed by a core) can be used to compute elements of one or more intermediate outputs. For example, if the thread is targeting... Index [i][j] The output at that point, then the row pooling circuit 120 uses this thread to use SLM from input[i][j] Obtain the next K elements and calculate their sum as an intermediate output, such as combining... Figure 5A Detailed description follows. Subsequently, the row pooling circuit 120 shifts the intermediate output to the SLM. For example, the row pooling circuit 120 generates the following intermediate output (e.g., Equation 3-5) such that when there are fewer than K remaining elements, only the remaining elements in that row are considered, as described below. Figure 5A Shown: input_v2 [ i ][ j ] = Equation 3 input_v2 [ i ][ j ] = ( input [ i ][ j ], input [ i ][ j +1], input [ i ][ j+2], Equation 4 …, input [ i ][ j + K 1-1]) and input_v2 [ i ][ j =From the same row input [ i ][ j The next equation after [the beginning] is 5. The sum of K elements In some examples, the apparatus includes means for performing row pooling operations. For example, the means for performing row pooling operations can be implemented by row pooler circuitry 120. In some examples, row pooler circuitry 120 can be implemented by programmable circuitry (e.g., Figure 9 Example programmable circuit 912) can be instantiated. For example, row pooling circuit 120 can be instantiated using... Figure 10 The example microprocessor 1000 executes machine-executable instructions (e.g., at least...). Figure 3 The line pooler circuit 120 can be instantiated by the machine-executable instructions implemented in block 315. In some examples, the line pooler circuit 120 can be instantiated by hardware logic circuitry, which can be an ASIC, XPU, or other hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 11 The FPGA circuit 1100 is used for implementation. Alternatively or additionally, the row pooler circuit 120 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the row pooler circuit 120 can be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGA, ASIC, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0031] The column pooling circuit 125 performs processing on the second dimension (e.g., column-based pooling operation). For example, the column pooling circuit 125 acquires one or more inputs. input_v2[H][W] One or more outputs are represented as input_ v3[H-K1+1][W-K2+1] In the examples disclosed herein, one or more threads associated with one or more cores can be used to determine elements of subsequent intermediate output. For example, if the thread is targeting... Index [i][j] The output at the point, the column pooling circuit 125 obtains the output from the SLM and input[i][j]The associated K subsequent elements (e.g., using column-major order) are summed to determine the corresponding output (e.g., equation 6-8), and the sum is divided by a constant value (K1 x K2), as described below. Figure 5B Shown: input_v3 [ i ][ j ] = ) / ( K 1 x K 2) Equation 6 input_v3 [ i ][ j ] = ( input_v2 [ i ][ j ], input_v2 [ i +1][ j Equation 7 input_v2 [ i +2][ j ], …, input_v2 [ i + K 2-1][ j ]) sum / ( K 1x K 2) input_v3 [ i ][ j = From the same column input [ i ][ j The next equation begins with 8. The sum of K elements As described in conjunction with row pooling circuit 120, if fewer than K elements remain, column pooling circuit 125 considers only the remaining elements in the column. In the example disclosed herein, intermediate outputs are stored by swapping dimensions to maintain memory access continuity. Therefore, intermediate outputs can be stored as... input_v2[W][H] Instead input_v2[H][W] Storing in this format introduces no overhead to the GPU while ensuring contiguous memory access for one or more subsequent dimensions. For two-dimensional (2D) matrices, this process is similar to storing in column-major order rather than row-major order, as combined with... Figure 6AAs shown in detail. Similarly, the processing order for all dimensions is changed (except for the last dimension). In the example disclosed in this paper, row-wise pooling is performed on the last dimension because the last dimension from the left is the original last dimension from the right, and the final output is stored in the correct order. Storing the output in the correct order does not introduce any overhead in the GPU. In the column-based pooling example, the current dimension is specified as the last dimension, as combined with... Figure 6B shown.

[0032] Furthermore, in the examples disclosed herein, custom stride and padding (e.g., expansion operations) can be used to determine how the convolution operation is applied to the input, which will affect the output size and / or feature extraction. Padding refers to the process of adding extra pixels to the boundaries of the input image, while stride refers to the number of pixels that a filter (e.g., a convolution kernel) moves or slides across the input image. In the examples disclosed herein, when one or more expansion operations have set values ​​(e.g., stride = 1 and padding = 0),... input_v3 The row pooling output is represented, and the row pooler circuit 120 moves the final output from the SLM to the L2 cache. For example, based on a custom step size, the row pooler circuit 120 includes elements that can be skipped or ignored during the row-based pooling operation. In addition to step size and padding, the methods and apparatus disclosed herein also cover other variations of pooling (e.g., average pooling (AvgPool), max pooling (MaxPool) for computing the maximum value of a local block of a feature map, min pooling (MinPool), adaptive feature pooling, etc.).

[0033] The methods and apparatus disclosed herein can be used for any type(s) of pooling variants that comply with the following: F (array[0:N]) = F (F(array[0:i]), F (array[i+1,N])), where F () represents the pooling computation function, where the sum is identified for AvgPool operations, the maximum value for MaxPool operations, and so on. For example, given Sum([1,5, 9, 2, 4]), pooling-based summation can be performed as Sum(Sum[1, 5], Sum[9, 2, 4]) = Sum(6,15) = 21, thus requiring only adjustments to the core operation that computes each element in the i-th dimension. Compared to operations performed using AvgPool, operations performed using MaxPool focus on identifying the maximum value of all elements that need to be computed to avoid dividing by the pooling size at the end of the operation (as described regarding AvgPool operations).

[0034] While the examples disclosed herein focus on the use of two-dimensional (2D) pooling, the methods and apparatus disclosed herein can also be applied to extensions of M-dimensional pooling. For example, 2D pooling involves processing (one or more) rows and then processing (one or more) columns, with the output of the preceding row processing used as the input to the column processing. Similarly, for M-dimensional pooling, processing can be performed sequentially on each dimension such that the input to the i-th dimension processing represents the output of the (i-1)-th dimension processing.

[0035] Considering the time and / or space complexity associated with computational operations, known methods for computing 2D pooling operations on GPUs (e.g., for deep learning frameworks) result in O(N²K²) time complexity and O((N-K+1)²) space complexity. In contrast, the examples disclosed in this paper achieve O(N²K) time complexity (e.g., a K-fold improvement) and O(N²) space complexity. While methods exist for computing AvgPool and MaxPool operations in O(N²) time, these are not optimal for GPUs (e.g., unlike CPUs associated with specific dimensions). Furthermore, there is currently a lack of a general method that can be used to perform all types of pooling variants while achieving O(N²) time complexity.

[0036] As previously stated, the methods and apparatus disclosed herein allow computation of any type of pooling variant on a GPU, reducing the number of computational steps(s) (e.g., by computing the common component only once) and maintaining a high number of parallel threads (e.g., by keeping threads independent by splitting(s) pooling operations(s) into two independent steps). For three-dimensional (3D) pooling (as opposed to the 2D pooling described above), known methods have a time complexity of O(N³K³), while the methods and apparatus disclosed herein aim to achieve a time complexity of O(N³K). Furthermore, for all D-dimensional pooling (e.g., denoted as O(K...),... D The method and apparatus disclosed in this paper improve the time complexity by O(K). D-1 / D) times, this improvement takes into account the need for pooling variants (such as AvgPool, MaxPool, etc.) to maintain a constant number of dimensions.

[0037] Furthermore, the number of SIMD instructions associated with 2D pooling and / or 3D pooling can be defined using the methods and apparatus disclosed herein. For 2D pooling-related operations using known methods, pooling is computed for each output element on a submatrix of size K×K in the input data. To improve the SIMD-based instructions associated with 2D pooling, K SIMD instructions are obtained via row pooler circuit 120 and / or column pooler circuit 125. In the examples disclosed herein, K consecutive elements associated with rows are concatenated to a single SIMD instruction (e.g., assuming K is less than or equal to the SIMD width). Similarly, the methods and apparatus disclosed herein introduce the use of two steps, such that a total of two SIMD instructions are executed for each element. For example, the number of SIMD instructions used by known methods can be expressed as N² × K, while the number of SIMD instructions in the examples disclosed herein can be expressed as N² × 2. In the examples disclosed herein, the number of SIMD instructions is reduced by a factor of K / 2 (e.g., defining the expected performance improvement). For D-dimensional pooling, the number of SIMD instructions is reduced by K. (D-1) / D times. Therefore, for 3D pooling, the improvement is determined to be K. 2 / 3 times. For example, when the kernel size is 11, the method and apparatus disclosed herein can achieve an improvement of 40.33 times.

[0038] Besides reducing SIMD-based instructions, memory accesses can also be reduced by a factor of K. For example, since pooling is a memory-constrained operation, reducing the number of memory accesses per element can significantly improve pooling-related performance. Using known methods, 2D pooling requires K×K input elements for each output element. Since this memory is only contiguous in the row-based dimension, K memory accesses are used for K contiguous memory segments, such as... Figure 7A As shown in detail. In the examples disclosed herein, the number of memory accesses is constant (e.g., representing the number of pooling dimensions). For 2D pooling, the examples disclosed herein rely on only two memory accesses per output element, such as... Figure 7B As shown in detail.

[0039] As previously described regarding row-based pooling operations, row pooler circuit 120 stores transposed data, converting columns into rows to enable contiguous memory accesses. In the example disclosed herein, reduced memory accesses contribute to improved execution of SIMD16 instructions. For M-dimensional pooling, the number of memory accesses associated with known methods can be defined as K. (D-1) The number of memory accesses associated with the methods and apparatus disclosed herein can be defined as D, where D represents the pooling dimension and K represents the pooling kernel size. Therefore, the methods and apparatus disclosed herein reduce the number of memory accesses by K. (D-1) / D times. In the examples published in this paper, the time complexity of the pooling method is improved by K times for GPU-based performance. (D-1) The number of memory accesses in the GPU is reduced by several times, and the number of computations and / or SIMD instructions is also reduced by up to 1000 times. (D-1) / D times. Accordingly, the methods and apparatus disclosed in this paper bring fundamental improvements to known pooling methods supported by GPU-based deep learning frameworks and / or libraries.

[0040] In some examples, the apparatus includes means for performing column pooling operations. For example, the means for performing column pooling operations can be implemented by row pooling circuitry 125. In some examples, column pooling circuitry 125 can be implemented by programmable circuitry (e.g., Figure 9 Example programmable circuit 912) can be instantiated. For example, column pooling circuit 125 can be instantiated by... Figure 10 The example microprocessor 1000 executes machine-executable instructions (e.g., at least...). Figure 3 The block 315 is instantiated using the machine-executable instructions implemented therein. In some examples, the column pooling circuit 125 can be instantiated by hardware logic circuitry, which may be an ASIC, XPU, or other hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 11 The FPGA circuit 1100 is used for implementation. Alternatively or additionally, the column pooling circuit 125 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the column pooling circuit 125 can be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGA, ASIC, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0041] Output generator circuit 130 generates output data, which has: (1) NM-dimensional output batch (B out )(For example{ Bout_1 B out_2 , …, B out_N-M}) and (2) M-dimensional output working data (O) {O1, O2, …, O M}, thus making O i = I i – K i+ 1 (e.g., where I corresponds to the input working data and K corresponds to the input pooling size). In the examples disclosed herein, the output generator circuit 130 provides one or more outputs related to row-based pooling operations (e.g., performed using row pooler circuit 120) and / or column-based pooling operations (e.g., performed using column pooler circuit 125). In some examples, the output generator circuit 130 outputs the number of memory accesses and / or the number of SIMD instructions related to 2D and / or 3D pooling operations. In some examples, the output generator circuit 130 provides one or more outputs related to an improvement factor (also known as an improvement multiple), such as... Figure 8 As shown. In some examples, the output generator circuit 130 applies an expansion operation (e.g., using a step or padding) as part of the pooling operation. In some examples, the output generator circuit 130 transfers the first intermediate data output and / or the second intermediate data output from the SLM to a buffer.

[0042] In some examples, the apparatus includes means for generating a final result. For example, the means for generating the final result may be implemented by an output generator circuit 130. In some examples, the output generator circuit 130 may be a programmable circuit (e.g., Figure 9 Example programmable circuit 912) can be instantiated. For example, output generator circuit 130 can be instantiated by... Figure 10 The example microprocessor 1000 executes machine-executable instructions (e.g., at least...). Figure 2 The output generator circuit 130 can be instantiated by the machine-executable instructions implemented in block 225. In some examples, the output generator circuit 130 can be instantiated by hardware logic circuitry, which can be an ASIC, XPU, or other hardware logic circuitry configured and / or constructed to perform operations corresponding to machine-readable instructions. Figure 11 The FPGA circuit 1100 is used for implementation. Alternatively or additionally, the output generator circuit 130 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the output generator circuit 130 can be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGA, ASIC, XPU, comparators, operational amplifiers, logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0043] Data storage device 140 can be used to store any information related to input identifier circuit 110, row pooler circuit 120, column pooler circuit 125 and / or output generator circuit 130. Figure 1The data storage device 140 shown in the example can be implemented using any memory, storage device, and / or storage disk used for storing data, such as flash memory, magnetic media, optical media, etc. Furthermore, the data stored in the data storage device 140 can be in any data format, such as binary data, comma-separated data, tab-separated data, Structured Query Language (SQL) structures, image data, etc.

[0044] Although Figure 1 The diagram illustrates the implementation. Figure 1 An example of the pooled actuator circuit 105, but Figure 1 One or more of the elements, processes, and / or devices shown may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other way. Additionally, the example input identifier circuit 110, the example row pooler circuit 120, the example column pooler circuit 125, the example output generator circuit 130, and / or more generally, Figure 1 The example pooling actuator circuit 105 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, any of the example input identifier circuit 110, example row pooler circuit 120, example column pooler circuit 125, example output generator circuit 130, and / or more generally, example pooling actuator circuit 105, can be implemented by combining machine-readable instructions (e.g., firmware or software) with programmable circuitry, processor circuitry, one or more analog circuits, one or more digital circuits, one or more logic circuits, one or more programmable processors, one or more programmable microcontrollers, one or more graphics processing units (GPUs), one or more digital signal processors (DSPs), one or more ASICs, one or more programmable logic devices (PLDs), vision processing units (VPUs), and / or one or more field-programmable logic devices (FPLDs) (e.g., FPGAs). Furthermore, Figure 1 The pooling actuator circuit 105 may include, in addition to Figure 1 Those other than or replacing those shown Figure 1 One or more of the elements, processes and / or devices shown, and / or may include any or all of more than one of the elements, processes and devices shown.

[0045] exist Figure 2-3The flowchart shown represents example machine-readable instructions that can be executed by programmable circuitry to implement and / or instantiate. Figure 1 The pooling actuator circuit 105, and / or the flowchart representing example operations, can be executed by programmable circuitry to implement and / or instantiate pooling actuator circuitry 105. Machine-readable instructions can be one or more executable programs or portions of one or more executable programs executable by programmable circuitry, such as those described below. Figure 9 The programmable circuitry 912 and / or machine-readable instructions shown in the example processor platform 1100 discussed below may be related to the following. Figure 10 and / or Figure 11 The examples discussed are programmable circuits (e.g., FPGAs) that perform one or more functions or portions of functions. In some examples, machine-readable instructions cause operations, tasks, etc., to be performed and / or executed in a real-world manner. As used herein, “automation” means without human intervention.

[0046] The program may be embodied as instructions (e.g., software and / or firmware) stored on one or more non-transitory computer-readable and / or machine-readable storage media, such as cache memory, magnetic storage devices or disks (e.g., floppy disks, hard disk drives (HDDs), etc.), optical storage devices or disks (e.g., Blu-ray discs, compact disks (CDs), digital versatile disks (DVDs), etc.), redundant arrays of independent disks (RAID), registers, ROM, solid-state drives (SSDs), SSD memory, non-volatile memory (e.g., electrically erasable programmable read-only memory (EEPROM), flash memory, etc.), volatile memory (e.g., any type of random access memory (RAM), etc.), and / or any other storage device or disk. Instructions on a non-transitory computer-readable and / or machine-readable medium may be programmed and / or executed by programmable circuitry located in one or more hardware devices, but the entire program and / or portions thereof may also be executed and / or instantiated by one or more hardware devices other than programmable circuitry, and / or embodied in dedicated hardware. Machine-readable instructions may be distributed across multiple hardware devices and / or executed by two or more hardware devices (e.g., server and client hardware devices). For example, client hardware devices may be implemented by endpoint client hardware devices (e.g., hardware devices associated with human and / or machine users) or by an intermediate client hardware device gateway (e.g., a radio access network (RAN)) that facilitates communication between a server and endpoint client hardware devices. Similarly, a non-transitory computer-readable storage medium may include one or more media. Additionally, although referenced... Figure 2-3 The flowchart shown is used to describe the example program, but the implementation can be used instead. Figure 1The example pooled actuator circuit 105 can be implemented in many other ways. For example, the execution order of the blocks in the flowchart can be changed, and / or some of the described blocks can be altered, eliminated, or combined. Additionally or alternatively, any or all blocks of the flowchart can be implemented by one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, comparator, operational amplifier, logic circuitry, etc.) configured to perform the corresponding operations without executing software or firmware. Programmable circuitry can be distributed across different network locations and / or local to one or more hardware devices (e.g., a single-core processor (e.g., a single-core CPU), a multi-core processor (e.g., a multi-core CPU, XPU, etc.)). As used herein, “programmable circuitry” includes any (one or more) types of circuitry that can be programmed to perform desired functions, such as CPUs, GPUs, VPUs, and / or FPGAs. Programmable circuitry may include one or more CPUs, one or more GPUs, one or more VPUs, and / or one or more FPGAs located in the same package (e.g., the same integrated circuit (IC) package or two or more separate housings); one or more CPUs, GPUs, VPUs, and / or one or more FPGAs within a single machine; multiple CPUs, GPUs, VPUs, and / or FPGAs distributed across multiple servers in a server rack; and / or multiple CPUs, GPUs, VPUs, and / or FPGAs distributed across one or more server racks. Additionally or alternatively, the programmable circuitry may include programmable logic devices (PLDs), general-purpose array logic (GAL) devices, programmable array logic (PAL) devices, complex programmable logic devices (CPLDs), simple programmable logic devices (SPLDs), microcontrollers (MCUs), programmable system-on-a-chip (PSoCs), and / or any combination of these in any of the contexts described above.

[0047] Machine-readable instructions described herein may be stored in one or more formats, including compressed formats, encrypted formats, segmented formats, compiled formats, executable formats, packaged formats, etc. Machine-readable instructions as described herein may be stored as data (e.g., computer-readable data, machine-readable data, one or more bits (e.g., one or more computer-readable bits, one or more machine-readable bits, etc.), bit streams (e.g., computer-readable bit streams, machine-readable bit streams, etc.), or data structures (e.g., as part of instructions, code, code representations, etc.). For example, machine-readable instructions may be segmented and stored on one or more storage devices, disks, and / or computing devices (e.g., servers) located in the same or different locations within a network or set of networks (e.g., in the cloud, in edge devices, etc.). Machine-readable instructions may require installation, modification, adaptation, updating, combination, supplementation, configuration, decryption, decompression, unpacking, distribution, reassignment, compilation, etc., to make them directly readable, interpretable, and / or executable by computing devices and / or other machines. For example, machine-readable instructions may be stored in multiple parts that are individually compressed, encrypted, and / or stored on separate computing devices, wherein when these parts are decrypted, decompressed, and / or combined, they form a set of computer-executable and / or machine-executable instructions that implement one or more functions and / or operations (which together may form a program such as that described herein).

[0048] In another example, machine-readable instructions may be stored in a state in which they can be read by programmable circuitry, but require the addition of libraries (e.g., dynamic link libraries (DLLs)), software development kits (SDKs), application programming interfaces (APIs), etc., to execute these machine-readable instructions on a specific computing device or other device. In another example, machine-readable instructions may need to be configured (e.g., storage settings, input data, recording network addresses, etc.) before they can be executed in whole or in part. Therefore, machine-readable, computer-readable, and / or machine-readable media as used herein may include instructions and / or (one or more) programs, regardless of the specific format or state of such machine-readable instructions and / or (one or more) programs.

[0049] The machine-readable instructions described in this article can be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, machine-readable instructions can be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0050] As mentioned above, Figure 2-3 Example operations can be implemented using executable instructions (e.g., computer-readable and / or machine-readable instructions) stored on one or more non-transitory computer-readable and / or machine-readable media. As used herein, the terms non-transitory computer-readable medium, non-transitory computer-readable storage medium, non-transitory machine-readable medium, and / or non-transitory machine-readable storage medium are explicitly defined to include any type of computer-readable storage device and / or storage disk, excluding propagation signals and transmission media. Examples of such non-transitory computer-readable medium, non-transitory computer-readable storage medium, non-transitory machine-readable medium, and / or non-transitory machine-readable storage medium include optical storage devices, magnetic storage devices, HDDs, flash memory, read-only memory (ROM), CDs, DVDs, caches, any type of RAM, registers, and / or any other storage device or storage disk in which information can be stored for any duration (e.g., long-term storage, permanent storage, short-term storage, temporary buffering, and / or caching of information). As used herein, the terms "non-transitory computer-readable storage device" and "non-transitory machine-readable storage device" are defined as any physical (mechanical, magnetic, and / or electrical) hardware that retains information for a period of time, excluding propagated signals and transmission media. Examples of non-transitory computer-readable storage devices and / or non-transitory machine-readable storage devices include any type of random access memory, any type of read-only memory, solid-state memory, flash memory, optical disc, hard disk, disk drive, and / or redundant array of independent disk (RAID) systems. As used herein, the term "device" refers to physical structures, such as mechanical and / or electrical equipment, hardware, and / or circuitry, which may or may not be configured to execute computer-readable instructions, machine-readable instructions, etc., and / or may or may not be manufactured to execute computer-readable instructions, machine-readable instructions, etc.

[0051] Figure 2This is a flowchart representing example machine-readable instructions and / or example operations 200, which can be implemented by programmable circuitry, instantiated, and / or executed. Figure 1 Example pooled actuator circuit 105. Figure 2 The machine-readable instructions and / or operations 200 begin at box 205, where the input identifier circuitry 110 determines whether to perform the pooling operation using a GPU-based architecture. In some examples, the input identifier circuitry 110 determines whether to continue with the pooling operation disclosed herein based on whether the pooling operation will be performed using the CPU and / or GPU. For pooling operations performed on a GPU, the input identifier circuitry 110 receives an n-dimensional input at box 210, which has: (1) an NM batch dimension, including an M working data dimension; and (2) a pooling size of M dimensions. In the examples disclosed herein, the input identifier circuitry 110 distributes the pooling operation workload (e.g., computational task) across one or more cores based on the input batch dimension at box 215. For example, the input identifier circuitry 110 allocates a set of data units (HxW) to each core for processing, such as in combination with Figure 1 As described. The pooling actuator circuit 105 then performs m-dimensional average pooling at block 220. For example, as in combination Figure 3 As described in detail, row pooler circuit 120 generates a first intermediate output by performing row-based pooling operations, and column pooler circuit 125 generates a second intermediate output by performing column-based pooling operations. Once all pooling operations are complete, output generator circuit 130 outputs data with updated NM batch dimensions and M working data dimensions at block 225.

[0052] Figure 3 This is a flowchart representing example machine-readable instructions and / or example operations 220 that can be implemented by example programmable circuitry running, instantiating, and / or executing. Figure 1 Example pooling actuator circuit 105 for performing m-dimensional average pooling. Figure 3 Machine-readable instructions and / or operations 220 begin at box 305, where row pooling circuitry 120 identifies the number of rows and / or columns that can be accommodated in shared local memory (SLM). Figure 3 In the example, row pooling circuit 120 identifies the input of the next K elements associated with a given dimension at box 310. For example, row pooling circuit 120 uses threads (e.g., instructions executed by the core) from... input[i][j]The next K elements are retrieved. Then, the row pooling circuit 120 generates a first intermediate output at box 315 by performing a summation on the next K elements. For example, the row pooling circuit 120 calculates the sum of elements as the resulting intermediate output (e.g., when the thread is targeting...). Index [i][j] (when outputting at the location), such as regarding Figure 1 and Figure 5A Detailed description. Next, the row pooling circuit 120 will output the first intermediate output (e.g., at block 320) at box 320. Figure 1 The one or more shown input_v2[i][j] The output is moved to the SLM. In some examples, the row pooling circuit 120 stores the first intermediate output at block 325 by swapping dimensions to allow for sequential memory access for subsequent dimensions. Accordingly, the example intermediate output... input_v2[H][W] Can be stored as input_v2 [W][H] .

[0053] Once the row-based pooling operation is complete, the column pooling circuit 125 marks the input of the next K elements at box 330. After retrieving the next K elements, the column pooling circuit 125 generates a second intermediate output at box 335 based on the summed elements divided by a constant value (e.g., the kernel size). For example, the column pooling circuit 125 obtains the input from the SLM... input[i][j] The associated K subsequent elements (e.g., based on column-major order) are summed to determine the corresponding output, and the sum is divided by a constant value (K1xK2), as shown in the following example. Figure 1 and Figure 5B As shown and described. Once the column-based pooling operation is completed at block 340, the column pooling circuit 125 moves one or more final outputs from the SLM to the cache at block 345.

[0054] Figure 4 The illustration shows a two-dimensional (2D) matrix (I) indicating an example input image 405 of a given size (e.g., N x M) and an example output image 415 of a pooling operation (e.g., (N - K1 + 1) x (M - K2 + 1)), with a convolution kernel size of K1 x K2. For all submatrices S of size K1 x K2 in the input I (e.g., the first submatrix 410), F(E) can be determined based on the mathematical operation (F) and the set of all elements (E) in the submatrix S. The resulting output F(E) represents... Index [i][j] The output at that location, where [i][j]It is the index of the top-left element of the submatrix S. Several variations can be represented based on the output F(E). For example, for the average pooling operation (AvgPool), F() represents the average of all elements; while for the maximum pooling operation (MaxPool), F() represents the maximum of all elements. In some examples, known pooling techniques involve iterating through all submatrices S in the input (e.g., of size K1 x K2). For a starting index of ( i, j The pooling is determined using the average of the AvgPool operation and the maximum of the MaxPool operation on the current submatrix. In some examples, the output elements may be set at indices of the output matrix (e.g., output image 415). i, j At position ), where the second submatrix 420 indicates the output of (one or more) pooling operations.

[0055] Figure 5A The illustration shows how to use it. Figure 1 The pooling executor circuit 105 performs a portion of the pooling operation 500, an example computation associated with one or more selected data rows (or rows thereof). Figure 5A In the example, the row pooling circuit 120 receives input 505 (e.g., input[H][W] And generate the first intermediate data 520 (e.g., input_v2[H][W] In the examples disclosed herein, row pooling circuit 120 (e.g., from...) input[i][j] The row pooling circuit 120 retrieves the first K elements 510 associated with the first row of input 505 and calculates the first sum of these K elements as an output (e.g., first output 525). Subsequently, the row pooling circuit 120 retrieves the second K elements 515 associated with the second row of input 505 and calculates the second sum of these K elements as an output (e.g., second output 530). Once the first intermediate output is ready, the row pooling circuit 120 moves the first intermediate output to shared local memory (SLM). For example, as per [reference to...] Figure 1 As described, the row pooling circuit 120 can... input_v2[i][j] Identified as ( input[i][j], input[i][j+1], input[i][j+2], …, input[i] [j+K1-1] ) and.

[0056] Figure 5B The illustration shows how to use it. Figure 1 The pooling executor circuit 105 performs a portion of the pooling operation 550, an example computation related to one or more selected data columns. Figure 5B In the example, the input received by the column pooling circuit 125 can be represented as the first intermediate output (e.g. input_v2[H][W] ), to generate a second intermediate output (such as input_v3[H-K2] [W-K1]For example, given an input of 555, the column pooling circuit 125 retrieves the input from the SLM and... input[i][j] The K related elements (e.g., K2 elements, 560) are used (e.g., column-major order), and the sum is determined as the first output (e.g., the sum of K2 elements in output 565, 570). The sum is then divided by a constant value (e.g., K1 x K2) to obtain the second output (e.g., the average value, 575). (See also: ...) Figure 1 An example of such an average calculation, as described, can be represented as: input_v3[i][j] = ( input_v2[i][j], input_ v2[i+1][j], input_v2[i+2][j], …, input_v2[i+K2-1][j] ) and / ( K1xK2 In some examples, the column pooling circuit 125 removes additional elements during computation, as shown in combination with output 580 (e.g., the modified output 585 is represented as...). input_v3[H-K2][W-K1] ).

[0057] Figure 6A The illustration shows how to use it. Figure 1 The pooling executor circuit 105 shown performs a pooling operation as part of an example dimension reordering 600 based on calculations associated with one or more selected data rows(s). This is done in conjunction with memory optimization and / or maintaining contiguous memory access. Figure 5A As part of the row pooling operation, row pooler circuit 120 stores one or more intermediate outputs by swapping dimensions. Figure 6A In the example, row pooling circuit 120 acquires... Figure 5A The initial intermediate outputs shown (e.g., based on the first output 525 and the second output 530, when the dimensions are not reordered) are such that input_v2[H][W] After the row pooling circuit 120 reorders the dimensions, the intermediate output (e.g., output 605) can be represented as follows: input_v2[W][H] ,like Figure 6A The output of dimension reordering using one or more examples is shown in 610 and 615.

[0058] Figure 6B The diagram illustrates the use of Figure 1 Example pooling operation 650 is performed by pooling executor circuit 105. During row-based and column-based pooling operations, the processing order of dimensions other than the last dimension processed is changed. The last dimension involves row-wise pooling because the last dimension from the left-hand side of the input represents the original last dimension from the right-hand side of the input, thus ensuring that the final output is stored in the correct order, which does not introduce any overhead in GPU-based architectures. Figure 6B In the example, the current dimension refers to the last dimension. For example, input 655 input_v2[W][H](With reordered dimensions) It includes a first group of K2 elements 660 and a second group of K2 elements 665. The resulting output 670 includes the sum of K2 elements represented by (one or more) reordered dimension outputs 675 and 680.

[0059] Figure 7A The illustration shows an example memory access 700 using a known pooling technique. As mentioned earlier, pooling is a memory-constrained operation, therefore reducing the number of memory accesses per element can significantly improve performance. Figure 7A In the example, matrix 702, with output 704, uses K×K input elements (similar to input element 708 of matrix 706). Specifically, 2D pooling uses K×K input elements for each output element. Figure 7A In the example, since there are K consecutive memory segments, a total of K memory accesses are required when using the known 2D pooling technique (such as memory accesses 712, 714, 716 for matrix 710).

[0060] In comparison, Figure 7B The illustration shows when using Figure 1 An example of reduced memory access 750 when the pooling executor circuit 105 performs pooling operations. Figure 7B In the example, the same output 702 produces a single memory access associated with a row-based pooling operation performed by row pooler circuit 120 (e.g., memory access 754 in matrix 752) and a single memory access associated with a column-based pooling operation performed by column pooler circuit 125 (e.g., memory access 762 in matrix 760). Therefore, the number of memory accesses is constant (e.g., representing the number of pooling dimensions), and a total of two memory accesses are used for each output element.

[0061] Figure 8 The example pooling operation results 800 are shown with and without Single Instruction Multiple Data (SIMD) based optimizations. The SIMD-based optimizations are achieved using GPU-based architectures (e.g., Intel with existing IPEX optimized pooling operations). ® Executed by Ponte Vecchio GPU. Figure 8In the example, the measurement results were obtained using an input data size of 128×128×32×32 (using NCHW format, where N represents the number of images, C represents the number of channels, H represents the height, and W represents the width). The pooling operation result 800 includes one or more example convolutional kernel sizes 802, example IPEX-based results 806, results 808 of the disclosed method (including results without and with SIMD optimization), an improvement factor result 810, and a theoretical improvement factor (K / 2) result 812. For examples without SIMD-based optimized implementations (such as using small convolutional kernel sizes no larger than 6), the improvement factor 810 is close to the theoretical improvement factor 812, i.e., K / 2. For larger convolutional kernel sizes, the observed improvement is approximately three to four times, as the generated SIMD instructions are serialized due to variable loop lengths (one or more). For example, the compiler cannot effectively optimize for larger filter sizes because loop iteration counting and alignment are unpredictable.

[0062] By incorporating SIMD optimizations using the methods and apparatus disclosed herein, thereby adjusting the loop structure to fit the SIMD width, a significant performance improvement (approximately K / 2x) consistent with theoretical predictions was observed. Significant performance improvements (e.g., an improvement factor of 810) were observed for all test convolutional kernel sizes. In the examples disclosed herein, the maximum workgroup-size memory was allocated on shared memory with no additional memory overhead. In production models, standard convolutional kernel sizes for pooling were 3×3, 5×5, 7×7, and 11×11, including variations using 2D and 3D pooling. Based on the observed improvements in pooling operations, the expected performance improvements for end-to-end inference using convolutional neural network (CNN) models are as follows: an overall improvement of 11.19% using AlexNet, 9.90% using VGG16, 13.70% using GoogleNet, and 13.48% using Inception-v3. For example, in most CNN models, convolutions consume up to 70% of inference time, followed by pooling operations (averaging up to 20% of inference time). Using the methods and apparatus disclosed herein for pooling operations, end-to-end inference performance can be improved by 9.90% to 13.70%.

[0063] Figure 9 This is a block diagram of an example programmable circuit platform 900, which is configured to perform and / or instantiate... Figure 2-3 Example machine-readable instructions and / or example operations to implement Figure 1 Example pooled actuator circuit 105. Programmable circuit platform 900 can be, for example, a server, personal computer, workstation, self-learning machine (e.g., neural network), mobile device (e.g., cellular phone, smartphone, such as iPad). TMTablet devices, personal digital assistants (PDAs), internet devices, DVD players, CD players, digital video recorders, Blu-ray players, game consoles, personal video recorders, set-top boxes, headphones (e.g., augmented reality (AR) headphones, virtual reality (VR) headphones, etc.) or other wearable devices, or any other type of computing and / or electronic device.

[0064] The illustrated programmable circuit platform 900 includes a programmable circuit 912. The illustrated programmable circuit 912 is hardware. For example, the programmable circuit 912 may be implemented by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers from any desired family or manufacturer. The programmable circuit 912 may be implemented by one or more semiconductor-based (e.g., silicon-based) devices. In this example, the programmable circuit 912 implements an input identifier circuit 110, a row pooler circuit 120, a column pooler circuit 125, and an output generator circuit 130.

[0065] The illustrated programmable circuit 912 includes local memory 913 (e.g., cache, registers, etc.). The illustrated programmable circuit 912 communicates via bus 918 with main memory, which includes volatile memory 914 and non-volatile memory 916. Volatile memory 914 can be implemented using synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of RAM device. Non-volatile memory 916 can be implemented using flash memory and / or any other desired type of memory device. Access to the illustrated main memory 914, 916 is controlled by memory controller 917. In some examples, memory controller 917 can be implemented by one or more integrated circuits, logic circuitry, microcontrollers from any desired family or manufacturer, or any other type of circuitry to manage the flow of data to and from main memory 914, 916.

[0066] The illustrated programmable circuit platform 900 also includes interface circuitry 920. Interface circuitry 920 can be implemented in hardware according to any type of interface standard, such as an Ethernet interface, a universal serial bus (USB) interface, a Bluetooth® interface, a near field communication (NFC) interface, a peripheral component interconnect (PCI) interface, and / or a peripheral component interconnect express (PCIe) interface.

[0067] In the illustrated example, one or more input devices 922 are connected to interface circuitry 920. The input devices 922 allow users (e.g., human users, machine users, etc.) to input data and / or commands into programmable circuitry 912. The input devices 922 may be implemented using, for example, audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, touchpads, trackballs, isopoint devices, and / or voice recognition systems.

[0068] One or more output devices 924 are also connected to the interface circuitry 920 of the illustrated example. The output devices 924 may be implemented, for example, by display devices (e.g., light-emitting diodes (LEDs), organic light-emitting diodes (OLEDs), liquid crystal displays (LCDs), cathode ray tube (CRT) displays, in-place switching (IPS) displays, touchscreens, etc.), haptic output devices, printers, and / or speakers. The interface circuitry 920 of the illustrated example thus typically includes a graphics driver card, a graphics driver chip, and / or graphics processor circuitry, such as a GPU.

[0069] The interface circuit 920 illustrated also includes communication devices, such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces, to facilitate data exchange with external machines (e.g., any kind of computing device) via network 926. Communication can be made via, for example, Ethernet connections, digital subscriber line (DSL) connections, telephone line connections, coaxial cable systems, satellite systems, line-of-sight wireless systems, cellular telephone systems, optical connections, and so on.

[0070] The illustrated programmable circuit platform 900 also includes one or more mass storage devices 928 for storing software and / or data. Examples of such mass storage devices 928 include magnetic storage devices (e.g., floppy disks, drives, HDDs, etc.), optical storage devices (e.g., Blu-ray discs, CDs, DVDs, etc.), RAID systems, and / or solid-state storage disks or devices (e.g., flash memory devices and / or SSDs).

[0071] can be Figure 2-3 The machine-executable instructions implemented by the machine-readable instructions 932 may be stored in a mass storage device 928, a volatile memory 914, a non-volatile memory 916, and / or at least one removable non-transitory computer-readable storage medium such as a CD or DVD.

[0072] Figure 10 yes Figure 9 A block diagram illustrating an example implementation of the programmable circuit 910. In this example, Figure 9 The programmable circuit 910 is implemented by the microprocessor 1000. For example, the microprocessor 1000 may be a general-purpose microprocessor (e.g., a general-purpose microprocessor circuit). The microprocessor 1000 executes... Figure 2-3 The flowchart contains part or all of the machine-readable instructions to effectively translate Figure 1 The circuit is instantiated as a logic circuit to perform operations corresponding to these machine-readable instructions. In some such examples, Figure 1 The circuitry is instantiated by the hardware circuitry of the microprocessor 1000 in conjunction with instructions. For example, the microprocessor 1000 can implement multi-core hardware circuitry (e.g., CPU, DSP, GPU, XPU, etc.). While it may include any number of example cores 1002 (e.g., one core), this example of the microprocessor 1000 is a multi-core semiconductor device including N cores. The cores 1002 of the microprocessor 1000 can operate independently or cooperate to execute machine-readable instructions. For example, machine code corresponding to firmware, embedded software, or software programs can be executed by one of the cores 1002, or by multiple cores of the cores 1002 at the same or different times. In some examples, the machine code corresponding to firmware, embedded software, or software programs is divided into threads and executed in parallel by two or more cores of the cores 1002. The software program may correspond to... Figure 2-3 The flowchart represents part or all of the machine-readable instructions and / or operations.

[0073] Core 1002 can communicate via a first example bus 1004. In some examples, the first bus 1004 can implement a communication bus to enable communication associated with one or more of cores 1002. For example, the first bus 1004 can implement at least one of an Inter-Integrated Circuit (I2C) bus, a Serial Peripheral Interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, the first bus 1004 can implement any other type of computing or electrical bus. Core 1002 can obtain data, instructions, and / or signals from one or more external devices via example interface circuitry 1006. Core 1002 can output data, instructions, and / or signals to one or more external devices via interface circuitry 1006. While the core 1002 of this example includes example local memory 1020 (e.g., a Level 1 (L1) cache that can be partitioned into an L1 data cache and an L1 instruction cache), the microprocessor 1000 also includes example shared memory 1010 (e.g., a Level 2 (L2) cache) that can be shared by the cores for high-speed access to data and / or instructions. Data and / or instructions can be transferred (e.g., shared) by writing to and / or reading from shared memory 1010. The local memory 1020 and shared memory 1010 of each core 1002 may include multi-level cache memory and main memory (e.g., Figure 9 The cache hierarchy (914, 916) is part of the main memory's storage device hierarchy. Typically, higher-level memories in this hierarchy exhibit lower access times and have smaller storage capacities compared to lower-level memories. Variations across the various levels of the cache hierarchy are managed by cache coherency strategies (e.g., reconciliation).

[0074] Each core 1002 may be referred to as a CPU, DSP, GPU, etc., or any other type of hardware circuitry. Each core 1002 includes control unit circuitry 1014, arithmetic and logic (AL) circuitry (sometimes called ALU) 1016, multiple registers 1018, an L1 cache 1020, and a second example bus 1022. Other structures may also be present. For example, each core 1002 may include vector unit circuitry, single instruction multiple data (SIMD) unit circuitry, load / store unit (LSU) circuitry, branch / jump unit circuitry, floating-point unit (FPU) circuitry, etc. Control unit circuitry 1014 includes semiconductor-based circuitry configured to control (e.g., coordinate) the movement of data within the corresponding core 1002. AL circuitry 1016 includes semiconductor-based circuitry configured to perform one or more mathematical and / or logical operations on the data within the corresponding core 1002. Some examples of AL circuitry 1016 perform integer-based operations. In other examples, AL circuit 1016 also performs floating-point operations. In still other examples, AL circuit 1016 may include a first AL circuit that performs integer-based operations and a second AL circuit that performs floating-point operations. In some examples, AL circuit 1016 may be referred to as an Arithmetic Logic Unit (ALU).

[0075] Register 1018 is a semiconductor-based structure used to store data and / or instructions, such as the results of one or more operations performed by the AL circuit 1016 of the corresponding core 1002. For example, register 1018 may include one or more vector registers, one or more SIMD registers, one or more general-purpose registers, one or more flag registers, one or more segment registers, one or more machine-specific registers, one or more instruction pointer registers, one or more control registers, one or more debug registers, one or more memory management registers, one or more machine check registers, and so on. Register 1018 can be as follows: Figure 10 The diagram shows the arrangement as a library group. Alternatively, register 1018 can be organized in any other arrangement, format, or structure, including distribution throughout core 1002 to reduce access time. The second bus 1022 can be implemented by at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.

[0076] Each core 1002 and / or more generally, the microprocessor 1000 may include additional and / or alternative structures as shown and described above. For example, one or more clock circuits, one or more power sources, one or more power gates, one or more cache home agents (CHAs), one or more converged / common mesh stops (CMSs), one or more shifters (e.g., one or more barrel shifters), and / or other circuitry may be present. The microprocessor 1000 is a semiconductor device manufactured to include a plurality of interconnected transistors to implement the above-described structures in one or more integrated circuits (ICs) contained in one or more packages.

[0077] The microprocessor 1000 may include one or more accelerators (e.g., acceleration circuitry, hardware accelerators, etc.) and / or work in conjunction with them. In some examples, accelerators are implemented by logic circuitry to perform certain tasks faster and / or more efficiently than a general-purpose processor. Examples of accelerators include ASICs and FPGAs, such as those discussed herein. GPUs, DSPs, and / or other programmable devices may also be accelerators. Accelerators may be on the microprocessor 1000 board, in the same chip package as the microprocessor 1000, and / or in one or more packages separate from the microprocessor 1000.

[0078] Figure 11 yes Figure 9 A block diagram illustrating another example implementation of the programmable circuit 912. In this example, the programmable circuit 912 is implemented by the FPGA circuit 1100. For example, the FPGA circuit 1100 can be implemented by an FPGA. For example, the FPGA circuit 1100 can be used to perform, for example, actions that would otherwise be possible through... Figure 10 The example microprocessor 1000 executes operations by executing corresponding machine-readable instructions. However, once configured, the FPGA circuitry 1100 instantiates the operations and / or functions corresponding to the machine-readable instructions in hardware, thus often executing operations / functions faster than a general-purpose microprocessor can execute the corresponding software.

[0079] More specifically, as described above Figure 10 The microprocessor 1000 (it is a general-purpose device that can be programmed to execute...) Figure 2-3 The flowchart represents part or all of the machine-readable instructions, but its interconnections and logic circuitry are fixed once manufactured. Figure 11 The example FPGA circuit 1100 includes interconnects and logic circuits that can be configured, constructed, programmed, and / or interconnected in different ways after manufacturing to instantiate, for example, with... Figure 10The flowchart represents part or all of the machine-readable instructions corresponding to the operations / functions. Specifically, FPGA circuit 1100 can be considered as an array of logic gates, interconnects, and switches. Switches can be programmed to change the way logic gates are interconnected, effectively forming one or more dedicated logic circuits (unless and until FPGA circuit 1100 is reprogrammed). The configured logic circuits enable logic gates to cooperate in different ways to perform different operations on data received by the input circuits. These operations can correspond to... Figure 2-3 The flowchart represents part or all of the instructions (e.g., software and / or firmware). Therefore, the FPGA circuit 1100 can be configured and / or constructed to effectively connect with... Figure 2-3 The flowchart instantiates some or all of the machine-readable instructions corresponding to the operations / functions, and performs the operations / functions corresponding to these software instructions in a dedicated manner similar to that of an ASIC. Therefore, the FPGA circuit 1100 performs the operations / functions corresponding to these software instructions in a dedicated manner. Figure 2-3 The speed at which some or all of the machine-readable instructions correspond to the operation / function can be faster than the speed at which a general-purpose microprocessor executes these instructions.

[0080] exist Figure 11 In some examples, FPGA circuit 1100 is configured and / or constructed in response to being programmed (and / or reprogrammed once or multiple times) based on a binary file. In some examples, the binary file can be compiled and / or generated based on instructions in a hardware description language (HDL), such as Lucid, the Very High Speed ​​Integrated Circuit (VHSIC) Hardware Description Language (VHDL), or Verilog. For example, a user (e.g., a human user, a machine user, etc.) can write code or programs corresponding to one or more operations / functions using HDL; the code / program can be translated into a low-level language as needed; and the code / program (e.g., code / program in a low-level language) can be (e.g., by a compiler, software application, etc.) converted into a binary file. In some examples, Figure 11 The FPGA circuit 1100 can access and / or load binary files to enable... Figure 11 The FPGA circuit 1100 is configured and / or constructed to perform one or more operations / functions. For example, the binary file may be transmitted via bit streams (e.g., one or more computer-readable bits, one or more machine-readable bits, etc.), data (e.g., computer-readable data, machine-readable data, etc.) and / or Figure 11The FPGA circuit 1100 is implemented using machine-readable instructions accessible to cause [access to / from] [the following]. Figure 11 The configuration and / or construction of the FPGA circuit 1100 or one or more of its components.

[0081] In some examples, the binary file is compiled, generated, transformed, and / or otherwise output from a unified software platform used to program the FPGA. For example, the unified software platform may translate first instructions (e.g., code or program) in a high-level language (e.g., C, C++, Python, etc.) corresponding to one or more operations / functions into second instructions in HDL corresponding to those one or more operations / functions. In some such examples, the binary file is compiled, generated, and / or otherwise output from the unified software platform based on the second instructions. In some examples, Figure 11 The FPGA circuit 1100 can access and / or load binary files to enable... Figure 11 The FPGA circuit 1100 is configured and / or constructed to perform one or more operations / functions. For example, the binary file may be transmitted via bit streams (e.g., one or more computer-readable bits, one or more machine-readable bits, etc.), data (e.g., computer-readable data, machine-readable data, etc.) and / or Figure 11 The FPGA circuit 1100 is implemented using machine-readable instructions accessible to cause [access to / from] [the following]. Figure 11 The configuration and / or construction of the FPGA circuit 1100 or one or more of its components.

[0082] Figure 11 The FPGA circuit 1100 includes example input / output (I / O) circuitry 1102 to obtain data from and / or output data to example configuration circuitry 1104 and / or external hardware 1106. For example, configuration circuitry 1104 may be implemented by interface circuitry that can obtain a binary file, which may be implemented as a bitstream, data, and / or machine-readable instructions, to configure FPGA circuitry 1100, or portions thereof. In some such examples, configuration circuitry 1104 may obtain the binary file from a user, a machine (e.g., hardware circuitry that can implement an Artificial Intelligence / Machine Learning (AI / ML) model to generate the binary file (e.g., programmable or dedicated circuitry)), and / or any combination of these. In some examples, external hardware 1106 may be implemented by external hardware circuitry. For example, external hardware 1106 may be... Figure 10 The microprocessor 1000 is implemented.

[0083] The FPGA circuit 1100 also includes an array of example logic gates 1108, multiple example configurable interconnects 1110, and example memory circuitry 1112. The logic gates 1108 and the configurable interconnects 1110 are configurable to instantiate and interact with... Figure 2-3 At least some of the machine-readable instructions correspond to one or more operations / functions, and / or other desired operations. Figure 11 The logic gate circuits 1108 shown are fabricated in blocks or groups. Each block includes semiconductor-based electrical structures that can be configured into logic circuits. In some examples, the electrical structures include logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that provide basic building blocks for the logic circuits. Each logic gate circuit 1108 contains electrically controllable switches (e.g., transistors) to enable the configuration of the electrical structures and / or logic gates to form a circuit that performs a desired operation / function. The logic gate circuits 1108 may include other electrical structures such as look-up tables (LUTs), registers (e.g., flip-flops or latches), multiplexers, etc.

[0084] The configurable interconnect 1110 illustrated is a conductive path, trace, via, or the like, which may include electrically controllable switches (e.g., transistors) whose states can be changed by programming (e.g., using an HDL instruction language) to activate or deactivate one or more connections between one or more logic gates 1108 to program the desired logic circuitry.

[0085] The illustrated storage circuit 1112 is configured to store the results of one or more operations performed by the corresponding logic gates. Storage circuit 1112 can be implemented using a register or similar device. In the illustrated example, storage circuit 1112 is distributed among logic gate circuits 1108 to facilitate access and improve execution speed.

[0086] Figure 11 The example FPGA circuit 1100 also includes example dedicated operation circuitry 1114. In this example, dedicated operation circuitry 1114 includes dedicated circuitry 1116, which can be invoked to implement common functions to avoid the need for field programming of these functions. Examples of such dedicated circuitry 1116 include memory (e.g., DRAM) controller circuitry, PCIe controller circuitry, clock circuitry, transceiver circuitry, memory, and multiplier-accumulator circuitry. Other types of dedicated circuitry may also be present. In some examples, FPGA circuitry 1100 may also include example general-purpose programmable circuitry 1118, such as example CPU 1120 and / or example DSP 1122. Other general-purpose programmable circuitry 1118 may be present additionally or alternatively, such as GPUs, XPUs, etc., which can be programmed to perform other operations.

[0087] Although Figure 10 and Figure 11 The diagram shows Figure 9 These are two example implementations of the programmable circuit 912, but many other schemes are envisioned. For example, the FPGA circuit may include an onboard CPU, such as... Figure 11 One or more example CPUs 1120. Therefore, Figure 9 The programmable circuit 912 can additionally be combined with at least Figure 10 Example microprocessor 1000 and Figure 11 The example FPGA circuit 1100 is implemented. In some such hybrid examples, Figure 11 One or more cores of 1102 can execute Figure 2-3 The flowchart represents the first part of the machine-readable instructions to perform one or more first operations / functions. Figure 11 The FPGA circuit 1100 can be configured and / or constructed to perform operations related to... Figure 2-3 The flowchart represents the second part of the machine-readable instructions corresponding to one or more second operations / functions, and / or the ASIC can be configured and / or constructed to perform the same operations / functions as those described above. Figure 2-3 The flowchart represents the third part of the machine-readable instructions, corresponding to one or more third operations / functions.

[0088] It should be understood that Figure 1 Some or all of the circuits can thus be instantiated at the same or different times. For example, Figure 10 One or more identical and / or different parts of the microprocessor 1000 may be programmed to execute one or more machine-readable instructions at the same and / or different times. In some examples, Figure 11 One or more identical and / or different portions of the FPGA circuit 1100 may be configured and / or constructed to perform operations / functions corresponding to one or more portions of machine-readable instructions at the same and / or different times.

[0089] In some examples, it can be instantiated, for example, in one or more threads that execute simultaneously and / or sequentially. Figure 2 Some or all of the circuits. For example, Figure 10 The microprocessor 1000 can execute machine-readable instructions in one or more threads that execute simultaneously and / or serially. In some examples, Figure 11 The FPGA circuit 1100 can be configured and / or constructed to perform operations / functions simultaneously and / or serially. Furthermore, in some examples, Figure 1 Some or all of the circuits can be Figure 10One or more virtual machines and / or containers are implemented and executed on the microprocessor 1000.

[0090] In some examples, Figure 9 The programmable circuit 912 can be housed in one or more packages. For example, Figure 12 Microprocessor 1200 and / or Figure 11 The FPGA circuit 1100 can be housed in one or more packages. In some examples, the XPU can be... Figure 9 The programmable circuit 912 is implemented and can be in one or more packages. For example, the XPU may include a CPU in a package (e.g., Figure 12 Microprocessor 1200, Figure 11 CPU 1120, etc.), and another packaged DSP (e.g., Figure 11 DSP 1122), GPU in another package, and FPGA in another package (e.g., Figure 11 FPGA circuit 1100).

[0091] exist Figure 12 The diagram illustrates a sample software distribution platform 1205, used to distribute software such as... Figure 9 Example machine-readable instructions 932, such as software, are distributed to other hardware devices (e.g., hardware devices owned and / or operated by a third party different from the owner and / or operator of the software distribution platform). Example software distribution platform 1205 may be implemented by any computer server, data facility, cloud service, etc., capable of storing software and transferring it to other computing devices. A third party may be a customer of the entity that owns and / or operates the software distribution platform 1205. For example, the entity owning and / or operating the software distribution platform 1205 may be the software (e.g., Figure 9 The developer, seller, and / or licensor of the example machine-readable instructions 932. A third party may be a consumer, user, retailer, OEM, etc., who purchases and / or licenses the software for use and / or resells and / or sublicenses it. In the illustrated example, the software distribution platform 1205 includes one or more servers and one or more storage devices. The storage devices store the machine-readable instructions 932, which may correspond to the instructions described above. Figure 2-3Example machine-readable instructions. One or more servers of the example software distribution platform 1205 communicate with the example network 1210, which may correspond to the Internet and / or any one or more of the example networks described above. In some examples, as part of a business transaction, one or more servers respond to a request to transfer software to a requesting party. Payment for the delivery, sale, and / or licensing of the software may be processed by one or more servers of the software distribution platform and / or by a third-party payment entity. These servers enable purchasers and / or licensors to download machine-readable instructions 932 from the software distribution platform 1205. For example, it may be compatible with... Figure 2-3 The software corresponding to the example machine-readable instructions 932 can be downloaded to the example programmable circuit platform 1100, which executes the machine-readable instructions 932 to implement... Figure 1 Pooled actuator circuit 105. In some examples, one or more servers of the software distribution platform 1205 periodically provide, transmit, and / or force software updates (e.g., Figure 9 The example machine-readable instruction 932 ensures that improvements, patches, updates, etc., are distributed and applied to the software at the end-user device. Although referred to as software above, the distributed "software" can also be firmware.

[0092] "Comprising" and "including" (and all their forms and tenses) are used herein as introductory terms. Thus, whenever a claim uses any form of "comprising" or "including" (e.g., including, containing, having, etc.) as a preamble or in any kind of claim recitation, it is understood that additional elements, terms, etc., may exist without falling outside the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term in, for example, the preamble of a claim, it is introductory in the same way that the terms "comprising" and "including" are introductory. The term "and / or" when used, for example, in the form of, say, A, B, and / or C, refers to any combination or subset of A, B, and C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, or (7) A and B and C. For the purposes of this document in the context of describing structures, components, items, objects and / or things, the phrase “at least one of A and B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, for the purposes of this document in the context of describing structures, components, items, objects and / or things, the phrase “at least one of A or B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. For the purposes of this document in the context of describing the execution or operation of processes, instructions, actions, activities, etc., the phrase “at least one of A and B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing the execution or operation of processes, instructions, actions, activities, etc., the phrase “at least one of A or B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.

[0093] As used herein, singular references (e.g., “a,” “an,” “first,” “second,” etc.) do not exclude pluralism. As used herein, the term “a” or “an” refers to one or more of that object. The terms “a” (or “an”), “one or more,” and “at least one” are used interchangeably herein. Furthermore, although listed separately, multiple means, elements, or actions may be implemented by, for example, the same entity or object. Moreover, while individual features may be included in different examples or claims, they may be combined, and inclusion in different examples or claims does not imply that the combination of features is infeasible and / or not advantageous.

[0094] As used herein, the phrase “communicate with”—including its variations—covers direct communication and / or indirect communication via one or more intermediate components, without requiring direct physical (e.g., wired) communication and / or continuous communication, but also includes selective communication at periodic intervals, scheduled intervals, non-periodic intervals, and / or one-off events.

[0095] As used herein, “programmable circuit” is defined as including (i) one or more special-purpose electrical circuits (e.g., special-purpose circuits (ASICs)) configured to perform one or more specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors), and / or (ii) one or more general-purpose semiconductor-based electrical circuits that are programmable with instructions to perform one or more specific functions and / or operations and include one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of programmable circuits include: programmable microprocessors, such as a central processing unit (CPU) capable of executing first instructions to perform one or more operations and / or functions; field-programmable gate arrays (FPGAs) that can be programmed with second instructions to configure and / or construct the FPGA to instantiate one or more operations and / or functions corresponding to the first instructions; graphics processing units (GPUs) capable of executing first instructions to perform one or more operations and / or functions; digital signal processors (DSPs) capable of executing first instructions to perform one or more operations and / or functions; XPUs; network processing units (NPUs); one or more microcontrollers capable of executing first instructions to perform one or more operations and / or functions; and / or integrated circuits, such as application-specific integrated circuits (ASICs). For example, an XPU can be implemented by a heterogeneous computing system that includes multiple types of programmable circuits (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more NPUs, one or more DSPs, etc., and / or any combination of these), and coordination technologies (e.g., one or more application programming interfaces (APIs) that can assign one or more computing tasks to any one or more of the multiple types of programmable circuits that are best suited to perform the one or more computing tasks).

[0096] As used herein, an integrated circuit is defined as one or more semiconductor packages containing one or more circuit elements, such as transistors, capacitors, inductors, resistors, current paths, diodes, and so on. For example, an integrated circuit can be implemented as one or more of ASICs, FPGAs, chips, microchips, programmable circuits, semiconductor substrates coupling multiple circuit elements, systems on chips (SoCs), and so on.

[0097] In summary, it should be understood that the example systems, methods, apparatuses, and artifacts disclosed herein improve GPU-based pooling computation by: (1) reducing the time complexity of 2D pooling by a factor of K (kernel size); (2) reducing memory accesses in the GPU; and (3) reducing the number of operands and / or single-instruction multiple-data (SIMD) instructions in 2D pooling by a factor of K / 2. For a specific element in a specific dimension, the methods and apparatuses disclosed herein perform pooling using the subsequent K elements and update the inputs(one or more) for the subsequent dimensions(one or more). In some examples, intermediate outputs are stored by swapping dimensions to ensure continuous memory accesses for subsequent dimensions. Therefore, the methods and apparatuses disclosed herein improve the performance of GPU-based 2D pooling operations by a factor of K / 2 and reduce the associated computations and / or memory accesses by a factor of K / 2. Thus, the examples disclosed herein achieve improvements in machine performance.

[0098] This document discloses example methods, apparatuses, systems, and artifacts for computing pooling operations on graphics processing units (GPUs). Further examples and combinations thereof include the following: Example 1 includes an apparatus comprising: interface circuitry; machine-readable instructions; and at least one processor circuitry of a graphics processing unit, the at least one processor circuitry being programmed via the machine-readable instructions to: distribute computational tasks among one or more cores based on a batch dimension; perform row pooling operations based on the batch dimension and a first pooling size to generate a first intermediate data output; perform column pooling operations based on the batch dimension and a second pooling size to generate a second intermediate data output; and generate a final output based on the first intermediate data output and the second intermediate data output.

[0099] Example 2 includes the apparatus described in Example 1, wherein one or more of the at least one processor circuitry are configured to: generate the first intermediate data output by identifying the sum of a first group of elements in a row.

[0100] Example 3 includes one or more of the apparatuses described in Examples 1-2, wherein one or more of the at least one processor circuitry is configured to: generate the second intermediate data output by identifying the average value of a second group of elements in a column.

[0101] Example 4 includes one or more of the apparatuses described in Examples 1-3, wherein one or more of the at least one processor circuitry is configured to: store the first intermediate data output by swapping the first dimension with the second dimension for continuous access to a shared local memory.

[0102] Example 5 includes one or more of the apparatuses shown in Examples 1-4, wherein the first dimension is the height of a two-dimensional matrix associated with the first intermediate data output or the second intermediate data output, and the second dimension is the width of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output.

[0103] Example 6 includes one or more of the apparatuses shown in Examples 1-5, wherein one or more of the at least one processor circuitry is configured to: transfer the first intermediate data output or the second intermediate data output from shared local memory to a cache.

[0104] Example 7 includes one or more of the apparatuses shown in Examples 1-6, wherein one or more of the at least one processor circuitry is configured to: apply an extended operation to the first intermediate data output or the second intermediate data output.

[0105] Example 8 includes one or more of the devices shown in Examples 1-7, wherein the expansion operation is one of step or padding.

[0106] Example 9 includes at least one non-transitory machine-readable medium, including machine-readable instructions that cause at least one processor circuitry of a graphics processing unit to at least: allocate computational tasks among one or more cores based on a batch dimension; perform a row pooling operation based on the batch dimension and a first pooling size to generate a first intermediate data output; perform a column pooling operation based on the batch dimension and a second pooling size to generate a second intermediate data output; and generate a final output based on the first intermediate data output and the second intermediate data output.

[0107] Example 10 includes at least one non-transitory machine-readable medium as described in Example 9, wherein the machine-readable instructions cause one or more of the at least one processor circuitry to: generate the first intermediate data output by identifying the sum of a first group of elements in a row.

[0108] Example 11 includes at least one non-transitory machine-readable medium as described in one or more of Examples 9-10, wherein the machine-readable instructions cause one or more of the at least one processor circuitry to: generate the second intermediate data output by identifying the average value of a second group of elements in a column.

[0109] Example 12 includes at least one non-transitory machine-readable medium as described in one or more of Examples 9-11, wherein the machine-readable instructions cause one or more of the at least one processor circuitry to: store the first intermediate data output by swapping a first dimension with a second dimension to enable continuous access to shared local memory.

[0110] Example 13 includes at least one non-transitory machine-readable medium as described in one or more of Examples 9-12, wherein the first dimension is the height of a two-dimensional matrix associated with the first intermediate data output or the second intermediate data output, and the second dimension is the width of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output.

[0111] Example 14 includes at least one non-transitory machine-readable medium as described in one or more of Examples 9-13, wherein the machine-readable instructions cause one or more of the at least one processor circuitry to: transfer the first intermediate data output or the second intermediate data output from shared local memory to a cache.

[0112] Example 15 includes at least one non-transitory machine-readable medium as described in one or more of Examples 9-14, wherein the machine-readable instructions cause one or more of the at least one processor circuitry to: apply an extended operation to the first intermediate data output or the second intermediate data output.

[0113] Example 16 includes at least one non-transitory machine-readable medium as described in one or more of Examples 9-15, wherein the expansion operation is one of stepping or padding.

[0114] Example 17 includes an apparatus comprising: means for distributing computational tasks among one or more cores based on a batch dimension; means for performing a row pooling operation based on the batch dimension and a first pooling size to generate a first intermediate data output; means for performing a column pooling operation based on the batch dimension and a second pooling size to generate a second intermediate data output; and means for generating a final output based on the first intermediate data output and the second intermediate data output.

[0115] Example 18 includes the device described in Example 17, wherein the means for performing row pooling operations generates the first intermediate data output by identifying the sum of a first group of elements in a row.

[0116] Example 19 includes the device described in one or more of Examples 17-18, wherein the means for performing column pooling operations generates the second intermediate data output by identifying the average value of a second group of elements in a column.

[0117] Example 20 includes the device described in one or more of Examples 17-19, wherein the means for performing row pooling operations stores the first intermediate data output by swapping the first dimension with the second dimension to enable continuous access to shared local memory.

[0118] Example 21 includes the device described in one or more of Examples 17-20, wherein the first dimension is the height of a two-dimensional matrix associated with the first intermediate data output or the second intermediate data output, and the second dimension is the width of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output.

[0119] Example 22 includes the device described in one or more of Examples 17-21, wherein one or more of the at least one processor circuitry is configured to: transfer the first intermediate data output or the second intermediate data output from shared local memory to a cache.

[0120] Example 23 includes the device described in one or more of Examples 17-22, wherein one or more of the at least one processor circuitry is configured to: apply an extended operation to the first intermediate data output or the second intermediate data output.

[0121] Example 24 includes the device described in one or more of Examples 17-23, wherein the expansion operation is one of step or padding.

[0122] The appended claims are hereby incorporated into this Detailed Description section by reference. While certain example systems, methods, apparatuses, and articles of manufacture are disclosed herein, the scope of this patent is not limited thereto. Rather, this patent covers all systems, methods, apparatuses, and articles of manufacture that fairly fall within the scope of the claims of this patent.

Claims

1. An apparatus comprising: Interface circuit; Machine-readable instructions; as well as At least one processor circuit of the graphics processing unit, said at least one processor circuit being programmed by the machine-readable instructions to: Distribute computing tasks across one or more cores based on the batch dimension; Row pooling is performed based on the batch dimension and the first pooling size to generate a first intermediate data output; Column pooling is performed based on the batch dimension and the second pooling size to produce a second intermediate data output; as well as The final output is generated based on the first intermediate data output and the second intermediate data output.

2. The apparatus of claim 1, wherein, One or more of the at least one processor circuitry are used to: generate the first intermediate data output by identifying the sum of the first group of elements in a row.

3. The apparatus as claimed in claim 1 or 2, wherein, One or more of the at least one processor circuitry are configured to: generate the second intermediate data output by identifying the average value of a second group of elements in a column.

4. The apparatus as claimed in claim 1, 2, or 3, wherein, One or more of the at least one processor circuitry are configured to: store the first intermediate data output by swapping the first dimension with the second dimension, in order to enable continuous access to the shared local memory.

5. The apparatus of claim 4, wherein, The first dimension is the height of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output, and the second dimension is the width of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output.

6. The apparatus as claimed in claim 1, 2, 3, 4 or 5, wherein, One or more of the at least one processor circuitry are configured to: transfer the first intermediate data output or the second intermediate data output from the shared local memory to a cache.

7. The apparatus as claimed in claim 1, 2, 3, 4, 5 or 6, wherein, One or more of the at least one processor circuitry are configured to: apply an extended operation to the first intermediate data output or the second intermediate data output.

8. The apparatus of claim 7, wherein, The expansion operation is at least one of step or padding.

9. At least one machine-readable medium, including machine-readable instructions, such that at least one processor circuitry of the graphics processing unit is used for at least: Distribute computing tasks across one or more cores based on the batch dimension; Row pooling is performed based on the batch dimension and the first pooling size to generate a first intermediate data output; Column pooling is performed based on the batch dimension and the second pooling size to produce a second intermediate data output; as well as The final output is generated based on the first intermediate data output and the second intermediate data output.

10. The at least one machine-readable medium as claimed in claim 9, wherein, The machine-readable instructions cause one or more of the at least one processor circuitry to: generate the first intermediate data output by identifying the sum of a first group of elements in a row.

11. The at least one machine-readable medium as claimed in claim 9 or 10, wherein, The machine-readable instructions cause one or more of the at least one processor circuitry to: generate the second intermediate data output by identifying the average value of a second group of elements in a column.

12. The at least one machine-readable medium as claimed in claim 9, 10, or 11, wherein, The machine-readable instructions cause one or more of the at least one processor circuitry to: store the first intermediate data output by swapping the first dimension with the second dimension, in order to enable continuous access to the shared local memory.

13. The at least one machine-readable medium as claimed in claim 12, wherein, The first dimension is the height of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output, and the second dimension is the width of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output.

14. The at least one machine-readable medium as described in claim 9, 10, 11, 12, or 13, wherein, The machine-readable instructions cause one or more of the at least one processor circuitry to: transfer the first intermediate data output or the second intermediate data output from the shared local memory to a cache.

15. At least one machine-readable medium as described in claims 9, 10, 11, 12, 13, or 14, wherein, The machine-readable instructions cause one or more of the at least one processor circuitry to: apply an extended operation to the first intermediate data output or the second intermediate data output.

16. The at least one machine-readable medium as described in claims 9, 10, 11, 12, 13, 14, or 15, wherein, The expansion operation is at least one of step or padding.

17. An apparatus comprising: A device for distributing computing tasks among one or more cores based on batch dimensions; A means for performing row pooling operations based on the batch dimension and the first pooling size to generate a first intermediate data output; A means for performing column pooling operations based on the batch dimension and the second pooling size to produce a second intermediate data output; as well as An apparatus for generating a final output based on the first intermediate data output and the second intermediate data output.

18. The device as claimed in claim 17, wherein, The apparatus for performing row pooling operations generates the first intermediate data output by identifying the sum of the first group of elements in a row.

19. The device as claimed in claim 17 or 18, wherein, The apparatus for performing column pooling operations generates the second intermediate data output by identifying the average value of a second group of elements in a column.

20. The device as claimed in claim 17, 18 or 19, wherein, The apparatus for performing row pooling operations stores the first intermediate data output by swapping the first dimension with the second dimension to enable continuous access to shared local memory.

21. A method comprising: Distribute computing tasks across one or more cores based on the batch dimension; Row pooling is performed based on the batch dimension and the first pooling size to generate a first intermediate data output; Column pooling is performed based on the batch dimension and the second pooling size to produce a second intermediate data output; as well as The final output is generated based on the first intermediate data output and the second intermediate data output.

22. The method of claim 21, further comprising: The first intermediate data output is generated by identifying the sum of the first group of elements in a row.

23. The method of claims 21 and 22, further comprising: The second intermediate data output is generated by identifying the average value of the second group of elements in a column.

24. The method of claim 21, 22 or 23, further comprising: The first intermediate data output is stored by swapping the first dimension with the second dimension to enable continuous access to the shared local memory.

25. The method as described in claim 21, 22, 23 or 24, wherein, The first dimension is the height of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output, and the second dimension is the width of the two-dimensional matrix associated with the first intermediate data output or the second intermediate data output.