Method for solving problem of not supporting three-dimensional (3D) operator on neural processing unit (NPU)
By using 2D operators to simulate 3D operators on NPUs, the problem of NPU not supporting 3D operators is solved, and the effect of effectively processing 3D data on hardware platforms that do not support 3D operators is achieved, expanding the application range of NPUs and maintaining high performance.
Patent Information
- Application Number
- CN202510199962.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
Many neural processing units (NPUs) and dedicated hardware platforms fail to natively support 3D operators, such as 3D convolution and 3D pooling, limiting the wide deployment of 3D CNNs, resulting in wasted computing resources and reduced processing speed.
The 3D convolution operator, 3D pooling operator and 3D batch normalization operator are simulated by using 2D convolution operators, 2D pooling operators and 2D batch normalization operators, and the spatial dimensions and temporal dimensions of 3D data are decoupled and processed respectively.
It realizes efficient processing of 3D data on NPUs that do not support 3D operators, expands the application range of NPUs, and maintains the same performance level as the original 3D operators.
Smart Images

Figure CN120146126A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a method for simulating 3D operators using two-dimensional (2D) operators, which relates to the technical field of deep learning and is a technical improvement for the inability of neural processing units (NPUs) to use 3D operators. Background Art
[0002] Three-dimensional convolutional networks (3D CNNs) have become a key technology for processing video sequences and three-dimensional image data. These technologies are widely used in multiple fields and particularly play an important role in enhancing the understanding and analysis of video content. Well-known implementations include I3D, R3D, S3D, Non-local, and SlowFast networks, etc. They capture information in the temporal and spatial dimensions by using three-dimensional convolutional operations, significantly improving the accuracy and efficiency of data processing.
[0003] The current technical challenge is that many neural processing units (NPUs) and dedicated hardware platforms do not natively support 3D operators such as 3D convolution and 3D pooling. This limitation not only hinders the wide deployment of 3D CNNs but also leads to waste of computing resources and a significant reduction in processing speed. Especially in application scenarios that require real-time or near-real-time data processing, this technical limitation becomes a key bottleneck. In view of this, there is a strong need to develop a new technical solution that can achieve equivalent 3D data processing capabilities on existing NPUs and hardware platforms that do not support 3D convolution operations. Such an improvement will directly affect processing speed and resource utilization, thus promoting the implementation and optimization of 3D CNN technology in a wider range of practical applications. Summary of the Invention
[0004] The object of the present invention is to overcome the deficiencies in the prior art and propose a method for simulating 3D convolution operators using two-dimensional (2D) convolution operators, which can solve the problem that NPUs do not support 3D operators.
[0005] To achieve the above object, the present invention adopts the following technical solutions: Simulate 3D convolution operators using 2D convolution operators; Simulate 3D pooling operators using 2D pooling operators; Simulate 3D batch normalization operators using 2D batch normalization operators.
[0006] When using a 2D convolution operator to simulate a 3D convolution operator, it is first necessary to decouple the spatial and temporal dimensions of the 3D convolution operator, and decompose the original 3D convolution operation into two successive 3D convolution kernels, namely a 3D spatial convolution kernel and a 3D temporal convolution kernel. Specifically, first perform feature extraction in the spatial dimension for each independent layer of the 3D data. In this process, a 3D spatial convolution kernel of size 1×H×W is required, where H and W represent the height and width of the image. To convert the 3D spatial convolution kernel into a 2D convolution kernel, its size can be converted from 1×H×W to a 2D convolution kernel of H×W, so that features can be effectively extracted in the spatial dimension. Then, perform feature extraction operations in the temporal dimension on each independent layer of the 3D data.
[0007] In this step, the size of the original 3D temporal convolution kernel is T×1×1, where T represents the temporal dimension. After replacing this 3D temporal convolution kernel with a 2D convolution kernel of T×1, features can be extracted in the temporal dimension.
[0008] When using a 2D pooling operator to simulate the 3D pooling operator during the input data dimension stacking, it is first necessary to stack the temporal dimension of the 3D data onto the batch dimension to form a new data format. Then, decompose the 3D pooling operation into two steps of 2D pooling operations. First, use a 2D pooling operator of size H×W to perform pooling processing on the spatial dimension of the input data. Second, use a 2D pooling operator of size T×1 to perform pooling processing on the temporal dimension. In these two steps, the type of 2D pooling operator used is consistent with the 3D pooling operator to ensure that the form of the operation matches the original 3D pooling operator.
[0009] When using a 2D batch normalization operator to simulate the 3D batch normalization operator during the input data dimension stacking process, it is first necessary to stack the temporal dimension of the 3D data with the channel dimension to form a new data format. Then, apply the 2D batch normalization operator to perform batch normalization on the processed data. This step normalizes the data in the spatial dimension to ensure the balance of data in each dimension during calculation. Finally, restore the data after batch normalization to the original 3D data format to restore the correct relationship between the temporal dimension and the channel dimension. Through this series of steps, the 2D batch normalization operator successfully simulates the function of the 3D batch normalization operator.
[0010] The beneficial effects of the present invention are as follows: The present invention proposes a method for using a two-dimensional (2D) operator to simulate a 3D operator, aiming to achieve the same performance level as the original 3D operator. Through this simulation technology, it is possible to effectively process 3D data on an NPU that does not support 3D operators, thereby expanding the application scope of the NPU. Description of the Drawings
[0011] Figure 1 is a decoupling schematic diagram of 3D operators; Figure 2 is a schematic diagram of the dimension stacking of 3D data during the pooling operation; Figure 3 is a schematic diagram of the dimension stacking of 3D data during batch normalization; Figure 4 is the structure of the 3D ResNets (R3D) model; Figure 5 is the accuracy curve of the R3D-18 model and the model that uses 2D operators to replace all operators in the R3D model; Figure 6 is the accuracy curve of the R3D-50 model and the model that uses 2D operators to replace all operators in the R3D model. Detailed implementation manners
[0012] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.
[0013] This embodiment introduces a method for solving the problem that a neural processing unit (NPU) does not support three-dimensional (3D) operators, including: using a 2D convolution operator to simulate a 3D convolution operator; using a 2D pooling operator to simulate a 3D pooling operator; using a 2D batch normalization operator to simulate a 3D batch normalization operator.
[0014] As Figure 1 , when using a 2D convolution operator to simulate a 3D convolution operator, it is first necessary to decouple the spatial dimension and the time dimension of the 3D convolution operator, and disassemble the original 3D convolution operation into two successive 3D convolution kernels, namely a 3D spatial convolution kernel and a 3D time convolution kernel; specifically, first perform feature extraction on the spatial dimension for each independent layer of the 3D data; in this process, a 3D spatial convolution kernel with a size of 1×H×W is required, where H and W represent the height and width of the image.
[0015] In order to convert the 3D spatial convolution kernel into a 2D convolution kernel, its size can be converted from 1×H×W to a 2D convolution kernel of H×W, so that features can be effectively extracted in the spatial dimension; then, perform feature extraction operations on the time dimension for each independent layer of the 3D data; in this step, the size of the original 3D time convolution kernel is T×1×1, where T represents the time dimension; after replacing the 3D time convolution kernel with a 2D convolution kernel of T×1, features can be extracted in the time dimension.
[0016] Figure 2Shows a schematic diagram of the superposition of input data dimensions when using a 2D pooling operator to simulate a 3D pooling operator; First, the time dimension of the 3D data is superimposed on the batch dimension to form a new data format; Then, the 3D pooling operation is decomposed into two steps of 2D pooling operations. In the first step, a 2D pooling operator with a size of H×W is used to perform pooling processing on the spatial dimension of the input data; In the second step, a 2D pooling operator with a size of T×1 is used to perform pooling processing on the time dimension. In these two steps, the type of 2D pooling operator used is consistent with the 3D pooling operator to ensure that the form of the operation matches the original 3D pooling operator.
[0017] Figure 3 Shows the dimension recovery process of the input data when using a 2D batch normalization operator to simulate a 3D batch normalization operator; First, the time dimension and the channel dimension of the 3D data are superimposed to form a new data format, and the superimposing method is the same as that shown in Figure 2 ; Then, a 2D batch normalization operator is applied to perform batch normalization processing on the processed data; This step normalizes the data in the spatial dimension to ensure the balance of data in each dimension during calculation; Finally, the data after batch normalization processing is restored to the original 3D data format to restore the correct relationship between the time dimension and the channel dimension, and the restoration process is as shown in Figure 3 ; Through this series of steps, the 2D batch normalization operator successfully simulates the function of the 3D batch normalization operator.
[0018] Figure 4 Shows the structure of the 3D ResNet (R3D) model, in which all 3D operators in the model (including 3D convolution operators, 3D pooling operators, and 3D batch normalization operators) are sequentially converted into corresponding 2D operators; The specific operation is as follows: First, the 3D convolution operator is replaced with a 2D convolution operator, and the conversion is carried out according to the time dimension and the spatial dimension to ensure that the features in each dimension can be extracted; Then, the 3D pooling operator is replaced with a 2D pooling operator, and the time and spatial dimensions in the 3D pooling operation are decoupled and processed separately using 2D pooling operators; Then, the 3D batch normalization operator is replaced with a 2D batch normalization operator. At this time, the time dimension and the channel dimension of the 3D data need to be superimposed together, and the 2D batch normalization operator is applied for processing, and finally the data is restored to the original 3D data format.
[0019] After completing these replacement operations, a new network model is obtained, and this model contains all 3D operators replaced by 2D operators; Subsequently, the initial R3D-50 model, R3D-18 model, and the new model after replacing 3D operators are trained and verified using the UCF-101 dataset, and the accuracy change curves are as shown in Figure 5 and Figure 6As shown; it can be seen from the figure that the accuracy of the R3D model using 2D operators instead of 3D operators has an error of about 1.5% compared with the original model, and the overall performance basically meets the requirements.
[0020] The above are only the embodiments of the present invention. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for solving the problem of not supporting three-dimensional (3D) operators on a neural processing unit (NPU), characterized in that: include: Use 2D convolution operator to simulate 3D convolution operator; Use 2D pooling operator to simulate 3D pooling operator; Use the 2D Batch Normalization operator to simulate the 3D Batch Normalization operator.
2. The method for solving the problem of not supporting three-dimensional (3D) operators on a neural processing unit (NPU) according to claim 1, characterized in that: Use 2D convolution operators to simulate 3D convolution operators, including: Decouple the spatial dimension and temporal dimension of the 3D convolution operator, and decompose the original 3D convolution operation into two consecutive 3D convolution kernels, namely the 3D spatial convolution kernel and the 3D temporal convolution kernel; First, perform feature extraction on each independent layer of the 3D data in the spatial dimension. In this process, a 3D spatial convolution kernel of size 1×H×W is used, where H and W represent the height and width of the image. In order to convert the 3D spatial convolution kernel into a 2D convolution kernel, its size can be converted from 1×H×W to a 2D convolution kernel of H×W, so that features can be effectively extracted in the spatial dimension; Next, feature extraction is performed on each independent layer of the 3D data in the time dimension. In this step, the size of the original 3D temporal convolution kernel is T×1×1, where T represents the time dimension. After replacing the 3D temporal convolution kernel with a T×1 2D convolution kernel, features can be extracted in the time dimension.
3. The method for solving the problem of not supporting three-dimensional (3D) operators on a neural processing unit (NPU) according to claim 1, characterized in that: Use 2D pooling operators to simulate 3D pooling operators, including: When using a 2D pooling operator to simulate a 3D pooling operator and the input data dimensions are superimposed, the time dimension of the 3D data needs to be superimposed on the batch dimension to form a new data format; Next, the 3D pooling operation needs to be decomposed into two 2D pooling operations. In the first step, a 2D pooling operator of size H×W is used to pool the spatial dimension of the input data. In the second step, a 2D pooling operator of size T×1 is used to pool the temporal dimension. In these two steps, the type of 2D pooling operator used is consistent with the 3D pooling operator to ensure that the form of the operation matches the original 3D pooling operator.
4. The method for solving the problem of not supporting three-dimensional (3D) operators on a neural processing unit (NPU) according to claim 1, characterized in that: Use the 2D Batch Normalization operator to simulate the 3D Batch Normalization operator, including: When using the 2D batch normalization operator to simulate the dimension superposition process of the input data of the 3D batch normalization operator, it is necessary to superimpose the time dimension and channel dimension of the 3D data to form a new data format; Next, the 2D batch normalization operator is applied to the processed data for batch normalization. This step normalizes the data in the spatial dimension to ensure the balance of data in each dimension during calculation. Finally, the batch normalized data is restored to the original 3D data format to restore the correct relationship between the time dimension and the channel dimension. Through this series of steps, the 2D batch normalization operator successfully simulates the function of the 3D batch normalization operator.