Data processing device, method and related products for executing neural network model

By folding the dimensions of the convolutional layer filters and rearranging the data in the width and height dimensions to the input channel dimension, the problem of redundant calculation in the convolution operation is solved, and the computing performance and resource utilization efficiency are improved.

CN114764608BActive Publication Date: 2025-09-26CAMBRICON TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011631707.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2025-09-26
Estimated Expiration
2041-03-10

AI Technical Summary

Technical Problem

In the existing technology, the computational performance of convolution operations is limited by the alignment requirements of various filter dimensions and hardware alignment requirements, resulting in redundant calculations and resource waste, especially when the input channel dimension is small.

Method used

By folding the dimensions of the convolutional layer filters, the data in the width and height dimensions are rearranged to the input channel dimension, and the convolution operation is optimized to reduce redundant calculations and improve computing performance.

Benefits of technology

Through dimension folding optimization, redundant calculations are reduced, and the computational performance of convolution operations and the efficiency of hardware resource utilization are improved, especially when the input channel dimension is small.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114764608B_ABST
    Figure CN114764608B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a data processing device, method, and related products for executing a neural network model. The data processing device can be included as a computing device in a combined processing device, and the combined processing device can also include an interface device and other processing devices. The computing device interacts with the other processing devices to jointly complete the computing operations specified by the user. The combined processing device can also include a storage device, which is connected to the computing device and the other processing devices respectively and is used to store data from the computing device and the other processing devices. The solution disclosed in the present disclosure optimizes the convolution operation of multidimensional arrays and improves the processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of data processing. More specifically, the present disclosure relates to a data processing device, a data processing method, a chip, and a board for executing a neural network model. Background Art

[0002] Deep learning has become a key branch of machine learning and is significantly driving the development of artificial intelligence (AI). Deep neural networks (DNNs), the core technology of deep learning, have been widely applied across numerous industries.

[0003] The convolutional layer is a commonly used hidden layer in neural network models, extracting features from input data through convolution operations. Neural network models incorporate numerous convolution operations, and their computational performance significantly impacts the overall performance of the neural network model. Convolution operations require both instruction alignment and hardware (e.g., parallel arithmetic units) alignment for each dimension of the convolutional layer's filters. Therefore, convolution operations need to be optimized to improve the computational performance of neural network models. Summary of the Invention

[0004] To at least address one or more of the technical issues mentioned above, the present disclosure proposes, in various aspects, a data processing scheme for executing a neural network model. By transforming the filters of the convolutional layer, the data processing scheme can effectively improve the computational performance of the convolution operation. The neural network model of the embodiments of the present disclosure can be applied to various fields, such as image processing, speech processing, text processing, etc. These processes can include, but are not limited to, recognition and classification.

[0005] In a first aspect, the present disclosure provides a data processing apparatus for executing a neural network model, comprising:

[0006] a storage circuit configured to store a folded filter of a convolutional layer of the neural network model, wherein the folded filter is obtained by dimensional folding of the original filter, wherein the dimensional folding includes rearranging data of a width dimension and / or a height dimension to an input channel dimension; and

[0007] Processing circuitry configured to:

[0008] Performing dimension folding on the input feature map to obtain a folded feature map; and

[0009] A convolution operation is performed on the folded feature map using the folded filter to obtain an output feature map.

[0010] In a second aspect, the present disclosure provides a chip comprising the data processing device of any embodiment of the aforementioned first aspect.

[0011] In a third aspect, the present disclosure provides a board comprising a chip according to any one of the embodiments of the second aspect.

[0012] In a fourth aspect, the present disclosure provides a method for executing a neural network model implemented by a data processing device, the data processing device including a storage circuit and a processing circuit, the method comprising:

[0013] The processing circuit performs dimension folding on the input feature map to obtain a folded feature map;

[0014] The processing circuit performs a convolution operation on the folded feature map using the folded filter of the convolution layer of the neural network model stored in the storage circuit to obtain an output feature map;

[0015] The folded filter is obtained by folding the original filter in the dimension, and the dimension folding includes rearranging data in the width dimension and / or height dimension to the input channel dimension.

[0016] Through the data processing device, chip, board and data processing method implemented by the data processing device as provided above, the scheme disclosed herein optimizes convolution operations by folding filters. The embodiments disclosed herein are particularly suitable for situations where the input channel dimension of the original filter is small. In conventional convolution operations, when the input channel dimension of the filter is small, due to the limitations of the vectorized alignment of the artificial intelligence chip instruction set, more redundant calculations will be caused. The embodiments disclosed herein can reduce redundant calculations as much as possible, avoid wasting computing resources, and improve the computing performance of convolution operations during hardware acceleration by folding the data of the convolution kernel width dimension and / or height dimension of the original filter onto the input channel dimension to meet the alignment requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0018] Figure 1 is a structural diagram showing a board according to an embodiment of the present disclosure;

[0019] Figure 2 is a structural diagram illustrating an integrated circuit device according to an embodiment of the present disclosure;

[0020] Figure 3is a schematic diagram showing the internal structure of a single-core computing device according to an embodiment of the present disclosure;

[0021] Figure 4 is a schematic diagram showing the internal structure of a multi-core computing device according to an embodiment of the present disclosure;

[0022] Figure 5 is a schematic diagram showing the internal structure of a processor core according to an embodiment of the present disclosure;

[0023] Figure 6 An exemplary convolution operation example to which embodiments of the present disclosure may be applied is shown;

[0024] Figure 7 An exemplary schematic diagram of a data processing scheme according to an embodiment of the present disclosure is shown;

[0025] Figure 8 shows a more detailed filter folding schematic diagram according to an embodiment of the present disclosure;

[0026] Figure 9 Schematic diagram showing the folding and padding of the convolution kernel according to the embodiment of the present disclosure;

[0027] Figure 10 A schematic diagram schematically illustrates the effect of the convolution step size on the effective multiple according to an embodiment of the present disclosure;

[0028] Figure 11 A schematic structural diagram exemplarily shows a data processing device that can implement the embodiments of the present disclosure; and

[0029] Figure 12 An exemplary flow chart of a data processing method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.

[0031] It should be understood that the terms "first," "second," and "third," etc., which may be used in the claims, specification, and drawings of this disclosure, are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0032] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0033] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0034] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0035] Figure 1 FIG. 1 is a schematic diagram showing the structure of a board 10 according to an embodiment of the present disclosure. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC), or system-on-chip, integrated with one or more combined processing devices. The combined processing device is an artificial intelligence computing unit that supports various deep learning and machine learning algorithms to meet the intelligent processing needs in complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in the field of cloud intelligence. A notable feature of cloud-based intelligent applications is the large amount of input data, which places high demands on the platform's storage and computing capabilities. Board 10 of this embodiment is suitable for cloud-based intelligent applications and has extensive off-chip storage, on-chip storage, and powerful computing capabilities.

[0036] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. Calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface formats, such as a PCIe interface.

[0037] Board 10 also includes a memory device 104 for storing data, which includes one or more storage units 105. Memory device 104 is connected to a control device 106 and chip 101 via a bus for data transmission. Control device 106 in board 10 is configured to control the state of chip 101. To this end, in one application scenario, control device 106 may include a microcontroller (MCU).

[0038] Figure 2 FIG. 1 is a block diagram showing the combined processing device in the chip 101 of this embodiment. Figure 2 As shown in , the combined processing device 20 includes a computing device 201 , an interface device 202 , a processing device 203 and a DRAM 204 .

[0039] The computing device 201 is configured to perform user-specified operations and is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.

[0040] Interface device 202 is used to transmit data and control instructions between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to a storage device on-chip of computing device 201. Furthermore, computing device 201 can obtain control instructions from processing device 203 via interface device 202 and write them to a control cache on-chip of computing device 201. Alternatively or optionally, interface device 202 can also read data from the storage device of computing device 201 and transmit it to processing device 203.

[0041] The processing device 203 is a general processing device that performs basic controls including but not limited to data handling, starting and / or stopping the computing device 201. Depending on the implementation method, the processing device 203 can be a central processing unit (CPU), a graphics processing unit (GPU) or one or more types of processors in other general and / or special processors. These processors include but are not limited to digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, with respect to the computing device 201 of the present disclosure, it can be regarded as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 201 and the processing device 203 are integrated and considered together, the two are regarded as forming a heterogeneous multi-core structure.

[0042] The DRAM 204 is used to store data to be processed. It is a DDR memory, typically 16G or larger, and is used to store data of the computing device 201 and / or the processing device 203 .

[0043] Figure 3 The single-core computing device 301 is used to process input data for computer vision, speech, natural language, data mining, etc. The single-core computing device 301 includes three modules: a control module 31, a computing module 32, and a storage module 33.

[0044] The control module 31 coordinates and controls the operations of the computing module 32 and the storage module 33 to complete deep learning tasks. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 retrieves instructions from the processing device 203, while the instruction decode unit 312 decodes the retrieved instructions and sends the decoded results as control information to the computing module 32 and the storage module 33.

[0045] The operation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations, supporting complex operations such as vector multiplication, addition, and nonlinear transformations. The matrix operation unit 322 is responsible for the core calculations of the deep learning algorithm, namely matrix multiplication and convolution.

[0046] The storage module 33 is used to store or transfer relevant data and includes a neuron RAM (NRAM) 331, a weight RAM (WRAM) 332, and a direct memory access (DMA) module 333. NRAM 331 stores input neurons, output neurons, and intermediate computational results; WRAM 332 stores the convolution kernels (i.e., weights) of the deep learning network; and DMA 333, which connects to DRAM 204 via bus 34, transfers data between the single-core computing device 301 and DRAM 204.

[0047] Figure 4 The figure shows the internal structure of a multi-core computing device 201. The multi-core computing device 41 employs a layered design. As a system-on-chip (SoC), the multi-core computing device 41 includes at least one cluster, each of which includes multiple processor cores. In other words, the multi-core computing device 41 is structured in a hierarchy of SoC-cluster-processor cores.

[0048] At the system-on-chip level, Figure 4 As shown, the multi-core computing device 41 includes an external storage controller 401 , a peripheral communication module 402 , an on-chip interconnect module 403 , a synchronization module 404 and multiple clusters 405 .

[0049] There can be multiple external storage controllers 401, and two are shown in the figure as an example. They are used to respond to access requests issued by the processor core and access external storage devices, such as Figure 2DRAM 204 in the chip, thereby reading data from outside the chip or writing data. The peripheral communication module 402 is used to receive control signals from the processing device 203 through the interface device 202, and start the computing device 201 to perform tasks. The on-chip interconnect module 403 connects the external storage controller 401, the peripheral communication module 402 and multiple clusters 405 to transmit data and control signals between each module. The synchronization module 404 is a global synchronization barrier controller (GBC) used to coordinate the work progress of each cluster and ensure information synchronization. Multiple clusters 405 are the computing cores of the multi-core computing device 41. Four are shown as an example in the figure. With the development of hardware, the multi-core computing device 41 of the present disclosure can also include 8, 16, 64, or even more clusters 405. Clusters 405 are used to efficiently execute deep learning algorithms.

[0050] At the cluster level, Figure 4 As shown, each cluster 405 includes multiple processor cores (IPU cores) 406 and a memory core (MEM core) 407 .

[0051] The figure shows four processor cores 406 as an example, but the present disclosure does not limit the number of processor cores 406. Figure 5 Each processor core 406 is similar to Figure 3 The single-core computing device 301 also includes three major modules: a control module 51, a computing module 52, and a storage module 53. The functions and structures of the control module 51, computing module 52, and storage module 53 are roughly the same as those of the control module 31, computing module 32, and storage module 33, and will not be repeated here. It should be noted that the storage module 53 includes an input / output direct memory access module (IODMA) 533 and a move direct memory access module (MVDMA) 534. The IODMA 533 controls memory access between the NRAM 531 / WRAM 532 and the DRAM 204 via the broadcast bus 409; the MVDMA 534 is used to control memory access between the NRAM 531 / WRAM 532 and the storage unit (SRAM) 408.

[0052] Back to Figure 4The storage core 407 is primarily used for storage and communication, namely, storing shared data or intermediate results between the processor cores 406, and performing communication between the cluster 405 and the DRAM 204, between the clusters 405, and between the processor cores 406. In other embodiments, the storage core 407 has scalar operation capabilities and is used to perform scalar operations.

[0053] The storage core 407 includes SRAM 408, a broadcast bus 409, a cluster direct memory access module (CDMA) 410, and a global direct memory access module (GDMA) 411. SRAM 408 acts as a high-performance data transfer station. Data reused between different processor cores 406 within the same cluster 405 does not need to be obtained from DRAM 204 by each processor core 406. Instead, it is transferred between the processor cores 406 via SRAM 408. The storage core 407 only needs to quickly distribute the reused data from SRAM 408 to multiple processor cores 406, thereby improving inter-core communication efficiency and significantly reducing on-chip and off-chip input / output access.

[0054] The broadcast bus 409, CDMA 410, and GDMA 411 are used for communication between the processor cores 406, communication between the clusters 405, and data transmission between the clusters 405 and the DRAM 204, respectively. Each of these will be described below.

[0055] Broadcast bus 409 facilitates high-speed communication between processor cores 406 within cluster 405. In this embodiment, broadcast bus 409 supports inter-core communication methods including unicast, multicast, and broadcast. Unicast refers to point-to-point data transmission (e.g., from one processor core to another), multicast transfers a copy of data from SRAM 408 to a specific number of processor cores 406, and broadcast transfers a copy of data from SRAM 408 to all processor cores 406, a special case of multicast.

[0056] The CDMA 410 is used to control memory access to the SRAM 408 between different clusters 405 within the same computing device 201 .

[0057] GDMA 411 works in conjunction with external memory controller 401 to control memory access from cluster 405's SRAM 408 to DRAM 204, or to read data from DRAM 204 to SRAM 408. As previously mentioned, communication between DRAM 204 and NRAM 431 or WRAM 432 can be achieved through two channels. The first channel directly connects DRAM 204 and NRAM 431 or WRAM 432 via IODAM 433. The second channel first transfers data between DRAM 204 and SRAM 408 via GDMA 411, and then transfers data between SRAM 408 and NRAM 431 or WRAM 432 via MVDMA 534. While the second channel may appear to require more components and a longer data flow, in some embodiments, the bandwidth of the second channel is significantly greater than the first channel. Therefore, communication between DRAM 204 and NRAM 431 or WRAM 432 through the second channel may be more efficient. The embodiments of the present disclosure can select a data transmission channel according to the hardware conditions.

[0058] In other embodiments, the functions of GDMA 411 and IODMA 533 can be integrated into the same component. For ease of description, this disclosure considers GDMA 411 and IODMA 533 as separate components. For those skilled in the art, as long as the functions implemented and the technical effects achieved are similar to those of this disclosure, they are within the scope of protection of this disclosure. Furthermore, the functions of GDMA 411, IODMA 533, CDMA 410, and MVDMA 534 can also be implemented by the same component.

[0059] A neural network model usually includes an input layer, a convolutional layer, an activation function, a pooling layer, a fully connected layer, etc., with a few layers and hundreds of layers. Each layer executes an operator. For example, the convolutional layer executes the convolution operator. The number of operators required depends on the number of layers.

[0060] Neural network model training involves inputting training samples to adjust the parameters of each layer, ensuring that the results calculated by the neural network model are as close as possible to the real-world results. Neural network model training involves forward propagation and backward propagation. Forward propagation, based on an existing model, passes the training samples through the layers of the neural network model, gradually extracting the input feature map into abstract features. Backward propagation, on the other hand, uses gradient descent and the chain rule to calculate the partial derivative of the loss function with respect to each parameter to update the parameters. Training is then repeated using the updated parameters, and this process is repeated multiple times until the forward propagation results meet expectations. Using the trained neural network model to perform forward operations on real-world inputs to complete a task is called inference.

[0061] Based on the aforementioned hardware environment, the disclosed embodiment provides a data processing solution for executing a neural network model, and more specifically, a solution for optimizing convolution operations in a neural network model.

[0062] Figure 6 An exemplary convolution operation example to which the present disclosure can be applied is shown. As shown in the figure, the convolution layer in the neural network model can perform feature extraction by applying a filter to the input feature map and performing convolution processing.

[0063] The figure illustrates an input feature map of size 6×6×3, which can represent three 6×6 feature maps (i.e., a 6×6×3 three-dimensional matrix), each representing three different features. In this example, the width W of the feature map is 6, and the height H is also 6. The number of input feature maps can also be referred to as the number of input channels Ci. For example, the input example in the figure has three feature maps, also known as three feature channels.

[0064] The figure also shows an example of a filter of size 2×3×3×3, which can represent two convolution kernels of size 3×3×3 (i.e., two three-dimensional matrices of size 3×3×3). Each convolution kernel has three different convolution kernels of size 3×3, corresponding to three different feature maps of the input. The number of stereo convolution kernels can be called the number of output channels Co, which is 2 in this example. In each stereo convolution kernel, the number of two-dimensional convolution kernels can be called the number of input channels Ci, which is consistent with the number of channels of the input feature map. Each two-dimensional convolution kernel has a corresponding width Kw and height Kh. In this example, Kw and Kh are both 3.

[0065] The convolution of the input feature map with the filter results in two 4×4 feature maps. The convolution of the input feature map with the upper stereo convolution kernel produces the upper 4×4 output feature map, while the convolution of the input feature map with the lower stereo convolution kernel produces the lower 4×4 output feature map. The value at each position in the output feature map is obtained by performing a two-dimensional convolution operation on the corresponding block of each input feature map with the corresponding convolution kernel, followed by the summation. For example, the figure shows that the value at position (0,0) in the upper output feature map is obtained by performing a two-dimensional convolution operation on the block outlined by the black cube in the input feature map with the upper stereo convolution kernel, resulting in three values, which are then summed. To obtain outputs at other positions, the convolution kernel can be moved within the input feature map. In the example shown, the convolution stride (Sx, Sy) is (1,1). When the convolution operation is performed after shifting one grid square to the right (width) or one grid square down (height), the values ​​at positions (0,1) or (1,0) in the upper output feature map are obtained, respectively.

[0066] As can be seen from the above description, a convolutional layer in a neural network has a set of input feature maps, containing a total of H×W×Ci information, where H and W are the height and width of the input feature maps, respectively, and Ci is the number of input feature maps, also known as the number of input channels. The convolutional layer has Ci×Co convolution kernels of size Kh×Kw, where Ci is the number of input channels, Co is the number of output feature maps (or output channels), and Kh and Kw are the height and width of the convolution kernels, respectively. The output feature map contains Ho×Wo×Co information, where Ho and Wo are the height and width of the output feature map, respectively, and Co is the number of output channels. Furthermore, the convolution operation involves the convolution stride (Sx, Sy), which affects the size of the output feature map.

[0067] In the embodiment of the present disclosure, the dimensions of the multidimensional data involved are characterized as (N, H, W, C) or (Co, H, W, Ci), which represents the storage order of the data in the memory. It can be understood that although the multidimensional data has multiple dimensions, because the layout of the memory is always one-dimensional, there is a corresponding relationship between the multidimensional data and the storage order on the memory. Multidimensional data is usually allocated in a continuous storage space, that is, the multidimensional data can be expanded in one dimension and stored in sequence on the memory. For example, in the embodiment of the present disclosure, sequential storage is performed in a low-dimensional (here Ci is the lowest dimension) priority manner. Adjacent dimensions refer to dimensions that are adjacent to each other in the dimensional information representation of multidimensional data. For example, W and Ci are adjacent. Adjacent dimensions can also be called continuous dimensions.

[0068] To improve memory access speed and fully utilize memory bandwidth, AI chip instruction sets typically require vectorized alignment. AI chip designs typically use the Ci dimension as the lowest dimension, which corresponds to the NHWC arrangement described above. Therefore, instruction alignment requires that the Ci dimension be aligned to a specified value, such as the instruction alignment value Aci, so that data is accessed in units of this instruction alignment value Aci. However, when the Ci dimension is small, this alignment restriction results in a large amount of redundant computation, leading to wasted resources.

[0069] In view of this, the disclosed embodiment provides a data processing solution for executing a neural network model based on the aforementioned hardware environment, which optimizes the Ci dimension in the convolutional layer to reduce resource waste while meeting the above-mentioned alignment constraints.

[0070] Figure 7 An exemplary schematic diagram of the data processing scheme of the disclosed embodiment is shown by a specific example. Assume that the alignment value required by the instruction alignment is Aci. Based on different instruction set designs, Aci can have different values, such as 32, 64, 128, etc. In the following example, Aci=64 is used as an example for description. According to the alignment requirement of the instruction, the Ci dimension of the filter needs to be aligned to Aci, that is, aligned to 64.

[0071] The left side of the figure shows the original filter of the convolution layer, which is represented as 64×4×4×4, that is, its output channel number Co is 64, the input channel number Ci is 4, and the size of each convolution kernel is 4×4. As can be seen from the figure, the Ci dimension of the original filter is much smaller than the alignment requirement of the instruction (64). According to the conventional processing method, the Ci dimension will be padded with zeros to align to 64. Aligning from 4 to 64 requires a lot of redundant calculations, resulting in a waste of resources.

[0072] The right side of the figure shows the folded filter according to the embodiment of the present disclosure, which is represented as 64×1×1×64, that is, its input channel number Ci' is 64, the output channel number is the same as the original filter, both are 64, and the size of each convolution kernel is 1×1. It can be seen that since the data in the width and height dimensions of the convolution kernel have been transferred to the input channel dimension, the input channel number of the folded filter has been aligned to the alignment requirement of the instruction (64), and no additional zero padding is required to align the Ci dimension, thus avoiding the redundant calculation caused by the zero padding.

[0073] In the folding process described above, the folded filter is the original filter obtained by dimensional folding. This dimensional folding process is based on the following considerations: if the original calculation method is used to achieve Ci alignment through zero padding, it will generate redundant calculations and waste computing resources. If the data of other dimensions is transferred to the Ci dimension and the Ci dimension is filled to the instruction alignment value, the waste of computing resources can be minimized and computing efficiency can be improved.

[0074] Figure 8 A more detailed filter folding diagram according to an embodiment of the present disclosure is shown. The figure shows the folding process of the original filter (64,4,4,4) in the above example. The Ci of the original filter is 4, and according to the instruction alignment requirement, it needs to be aligned to 64, so N is required. total = 64 ÷ 4 = 16-fold folding. In the example in the figure, the total folding factor of 16 is distributed into 4-fold folding of the convolution kernel width dimension and 4-fold folding of the convolution kernel height dimension.

[0075] As shown in the figure, preferably, folding along the width W dimension can be performed first. 801 in the figure shows that a single layer of W-dimensional data can be divided into four folds or segments, which are then sequentially arranged along the input channel Ci dimension. When the original filter 800 is completely folded along the W dimension, an intermediate filter shown in 810 is obtained, whose dimensions can be expressed as (64, 4, 1, 16).

[0076] Next, based on the intermediate filter 810, a highly H-dimensional folding can be performed. As shown in the figure, a 4-fold folding is also performed in the H dimension. The H-dimensional data of the intermediate filter 810 is divided into four folds or four segments, and then arranged in sequence on the input channel Ci dimension. In this folding process, each of the four segments on the H dimension corresponds to the data previously folded on the single-layer W dimension. After complete folding on the H dimension, the final folded filter shown in 820 can be obtained, and its dimensions can be expressed as (64, 1, 1, 64). The input channel Ci dimension of the final folded filter is 64, which meets the alignment requirements of the instruction.

[0077] The above describes a folding method in which the width W dimension is first folded, followed by the H dimension. Those skilled in the art will appreciate that the H dimension can also be folded first, followed by the W dimension. However, since the H dimension is not adjacent to the Ci dimension, additional processing, such as a data transposition operation, is required compared to folding the W dimension first.

[0078] As can be seen from the previous folding process, the dimension of the input channel Ci of the folded filter has changed. Therefore, the input feature map also needs to undergo the same type of folding. Because the input feature map and the filter are folded at the same magnification, the output of the convolution operation between the folded input feature map and the folded filter is the same as the output of the convolution operation before the folding, and no further processing is required.

[0079] Further, from Figure 8 As can be seen from the folding process, the number of input channels of the folded filter increases exponentially compared to the number of input channels of the original filter. Therefore, the folding scheme of the disclosed embodiments is particularly suitable for situations where the number of input channels Ci of the original filter is small, for example, Ci does not exceed the first threshold Th1, and the first threshold Th1 is less than the instruction alignment value Aci. In some embodiments, the first threshold Th1 ≤ Aci / 2. Practical results show that the smaller Ci, the greater the potential for improvement compared to existing algorithms.

[0080] from Figure 8 It can also be seen in the folding process that based on the various parameters of the original filter and the instruction alignment requirements, the various parameters of the folding filter can be determined.

[0081] First, the total folding factor N can be determined based on the number of input channels Ci of the original filter and the instruction alignment value Aci. total .

[0082] In some embodiments, the total folding factor N can be determined as follows total :

[0083] N total =Aci / Cia (1)

[0084] Where Cia is Ci aligned to the nearest Aci / 2 n The value of , n is a natural number.

[0085] As mentioned above, the embodiment of the present disclosure aligns Ci to the instruction alignment value Aci by folding in multiples. n When N is multiple, the two can be directly divided to get the total number of times to be folded. For example, in the above example, Ci=4, so N total =64 / 4=16 times. When Aci is not 2 times Ci n When times, you need to align Ci to the nearest Aci / 2 n For example, if Aci is 64, then Aci / 2 nIncluding 32, 16, 8, 4 and 2, Ci needs to be aligned to the nearest value among these numbers. For example, if Ci = 3, it needs to be aligned to 4; if Ci = 5, it needs to be aligned to 8, and so on. After alignment, remove Aci and you can get the total folding multiple N total .

[0086] Then, after determining the total folding multiple N total Afterwards, according to the previous considerations, it can be split into the convolution kernel width folding multiple Nw and the convolution kernel height folding multiple Nh. The multiple splitting can be performed according to various different rules to achieve different advantages.

[0087] In one embodiment, the folding multiple can be evenly divided into the width direction and the height direction of the convolution kernel. Figure 7 and Figure 8 In the example, the total folding multiple of 16 is evenly divided into 4 times the width and 4 times the height.

[0088] In another embodiment, the folding factor can be preferentially split along the convolution kernel width W. As mentioned above, in the data placement order NHWC, the W and Ci dimensions are continuous. Therefore, folding along the W dimension is relatively simple to implement, requiring only adjustment of the filter's dimensional representation, or dimensional reorganization, without any other processing. Therefore, preferentially splitting the folding factor along the convolution kernel width W simplifies processing.

[0089] The splitting of the total fold is also affected by the following two factors.

[0090] On the one hand, depending on the convolution kernel size of the original filter, if the values ​​of the various dimensions of the convolution kernel (Kw and Kh) are not sufficient to achieve the desired folding ratio, other measures need to be taken.

[0091] For example, in the above example, if the width W direction is prioritized, the total folding factor of 16 can be split into 16 times the width and 1 times the height. However, since the convolution kernel width Kw of the original filter is only 4, it supports a maximum of 4 folds. Therefore, the allocation of folding factors needs to be adjusted. For example, the splitting factor of the width W dimension can be reduced, for example, splitting it into 4 times the width and 4 times the height.

[0092] For example, if the convolution kernel width dimension and / or the convolution kernel height dimension cannot be 2 n Folding, for example, if Kw or Kh is an odd number and cannot be divided by 2, then folding by 2 is not possible; if it cannot be divided by 4, then folding by 4 is not possible. In this case, in order to fold the convolution kernel, it is necessary to fill it according to the folding multiple.

[0093] Figure 9The following diagram illustrates folding and padding a convolution kernel according to an embodiment of the present disclosure. The example in the figure assumes that the convolution kernel size of the filter to be folded is 3×3, but both the width W dimension and the height H dimension need to be folded four times. In this case, both the W and H dimensions need to be padded to multiples of 4.

[0094] Figure 9 The padding in the H and W directions is split and illustrated. The original single-layer convolution kernel 901 is 3 in the W direction. In order to perform four folds, it needs to be padded to 4, as shown in 911, where the white squares represent the padding blocks. Similarly, in the H direction, the original single-layer convolution kernel 901 is 3 in the H direction. In order to perform four folds, it needs to be padded to 4, as shown in 912, where the white squares represent the padding blocks. When padding is performed on both the H and W directions at the same time, after the final padding of the original convolution kernel 900, a convolution kernel 910 can be obtained, whose convolution kernel size becomes 4×4, and both the width and height can be folded four times.

[0095] from Figure 9 It can be seen that the padding operation on the convolution kernel will introduce invalid values, resulting in invalid redundant calculations.

[0096] Therefore, in some embodiments, in order to achieve folding, the total folding multiple can be split into a size that makes the padding amount caused by the alignment of the folding multiples as small as possible. For example, assuming that the original convolution kernel size is 1×6 and the total folding multiple is 16, if it is split into 4 times in the H direction and 4 times in the W direction, the original convolution kernel needs to be padded to 4×8, and 26 zero-padding values ​​will be added to each layer; if it is split into 2 times in the H direction and 8 times in the W direction, the original convolution kernel only needs to be padded to 2×8, and 10 zero-padding values ​​will be added to each layer; if it is split into 1 times in the H direction and 16 times in the W direction, the original convolution kernel needs to be padded to 1×16, and 10 zero-padding values ​​will also be added to each layer. In the case of the same padding amount, it is preferred to allocate the folding multiple to the W direction, for example, choosing to split it into 1 times in the H direction and 16 times in the W direction.

[0097] On the other hand, the actual effective multiple of the filter folding scheme of the disclosed embodiments is also related to the convolution step size. When the convolution step size is not divisible by the corresponding folding multiple, the folding method of the convolution kernel remains unchanged, but the folding of the input feature map will overlap, resulting in a certain amount of redundant calculation in the convolution operation.

[0098] Figure 10The schematic diagram of the effect of the convolution step size on the effective multiple according to the embodiment of the present disclosure is schematically shown. The figure takes the H direction as an example to illustrate the folding method of the input feature map containing overlapping areas. As shown in the figure, when the convolution step size Sy in the H direction is Sy=4, the H direction is folded four times with α and β as two consecutive blocks of data, and there is no overlapping area. However, when the convolution step size Sy in the H direction is Sy=2, the H direction is folded four times with γ and δ as two consecutive blocks of data, and there is an overlapping area. The size of the overlapping area is the folding multiple in that direction minus the convolution step size in that direction.

[0099] At this time, the folding multiples in the H and W directions can be flexibly configured. For example, the folding multiples are allocated to higher dimensions first to avoid redundant calculations in lower dimensions. For example, for the NHWC data arrangement order, the H dimension is higher than the W dimension. If Sx = 2, Sy = 2, N total =16, the W dimension can be allocated 20% and the H dimension 80%. In this way, the redundant calculations caused by the overlapping area are distributed in the H dimension. In this case, when performing these redundant calculations, due to the higher H dimension, more data needs to be read for each operation, which helps improve data access I / O efficiency.

[0100] In summary, in the embodiments of the present disclosure, the folding multiples of the H dimension and the W dimension can be flexibly configured based on the various rules described above to avoid redundant calculations as much as possible and maximize the improvement of computing efficiency.

[0101] In some embodiments, the total folding multiples may be split as follows: The folding multiples in the W direction may be determined first, for example, the total folding multiples N may be split in an average manner. total To determine the folding multiple Nw in the W direction:

[0102]

[0103] For example, let's take the original filter (64, 6, 6, 4), convolution step (Sx, Sy) = (2, 2), and instruction alignment value Aci = 64 as an example.

[0104] Then, the size in the W direction and the convolution step size can be determined accordingly.

[0105] As needed, the convolution kernel width k of the original filter can be first w The multiple of the folding factor Nw aligned to the W direction, denoted as k wa Then, calculate the size k of the folded convolution kernel W dimension w ′:

[0106] k w ′=k wa / Nw (3)

[0107] Continuing with the above example, since k w is 6, and Nw is 4, so a padding operation is required. At this time, k wa =8, then

[0108] Then, the convolution step size S in the W direction of the folded filter can be determined as follows x ′:

[0109]

[0110] Continuing with the above example, this time Therefore S′ xx = 1. This means that if four folds are performed in the W dimension, there will be overlapping areas, which will introduce redundant calculations. The actual effective folding ratio in the W dimension is S x / S x ′=2 / 1=2, not 4 times corresponding to a quarter fold.

[0111] At this point, the folding ratio can be adjusted. For example, the maximum effective folding ratio of 2 can be achieved in the W direction without overlapping areas, and the remaining required folding ratios can be allocated to the H direction. Therefore, the folding ratio in the H direction can be calculated based on the effective folding ratio in the W direction as follows:

[0112] Nh=N total / (S x / S x ′) (5)

[0113] In the above example, Nh=16 / 2=8, so the folding ratio in the H direction is 8 times.

[0114] At this point, the folding ratio in the W direction can be updated accordingly:

[0115] Nw=N total / Nh (6)

[0116] In the above example, Nw=16 / 8=2, which is equal to its maximum effective folding ratio. With this folding ratio allocation method, there will be no overlapping area in the W direction, but there will be more overlap in the H direction.

[0117] After determining the folding multiples Nw and Nh in the W and H directions, the remaining parameters in each dimension can be calculated accordingly, including dimension size, convolution step size, etc.

[0118] Similar to the calculation of the W dimension and the convolution step size described above, the convolution kernel width k of the original filter can be first reduced to wThe multiple of the folding factor Nw aligned to the W direction, denoted as k wa ; Increase the convolution kernel height k of the original filter h The multiple of the folding factor Nh aligned to the H direction is denoted as k ha Then, the size k of the folded convolution kernel W is calculated as follows w ′ and the size k of the H dimension h ′:

[0119] k w ′=k wa / Nw (7)

[0120] k h ′=k ha / Nh (8)

[0121] Continuing with the above example, since k w is 6, and Nw is 2, which is divisible by an integer, so no padding operation is required. At this time, k wa =6, then k′ w =6 / 2=3. For the H direction, due to k h is 6, and Nh is 8, which is not divisible, so a padding operation is required. At this time, k ha =8, then k h ′=8 / 8=1.

[0122] Then, the convolution step size S in the W direction of the folded filter can be determined as follows x The convolution step size S in the ′ and H directions y ′:

[0123]

[0124]

[0125] For the above example, S′ x =S x / Nw=2 / 2=1; due to S y / Nh=2 / 8<1, therefore, S′ y =1.

[0126] In summary, it is described how to design various parameters of the folding filter according to the instruction alignment requirements.

[0127] The above describes a scheme for generating a folded filter according to an embodiment of the present disclosure. In some embodiments, the folded filter can be generated offline. For example, in the process of reasoning using a neural network model, a pre-arranged, offline-generated folded filter can be used to perform a convolution operation with an input feature map that is similarly folded online to perform the reasoning process. In other embodiments, the folded filter can be generated online. For example, in the process of training a neural network model, the filter of the convolution layer can be folded online, and the training data can be similarly folded online, and then the two can perform a convolution operation to perform the training process.

[0128] Regardless of the process in which the folded filter of the disclosed embodiment is used, aligning the Ci dimensions through folding can greatly optimize the computational complexity of the convolution operation. The following compares the performance of the solution of the disclosed embodiment with that of existing convolution operations in terms of convolution computation complexity.

[0129] Let P represent the amount of convolution computation, A ci Indicates the value after Ci alignment, A CO Indicates the value after Co alignment,

[0130] Then P=n*A ci *A CO *H o *W o *k w *k h (11)

[0131] Hardware using the NHWC dimension order needs to be aligned to A because the Ci dimension is the lowest dimension and the vector instruction alignment requirement ci , so A ci It is aimed at the requirement of vector instruction alignment; artificial intelligence computing acceleration hardware usually has multiple parallel high-performance convolution computing units, so A CO It targets the requirement of dimensional alignment of the convolution kernel Co, and its value is the number of high-performance parallel computing units.

[0132] Before optimization, the computational complexity of the existing convolution operation is:

[0133] P before =n*A ci *A CO *H o *W o *k w *k h (12)

[0134] After optimization using the solution of the disclosed embodiment, the computational complexity of the convolution operation is:

[0135] P after=n*A ci *A CO *H o *W o *k′ w *k h '

[0136] =n*A ci *A CO *H o *W o *alignTo(k w ,Nw)*alignTo(k h ,Nh) / (Nw*Nh) (13)

[0137] The performance optimization rate after Ci folding is:

[0138]

[0139] Reference Figure 7-Figure 8 Taking the example described above as an example,

[0140] Before optimization, P before =1*64*64*ho*wo*4*4 (Sx,Sy=4,4)

[0141] After optimization, P after =1*64*64*ho*wo*1*1 (Sx',Sy'=1,1)

[0142] =1*64*64*ho*wo*4*4*(1 / 16)

[0143] The convolution computation amount of the reduced convolution computation unit is 93.75%.

[0144] I understand. Figure 7-Figure 8 The example above illustrates an ideal folding scenario, where the convolution stride fully meets the folding requirements and the kernel size and input feature map size do not need to be aligned by the folding factor. In practice, depending on the specific parameter values, the actual optimization rate may be lower than the peak optimization rate of 93.75%.

[0145] From the above comparison of computational complexity, it can be seen that the folded filter solution provided by the disclosed embodiment can effectively save the computational complexity of the convolution unit, thereby improving the computational performance of the convolution operation.

[0146] The disclosed embodiments also provide a data processing device for executing a neural network model, and a method for executing a neural network model implemented by the data processing device.

[0147] Figure 11The following is a schematic structural diagram of a data processing device that can implement the embodiment of the present disclosure. Figure 11 As shown, the data processing device 1100 includes a processing circuit 1110 and a storage circuit 1120 .

[0148] The processing circuit 1110 is responsible for processing various functions on the data processing device 1100, including but not limited to control, decoding, calculation, etc. The processing circuit 1110 may include, for example Figure 3 The control module 31 and / or the operation module 32 in.

[0149] In some embodiments, the processing circuit 1110 can be configured to perform dimensionality folding on the input feature map to obtain a folded feature map; and then perform a convolution operation on the folded feature map using the folded filter of the disclosed embodiment to obtain an output feature map. The dimensionality folding of the input feature map is the same as that of the folded filter, so that the output feature map is the same as the result obtained by the original operation, without requiring any additional processing.

[0150] The storage circuit 1120 can be used to store or carry relevant data, which can be, for example, Figure 3 or Figure 5 The various RAMs shown are also referred to as on-chip caches. In some embodiments, the storage circuit 1120 can be configured to store the folded filters of the convolutional layer of the neural network model. The folded filters are obtained by dimensional folding the original filters according to the embodiments of the present disclosure.

[0151] Dimension folding processing may include rearranging data in the width dimension and / or the height dimension to the input channel dimension. For example, for data rearrangement or dimension folding in the width dimension, the processing circuit 1110 may implement it by dimension reorganization. For data rearrangement or dimension folding in the height dimension, the processing circuit 1110 may implement it by dimension transposition.

[0152] In some embodiments, the data processing device 1100 can be configured to perform a training process for a neural network model. In this case, the processing circuit 1110 can be configured to perform the folding process of the disclosed embodiment on the filters of the convolutional layer of the neural network model and the training data online during training. The resulting folded filter is then used to perform a convolution operation on the folded training data to perform the training process. The specific folding process performed by the processing circuit 1110 can be referred to the description above and will not be repeated here.

[0153] In other embodiments, the data processing device 1100 may be configured to perform an inference process on a neural network model. In this case, the processing circuit 1110 may be configured to first perform a dimensionality folding process on the input neurons, and then directly use the folding filter stored in the storage circuit 1120 to perform a convolution operation on the folded input neurons to perform the inference process. The dimensionality folding method of the input neurons is consistent with the dimensionality folding method of the stored folding filter.

[0154] Figure 12 An exemplary flow chart of a data processing method according to an embodiment of the present disclosure is shown.

[0155] As shown, the data processing method 1200 includes step 1210, where the processing circuit performs dimension folding on the input feature map to obtain a folded feature map. Next, at step 1220, the processing circuit performs a convolution operation on the folded feature map using the folded filter of the convolution layer of the neural network model stored in the storage circuit to obtain an output feature map.

[0156] The folded filter stored in the storage circuit is the original filter obtained by dimensional folding. In the disclosed embodiment, dimensional folding can include rearranging data in the width dimension and / or height dimension to the input channel dimension. In the above process, the dimensional folding method of the input feature map is consistent with the dimensional folding method of the stored folded filter.

[0157] Those skilled in the art will appreciate that the filter folding process of the embodiment of the present disclosure described above in conjunction with the accompanying drawings can also be applied to Figure 11 data processing device and Figure 12 Therefore, the data processing method will not be described again.

[0158] The present disclosure also provides a chip, which may include the data processing device of any embodiment described above in conjunction with the accompanying drawings. Furthermore, the present disclosure also provides a board, which may include the aforementioned chip.

[0159] According to different application scenarios, the electronic devices or devices disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, driving recorders, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, automatic driving terminals, vehicles, household appliances, and / or medical equipment. The vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs. The electronic devices or devices disclosed herein may also be applied to the Internet, Internet of Things, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, medical care and other fields. Furthermore, the electronic devices or devices disclosed herein may also be used in cloud, edge, terminal and other application scenarios related to artificial intelligence, big data and / or cloud computing. In one or more embodiments, electronic devices or apparatuses with high computing power according to the disclosed solution can be applied to cloud devices (such as cloud servers), while electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or edge device are compatible with each other, so that according to the hardware information of the terminal device and / or edge device, appropriate hardware resources can be matched from the hardware resources of the cloud device to simulate the hardware resources of the terminal device and / or edge device, so as to complete the unified management, scheduling and collaborative work of end-to-end or cloud-edge-to-end.

[0160] It should be noted that, for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art will understand that the scheme of the present disclosure is not limited by the order of the actions described. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art will understand that some of the steps therein can be performed in other orders or simultaneously. Further, those skilled in the art will understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved therein are not necessarily necessary for the implementation of one or more schemes of the present disclosure. In addition, depending on the different schemes, the description of some embodiments of the present disclosure also has different emphases. In view of this, those skilled in the art will understand that the parts that are not described in detail in a certain embodiment of the present disclosure may also refer to the relevant descriptions of other embodiments.

[0161] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented in other ways not disclosed herein. For example, with respect to the various units in the electronic device or device embodiments described above, this document divides them based on the consideration of logical functions, and there may be other ways of division in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. With respect to the connection relationship between different units or components, the connection discussed above in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.

[0162] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network elements. In addition, according to actual needs, some or all of the units may be selected to achieve the purpose of the solution described in the embodiments of this disclosure. In addition, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically separately.

[0163] In some implementation scenarios, the above-mentioned integrated unit can be implemented in the form of a software program module. If implemented in the form of a software program module and sold or used as an independent product, the integrated unit can be stored in a computer-readable memory. Based on this, when the scheme of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in a memory, which may include several instructions to enable a computer device (such as a personal computer, a server or a network device, etc.) to perform some or all of the steps of the method described in the embodiment of the present disclosure. The aforementioned memory may include, but is not limited to, various media that can store program code, such as a USB flash drive, a flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0164] In some other implementation scenarios, the above-mentioned integrated unit can also be implemented in the form of hardware, that is, a specific hardware circuit, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include but is not limited to physical devices, and the physical devices may include but are not limited to devices such as transistors or memristors. In view of this, the various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM and RAM, etc.

[0165] Although a plurality of embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may conceive of many modifications, changes, and alternatives without departing from the ideas and spirit of the present disclosure. It should be understood that in practicing the present disclosure, various alternatives to the embodiments of the present disclosure described herein may be adopted. The appended claims are intended to define the scope of protection of the present disclosure and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A data processing device for executing a neural network model, comprising: a storage circuit configured to store a folded filter of a convolutional layer of the neural network model, wherein the folded filter is obtained by dimensional folding of the original filter, wherein the dimensional folding includes rearranging data of a width dimension and / or a height dimension to an input channel dimension; as well as Processing circuitry configured to: Performing dimension folding on the input feature map to obtain a folded feature map; as well as performing a convolution operation on the folded feature map using the folded filter to obtain an output feature map; The input channel dimension of the original filter does not exceed a first threshold value A1, and the input channel dimension of the folded filter is equal to a second threshold value Aci, wherein the first threshold value A1 is less than the second threshold value Aci, and the processing circuit is configured to perform the dimension folding as follows: Determine the total folding multiple N based on the input channel dimension Ci of the multidimensional data to be folded and the second threshold Aci total ; Split the total folding multiple Ntotal into a width dimension folding multiple Nw and a height dimension folding multiple Nh; Determine the width and height dimensions of the folded multidimensional data based on Nw, Nh, and the width and height dimensions of the multidimensional data to be folded; as well as Based on Nw, Nh and the original convolution step size of the convolution operation, the folded convolution step size of the convolution operation is determined.

2. The data processing apparatus according to claim 1, wherein the processing circuit is further configured to determine the total folding factor N as follows: total : N total =Aci / Cia, where Cia is Ci aligned to the nearest Aci / 2 n The value of , n is a natural number.

3. The data processing device according to any one of claims 1 to 2, wherein the processing circuit is further configured to split the total folding multiple N according to any one of the following rules or a combination of rules: total : Split into width dimension first; Split evenly into width and height dimensions; Split into sections that minimize the amount of padding caused by fold alignment; or The convolution step size in the width dimension is divisible by the folding factor of the width dimension.

4. The data processing apparatus according to claim 1 , wherein the processing circuit is further configured to determine the size of the width dimension and the size of the height dimension of the folded multidimensional data as follows: k w ′=k wa / Nw (7) k h ′=k ha / Nh (8) in, k w ′, k h ′ are the width dimension and height dimension of the folded multidimensional data, respectively, k wa is the width dimension k of the multidimensional data to be folded w Align to the nearest width dimension fold multiple Nw value, k ha is the height dimension k of the multidimensional data to be folded h Align to the nearest height dimension with the fold multiplier Nh.

5. The data processing apparatus according to claim 1 , wherein the processing circuit is further configured to determine a convolution step size of the folded multidimensional data as follows: in, s x 、S y are the convolution step size of the original width dimension and the convolution step size of the height dimension of the convolution operation, S x ′、S y ′ are the convolution stride of the width dimension and the convolution stride of the height dimension after folding of the convolution operation respectively. 6 . The data processing apparatus according to claim 1 , wherein the second threshold value Aci is determined based on an instruction alignment requirement, and the first threshold value A1 ≤ Aci / 2.

7. The data processing apparatus according to claim 1 , wherein the processing circuit is further configured to: said dimensional folding in the width dimension is achieved by dimensional reorganization; and / or The dimensional folding in the height dimension is achieved by dimensional transposition. 8 . The data processing apparatus according to claim 1 , wherein the output channel dimension size of the original filter is equal to the output channel dimension size of the folded filter. 9 . The data processing apparatus according to claim 1 , wherein the folding filter is generated offline or online.

10. A chip, characterized in that: The chip includes the data processing device according to any one of claims 1-9.

11. A board, characterized in that: The board includes the chip according to claim 10.

12. A method for executing a neural network model, implemented by a data processing device, the data processing device comprising storage circuitry and processing circuitry, the method comprising: The processing circuit performs dimension folding on the input feature map to obtain a folded feature map; The processing circuit performs a convolution operation on the folded feature map using the folded filter of the convolution layer of the neural network model stored in the storage circuit to obtain an output feature map; The folded filter is obtained by folding the original filter in the dimension, and the dimension folding includes rearranging data of the width dimension and / or the height dimension to the input channel dimension; The input channel dimension of the original filter does not exceed a first threshold value A1, the input channel dimension of the folded filter is equal to a second threshold value Aci, wherein the first threshold value A1 is less than the second threshold value Aci, and further, the processing circuit performs the dimension folding as follows: Determine the total folding multiple N based on the input channel dimension Ci of the multidimensional data to be folded and the second threshold Aci total ; Split the total folding multiple Ntotal into a width dimension folding multiple Nw and a height dimension folding multiple Nh; Determine the width and height dimensions of the folded multidimensional data based on Nw, Nh, and the width and height dimensions of the multidimensional data to be folded; as well as Based on Nw, Nh and the original convolution step size of the convolution operation, the folded convolution step size of the convolution operation is determined.

13. The method according to claim 12, further comprising: The processing circuit further determines the total folding factor N as follows total : N total =Aci / Cia, where Cia is Ci aligned to the nearest Aci / 2 n The value of , n is a natural number.

14. The method according to any one of claims 12-13, further comprising: The processing circuit splits the total folding times N according to any one of the following rules or a combination of rules: total : Split into width dimension first; Split evenly into width and height dimensions; Split into sections that minimize the amount of padding caused by fold alignment; or The convolution step size in the width dimension is divisible by the folding factor of the width dimension.

15. The method according to claim 12, further comprising: The processing circuit determines the width dimension size and the height dimension size of the folded multidimensional data as follows: k w ′=k wa / Nw (7) k h ′=k ha / Nh (8) Among them, k w ′, k h ′ are the width dimension and height dimension of the folded multidimensional data, respectively, k wa is the width dimension k of the multidimensional data to be folded w Align to the nearest width dimension fold multiple Nw value, k ha is the height dimension k of the multidimensional data to be folded h Align to the nearest height dimension with the fold multiplier Nh.

16. The method according to claim 12, further comprising: The processing circuit determines the convolution step size of the folded multidimensional data as follows: Among them, S x 、S y are the convolution step size of the original width dimension and the convolution step size of the height dimension of the convolution operation, S x ′、S y ′ are the convolution stride of the width dimension and the convolution stride of the height dimension after folding of the convolution operation respectively. The method according to claim 12 , wherein the second threshold value Aci is determined based on an instruction alignment requirement, and the first threshold value A1 ≤ Aci / 2.

18. The method of claim 12, further comprising: The processing circuit implements the dimensional folding in the width dimension by dimensional reorganization; and / or The processing circuit implements the dimension folding in the height dimension by dimension transposition.

19. The method of claim 12, wherein the output channel dimension size of the original filter is equal to the output channel dimension size of the folded filter.

20. The method of claim 12, wherein the folding filter is generated offline or online.

Citation Information

Patent Citations

  • Method and device for performing operation of convolutional layer in convolutional neural network

    CN107844827A