For data processing devices, data processing methods and related products

By processing the input data in blocks and adapting to the processing capabilities of the hardware computing device, the problem of limited computing performance of convolutional neural networks is solved, and more efficient computing efficiency and speed are achieved.

CN113850380BActive Publication Date: 2025-09-23ANHUI CAMBRICON INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111129610.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2025-09-23
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

The computing performance of existing convolutional neural networks is limited by convolution operations of different scales and types, making it difficult to fully utilize the hardware advantages of deep learning processors, resulting in low computing efficiency.

Method used

Through the data processing device and method, block instructions are used to split and store input data, adapt convolution operation hardware, and thus improve computing efficiency.

Benefits of technology

The computational efficiency of convolution operations is improved, the parallel processing capabilities of hardware are fully utilized, power consumption is reduced and the computing speed is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113850380B_ABST
    Figure CN113850380B_ABST
Patent Text Reader

Abstract

This disclosure discloses a data processing device, a data processing method using the data processing device to execute block instructions, and related products. The data processing device can be included as a computing device in a combined processing device, and the combined processing device can also include an interface device and other processing devices. The computing device interacts with the other processing devices to jointly complete the computing operations specified by the user. The combined processing device can also include a storage device, which is connected to the computing device and the other processing devices respectively and is used to store data from the computing device and the other processing devices. The solution disclosed in this disclosure realizes the data splitting and storage in small convolution operations, thereby improving the operation processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of data processing, and more specifically, to a data processing device, a data processing method, a chip, and a board for executing block instructions on data using the data processing device. Background Art

[0002] Deep learning has become a key branch of machine learning and is significantly driving the development of artificial intelligence (AI). Deep neural networks (DNNs), the core technology of deep learning, have been widely applied across numerous industries.

[0003] Neural networks are one of the most critical technologies in artificial intelligence and deep learning, with convolutional neural networks (CNNs) being the most important type of network. The most critical computation in a CNN is the convolution operation in the convolution layer. The convolution layer extracts features from the input data. Through multiple layers of convolution, complex features are extracted to ensure the network has sufficient expressive power and generalization capabilities. Neural network models incorporate a large number of convolution operations of various types, and their computational performance significantly impacts the overall performance of the neural network model. When neural network models are applied to different fields, such as speech recognition, machine translation, and image processing, the dimensions of the corresponding input feature maps and weights may vary. To fully leverage the hardware advantages of deep learning processors, optimization is required for convolution operations of varying sizes and types to improve the computational performance of neural network models. Summary of the Invention

[0004] To at least address one or more of the above-mentioned technical issues, the present disclosure provides, in various aspects, a data processing device that, by executing block instructions on data, can adapt data of various dimensions to the hardware of the convolution operation, thereby improving the computational efficiency of the convolution operation. The convolution operation of the embodiments of the present disclosure can be an operation in various neural network models, which can be applied to various fields such as image processing, speech processing, text processing, etc. These processes can include, but are not limited to, recognition and classification.

[0005] In a first aspect, an embodiment of the present disclosure provides a data processing device, comprising a control circuit, a first storage circuit, and a second storage circuit, wherein: the first storage circuit is used to store data before processing; the second storage circuit is used to store data after processing; and the control circuit is used to configure and execute a block instruction to split the input data stored in the first storage circuit in a first-dimensional storage order into split units and store them as output data on the second storage circuit, wherein on the second storage circuit, the data is stored in each split unit in a second-dimensional storage order, and the data is stored between the split units in a third-dimensional storage order.

[0006] In a second aspect, an embodiment of the present disclosure provides a chip comprising the data processing device of the first aspect.

[0007] In a third aspect, an embodiment of the present disclosure provides a board comprising the chip of the aforementioned second aspect.

[0008] In a fourth aspect, an embodiment of the present disclosure provides a data processing method for executing block instructions on input data using the data processing device of the first aspect.

[0009] Through the data processing device, chip, board and data processing method for executing block instructions by the data processing device as provided above, the scheme of the disclosed embodiment performs block processing on the data in various convolution splitting schemes to adapt to the processing capabilities of the hardware computing device, thereby making full use of the parallel processing capabilities of multiple slave processing circuits, and effectively improving the computational efficiency of the convolution operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0011] Figure 1 A structural diagram of a board according to an embodiment of the present disclosure is shown;

[0012] Figure 2 A structural diagram showing a combined processing device according to an embodiment of the present disclosure;

[0013] Figure 3a A schematic diagram showing the internal structure of a processor core of a single-core computing device according to an embodiment of the present disclosure;

[0014] Figure 3b A simplified schematic diagram showing the internal structure of a multi-core computing device according to an embodiment of the present disclosure;

[0015] Figure 4An example of an exemplary convolution operation principle to which the present disclosure embodiments can be applied is shown;

[0016] Figure 5 shows a schematic structural block diagram of a computing device according to an embodiment of the present disclosure;

[0017] Figure 6 An exemplary data storage order according to an embodiment of the present disclosure is shown;

[0018] Figures 7a-7c Several exemplary grouping modes according to embodiments of the present disclosure are shown;

[0019] Figure 8 shows an exemplary splitting diagram of an input feature map according to an embodiment of the present disclosure;

[0020] Figure 9 A schematic diagram illustrating splitting and storage of the Forward4 solution according to an embodiment of the present disclosure is shown;

[0021] Figure 10 A schematic diagram showing the division of output points of the operation circuit in the Forward4 solution according to an embodiment of the present disclosure is shown;

[0022] Figure 11 A schematic diagram illustrating a single operation in the Forward4 solution according to an embodiment of the present disclosure is shown;

[0023] Figure 12 A schematic diagram of sliding convolution in the Forward4 solution according to an embodiment of the present disclosure is shown;

[0024] Figure 13 A schematic diagram showing the output data format of the Forward4 solution according to an embodiment of the present disclosure;

[0025] Figure 14 Shows the overall data handling process according to the embodiment of the present disclosure;

[0026] Figure 15 A schematic conceptual diagram illustrating Trans Tiling according to an embodiment of the present disclosure;

[0027] Figure 16 A schematic diagram showing the front and rear meter configuration; and

[0028] Figure 17 A schematic diagram illustrating executing a blocking instruction on neuron data according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0029] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.

[0030] It should be understood that the terms "first," "second," "third," and "fourth," etc., which may appear in the claims, specification, and drawings of this disclosure, are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0031] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0032] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context.

[0033] Figure 1 FIG. 1 is a schematic diagram showing the structure of a board 10 according to an embodiment of the present disclosure. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC), or system-on-chip, integrated with one or more combined processing devices. The combined processing device is an artificial intelligence computing unit that supports various deep learning and machine learning algorithms to meet the intelligent processing needs in complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in the field of cloud intelligence. A notable feature of cloud-based intelligent applications is the large amount of input data, which places high demands on the platform's storage and computing capabilities. Board 10 of this embodiment is suitable for cloud-based intelligent applications and has extensive off-chip storage, on-chip storage, and powerful computing capabilities.

[0034] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. Calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface formats, such as a PCIe interface.

[0035] Board 10 also includes a memory device 104 for storing data, which includes one or more storage units 105. Memory device 104 is connected to a control device 106 and chip 101 via a bus for data transmission. Control device 106 in board 10 is configured to control the state of chip 101. To this end, in one application scenario, control device 106 may include a microcontroller (MCU).

[0036] Figure 2 FIG. 1 is a block diagram showing the combined processing device in the chip 101 of this embodiment. Figure 2 As shown in , the combined processing device 20 includes a computing device 201 , an interface device 202 , a processing device 203 and a storage device 204 .

[0037] The computing device 201 is configured to perform user-specified operations and is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.

[0038] Interface device 202 is used to transmit data and control instructions between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to a storage device on-chip of computing device 201. Furthermore, computing device 201 can obtain control instructions from processing device 203 via interface device 202 and write them to a control cache on-chip of computing device 201. Alternatively or optionally, interface device 202 can also read data from the storage device of computing device 201 and transmit it to processing device 203.

[0039] The processing device 203 is a general processing device that performs basic controls including but not limited to data handling, starting and / or stopping the computing device 201. Depending on the implementation, the processing device 203 can be a central processing unit (CPU), a graphics processing unit (GPU) or one or more types of processors in other general and / or special processors, these processors include but are not limited to digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, only with respect to the computing device 201 disclosed in the present invention, it can be regarded as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 201 and the processing device 203 are integrated and considered together, the two are regarded as forming a heterogeneous multi-core structure.

[0040] The storage device 204 is used to store data to be processed, which may be DRAM or DDR memory, and is typically 16G or larger in size, for storing data of the computing device 201 and / or the processing device 203 .

[0041] Figure 3a The figure shows the internal structure of the processing core when the computing device 201 is a single-core device. The computing device 301 is used to process input data such as computer vision, speech, natural language, and data mining. The computing device 301 includes three modules: a control module 31, a calculation module 32, and a storage module 33.

[0042] The control module 31 coordinates and controls the operations of the computing module 32 and the storage module 33 to complete deep learning tasks. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 retrieves instructions from the processing device 203, while the instruction decode unit 312 decodes the retrieved instructions and sends the decoded results as control information to the computing module 32 and the storage module 33.

[0043] The operation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations, supporting complex operations such as vector multiplication, addition, and nonlinear transformations. The matrix operation unit 322 is responsible for the core calculations of the deep learning algorithm, namely matrix multiplication and convolution.

[0044] The storage module 33 is used to store or transfer relevant data and includes a neuron RAM (NRAM) 331, a weight RAM (WRAM) 332, and a direct memory access (DMA) module 333. NRAM 331 stores input neurons, output neurons, and intermediate computational results; WRAM 332 stores the convolution kernels (i.e., weights) of the deep learning network; and DMA 333, which connects to DRAM 204 via bus 34, transfers data between the computing device 301 and DRAM 204.

[0045] Figure 3b The figure shows a simplified schematic diagram of the internal structure of a multi-core computing device 201. A multi-core computing device can be abstracted using a hierarchical hardware model. As shown in the figure, the multi-core computing device can be abstracted into four levels, namely, the card level (Card) 350, the chip level (Chip) 360, the processor cluster level (Cluster) 370, and the processor core level (Core) 380. The embodiments disclosed herein mainly involve the data transmission of the storage unit and the computing unit part. Therefore, the drawings and descriptions briefly illustrate and introduce the relevant computing structure, omitting other parts.

[0046] At the board level, each board contains local DDR storage, and each processor chip serves as a computing and control unit.

[0047] At the chip level, each processor chip contains multiple multiprocessors as computing units.

[0048] At the computing cluster level, each multiprocessor includes multiple accelerator cores as control and computing units, and a shared memory SRAM as a storage unit.

[0049] At the processor core level, each accelerator core contains local storage and an array of local processing units. NFU stands for Neuron Function Unit, which is used for convolutional computations.

[0050] In this multi-core computing device, the storage model includes the global memory of the board, SRAM (shared memory) on the Cluster, NRAM, WRAM and registers on the Core, etc. In order to obtain better performance, the data movement between the storage levels below the Card and the balance between memory access / computation can be explicitly controlled. SRAM is contained in the memory processing unit MPU (Memory Process Unit Core, abbreviated as MPU, or Mem Core). Core refers to the intelligent processing core (Intelligent Process Unit Core, abbreviated as IPU Core or Core) in the multi-core computing device. 1 IPU Core contains NRAM, WRAM, NFU, etc. Cluster refers to a processor cluster or a computing cluster. Usually a multi-core computing device contains several Clusters, and a Cluster contains 1 Mem Core + N IPU Cores.

[0051] Example convolution operation types

[0052] The convolution layer in a neural network model can perform convolution operations, extracting features by applying convolution kernels (also known as filters, weights, etc.) to the input feature map (also known as input data, neurons, or input neurons). A convolution layer can contain multiple convolution kernels, with each element of the convolution kernel corresponding to a weight coefficient and a bias. The disclosed embodiments can be applied to data splitting for various convolution operations.

[0053] In conventional 3D convolution operations, assuming that the input feature map tensor shape in the convolution layer is expressed as X[N HiWi Ci], the convolution kernel tensor shape is expressed as K[Co Kh Kw Ci], and the output result is Y[N HoWo Co], then the simplified mathematical calculation formula of the convolution operation can be expressed as follows:

[0054]

[0055] In the above formula, X is the input data, Y is the output data, K is the convolution kernel, Kh and Kw are the length and width of K, and sh and sw are the strides in the length and width directions. The formula ignores bias, padding, and dilation, and assumes that the input data X is padded and the convolution kernel is dilated. The formula ignores the N and C dimensions. The forward calculation of the neural network model is independent in the N dimension and fully connected in the C dimension. When operating, the convolution kernel sweeps over the input features at a certain stride, performing matrix multiplication on the input features within the convolution window and adding the bias.

[0056] Figure 4 An example of an exemplary conventional 3D convolution operation principle to which the disclosed embodiments can be applied is shown.

[0057] The figure exemplifies four-dimensional input data X of size [N HiWiCi], which can be represented as N Hi×Wi×Ci 3D rectangles 410. The figure also exemplifies a four-dimensional convolution kernel K of size [Co Kh Kw Ci], which can be represented as Co Kh×Kw×Ci 3D convolution kernels 420. The convolution of the input data X with the convolution kernel K results in output data Y, which is four-dimensional data of size [N Ho Wo Co] and can be represented as N Ho×Wo×Co 3D rectangles 430.

[0058] The figure also specifically shows an example of a convolution operation, where the input data is a 6×6×3 input feature map 440, omitting the N dimension; the convolution kernel is a 3×3×3 stereo convolution kernel 450, targeting a single Co; and the output data is a 4×4 output feature map 460. The specific operation process is as follows:

[0059] The convolution kernel 450 scans the input feature map 440 at a certain step size, performs matrix element multiplication and summation on the input features within the convolution window 470, and superimposes the bias. That is, the value at each position in the output feature map 460 is obtained by performing a two-dimensional convolution operation on the corresponding block of each input feature map and the corresponding convolution kernel, and then adding them together. For example, the figure shows that the value of the (0,0) position on the output feature map 460 (that is, the convolution output point) is obtained by performing a two-dimensional convolution operation on the convolution window 470 framed by the black cube in the input feature map and the three-dimensional convolution kernel 450 to obtain three values, and then adding them together to obtain the final value.

[0060] To obtain outputs at other locations, the convolution kernel 450 can be moved on the input feature map 440, that is, the convolution window at the convolution output point can be moved. In the example shown in the figure, the convolution stride (Sx, Sy) is (1,1). When the convolution operation is performed after shifting one grid rightward (widthwise) or downward (heightwise), the value at position (0,1) or (1,0) on the output feature map 460 can be obtained, respectively.

[0061] As can be seen from the above description, in a convolutional layer of a neural network, there are N sets of input feature maps, each containing Hi × Wi × Ci information, where Hi and Wi are the height and width of the input feature map, respectively, and Ci is the number of input feature maps, also known as the number of input channels. The convolutional layer has Ci × Co convolution kernels of size Kh × Kw, where Ci is the number of input channels, Co is the number of output feature maps (or output channels), and Kh and Kw are the height and width of the convolution kernel, respectively. The output feature map contains Ho × Wo × Co information, where Ho and Wo are the height and width of the output feature map, respectively, and Co is the number of output channels. Furthermore, the convolution operation involves the convolution stride (Sx, Sy), which affects the size of the output feature map.

[0062] In this article, input feature map, input data, neurons or input neurons can be used interchangeably; convolution kernel, filter or weight can be used interchangeably. In addition, H (height) and Y dimension can be used interchangeably, and W (width) and X dimension can be used interchangeably. Accordingly, the H dimension of the input feature map can be expressed as Hi or Yi, the H dimension of the output feature map can be expressed as Ho or Yo, and the W dimension is expressed similarly. In the embodiment of the present disclosure, each convolution output point has a corresponding convolution window, and the shape of the convolution window is equal to the shape of the convolution kernel. The value of each convolution output point corresponds to the result of the bitwise multiplication and accumulation of the input feature map and the weight within its convolution window.

[0063] Exemplary computing devices / data processing devices

[0064] In the disclosed embodiments, a master-slave computing device can be used to implement the convolution operation. Furthermore, different data paths can be configured for the input feature map and the convolution kernel to improve memory access efficiency.

[0065] Figure 5 FIG3 shows a schematic structural block diagram of a computing device 500 according to an embodiment of the present disclosure. It can be understood that this structure can be regarded as a refinement of the internal structure of the computing module of a single processing core in FIG3 , or as a functional division block diagram based on the computing modules of multiple processing cores shown in FIG3 . Figure 5 As shown, the computing device 500 of the embodiment of the present disclosure can be configured to perform various types of convolution operations, and may include a master processing circuit (MA) 510 and multiple slave processing circuits (SL) 520. The figure shows 16 slave processing circuits SL0 to SL15. Those skilled in the art will appreciate that the number of slave processing circuits may be greater or less, depending on the specific hardware configuration, and the embodiment of the present disclosure is not limited in this respect.

[0066] The master processing circuit and the slave processing circuit, as well as multiple slave processing circuits, can communicate with each other via various connections. In different application scenarios, the connection between multiple slave processing circuits can be hard-wired or logically configured according to, for example, microinstructions, to form a topology of multiple slave processing circuit arrays. The disclosed embodiments are not limited in this respect. The master processing circuit and the slave processing circuits can cooperate with each other to achieve parallel computing processing.

[0067] To support computational functions, the master and slave processing circuits may include various computational circuits, such as vector operation units (VAUs) and matrix operation units (MAUs). The VAUs perform vector operations, supporting complex operations such as vector multiplication, addition, and nonlinear transformations. The MAUs are responsible for the core computations of deep learning algorithms, such as matrix multiplication and convolution.

[0068] The slave processing circuit may be configured to, for example, perform intermediate operations on corresponding data in parallel according to the operation instruction to obtain multiple intermediate results, and transmit the multiple intermediate results back to the master processing circuit.

[0069] By setting the computing device 500 into a master-slave structure (for example, a one-master-multiple-slave structure, or a multi-master-multiple-slave structure, and the present disclosure has no limitations in this regard), for the calculation instructions of the forward operation, the data can be split according to the calculation instructions, so that the parts with larger calculation amount can be calculated in parallel through multiple slave processing circuits to improve the calculation speed, save calculation time, and thus reduce power consumption.

[0070] In some embodiments of the present disclosure, by utilizing different data paths to transmit input feature maps and weights, multiple multiplexing methods of input feature maps and weights can be supported, thereby reducing the amount of data memory access during calculations and improving processing efficiency.

[0071] Specifically, the computing device 500 may further include a first storage circuit 530 and a second storage circuit 540 for respectively storing data transmitted via different data channels.

[0072] The first storage circuit 530 can be used to store multicast data. This means that the data in the first storage circuit will be transmitted to multiple slave processing circuits via a broadcast bus, and these slave processing circuits will receive the same data. It will be understood that both broadcast and multicast can be implemented via a broadcast bus. Multicast refers to a communication method that transmits one copy of data to multiple slave processing circuits; broadcast, on the other hand, is a communication method that transmits one copy of data to all slave processing circuits and is a special case of multicast. Because both multicast and broadcast correspond to one-to-many transmission methods, this document does not specifically distinguish between the two. Broadcast and multicast may be collectively referred to as multicast, and those skilled in the art will be able to clarify their meaning based on the context.

[0073] The second storage circuit 540 may be used to store distribution data, that is, the data in the second storage circuit will be transmitted to different slave processing circuits respectively, and each slave processing circuit receives different data.

[0074] By providing the first storage circuit and the second storage circuit separately, it is possible to support the transmission of data to be operated in different transmission modes, thereby reducing the amount of data memory access by multiplexing multicast data among multiple slave processing circuits.

[0075] In some embodiments, the master processing circuit may determine one of the input feature map and the convolution kernel as multicast data and store it in a first storage circuit, so that the data can be broadcast to multiple scheduled slave processing circuits during the operation. Correspondingly, the master processing circuit may determine the other of the input feature map and the convolution kernel as distributed data and store it in a second storage circuit. This distributed data can be distributed to the corresponding slave processing circuits before the operation.

[0076] Figure 5 A schematic diagram of the internal structure of a slave processing circuit SL according to an embodiment of the present disclosure is also shown. As shown, each slave processing circuit 520 may include multiple computing circuits CU 521, a first buffer circuit 522, and a second buffer circuit 523. Four computing circuits CU0-CU3 are shown. Those skilled in the art will appreciate that the number of computing circuits may be greater or less, depending on the specific hardware configuration, and the present disclosure is not limited in this respect.

[0077] In some embodiments, the first buffer circuit 522 can be used to cache the weights or input feature maps assigned to the slave processing circuit. Correspondingly, the second buffer circuit 523 can be used to cache the input feature maps or weights assigned to the slave processing circuit. Both buffer circuits are used to select data to participate in the operation. The data of the first buffer circuit 522 can be, for example, multiple data rows from the first storage circuit 530 or the second storage circuit 540, and correspondingly, the data of the second buffer circuit 523 can be, for example, multiple data rows from the second storage circuit 540 or the first storage circuit 530. Depending on the specific multiplexing method, these data rows can be distributed to the corresponding operation circuit CU 521 or broadcast to all CUs 521 in the slave processing circuit 520 during the operation.

[0078] Each operation circuit CU 521 is configured to perform a bitwise multiplication-accumulation operation on a data row selected from the first buffer circuit and a data row selected from the second buffer circuit during each calculation.

[0079] By providing a first buffer circuit and a second buffer circuit respectively, it is possible to support transmission of data to be operated in different transmission modes, thereby reducing the amount of data memory access by multiplexing data as much as possible between multiple operation circuits within a single slave processing circuit.

[0080] The slave processing circuit 520 may further include a third buffer circuit 524 for buffering the operation results of each operation circuit CU 521 .

[0081] Understandably, although Figure 5 The various processing circuits and storage circuits are shown as separate modules, but depending on the configuration, the storage circuits and processing circuits can also be combined into one module. For example, the first storage circuit 530 can be combined with the master processing circuit 510, and the second storage circuit 540 can be shared by multiple slave processing circuits 520, and an independent storage area is allocated to each slave processing circuit to speed up access. The embodiments of the present disclosure are not limited in this regard. In addition, in the computing device, the master processing circuit and the slave processing circuit can belong to different modules of the same processor or chip, or to different processors, and the present disclosure is not limited in this regard.

[0082] Exemplary Data Splitting and Storage

[0083] In the embodiment of the present disclosure, the dimensions of the multidimensional data involved are characterized as (N, H, W, C) or (Co, H, W, Ci), which represents the storage order of the data in the memory. It can be understood that although the multidimensional data has multiple dimensions, because the layout of the memory is always one-dimensional, there is a corresponding relationship between the multidimensional data and the storage order on the memory. Multidimensional data is usually allocated in a continuous storage space, that is, the multidimensional data can be expanded in one dimension and stored in sequence on the memory. For example, in the embodiment of the present disclosure, the initial input feature map can be stored sequentially in a low-dimensional (here C / Ci is the lowest dimension) priority manner; and in order to optimize the convolution operation, the storage order of the input feature map can be adjusted during the operation, as will be described in detail later. Adjacent dimensions refer to dimensions that are adjacent to each other in the dimensional information representation of multidimensional data. For example, W and Ci are adjacent. Adjacent dimensions can also be called continuous dimensions.

[0084] In intelligent processors, the primary hardware arithmetic unit is the vector multiplication-addition unit (MAU), driven by computing power requirements and area and power consumption considerations. Implementing support for various convolution algorithms in hardware design essentially involves maximizing the M&A operations within the algorithms and efficiently exchanging input and output data between on-chip RAM (such as the NRAM and WRAM shown in Figure 3) and the AMU through data paths.

[0085] Hardware stores data in cache rows (cache lines). Read, write, and compute operations are most efficient when aligned across the entire row. Therefore, to fully utilize bandwidth and adapt to the memory access requirements of the operator array, data is typically vectorized and aligned. Artificial intelligence chip designs typically use the Ci dimension as the lowest dimension, as described above in the NHWC arrangement order. Data along the Ci dimension is continuous. Therefore, vectorized alignment requires that the Ci dimension be aligned to a specified value, such as the alignment value M, so that accesses are performed in units of this alignment value M, which is also referred to as the maximum single hardware operation. Depending on the hardware design, M can have different values, such as 64 bits, 128 bits, 256 bits, or 512 bits. The input port size of the operator array is typically related to M. For example, if the input data bit width is symmetrical, the input port size of the operator array is typically twice M, meaning that input feature map data and weight data of the alignment value M can be processed simultaneously. Larger Ci dimensions of the input feature map make it easier to meet these alignment requirements.

[0086] When the Ci dimension of the input feature map is small, for example, smaller than the size of a cache line, the Ci dimension needs to be padded to one row of data (for example, 512 bits), that is, filled with invalid data 0. This padding will cause a lot of redundant calculations, resulting in resource waste and reduced computational efficiency.

[0087] In the disclosed embodiments, a convolution operation scheme is proposed that can determine the corresponding convolution splitting scheme based on the size of the lowest storage dimension (e.g., Ci) of the input feature map, where the convolution splitting scheme at least indicates the shape of the split unit of the data to be operated on. The amount of data contained in a split unit does not exceed the maximum single operation capacity of the hardware.

[0088] In some embodiments, the amount of data contained in a split unit can be set to the hardware's one-time processing alignment value M, so that calculation processing is performed in units of split units, which can fully utilize the computing power of the hardware and avoid or reduce invalid calculations.

[0089] In the exemplary description of this disclosure, it is assumed that M = 512 bits = 64 bytes, the data type can be Int8, Int16, Float16, or Float32, and the input feature map is consistent with the data type of the convolution kernel. Since the data type requires a width of at least 1 byte and the minimum unit of operation processing is one data, various calculations are performed in bytes in the following examples, such as M = 64 bytes, Ci = 28 bytes, etc., where units are sometimes omitted for brevity.

[0090] When the data volume of a split unit is equal to M, the data block shape of each split unit is blockC*blockY*blockX. There are many possible situations. Table 1 lists some of them:

[0091]

[0092] Table 1. Data block shape

[0093] As can be seen from Table 1, some data block shapes have equal X and Y dimensions (as shown in the dark rows), which can simplify subsequent operations. Therefore, in the disclosed embodiments, this data block shape can be preferably used to split the data to be operated.

[0094] For simplicity, the 64B×1×1 splitting scheme is called Forward64, the 16B×2×2 splitting scheme is called Forward16, the 4B×4×4 splitting scheme is called Forward4, the 4B×4×4 splitting scheme used for depthwise convolution operations is called Forward1, the 4B×4×4 splitting scheme used for reverse depthwise convolution operations is called Update1, and the 4B×4×4 splitting scheme used for cross-product convolution operations is called Update4. Except for Forward64, these splitting schemes are suitable for scenarios where the channel C in convolution calculations is relatively small, and therefore can also be collectively referred to as small convolutions. In these small convolution splitting schemes, a split unit includes data of the lowest storage dimension and at least one other storage dimension, and the total amount of data for a split unit does not exceed the maximum single-time computation capacity of the hardware.

[0095] Different convolution splitting schemes can be applied to different computing scenarios, thereby achieving different degrees of performance optimization.

[0096] After the splitting scheme is determined, the input feature map and convolution kernel can be split into multiple corresponding splitting units according to the determined convolution splitting scheme and their dimensional storage order can be converted so that the data in a splitting unit is continuously stored as a data row, thereby facilitating subsequent reading and processing in units of splitting units (data rows).

[0097] In some embodiments, for three-dimensional or four-dimensional neurons or weight data, all of them are divided into data blocks of size blockC*blockY*blockX (Uc×Uy×Ux), and each data block is continuously stored on a row of, for example, M=64B, so that when a row of data is read, the data of a data block is actually taken out.

[0098] Specifically, one or more split units can be read from the data to be calculated stored in the first dimensional storage order, with the split units as units, in a first reading order, and the read split units can be stored in the corresponding storage circuit, wherein the data in each split unit is stored in the second dimensional storage order, and the data between the split units is stored in the third dimensional storage order.

[0099] Figure 6 An exemplary data storage order according to an embodiment of the present disclosure is shown.

[0100] As shown in the figure, 610 represents the storage of the 4D tensor to be calculated, which consists of N 3D sub-tensors, with N being the highest dimension, i.e., the first dimension of the 4D tensor, stored in the order NHWC. Note that H and Y, W and X are used interchangeably in this article. Each sub-tensor is divided into smaller data blocks or split units, with the number of data blocks in each dimension being C / Y / X respectively.

[0101] The middle figure 620 shows the storage method of each sub-tensor, and each data block is stored as a continuous 64Byte, that is, a row. When the order of reading the data blocks is different, the order between the rows will also change accordingly. In the example in the figure, the data blocks are read in the direction of C first, then X, and finally Y, that is, the first reading order is YXC, then the rows are stored in the order of Y*X*C, that is, the third dimension storage order is YXC or HWC. In this example, the third dimension storage order is the same as the first dimension storage order. It can be understood that other reading orders can also be used, which will cause the third dimension storage order to be different from the first dimension storage order, and they will not be listed here one by one.

[0102] The diagram 630 on the right represents the order within each row, that is, the data order within each data block, and its shape is blockC*blockY*blockX. In this case, the second dimension storage order is CYX or CHW.

[0103] Example Grouping Operation

[0104] Small convolution uses block form. Its advantage over traditional convolution is that the alignment in the Ci direction only needs to satisfy the alignment of the block in the Ci direction. In this small channel scenario, the weights (co*Kh*kw*ci) are generally small, Kh and Kw are usually in the single digit, and co and ci are similar. Figure 5In the computing device / data processing device described above, the second storage circuit (e.g., WRAM 332 in FIG. 3 ) typically has a larger storage space than the first storage circuit (e.g., NRAM 331 in FIG. 3 ). Therefore, to fully utilize the on-chip computing space, most small convolution schemes, such as Forward4 and Forward1, employ a scheme where the storage locations of neurons and weights are swapped with those of normal convolutions. That is, neurons are stored in the second storage circuit WRAM, and weights are stored in the first storage circuit NRAM.

[0105] The calculation of convolution is that each input feature map needs to be multiplied and added with each Co convolution kernel, so as to output Co output feature maps. However, the on-chip space is not necessarily able to store all sizes of convolution kernels and input feature maps at the same time. Therefore, for the hardware, there is a series of operations of repeatedly loading input feature data or weight data. How to balance the repeated loading of input feature data or weight data will have a certain impact on the efficiency of the calculation. In actual operations, in order to reduce frequent off-chip memory accesses, there is a problem of splitting strategies for neurons and weights. In some embodiments, different splitting methods can be adopted according to the scale characteristics of the data involved in the operation.

[0106] According to the convolution operation principle described above, the operation results on the Co dimension (C dimension for depthwise convolution) do not need to be accumulated, so the operations on different Co can be performed relatively independently on different operation circuits. In small convolution scenarios, usually the size of the output channel Co dimension of the convolution kernel in a single round of operation does not exceed the number of scheduled slave processing circuits, so the operation of a single Co needs to be completed by one or more slave processing circuits. More generally, even when the Co dimension is large, it can be achieved by splitting it into multiple rounds of operations, where the size of Co processed in each round of operation does not exceed the number of scheduled slave processing circuits. Therefore, in one example, the number of operation rounds required to complete the convolution operation and the number of Co processed in each round of operation or the corresponding grouping mode can be determined first based on the size of the output channel Co dimension of the convolution kernel and the number of schedulable slave processing circuits Ns.

[0107] Regardless of the allocation method, in a single round of operation, Co may be allocated in two ways: multiple slave processing circuits process one Co value, or a single slave processing circuit processes one or more Co values. Specifically, in a single operation round that processes Nco output channels, each Rs SL constitutes a slave processing circuit group SLB, which processes the convolution kernel corresponding to the same output Co value, Rs = [Ns / Nco], that is, the same convolution kernel is reused on the Rs SLs within the same SLB, and Rs represents the number of times the convolution kernel is reused between the slave processing circuits. Correspondingly, the input feature map can be reused between each slave processing circuit group SLB, Rn = [Ns / Rs], which represents the number of times the input feature map is reused between the slave processing circuits.

[0108] Alternatively or additionally, when each slave processing circuit processes convolution kernels corresponding to rn Co values, rn = [Nco / Ns], the input feature map processed by each slave processing circuit can be reused for rn convolution kernels, where rn represents the number of times the input feature map is reused within a single slave processing circuit. Factors such as hardware buffer space limitations (e.g. Figure 5 The size of the first buffer circuit and the second buffer circuit in the processing circuit are used to determine the maximum number of convolution kernel multiplexing times rs and the maximum number of input feature map multiplexing times rn that can be applied within a single slave processing circuit.

[0109] Taking into account the cache size limitations and reuse benefits in the hardware circuit, in some embodiments of the present disclosure, the case where a slave processing circuit processes multiple Co values ​​in a single round of operation is temporarily not considered, but only the case where one or more slave processing circuits only process one Co value in a single round of operation is considered.

[0110] Different grouping modes can be used depending on the number of slave processing circuits SL that process the same Co value in a single round of operation. It is understood that it is preferable to evenly distribute the available slave processing circuits SL to balance the computing power. For example, every 2 SLs are grouped together, so that 16 SLs can process 8 Co values ​​at the same time; or every 4 SLs are grouped together, so that 16 SLs can process 4 Co values ​​at the same time; and so on. Figure 5 In the computing device described above, the second storage circuit WRAM has 16 storage areas, which are respectively allocated to 16 slave processing circuits SL. Furthermore, every 4 blocks can be combined into a storage block, which is allocated to the corresponding slave processing circuit group SLB. Therefore, in some embodiments, for Figure 5 The computing device shown includes Ns = 16 SLs, and the following grouping modes can be selected: Group 1 mode, Group 4 mode, and Group 16 mode. Those skilled in the art will understand that different grouping modes can be used depending on the value of Ns. Each grouping mode can refer to the three representative grouping modes given above for corresponding processing.

[0111] In some embodiments, the grouping pattern described above can be uniformly represented as GroupN, which means that all slave processing circuits SL scheduled in the current round of operations are divided into N groups. Each slave processing circuit group SLB processes the same Co value, and different slave processing circuit groups SLB process different Co values. For a total of 16 schedulable SLs, N can be 1, 4, or 16, corresponding to Group1, Group4, and Group16, respectively.

[0112] Figures 7a-7c Several exemplary grouping modes according to embodiments of the present disclosure are shown. Figure 7a shows the Group1 mode, Figure 7b Group 16 mode is shown, Figure 7c A Group 4 mode is shown.

[0113] like Figure 7a As shown, Group 1 mode means that all 16 schedulable SLs belong to a group and jointly process a single Co value. For example, SL0 through SL15 belong to group G0. Consequently, operations on a single output channel are distributed across the 16 SLs. In this mode, the convolution kernel 720 for that output channel is preferentially broadcast to each SL, while the input feature map 710 is split and distributed to each SL, thereby improving memory access efficiency.

[0114] In one embodiment, the convolution kernel can be stored in Figure 5 The input feature map is stored in the first storage circuit 530 for transmission via a broadcast channel. The input feature map can be divided according to the XY direction of the output feature map and stored in the second storage circuit 540 for distribution to different SLs. Thus, all SLs jointly calculate an output feature map Co. The division and storage of the input feature map will be described in detail later in conjunction with the accompanying drawings.

[0115] like Figure 7b As shown, the Group16 mode divides all 16 schedulable SLs into 16 groups, with one SL per group, and each SL processes a different Co value. For example, SL0 belongs to group G0, SL1 belongs to group G1, and so on, until SL15 belongs to group G15. In this mode, the same input feature map 730 can be reused across the 16 SLs. Therefore, it is prioritized to broadcast the input feature map 730 to each SL, while the convolution kernels 740 corresponding to different Co values ​​are distributed to the corresponding SLs.

[0116] In one embodiment, the input feature map can be replicated 16 times and stored in 16 storage areas allocated to 16 slave processing circuits on the second storage circuit. The convolution kernel is then divided according to Co, with one SL corresponding to one Co. 16 Cos are processed at a time, stored in the first storage circuit, and distributed to different SLs in a unicast manner. As a result, all SLs calculate output feature maps with different Cos for the same input feature map.

[0117] like Figure 7c As shown, Group 4 mode divides all 16 schedulable SLs into four groups, each processing one Co value. Each SL group (abbreviated as SLB) includes SLs equal to Rs = Ns / 4 = 4. For example, SL0-SL3 belong to group G0, SL4-SL7 belong to group G1, SL8-SL11 belong to group G2, and SL12-SL15 belong to group G3. This mode is between Group 1 and Group 16, so either the convolution kernel or the input feature map can be determined as multicast data, while the other can be determined as distributed data.

[0118] In one embodiment, the convolution kernels can be divided into 4 groups according to Co and stored in Figure 5 The input feature map is then divided into four parts along the XY direction of the output feature map, and four copies are stored in the second storage circuit 540 for distribution to the four SLBs. Each SLB receives the same input feature map, which is then distributed to the four SLs within the SLB according to the four copies. Thus, all SLs in each SLB jointly calculate the output feature map of Co, and the four SLBs each process a different Co.

[0119] like Figure 7c As shown, the convolution kernels are divided into four groups, each grouped by Co, with an interval of 1. For example, when Co = 12, the four groups are divided into Co, respectively {0, 4, 8}, {1, 5, 9}, {2, 6, 10}, and {3, 7, 11}. Each time, a Co is sent from each group. For example, the first time Co = 0 to 3 is sent, each Co corresponds to an SLB, and the four SLs within an SLB share the same weights. The second time, Co = 4 to 7 is sent, and so on. As a result, after each round of calculation, the Co dimensions of the calculation results output by each SLB are continuous.

[0120] When using the small convolution splitting operation scheme such as Forward4, in order to support the above three modes at the same time, the neurons can be uniformly stored in the second storage circuit WRAM and the weights can be stored in the first storage circuit NRAM.

[0121] Example split of input feature map

[0122] From the previous description, it can be seen that when multiple SLs jointly process a Co value, the input feature map needs to be split among these multiple SLs. For example, the Group1 grouping mode needs to split the input feature map into 16 parts, while the Group4 grouping mode needs to split the input feature map into 4 parts.

[0123] In order to ensure that the split input feature map can share the convolution kernel, it can be divided according to the Ho / Wo direction of the output feature map, so as to map back to the division of the input feature map. In some embodiments, the input feature map can be divided as follows between the Rs slave processing circuits SL included in each slave processing circuit group: according to the size of the corresponding output feature map, the output feature map is evenly divided into Rs output feature blocks of the same shape in the XY dimension (that is, the Ho / Wo dimension); and according to the input feature map area required to calculate each output feature block, the input feature map is divided into Rs input feature blocks in the XY dimension (that is, the Hi / Wi dimension) to be allocated to the Rs slave processing circuits. It can be understood that depending on the convolution kernel size and the convolution step size, the input feature maps corresponding to adjacent output points on the output feature map may overlap.

[0124] Figure 8 FIG2 shows an exemplary splitting diagram of an input feature map according to an embodiment of the present disclosure. In this example, the input feature map is divided into 16 parts and distributed on 16 SLs, corresponding to the Group 1 mode.

[0125] In the figure, 810 represents the output feature map of a single Co. This is divided into 16 output feature blocks of the same shape in a 4×4 pattern in the X and Y directions, and assigned to SL0 through SL15. These 16 output feature blocks are then mapped onto the input feature map 820, obtaining the 16 input feature map regions required to calculate each of these 16 output feature blocks. This is also done by dividing the input feature map in the X and Y directions. These 16 input feature map regions are then assigned to the 16 slave processing circuits SL.

[0126] According to the above description, the input feature map will be split into split units according to the determined convolution splitting scheme. Therefore, in the above embodiment, the input feature map is divided into blocks so that each divided input feature map block is a multiple of the XY dimension of the split unit in the XY direction, that is, it can be aligned according to the split unit in the XY direction. For example, when a 4×4×4 convolution splitting scheme is selected, each input feature map block is aligned as 4×4; and when a 16×2×2 convolution splitting scheme is selected, each input feature map block is aligned as 2×2.

[0127] If the output feature map is not aligned according to the split unit (for example, 4×4 or 2×2), it is necessary to pad the input feature map accordingly (for example, with 0s) so that the actual calculated output XY is aligned according to the split unit (for example, 4×4 or 2×2) and the input XY is also aligned according to the split unit (for example, 4×4 or 2×2).

[0128] Those skilled in the art will appreciate that the output feature map may also be split according to other rules in the XY direction, for example, split into 16 output feature blocks of the same shape in a 1×16 manner, and assigned to SL0 to SL15 respectively. The disclosed embodiment is not limited in this respect. In addition, it will be appreciated that although the foregoing description is made in conjunction with the splitting between slave processing circuits, this splitting method may also be applied to splitting in other scenarios, such as the splitting between computing circuits CU within a single slave processing circuit SL, and the disclosed embodiment is not limited in this respect.

[0129] Example convolution operation process within a single slave processing circuit

[0130] After the data to be operated is split and stored accordingly, multiple slave processing circuits can be scheduled to perform convolution operations on the corresponding data rows of the input feature map and the convolution kernel. Then, according to the convolution splitting scheme, the operation results returned by the multiple slave processing circuits can be spliced ​​to obtain the output feature map of the convolution operation of the input feature map and the convolution kernel. Specifically, multiple operation circuits CU and various buffer circuits in the slave processing circuit can be used (see Figure 5 ) to perform the specific convolution operation process. Depending on the space available for the buffer circuit within the processing circuit and the computing power limitations of the operation circuit, multiple operation cycles are usually required to complete the required operation in each round of operation.

[0131] As can be seen from the previous description, in the conventional 3D convolution operation scenario, all the operation circuits in a single slave processing circuit calculate an output feature map or a partial output feature map corresponding to the same output channel Co. Depending on the buffer space size of the first buffer circuit and the second buffer circuit in the slave processing circuit SL and the processing capability of the operation circuit CU (such as internal registers, etc.), the slave processing circuit may not be able to calculate the output feature map assigned to it at one time. Therefore, the output feature block can be divided into units based on the single operation capability of the operation circuit (for example, a single calculation of Nop output points or partial sums), and each output feature block corresponds to all N schedulable output points in a single SL. CU The single operation capability of the operation circuit (N CU *Nop output points). For example, Figure 5Taking the example where each SL includes 4 CUs, assuming that each CU can calculate Nop=4 output points or partial sums of output points at a time, a single SL can calculate 4*4=16 output points (or partial sums) at a time. Therefore, the output feature map can be divided into output feature blocks according to the alignment of 16 output points in the XoYo dimension, and each output feature block can be calculated one by one. It can be understood that these 16 output points can be in the form of 4*4 or 1*16, and the disclosed embodiment has no limitation in this regard.

[0132] When calculating the output feature blocks of each partition, we can further CU The output points of the output feature block are divided among the operation circuits to determine the processing objects of each operation circuit. Then, according to the division of the output points, N split units can be used as sliding windows to select N from the first buffer circuit. CU The input feature data rows are distributed to N CU An operation circuit selects the corresponding weight data from the second buffer circuit and broadcasts it to N CU A computation circuit is used to parallelize the computation of the output points corresponding to multiple sliding windows by reusing the weight data. Nk sliding selections are performed, where Nk is determined by the smaller value of the convolution kernel size in the X and Y dimensions and the maximum convolution kernel size supported by the processing circuit in a single operation in the current convolution split mode.

[0133] In some embodiments, when performing a conventional three-dimensional convolution operation, the corresponding weight data can be selected as follows: 1 / Nop weight rows are selected from the second buffer circuit according to the sliding method corresponding to that in the first buffer circuit, and Nop-1 copies are copied to expand it into an extended weight row, which is broadcast to N slave processing circuits. CU An operational circuit.

[0134] At this time, each operation circuit can perform bitwise multiplication and accumulation in units of 1 / Nop data rows for an input feature row from the first buffer circuit and an extended weight data row from the second buffer circuit during each sliding selection calculation to obtain Nop partial sums; and accumulate the Nk*Nop partial sums obtained from the Nk sliding selection calculations according to the corresponding convolution output points to obtain and output Nop operation results.

[0135] When outputting the output points of the arithmetic circuits within the slave processing circuit, the slave processing circuit can output the output points calculated by the multiple arithmetic circuits within the slave processing circuit in a specific order based on the output point division method, so that the output points outputted continuously are continuous in the X and / or Y dimensions, facilitating subsequent processing. In some embodiments, the master processing circuit can further store the operation results returned by each slave processing circuit in a fourth dimensional storage order. Depending on the circumstances, the master processing circuit can also convert the operation results into a desired dimensional storage order for storage.

[0136] There are many ways to divide the output points between the operation circuits, and the corresponding sliding selection convolution process and the output order of the output points are also different.

[0137] The following describes in detail the entire process of data splitting, storage, convolution sliding, and calculation output in combination with the Forward4 solution.

[0138] Shape description of input neurons and weights of the Forward4 scheme

[0139] In Forward4, the block size is 4B × 4 × 4. The block shape varies slightly depending on the data type. Table 2 shows the block shapes used by Forward4 for different data types.

[0140]

[0141] Table 2. Data block shapes of Forward4 under different data types

[0142] Figure 9 FIG2 shows a schematic diagram of splitting and storing the Forward4 solution according to an embodiment of the present disclosure. For simplicity, the example in the figure assumes that the data type is Int8.

[0143] The figure 910 shows the original data to be calculated (which can be neurons or weights), which is stored in the HWC order. The figure also shows four data blocks 911-914 in which the original data to be calculated is split according to the splitting unit, each data block including 4×4×4=64 data.

[0144] Figure 920 shows the format of the split data for easier reading. As can be seen, the original data blocks (e.g., 911-914) are arranged into a row (e.g., 921-924) along the C dimension. Within each row, the data is stored in CHW order. For example, for data row 921, the 16 data points with C=0 are stored first, followed by the 16 data points with C=1, then the 16 data points with C=2, and finally the 16 data points with C=3.

[0145] Specifically, for neurons, the data needs to be arranged from [1Hi Wi Ci] to:

[0146] [1*Hi / 4*Wi / 4*Ci / 4*(4×4×4)], the shape of this seven-dimensional tensor.

[0147] For the weights, the data needs to be arranged from [Co Kh Kw Ci] to:

[0148] [Co*Kh / 4*Kw / 4*Ci / 4*(4×4×4)], the shape of this seven-dimensional tensor.

[0149] As described above, the Forward4 solution supports multiple grouping modes. For neurons, depending on the grouping mode and the HoWo splitting method within the group, the seven-dimensional shape of the block format described above is ultimately split into each storage area of ​​the second storage circuit, which can be slightly different.

[0150] Assume that the original input neuron size is: [1*hi*wi*ci]

[0151] In Group 1 mode, the number of input neurons varies depending on the HoWo splitting method:

[0152] Ho*Wo 4*4 split: 16[hi / (4*4),wi / (4*4),ci / 4*(4*4*4)]

[0153] Ho*Wo 1*16 split: 16[hi / (4),wi / (4*4*4),ci / 4*(4*4*4)]

[0154] In the above 4*4 split, 16 represents 16 slave processing circuits (SLs). The final 4*4*4 (CHW) represents the CHW block split across three dimensions. The first 4 in the division of hi and wi indicates splitting hi*wi into 16 parts and distributing them to the 16 SLs. The second 4 indicates folding hi and wi toward ci. The same applies to the 1*16 split.

[0155] In Group4 grouping mode, the number of input neurons varies depending on the HoWo splitting method:

[0156] Ho*Wo 1*4 split: 4*4*[hi / (1*4),wi / (4*4),ci / 4*(4*4*4)]

[0157] For a slave processing circuit SL: hi / (1*4), wi / (4*4), ci / 4*(4*4*4)

[0158] In the above representation, the first 4 indicates 4 SLBs, and the neuron is replicated 4 times. The second 4 indicates that the neuron is split on the 4 SLs of an SLB. The last 4*4*4 represents the CHW BLOCK split by three dimensions.

[0159] In Group 16 mode, the input neurons do not need to be split, and the number of neurons is as follows:

[0160] 16*[hi / 4,wi / 4,ci / 4*(4*4*4)]

[0161] The above 16 means that the neurons are replicated on 16 SLs. The last 4*4*4 represents the CHW BLOCK split from three dimensions. Both hi and wi are divided by 4, which means folding hi and wi to ci direction.

[0162] Output point splitting between operation circuits in the Forward4 scheme

[0163] When multiple computing circuits CU within a single slave processing circuit SL jointly process a Co value, the output points need to be split among the multiple CUs.

[0164] Figure 10 FIG. 4 shows a schematic diagram of allocating interval output points to each operation circuit in the Forward4 scheme according to some embodiments of the present disclosure. CU The output feature block is evenly divided into Nop output feature sub-blocks of the same shape between the operation circuits. Each output feature sub-block includes N CU output points, divided into N CU operation circuits. For example, the figure takes the example of each SL including 4 CUs and each CU being able to calculate Nop=4 output points or partial sums at a time, showing that the output feature block 1010 includes 4*4 output points, and each output feature sub-block 1011~1014 divided evenly includes 2*2 output points. In each output feature sub-block, these 2*2 output points are assigned to 4 operation circuits. Thus, each operation circuit calculates one output point in each of the 4 output feature sub-blocks. The output points assigned to 4 different operation circuits CU0~CU3 are shown with different backgrounds in the figure. It can be seen from the figure that in each calculation, each operation circuit calculates multiple output points spaced in the X and / or Y dimensions on the output feature map.

[0165] Based on the above output point division, when performing convolution operation by sliding selection, N can be selected from the first buffer circuit corresponding to the output point position of each output feature sub-block according to the data required for calculating the output feature sub-block. CU Calculations are performed on each data row. For example, when initially selecting the input feature data, four input data rows can be selected from the corresponding input feature blocks based on the four input feature blocks required to calculate the four output points within the first output feature sub-block 1011, and distributed to the four calculation circuits. It will be understood that since these four output points are continuous in the X and / or Y directions, the interval or step size of the four simultaneously selected input data rows in the X and / or Y directions is 1.

[0166] When selecting weight data, the corresponding weight data can be selected from the second buffer circuit and broadcast to N CUArithmetic circuits are provided, thereby implementing parallel calculation of output points corresponding to multiple arithmetic circuits by multiplexing weight data. Furthermore, in some embodiments, in order to fully utilize the computing power (e.g., multiplication and addition units) within the arithmetic unit CU, for example, to calculate Nop output points or partial sums at a time, weight multiplexing can be performed within a single input data row, thereby simultaneously calculating Nop output points or partial sums.

[0167] For example, when selecting weight data, we can take only 1 / Nop weight rows, copy them Nop-1 times to expand them into 1 weight row, and this expanded weight row includes Nop identical 1 / Nop weight rows. The expanded weight row can also be broadcast to N CU A plurality of operation circuits are provided, so that while the weights are multiplexed between the multiple operation circuits, the weights are multiplexed with a smaller granularity (for example, 1 / Nop row) between the calculations of the Nop output points of a single operation circuit.

[0168] Therefore, by taking N CU Input feature data rows, take 1 / Nop weight rows and copy them to expand them into 1 weight row, and N weights can be calculated each time. CU *Nop output points or partial sums. When the calculation result is a partial sum, multiple sliding operations can be used to calculate the partial sum multiple times. The partial sums are accumulated according to the output points to obtain the final result.

[0169] According to the division method of the output points, the number of sliding times and sliding steps of the convolution operation can be determined. Figure 10 The partitioning method is as follows: the number of slides Nk = ceil(Kx / 2) * ceil(Ky / 2), where Kx and Ky are the smaller values ​​of the convolution kernel size in the X and Y dimensions and the maximum convolution kernel size supported by the slave processing circuit in a single operation in the current convolution splitting mode, and the sliding step size = 2. The maximum convolution kernel size supported by the slave processing circuit in a single operation is determined by, for example, at least the spatial size of the first buffer circuit and the second buffer circuit. It can be understood that when the convolution kernel exceeds the maximum convolution kernel size, it is necessary to split according to the maximum convolution kernel size in the Kx and Ky directions.

[0170] Convolution sliding process in the Forward4 scheme

[0171] Figure 11 Figure 1 shows a schematic diagram of a single operation in the Forward4 scheme according to an embodiment of the present disclosure. In this example, the first buffer circuit 1110 has a size of 3×3×64B, which means it can cache up to 9 rows of data. The second buffer circuit 1120 has a size of 2×2×64B, which means it can cache up to 4 rows of data. To be consistent with the split unit, the storage within the buffer circuit in the figure is also shown in units of split units.

[0172] The figure shows the operation process of the first sliding data acquisition. According to the method corresponding to the division of the output point, the split unit is used as a sliding window to slide and select N from the first buffer circuit. CU Input feature rows are sent to N CU arithmetic circuits for calculation; select 1 / Nop weight rows from the second buffer circuit according to the sliding method corresponding to the first buffer circuit, where Nop is the maximum number of convolution output points that can be calculated in a single operation circuit, copy Nop-1 copies to expand it into an extended weight row, and broadcast it to N slave processing circuits. CU An operational circuit.

[0173] Specifically, in Figure 5 In the computing device shown, N CU = 4, Nop = 4. When dividing the output points, each operation circuit calculates 2×2 output points with an interval of 1 in both the X and Y dimensions for division in each calculation.

[0174] As shown in the figure, one input feature data row is selected from the first buffer circuit 1110 at the starting position and at a position shifted by 1 in the X and / or Y directions, for a total of four input feature data rows, which are correspondingly sent to the four computing circuits 1140 in the slave processing circuit SL. A quarter of a weight data row, i.e., data of size 2×2, is selected from the second buffer circuit 1120 at the starting position, replicated three times, and expanded into an extended weight data row 1130, which is then broadcast to the four computing circuits 1140 in the SL.

[0175] During each calculation, each operation circuit performs bitwise multiplication and accumulation on an input feature row from the first buffer circuit and an extended weight row from the second buffer circuit in units of 1 / Nop data rows to obtain Nop partial sums.

[0176] As shown in the figure, four computation circuits 1140 perform bitwise multiplication and accumulation operations on the distributed input feature data rows and the broadcasted expanded weight data rows, producing computation results 1150. The results in 1150 with different background colors represent those obtained by different computation circuits 1140. It can be seen that in each computation, a CU calculates the partial sum of four output points, resulting in a total of 4×4 partial sums across the four CUs. It can be seen that the output points calculated by each CU are not adjacent in the XoYo dimension of the output feature map.

[0177] Next, the first and second buffer circuits perform synchronous sliding fetches to perform the next calculation. Nk sliding fetches are performed, where Nk = ceil(Kx / 2) * ceil(Ky / 2), where Kx and Ky are the smaller of the convolution kernel size in the X and Y dimensions or the maximum convolution kernel size supported by the processing circuit in the current convolution split mode. Accordingly, the computation circuit accumulates the Nk*Nop partial sums obtained from the Nk sliding fetches according to the corresponding convolution output points, yielding Nop computation results.

[0178] In some embodiments, in Forward4 mode, the maximum convolution kernel size supported by a single operation of the slave processing circuit is 8×8.

[0179] Figure 12 A schematic diagram of the sliding convolution process in the Forward4 scheme according to an embodiment of the present disclosure is shown. This example takes a 9×9 input feature map and a 5×5 convolution kernel as an example. The convolution step size is 1, and the output feature map size is 5×5. The input feature map needs to be aligned to 12×12, divided into 9 blocks of 4×4×4 (C×H×W) in size, and stored in the first buffer circuit, shown as 1210 in the figure, where the C dimension is omitted. The convolution kernel 5×5 needs to be aligned to 8×8, and the aligned part is padded with 0, and stored in the second buffer circuit, shown as 1220 in the figure, and the C dimension is also omitted. Each time calculation is performed, a 2×2 block in the convolution kernel is selected and copied 4 times, which just corresponds to the 4×4 block of the input feature map. The copy operation can be implemented by hardware.

[0180] The selection range of the input feature map and the convolution kernel in the first buffer circuit and the second buffer circuit at each slide is as follows: Figure 12 As shown, there are nine images, representing nine slides. Block 1210 represents the input feature map in the first buffer circuit, and the four dashed boxes indicate the regions selected for distribution to the four CUs. Block 1220 represents the convolution kernel in the second buffer circuit, and the dashed boxes represent the selected 1 / 4 row, which is replicated three times to form a single row and then broadcast to the four CUs. Number of slides Nk = ceil(Kx / 2) * ceil(Ky / 2) = 9.

[0181] During each calculation, each CU performs bitwise multiplication and accumulation on an input feature data row from the first buffer circuit and an extended weight data row from the second buffer circuit in units of 1 / 4 data row to obtain 4 partial sums; and in the current calculation round, the Nk partial sums corresponding to the same convolution output point obtained in Nk calculations are accumulated to obtain and output 4 calculation results.

[0182] Specifically, for Figure 12For each image in the image, the number of CUs Ncu=4, and each CU calculates Nop=4 output points or partial sums at a single time. The partial sum is the result of the bitwise multiplication and accumulation of 1 / 4 data rows, that is, each output point is a standard convolution of 4×2×2(Ci×Y×X). After sliding Nk=ceil(Kx / 2)*ceil(Ky / 2)=9 times, the accumulation is completed in the Y×X direction, and finally a complete 4×4(Y×X) output is obtained in 1 SL. In this mode, a single calculation only supports the case where the convolution kernel is no larger than 8×8. For larger convolution kernels, it is necessary to split it into 8×8 in the Kx and Ky directions, and the splitting operation can be performed according to the same principle as above.

[0183] It can be understood that when Ci>4, it is necessary to traverse in the Ci direction, switching inputs and weights at the same time, until the complete output is calculated. When the Xo / Yo calculated by each CU is greater than 4, it is necessary to slide along the Xo / Yo direction to read different input neurons and weights. Those skilled in the art can similarly deduce the calculation process based on the above description, and will not be repeated here.

[0184] Output shape description in Forward4 scheme

[0185] As can be seen from the previous output point division method and sliding convolution process, the results of the sliding mode output are not arranged in the normal order of traditional convolution output data. Therefore, during the output process, each slave processing circuit SL can convert the operation results of its internal operation circuit CU into a specified format, such as the format of Nco×Uy×Ux. In some embodiments, each slave processing circuit can output partial operation results of its internal partial operation circuit each time, and the partial operation results are continuous in the X and / or Y dimensions of the output feature map. The master processing circuit can further store the operation results returned from each slave processing circuit in a fourth dimensional storage order. Depending on the situation, the master processing circuit can also convert the operation results into the desired dimensional storage order for storage.

[0186] When the grouping mode and / or the splitting method of the input feature map within a single SLB (that is, the splitting method based on the HoWo of the output feature map) are different, the output data format is slightly different.

[0187] Figure 13 The output data format of the Forward4 solution according to one embodiment of the present disclosure is shown in FIG. In this embodiment, the grouping mode is Group 1, and the input feature map in a single SLB (including 16 SLs) is split according to Ho×Wo=1×16.

[0188] Figure 1310 shows the original output of one SL. As can be seen from the figure, each SL outputs a 1×1×4 (Co×Y×X) area each time, that is, it outputs part of the calculation results of its internal calculation circuit each time, for example, 2 calculation results of each of the 2 CUs (see Figure 10 ), this part of the operation results are continuous in the X and / or Y dimensions of the output feature map, for example, in the same row ( Figure 13 ) or the same column. This returns a 1×4×4 (Co×Y×X) region four times in a row, representing the four computation results for each of the four CUs. Different SLs output different regions of the output feature map for the same Co. After outputting all 4×4 regions of Co, further output switches to different output points.

[0189] Figure 1320 shows the data structure for storing and outputting 16 SLs. As shown, after being written to a storage circuit (e.g., the first storage circuit), the final output data is formatted as Yo*Xo*Co*4*16*4, where Yo and Xo represent the number of blocks in the output feature map that each SL is partitioned into, and 16 represents the number of blocks divided across the 16 SLs. In some implementations, further pendulum operations can be performed to convert the data into other desired formats, as needed.

[0190] As mentioned above, the output data format may be slightly different depending on the grouping mode and / or the way the input feature maps are split between multiple SLs within a single SLB. Assuming the original output size is:

[0191] 1*ho*wo*co

[0192] Then, the output data shape of Group1 when Ho*Wo is split into 4*4 is:

[0193] ho / (4*4)*wo / (4*4)*co / group*(4*16*4)

[0194] In the above formula, (4*16*4) is the basic output block of forward4, with directions corresponding to h*c*w. The 16 represents the division of the same co in the 16 SLs into ho and wo. Ho and wo are divided by 4 twice. The first 4 indicates a 4×4 split when storing data in the SL, and the second 4 indicates data block folding in the h and w directions. In Group1 mode, group = 1.

[0195] The output data shape of Group1 when Ho*Wo is split according to 1*16 is:

[0196] ho / (4)*wo / (4*16)*co / group*(4*16*4)

[0197] In the above formula, (4*16*4) is the basic output block of forward4, and the directions correspond to h*c*w respectively, where 16 represents the division of ho and wo for the same co on 16 SLs; in Group1 mode, the above group = 1.

[0198] As can be seen, in Group 1, the 16 SLs equally divide the Yo*Xo dimension of an output feature map. The data in the row dimension SL during output corresponds one-to-one to the way the 16 SLs equally divide the output neurons in the Yo*Xo direction. This scenario is suitable for input neurons with large Y*X values ​​and small Co values.

[0199] The output data shape of Group4 is:

[0200] ho / (2*4)*wo / (2*4)*co / group*(4*16*4)

[0201] In the above formula, (4*16*4) has the same meaning as above, except that 16 represents the division of wo outputs of 4 cos on 4 SLs. In Group 4 mode, group = 4.

[0202] The output data shape of Group16 is:

[0203] ho / 4*wo / 4*co / group*(4*16*4)

[0204] In the above, (4*16*4) has the same meaning as above, except that 16 represents the output division of 16 COs on 16 SLs. In Group 16 mode, group = 16.

[0205] Since Group has different split categories in the H*W direction, the 16 in the 4*16*4 mentioned above has differences in specific splits. Since Forwrd4 is calculated based on 4B*4*4 blocks, it is inevitable that there will be alignment restrictions during calculation. According to different Group modes, different H*W splitting methods of the same Group mode will eventually result in different alignment restrictions during calculation. In the calculation of alignment, you can first determine the alignment restrictions of ho*wo based on the splitting method of the output feature map, and then reversely infer hi*wi from ho*wo. Since the input neurons need to be arranged in the form of split unit blocks, they need to be aligned again. The above alignment restrictions can be summarized in the following Table 3:

[0206]

[0207] Table 3. Alignment constraints

[0208] In summary, during output, the hardware can automatically output neurons in the form of 4*16*4 (Y*SL*X) dimensions within rows and Y*X*C dimensions between rows. The same applies to larger convolution kernels.

[0209] Biased shape description in Forward4 scheme

[0210] Bias is the bias after the convolution calculation is completed. The original format of the bias is: [1 1co].

[0211] Since the data output by Forward4 is in the format of ho*wo*co / group*(4*16*4), if you need to add an offset to the data directly output by Forward4 on the chip, you need to change the basic shape of the offset. The layout of the offset in the on-chip space is related to the grouping mode. Specifically, the number of offsets in various grouping modes is as follows:

[0212] In Group 1 mode, the number of biased pendulums is: [1 1co*64]

[0213] Here, 64 means that a single offset is copied 64 times and placed consecutively.

[0214] In Group 4 mode, the number of offset pendulums is: [1 1co*16]

[0215] Here, 16 means that a single offset is copied 16 times and placed consecutively.

[0216] In Group 16 mode, the number of offset pendulums is: [1 1co*4]

[0217] Here, 4 means that a single offset is copied 4 times and placed consecutively.

[0218] Data transfer process

[0219] From the previous description of the small convolution operation scheme, we can see that the input neurons and weights need to be split and stored in dimension transformation, and the output neurons also need to undergo certain dimension transformation. Figure 3bRegarding the hardware structure of a multi-core computing device, for the sake of hardware IO efficiency, the input data needs to be read from the global memory first, and then stored in the shared storage SRAM after loading the data. As mentioned earlier, Forward4 needs to split neurons. Taking into account the alignment factor, its splitting feature determines that Forward4 has more computational advantages when processing input feature maps that are relatively large and have a relatively small number of channels. Therefore, when designing the hardware of Forward4, larger neurons can be stored on WRAM and relatively smaller weights can be stored in NRAM. At the same time, since the weights and neuron data need to be arranged in the form of blocks described above, the neurons stored on WRAM also need to pass through NRAM once to perform a shape transformation of the tensor data.

[0220] Figure 14 The overall data handling process according to the embodiment of the present disclosure is shown.

[0221] As shown in the figure, weights are read from off-chip storage, such as DDR, via the global direct memory access module (GDMA) into SRAM. Hardware-side alignment and padding operations are performed on the SRAM. Tiling instructions are used as data is transferred from SRAM to NRAM, completing both the data transfer process and the data dimension transformation and alignment.

[0222] The process of transferring neurons is similar to that of weights, except that after being transferred to NRAM via the block instruction, they also need to be transferred to WRAM. Because neurons are calculating, as the convolution kernel slides, there is a large amount of data overlap, which greatly reduces the efficiency of data transfer. To solve this problem, some embodiments of this disclosure use the img2col instruction to distribute data, which will be described in detail below.

[0223] Output data can be stored back to NRAM and can also be transformed into data dimensions and moved to SRAM using block instructions. It can then be stored back to off-chip DDR memory via GDMA.

[0224] Exemplary principle of block instructions

[0225] Data dimensionality change and manipulation refers to the process of arranging tensor data of a specific shape into the desired specific shape. Data movement refers to the reading and writing of data between different memory spaces. As mentioned earlier, the Forward4 convolution operation scheme requires that the neurons and weights used for convolution operations are arranged and aligned according to a specific block pattern. In addition, the output data is also output according to the specific Forward4 output format. This requires that the tensor data be arranged in blocks before calculation and then restored to its normal tensor shape after the calculation is completed.

[0226] In the disclosed embodiments, the process of transferring input neuron, weight, and bias data from SRAM to NRAM, and the process of transferring output data from NRAM to SRAM, all utilizes tiling instructions to complete this transfer operation. This transfer process requires completing the basic data transfer process, as well as the data dimensionality change and placement process to meet the computational requirements.

[0227] The Deform instruction family provides the ability to transform data shapes and convert data types for IO data paths, primarily including functions such as TRANS, MOVE, and ROTATE. The mode for implementing the transposition function in this instruction family is named Trans Tiling, and its primary purpose is to provide performance support for various shape transformations of small convolutions. Deform divides a 3D data block into two layers: an inner and an outer layer. The inner layer has three dimensions (corresponding to parameters n0-2 in the instruction). The lowest dimension is in bytes, while the next lowest and highest dimensions are unitless and represent the number of units in the previous layer. The outer layer also has three dimensions (corresponding to parameters n3-n5 in the instruction), each representing a multiple of the corresponding inner layer dimension.

[0228] When implementing the small convolution splitting scheme, the input data stored in the first dimension storage order (such as HWC) needs to be split, dimensionally converted and stored in split units. Each split unit is stored in the second dimension storage order (such as CHW), and the split units are stored in the third dimension storage order (such as HWC).

[0229] Figure 15 FIG2 is a schematic conceptual diagram of Trans Tiling according to an embodiment of the present disclosure.

[0230] The left figure shows the input data before deformation. It can be seen that the three-dimensional input data is described using six dimensions, n0 and n3 correspond to the first dimension of the original three-dimensional data (e.g., the lowest dimension), n1 and n4 correspond to the second dimension of the original three-dimensional data (e.g., the second lowest dimension), and n2 and n5 correspond to the third dimension of the data block (e.g., the highest dimension). In the example in the figure, the inner layer of the input data corresponds to the split unit. Taking the Forward4 scheme as an example, the inner layer data block of the input data is a 4B×4×4 data block, where n0=4B, n1=n2=4.

[0231] The right image shows the deformed output data. Three-dimensional output data is also described using six dimensions. In this case, the inner layer of the output data corresponds to the deformed split unit. In the forward scheme, the inner layer of the output data is a 64B × 1 × 1 data block, where n0 = 64B and n1 = n2 = 1.

[0232] Trans Tiling also features inline shuffles, including pre-tiling inline shuffles based on a pretable and post-tiling inline shuffles based on a posttable. The pretable shuffles the data in the n0 tiling input, while the posttable shuffles the data in the n0 tiling output. Without considering the table's flags, the pretable and posttable are essentially arrays representing the positions of 64 bytes of data.

[0233] Figure 16 A schematic diagram of the front and rear matching tables is shown.

[0234] As shown in the figure, the front and back tables respectively represent the rearranged position of a row of data in the n0 dimension of the input or output, which consists of 64 bytes. Each 8-bit byte includes a 6-bit index bit, which records the order of the byte data of bytes 0 to 63 in the original data; a 1-bit zero_en bit, which indicates whether it is set to 0. If this bit is 1, it is forced to write 0, and the [5,0] bit is invalid; and a 1-bit mask bit, which indicates whether the data in this bit is valid.

[0235] By using the front and back tables, the data of n0 of the input data of the block instruction can be rearranged when necessary, and / or the data of n0 of the output data of the block instruction can be rearranged.

[0236] Table 4 shows the meaning of the various parameters of the block instruction. Assuming that the bit width of the data to be blocked is dwidth, in bytes, the amount of data that the block instruction atomically operates on is called the block width T, in bytes. Among the parameters of the block instruction, 11 parameters, n0 to n5 and s1 to s5, are required to describe the tensor shape of the inner and outer data. Among them, n0 to n2 and s1 to s2 are parameters describing the inner layer, and n3 to n5 and s3 to s5 are parameters describing the outer layer.

[0237]

[0238] Table 4. Meaning of parameters of block instructions

[0239] The tensor descriptions before and after the execution of a block instruction require a set of parameters for each input tensor, described by 22 parameters: in0-in5, is1-is5, on0-on5, and os1-os5. Block instructions support a variety of block widths T, such as 1B, 2B, 4B, 6B, 8B, 16B, and 32B. Depending on the specific block task, the corresponding value can be set. Therefore, block instructions also include the block width T parameter.

[0240] When using block instructions, there are some basic usage restrictions or constraints. These restrictions include, for example: in0, in1, in2, on0, on1, on2 <= 64; n0 performance requires 64B alignment; in0 = on1*on2*T, on0 = in1*in2*T; in3*in4*in5 = on3*on4*on5; T <= 32B; front and back table = 64B.

[0241] In addition, the block instruction cannot be operated in situ, that is, two storage areas are required. Therefore, in the embodiment of the present disclosure, a data processing device is provided, including a control circuit, a first storage circuit and a second storage circuit. The first storage circuit is used to store data before the block instruction is executed; the second storage circuit is used to store data after the block instruction is executed. The control circuit is used to configure and execute the block instruction. In some embodiments, the data processing device can be, for example Figure 3b In the illustrated multi-core computing device, the control circuit is, for example, a processor core within the processor cluster, the first storage circuit is, for example, a shared memory SRAM within the processor cluster, and the second storage circuit is, for example, an NRAM within the processor core. When executing block instructions for different data (input neurons, weights, output neurons, etc.), the required dimensional changes and handling processes are also different, necessitating the design of different block instruction parameter configuration schemes.

[0242] General scheme for block instructions of neurons

[0243] According to the description of small convolution operation schemes such as Forward4 in the previous article, for input neurons, the role of the block instruction is to split, convert and store the input neurons stored in the first dimension storage order (such as HWC) in units of split units during the process of transferring the input neurons from SRAM to NRAM. The second dimension storage order (such as CHW) is used to store the input neurons in each split unit, and the third dimension storage order (such as HWC) is used to store the input neurons between the split units. The shape of the split unit is CHW=U Ci ×U H ×U W The alignment value M required by the block instruction is U Ci multiples of .

[0244] Specifically, for the neuron data in the Forward4 scheme, the data needs to be arranged from [1*hi*wi*ci] to:

[0245] [1*hi / 4*wi / 4*ci / 4*(4*4*4)]

[0246] Figure 17 A schematic diagram of executing a block instruction on neuron data according to an embodiment of the present disclosure is shown.

[0247] The left figure shows the neuron data before block processing (i.e., the input tensor of the block instruction). It can be seen that the three-dimensional neuron data [hi*wi*ci] (N dimension is omitted here) is divided into two layers, inner and outer, each described by three dimensions. The in0 dimension of the inner layer data block 1701 is aligned to the first alignment value according to the restriction of the block instruction, for example, M=64B; the in1 dimension can be set to U according to the shape of the split unit. W , in this example it is 4; the in2 dimension can also be set to U according to the shape of the split unit H , which is 4 in this example. After the inner data block is determined, the sizes of the three outer dimensions in3, in4, and in5 can also be determined accordingly, and their sizes are equal to the number of inner data blocks in the corresponding dimensions.

[0248] The right figure shows the neuron data after block processing (that is, the output tensor of the block instruction). It can be seen that the shape of the neuron data at this time becomes [hi / 4*wi / 4*(ci*16)], which is also divided into two layers, inner and outer, each using three dimensions to describe. Since the neuron data needs to be split according to the split unit, the block width T can be set to U according to the constraints of the block instruction. Ci , that is, the amount of data for one atomic operation is U Ci, so that it is easy to adjust the storage order in units of split units. At this time, the inner data block 1702 can correspond to the inner data block 1701 of the input tensor, but the shape is M×U H ×U W becomes (M*U H *U W )×1×1, shown in the figure as a large strip composed of 16 thin strips. The inner data block's on0 dimension is set to in1*in2*T=M, or 64 bytes, according to the constraints of the block instruction; the on1 dimension is set to in0 / T=M / T, which is 16 in this example; and the on2 dimension is set to in0 / T / on1=1. After determining the inner data block, the sizes of the three outer dimensions, on3, on4, and on5, can also be determined accordingly. Their sizes are equal to the number of inner data blocks containing output tensors in the corresponding dimensions.

[0249] From the above block instruction execution process, we can see that although the input neurons can be arranged from [1*hi*wi*ci] to [1*hi / 4*wi / 4*ci / 4*(4*4*4)], the lowest-dimensional 4*4*4 split unit block is still in HWC order, not CHW order. In order to achieve CHW order within the lowest-dimensional split unit block, it is necessary to use the post-matching table described above.

[0250] In some embodiments, the control circuitry in the data processing apparatus may be further configured to configure the blocking instruction as follows: set a post-matching table for the blocking instruction to rearrange the inner lowest-dimensional data of the output tensor of the blocking instruction according to the post-matching table. Specifically, the control circuitry may be further configured to set the post-matching table as follows: convert the inner lowest-dimensional on0 data of the output tensor arranged in a first-dimensional storage order (e.g., HWC) to an arrangement in a second-dimensional storage order (e.g., CHW).

[0251] The post-matching table at this point only obtains the output tensor on0, that is, the writing order of the input in1*in2*T=M=64B data. The writing order of this 64B data is related to the data bit width dwidth. In some embodiments, the post-matching table can be configured according to the logic shown in the pseudo code of Table 5 below.

[0252]

[0253]

[0254] Table 5. Pseudocode for tabulation after neuron block instruction

[0255] In actual operations, the shape of neurons changes dynamically, meaning the size of ci is arbitrary. Because the block instruction requires the lowest dimension n0 of the inner data block to be aligned to M = 64B, the entire block processing needs to be executed in two stages. The first stage is the 64B-aligned integer stage, and the second stage is the remainder stage that does not reach 64B.

[0256] Therefore, in some embodiments, the control circuit in the data processing device may divide the input data (e.g., neurons) into integer segments and remainder segments according to the input channel Ci dimension, wherein the Ci dimension of the integer segment is aligned to the alignment value M, and the Ci dimension of the remainder segment is smaller than M. Subsequently, a first blocking instruction may be configured and executed for the integer segment, and a second blocking instruction may be configured and executed for the remainder segment.

[0257] It's understandable that depending on the value of ci, there may be only integer segments, only remainder segments, or both integer and remainder segments. Assume that the length of the 64B-aligned integer segment in ci is ci_full, and the length of the unaligned remainder segment is ci_rem. For example, for an INT8 neuron of type 1*256*256*96, with ci = 96, then ci_full = 64 and ci_rem = 32.

[0258] Table 6 shows the shape change of neuron data before and after executing the block instruction.

[0259]

[0260] Table 6. Shape changes of neuron data before and after block processing

[0261] Note that the shape assumptions in Table 6 are parameters that have been aligned according to the alignment restrictions in Table 3 regarding different Forward4 Group modes and different H*W splitting methods.

[0262] For the integer segment part, you can refer to the previous text and combine Figure 17 Describes the content configuration of the first block directive.

[0263] In one embodiment, the parameters of the input tensor of the first block instruction can be configured as follows: the inner lowest dimension size in0 of the input tensor in the first block instruction is set to M, and the inner lowest dimension size in1 is set to U W , the inner highest dimension size in2 is set to U H ; and according to the size of each dimension of the integer segment of the input data, set the size values ​​in3, in4 and in5 of the three outer dimensions of the input tensor in the first block instruction, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks containing the input tensor on the corresponding dimensions.

[0264] Additionally, in one embodiment, the parameters of the output tensor of the first block instruction can be configured as follows: the inner lowest dimension size on0 of the output tensor in the first block instruction is set to in1*in2*T, the inner low dimension size on1 is set to M / T, and the inner highest dimension size on2 is set to 1; and according to the size of each dimension of the integer segment of the input data, the size values ​​on3, on4 and on5 of the three outer dimensions of the output tensor in the first block instruction are set, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks containing the output tensor in the corresponding dimensions.

[0265] In addition to the dimension parameters, the dimension step size also needs to be set. In some embodiments, the step size of the storage space of adjacent data points in the five dimensions other than the lowest inner dimension can be set based on the dimensional size of the six dimensions of the input tensor and the output tensor and the size of each dimension of the input data before processing. In some embodiments, the block width T can be set to U according to the constraints of the block instruction and the shape transformation of the split unit before and after processing. Ci .

[0266] In one example, when the Forward4 solution is used, the split unit is U Ci ×U H ×U W =4B×4×4, M=64B, the first block instruction for the integer segment can be configured according to the following Table 7.

[0267]

[0268]

[0269] Table 7. Parameter configuration scheme for neuron integer segment block instructions

[0270] Where ci, hi, and wi represent the number of data in the Ci, H, and W dimensions of the input data, respectively. dwidth represents the data bit width. ci_full represents the number of data in the Ci dimension of the integer segment. B represents bytes. T represents the block bit width. is1 to is5 represent the five-dimensional strides of the input tensor. os1 to os5 represent the five-dimensional strides of the output tensor.

[0271] For the remainder segment part, the second blocking instruction can be configured based on a slight adjustment of the integer segment part.

[0272] In one embodiment, the second block instruction can be configured as follows: according to the Ci dimension size of the integer segment, an input tensor bias and an output tensor bias of the second block instruction executed for the remainder segment are set, wherein the input tensor bias represents the offset of the remainder segment before processing relative to the starting storage address of the input data, and the output tensor bias represents the offset of the remainder segment after processing relative to the starting storage address of the output data. By setting the input tensor bias and the output tensor bias of the second block instruction, the input tensor address and the output tensor address of the second block instruction can be adjusted after taking into account the memory storage space of the integer segment portion.

[0273] In one embodiment, the parameters of the input tensor of the second block instruction can be configured as follows: the inner lowest dimension size in0 of the input tensor in the second block instruction is set to R, where R is the Ci dimension size of the remainder segment, and the inner lowest dimension size in1 is set to U W , set the inner highest dimension size in2 to U H ; and according to the size of each dimension of the remainder segment of the input data, set the size values ​​in3, in4 and in5 of the three outer dimensions of the input tensor in the second block instruction, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks containing the input tensor on the corresponding dimensions.

[0274] Additionally, in one embodiment, the parameters of the output tensor of the second block instruction can be configured as follows: the inner lowest dimension size on0 of the output tensor in the second block instruction is set to in1*in2*T, the inner lowest dimension size on1 is set to R / T, and the inner highest dimension size on2 is set to 1; and according to the size of each dimension of the remainder segment of the input data, the size values ​​on3, on4 and on5 of the three outer dimensions of the output tensor in the second block instruction are set, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks containing the output tensor on the corresponding dimensions.

[0275] Similarly, the control circuit can set the step size of adjacent data points in the storage space in the five dimensions except the lowest inner dimension based on the dimensional sizes of the six dimensions of the set input tensor and output tensor and the dimensional sizes of the input data before processing.

[0276] In one example, when the Forward4 solution is used, the split unit is U Ci ×U H ×U W =4B×4×4, M=64B, the second block instruction for the remainder segment can be configured according to the following Table 8.

[0277]

[0278]

[0279] Table 8. Parameter allocation scheme for neuron remainder segment block instructions

[0280] Wherein, ci, hi, and wi represent the number of data in the Ci, H, and W dimensions of the input data respectively, dwidth represents the data bit width, ci_rem represents the number of data in the Ci dimension of the remainder segment, B represents bytes, T represents the block bit width, is1~is5 represent the five-dimensional step size of the input tensor, and os1~os5 represent the five-dimensional step size of the output tensor.

[0281] Therefore, the disclosed embodiments provide a block processing solution for neuron data. When the neuron data has an arbitrary shape, a two-stage block processing method can be used to arrange the neuron data of arbitrary shape from [1*hi*wi*ci] to [1*hi / 4*wi / 4*ci / 4*(4*4*4)].

[0282] General scheme for block instructions of weights

[0283] The block processing of weight data is similar to that of neuron data. Specifically, for the weight data in the Forward4 scheme, the data needs to be arranged from [co*kh*hw*ci] to:

[0284] [co*kh / 4*kw / 4*ci / 4*(4*4*4)]

[0285] Unlike neuron data, weight data has an additional co dimension. Since the co dimension and the kh dimension are continuous, the co dimension can be merged into the kh dimension.

[0286] Table 9 shows the shape change of the weight data before and after executing the block instruction.

[0287]

[0288] Table 9. Shape changes before and after weight data block processing

[0289] Note that the shape assumptions in Table 9 are parameters that have been aligned according to the alignment restrictions in Table 3 regarding different Forward4 Group modes and different H*W splitting methods.

[0290] By incorporating the co dimension into the kh dimension, the weight data can still be processed in blocks using the scheme described above for neuron data. Specifically, in some embodiments, a two-stage processing scheme can be used for weight data of any size, namely integer segment block processing and remainder segment block processing. Detailed block instruction configuration schemes can be found in the description above.

[0291] In one example, when the Forward4 solution is used, the split unit is U Ci ×U H ×U W =4B×4×4, M=64B, the first block instruction for the integer segment of the weight data can be configured according to the following Table 10.

[0292]

[0293] Table 10. Weight integer segment block instruction parameter allocation scheme

[0294] Alternatively or additionally, in one example, when the Forward4 solution is used, the split unit is U Ci ×U H ×U W =4B×4×4, M=64B, the second block instruction for the remainder segment of the weight data can be configured according to the following Table 11.

[0295]

[0296] Table 11. Weight remainder segment block instruction parameter allocation scheme

[0297] The disclosed embodiments also provide a data processing method for executing block instructions using the aforementioned data processing device. Those skilled in the art will appreciate that the steps of the method for executing block instructions correspond to the various features of the computing device described above in conjunction with the accompanying drawings. Therefore, the features described above also apply to the steps of the method and will not be repeated here.

[0298] The present disclosure also provides a chip, which may include the data processing device of any of the embodiments described above in conjunction with the accompanying drawings. Furthermore, the present disclosure also provides a board, which may include the aforementioned chip.

[0299] According to different application scenarios, the electronic devices or devices disclosed herein may include servers, cloud servers, server computing clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, driving recorders, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, automatic driving terminals, vehicles, household appliances, and / or medical equipment. The vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs. The electronic devices or devices disclosed herein may also be applied to the Internet, Internet of Things, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, medical care and other fields. Furthermore, the electronic devices or devices disclosed herein may also be used in cloud, edge, terminal and other application scenarios related to artificial intelligence, big data and / or cloud computing. In one or more embodiments, electronic devices or apparatuses with high computing power according to the disclosed solution can be applied to cloud devices (such as cloud servers), while electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or edge device are compatible with each other, so that according to the hardware information of the terminal device and / or edge device, appropriate hardware resources can be matched from the hardware resources of the cloud device to simulate the hardware resources of the terminal device and / or edge device, so as to complete the unified management, scheduling and collaborative work of end-to-end or cloud-edge-to-end.

[0300] It should be noted that, for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art will understand that the scheme of the present disclosure is not limited by the order of the actions described. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art will understand that some of the steps therein can be performed in other orders or simultaneously. Further, those skilled in the art will understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved therein are not necessarily necessary for the implementation of one or more schemes of the present disclosure. In addition, depending on the different schemes, the description of some embodiments of the present disclosure also has different emphases. In view of this, those skilled in the art will understand that the parts that are not described in detail in a certain embodiment of the present disclosure may also refer to the relevant descriptions of other embodiments.

[0301] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented in other ways not disclosed herein. For example, with respect to the various units in the electronic device or device embodiments described above, this article splits them based on the consideration of logical functions, and there may be other ways of splitting them in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions in the units or components can be selectively disabled. With respect to the connection relationship between different units or components, the connection discussed above in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.

[0302] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network elements. In addition, according to actual needs, some or all of the units may be selected to achieve the purpose of the solution described in the embodiments of this disclosure. In addition, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically separately.

[0303] In some other implementation scenarios, the above-mentioned integrated units can also be implemented in the form of hardware, that is, as specific hardware circuits, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include but is not limited to physical devices, and the physical devices may include but are not limited to devices such as transistors or memristors. In view of this, the various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM and RAM, etc.

[0304] The above is a detailed introduction to the embodiments of the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, for those skilled in the art, based on the ideas of the present disclosure, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

Claims

1. A data processing device comprising a control circuit, a first storage circuit, and a second storage circuit, wherein: The first storage circuit is used to store data before processing; The second storage circuit is used to store processed data; as well as The control circuit is used to configure and execute a block instruction to split the input data stored in the first storage circuit according to the first dimension storage order into split units and store the split data as output data on the second storage circuit, wherein on the second storage circuit, the data is stored in each split unit according to the second dimension storage order, and the data is stored between the split units according to the third dimension storage order; The amount of data contained in the split unit is set to the one-time processing alignment value of the hardware; the input data is neuron data including three dimensions of H, W, and C, the storage order of the first dimension is HWC, the storage order of the second dimension is CHW, and the storage order of the third dimension is HWC, where C represents the input channel dimension, H represents the height dimension, and W represents the width dimension.

2. The data processing apparatus according to claim 1 , wherein the control circuit is further configured to configure the block instruction as follows: A post-matching table of the block instruction is set to rearrange the inner lowest dimensional data of the output tensor of the block instruction according to the instruction of the post-matching table.

3. The data processing device according to claim 2, wherein the control circuit is further configured to set the post-matching table as follows: The inner lowest dimensional data of the output tensor arranged in the first dimensional storage order is converted into the inner lowest dimensional data arranged in the second dimensional storage order.

4. The data processing apparatus according to claim 2, wherein the control circuit is further configured to: Divide the input data into an integer segment and a remainder segment according to the input channel Ci dimension, wherein the Ci dimension of the integer segment is aligned to the alignment value M, and the Ci dimension of the remainder segment is smaller than the M; configuring and executing a first blocking instruction for the integer segment; and A second block instruction is configured and executed for the remainder segment.

5. The data processing device according to claim 4, wherein the shape of the split unit is CHW=U Ci ×U H ×U W , where C represents the input channel dimension, H represents the height dimension, W represents the width dimension, and M is U Ci multiples of . The data processing apparatus according to claim 5 , wherein: When the integer segment exists, the control circuit is further configured to configure a first block instruction for the integer segment as follows: Set the inner lowest dimensional size in0 of the input tensor in the first block instruction to M, and the inner lowest dimensional size in1 to U W , the inner highest dimension size in2 is set to U H ;as well as According to the sizes of each dimension of the integer segment of the input data, the size values ​​in3, in4 and in5 of the three outer dimensions of the input tensor in the first block instruction are set, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks contained in the corresponding dimensions.

7. The data processing apparatus according to claim 6, wherein the control circuit is further configured to configure the first blocking instruction for the integer segment as follows: The inner lowest dimension size on0 of the output tensor in the first block instruction is set to in1*in2*T, the inner lowest dimension size on1 is set to M / T, and the inner highest dimension size on2 is set to 1, where T is the block width of the block instruction, which represents the amount of data for one atomic operation of the block instruction; and According to the size of each dimension of the integer segment of the input data, the size values ​​on3, on4 and on5 of the three outer dimensions of the output tensor in the first block instruction are set, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks contained in the corresponding dimensions.

8. The data processing apparatus according to claim 7, wherein the control circuit is further configured to configure the first blocking instruction for the integer segment as follows: Based on the size of the six dimensions of the input tensor and output tensor and the size of each dimension of the input data before processing, set the stride of adjacent data points in the storage space in the five dimensions except the lowest inner dimension.

9. The data processing apparatus according to claim 8, wherein the control circuit is further configured to configure the first blocking instruction for the integer segment as follows: According to the constraints of the block instruction and the shape transformation of the split unit before and after processing, the block width T is set to U Ci .

10. The data processing apparatus according to claim 4, wherein: When the remainder segment exists, the control circuit is further configured to configure the second block instruction as follows: According to the Ci dimension size of the integer segment, the input tensor bias and output tensor bias of the second block instruction executed on the remainder segment are set, wherein the input tensor bias represents the offset of the remainder segment before processing relative to the starting storage address of the input data, and the output tensor bias represents the offset of the remainder segment after processing relative to the starting storage address of the output data.

11. The data processing apparatus according to claim 10 , wherein the control circuit is further configured to configure a second blocking instruction for the remainder segment as follows: Set the inner lowest dimensional size in0 of the input tensor in the second block instruction to R, where R is the Ci dimension size of the remainder segment, and set the inner low dimensional size in1 to U W , set the inner highest dimension size in2 to U H ;as well as According to the size of each dimension of the remainder segment of the input data, the size values ​​in3, in4 and in5 of the three outer dimensions of the input tensor in the second block instruction are set, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks contained in the corresponding dimensions.

12. The data processing apparatus according to claim 11 , wherein the control circuit is further configured to configure a second blocking instruction for the remainder segment as follows: Set the inner lowest dimension size on0 of the output tensor in the second block instruction to in1*in2*T, the inner lowest dimension size on1 to R / T, and the inner highest dimension size on2 to 1; and According to the size of each dimension of the remainder segment of the input data, the size values ​​on3, on4 and on5 of the three outer dimensions of the output tensor in the second block instruction are set, where the size values ​​of the three outer dimensions respectively represent the number of inner data blocks contained in the corresponding dimensions.

13. The data processing apparatus according to claim 12, wherein the control circuit is further configured to configure a second blocking instruction for the remainder segment as follows: Based on the size of the six dimensions of the input tensor and output tensor and the size of each dimension of the input data before processing, set the stride of adjacent data points in the storage space in the five dimensions except the lowest inner dimension.

14. The data processing apparatus according to claim 9, wherein when U Ci ×U H ×U W =4B×4×4, M=64B, the parameters of the first block instruction are configured according to Table 1: Table 1 Where ci, hi, and wi represent the number of data in the Ci, H, and W dimensions of the input data, respectively; dwidth represents the data bit width; ci_full represents the number of data in the Ci dimension of the integer segment; B represents bytes; T represents the block bit width; is1 to is5 represent the five-dimensional step size of the input tensor; and os1 to os5 represent the five-dimensional step size of the output tensor.

15. The data processing apparatus according to claim 13, wherein when U Ci ×U H ×U W =4B×4×4, M=64B, the parameters of the second block instruction are configured according to Table 2: Table 2 Wherein, ci, hi, and wi represent the number of data in the Ci, H, and W dimensions of the input data respectively, dwidth represents the data bit width, ci_rem represents the number of data in the Ci dimension of the remainder segment, B represents bytes, T represents the block bit width, is1~is5 represent the five-dimensional step size of the input tensor, and os1~os5 represent the five-dimensional step size of the output tensor.

16. A data processing device according to any one of claims 1 to 15, wherein the input data is weight data including four dimensions: Ci, Co, Kh, and Kw, wherein the Co dimension and the Kh dimension are merged into the H dimension, the Ci dimension corresponds to the C dimension, and the Kw dimension corresponds to the W dimension. The storage order of the first dimension is HWC, the storage order of the second dimension is CHW, and the storage order of the third dimension is HWC, wherein Ci represents the weight input channel dimension, Co represents the weight output channel dimension, Kh represents the weight height dimension, Kw represents the weight width dimension, C represents the input channel dimension, H represents the height dimension, and W represents the width dimension.

17. A chip comprising the data processing device according to any one of claims 1 to 16.

18. A board comprising the chip according to claim 17.

19. A data processing method comprising executing a block instruction on input data using the data processing device according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Operation method and device, and related product

    CN111695682A