Ai processor, data processing method, and computer device

By storing the operation data in the target data format in the first memory of the AI ​​processor, the problem of low bus bandwidth utilization in the prior art is solved, and more efficient matrix operation is achieved.

WO2025107800A1PCT designated stage expired Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/115871
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-08-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing AI processors perform matrix operations, the bus bandwidth utilization rate is low, resulting in low matrix computing efficiency.

Method used

By storing operation data in the target data format in the first memory, the data dimensions match the array dimensions of the pulsating array in the matrix computing engine, so that the data handling engine can carry data according to the bus bit width.

Benefits of technology

This improves the bandwidth utilization rate of data handling during matrix computing, thereby improving the efficiency of matrix computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115871_30052025_PF_FP_ABST
    Figure CN2024115871_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application belong to the technical field of processors. Disclosed are an AI processor, a data processing method, and a computer device. The AI processor comprises: a first memory, which is used for storing first operation data and second operation data; a data transfer engine, which is used for reading first operation sub-data and second operation sub-data from the first memory on the basis of a bus width, and transferring by means of a bus the first operation sub-data and the second operation sub-data into a second memory in a matrix operation engine; and the matrix operation engine, which is used for performing, by means of a matrix operation unit, matrix operation on the first operation sub-data and the second operation sub-data which are in the second memory, so as to obtain sub-operation results, wherein the matrix operation engine is further used for accumulating the sub-operation results by means of an accumulator, so as to obtain a matrix operation result of the first operation data and the second operation data. By adopting the solution provided in the embodiments of the present application, the bandwidth utilization rate during data transfer is increased.
Need to check novelty before this filing date? Find Prior Art

Description

AI processor, data processing method, and computer device

[0001] This application claims priority to the Chinese patent application filed on November 22, 2023, with application number 202311579613.0 and invention name “AI processor, data processing method and computer device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of processor technology, and in particular to an AI processor, a data processing method, and a computer device. Background Art

[0003] The core computations in deep learning algorithms primarily include convolution and matrix computations. Convolution can be equivalent to matrix computations. Therefore, accelerating matrix computations can accelerate deep learning algorithms.

[0004] In related technologies, matrix calculation speeds are increased by implementing a matrix acceleration engine, and input data is moved based on the array dimensions of the systolic array within the matrix acceleration engine. Because the data dimensions represented by the data formats used for different input data often do not match the array dimensions of the systolic array, direct data movement based on the systolic array dimensions often results in low bus bandwidth utilization during bus data transfer.

[0005] Summary of the Invention

[0006] The present invention provides an AI processor, a data processing method, and a computer device. The technical solution is as follows:

[0007] In one aspect, an embodiment of the present application provides an AI processor, comprising: a matrix operation engine, a data handling engine, and a first memory, wherein the matrix operation engine and the data handling engine are connected via a bus;

[0008] The first memory is used to store first operation data and second operation data, wherein the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine;

[0009] The data transfer engine is configured to read the first operator data and the second operator data from the first memory based on a bus bit width, and transfer the first operator data and the second operator data to the second memory within the matrix operation engine via the bus, wherein the bus bit width is greater than the data bit width corresponding to the array dimension;

[0010] The matrix operation engine is configured to perform a matrix operation on the first operator data and the second operator data in the second memory through a matrix operation unit to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension;

[0011] The matrix operation engine is further configured to accumulate the sub-operation results through an accumulator to obtain a matrix operation result of the first operation data and the second operation data.

[0012] On the other hand, an embodiment of the present application provides a data processing method for an AI processor, wherein the AI ​​processor includes a matrix operation engine, a data transfer engine, and a first memory, wherein the matrix operation engine and the data transfer engine are connected via a bus; the method includes:

[0013] storing first operation data and second operation data in the first memory, wherein the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine;

[0014] Based on a bus bit width, reading first operator data and second operator data from the first memory by the data transfer engine, and transferring the first operator data and the second operator data to a second memory within the matrix operation engine via the bus, wherein the bus bit width is greater than a data bit width corresponding to the array dimension;

[0015] performing a matrix operation on the first operator data and the second operator data in the second memory by a matrix operation unit in the matrix operation engine to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension;

[0016] The sub-operation results are accumulated by an accumulator in the matrix operation engine to obtain a matrix operation result of the first operation data and the second operation data.

[0017] On the other hand, an embodiment of the present application provides a computer device, which includes a memory and an AI processor as described in the above aspects, wherein the memory stores at least one instruction, and the at least one instruction is used to be executed by the AI ​​processor.

[0018] In an embodiment of the present application, by storing the first operation data and the second operation data in the target data format in the first memory, the data dimensions of the first operation data and the second operation data match the array dimensions of the systolic array in the matrix operation unit in the matrix operation engine, so that the data transfer engine can read the first operator data and the second operator data from the first memory according to the bus bit width, and transfer the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and then the matrix operation engine performs a matrix operation on the first operator data and the second operator data through the matrix operation unit to obtain a sub-operation result, and accumulates the sub-operation results through the accumulator to obtain a matrix operation result of the first operation data and the second operation data. Using the AI ​​processor provided by the embodiment of the present application, by storing the operation data in the target data format in the first memory, it is possible to read the operation data according to the bus bit width. Since the bus bit width is greater than the data bit width corresponding to the array dimension, compared with reading data according to the data bit width corresponding to the array dimension, the AI ​​processor provided by the embodiment of the present application can improve the bandwidth utilization of data transfer during the matrix operation, thereby improving the efficiency of the matrix operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] FIG1 shows a schematic structural diagram of an AI processor provided by an exemplary embodiment of the present application;

[0020] FIG2 shows a schematic diagram of a systolic array operation provided by an exemplary embodiment of the present application;

[0021] FIG3 is a schematic diagram showing the logical expression of convolutional network data and the physical storage of data in different formats in related technologies;

[0022] FIG4 is a schematic diagram showing a physical storage form of a first target data format provided by an exemplary embodiment of the present application;

[0023] FIG5 is a schematic diagram showing data processing by a systolic array based on a general data format in the related art;

[0024] FIG6 is a schematic diagram showing data processing by a systolic array based on a target data format according to an exemplary embodiment of the present application;

[0025] FIG7 is a schematic diagram showing data processing by a systolic array based on a general data format in another related art;

[0026] FIG8 is a schematic diagram showing data processing by a systolic array based on a target data format provided by another exemplary embodiment of the present application;

[0027] FIG9 shows a schematic diagram of data format conversion provided by an exemplary embodiment of the present application;

[0028] FIG10 shows a flow chart of a data processing method provided by an exemplary embodiment of the present application;

[0029] FIG11 shows a structural block diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0031] It should be understood that the term "several" in this document refers to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exists simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0032] For ease of understanding, the nouns involved in the embodiments of this application are explained below.

[0033] Operational data: refers to the data involved in the operation. In the embodiments of the present application, the operational data is used for convolution operations or matrix multiplication operations. In the convolution operation process, the operational data includes the convolution input data and the convolution weight data, which can also be called the convolution kernel data; in the matrix multiplication process, the operational data includes the left matrix input data and the right matrix input data.

[0034] Pulse Sequence: The array in the Matrix Arithmetic Logic Unit (MALU) that implements matrix multiplication. The core concept of a systolic array is to partition the matrix into blocks and then perform the entire matrix multiplication through a series of data moves and local multiplications. The inputs to a systolic array include left and right inputs, with the number of rows denoted by MALU_K and the number of columns denoted by MALU_N.

[0035] Dimensionality: refers to the number of data in a certain dimension, such as the number of rows, columns, height, width, channels, etc. When the data format is HWC, if the data is 3×4×5, the dimension of the data in the H (height) dimension is 3, the dimension in the W (width) dimension is 4, and the dimension in the C (channel) dimension is 5.

[0036] Deep learning methods are widely used in various fields, including image processing, video processing, speech processing, and content generation. The core characteristics of deep learning are high computational complexity and a large number of parameters. Deep learning networks can be divided into three categories: Convolutional Neural Networks (CNNs), Transformers, and Vision Transformers (ViTs). The core computations in these three types of networks are primarily convolutional and matrix operations, which can be equivalent to matrix operations. Therefore, accelerating deep learning algorithms lies in accelerating matrix operations. How to perfectly integrate these core matrix operations with other commonly used vector operations is a key issue in AI processor design.

[0037] Optionally, the AI ​​processor can be a processor based on neural network algorithms and acceleration, such as a neural network processing unit (NPU).

[0038] In the related art, an AI processor is integrated into a general deep learning framework, and matrix operations are performed through the matrix operation engine in the AI ​​processor. The matrix operation unit in the matrix operation engine is generally structured in the form of a two-dimensional systolic array. Therefore, in order to use the matrix operation unit to perform data operations on the operation data, the related art generally transfers the operation data according to the array dimension of the systolic array in the matrix operation unit. However, for buses with a bus width greater than the data width corresponding to the array dimension, the bus bandwidth is often underutilized during data transfer, resulting in low bus bandwidth utilization.

[0039] In an embodiment of the present application, by storing the first operation data and the second operation data in the target data format (the data dimension represented matches the array dimension of the systolic array in the matrix operation unit) in the first memory, the first operator data and the second operator data can be directly read from the first memory according to the bus bit width through the data transfer engine, and the first operator data and the second operator data can be transferred to the second memory inside the matrix operation engine through the bus, thereby fully improving the utilization of the bus bit width and optimizing the data transfer process.

[0040] Please refer to Figure 1, which shows a structural diagram of an AI processor provided by an exemplary embodiment of the present application. The AI ​​processor 100 mainly includes a matrix operation engine 110, a data transfer engine 120 and a first memory 130, wherein the matrix operation engine 110 and the data transfer engine 120 are connected via a bus.

[0041] The first memory 130 is used to store the first operation data and the second operation data. The first operation data and the second operation data are in a target data format. The data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine.

[0042] In the related art, the AI ​​processor stores the computational data in the dimensional expression mode unique to the deep learning network through the first memory, resulting in insufficient bandwidth utilization during data transfer through the data transfer engine. For example, in PyTorch, the convolutional network data format is NCHW, and the weight format is CoCiKhKw; in TensorFlow, the convolutional network data format is NHWC, and the weight format is KhKwCiCo. Among them, N is the number of batches, C is the number of data channels, H is the height, W is the width, Kh is the height of the convolution kernel, Kw is the width of the convolution kernel, Ci is the number of input channels, and Co is the number of output channels.

[0043] In the embodiment of the present application, the first operation data and the second operation data in the first memory are both stored in the target data format. The original data formats of the first operation data and the second operation data may be the target data format, or the first operation data and the second operation data in the original data format may be format-converted and then stored in the first memory in the target data format.

[0044] Optionally, the architecture of the matrix operation engine is a multi-dimensional systolic array composed of multiple physical matrices of multiply accumulate operations (MACs), which is used to perform computational processing on a series of matrix operations of a convolutional neural network.

[0045] Optionally, the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine. This may mean that the data dimension of each data channel represented by the target data format is equal to the array dimension of the systolic array, or that the data dimension of some data channels among the data channels represented by the target data format is equal to the array dimension of the systolic array.

[0046] Optionally, the array dimension of the systolic array can be the number of array rows or the number of array columns. When the computational data is convolutional data, the data dimension represented by the target data format can be the number of data channels; when the computational data is matrix data, the data dimension represented by the target data format can be the number of matrix rows or the number of matrix columns.

[0047] For example, the number of data channels represented by the target data format is equal to the array dimension of the systolic array, or the number of data rows or columns represented by the target data format is equal to the array row number or column number of the systolic array.

[0048] Illustratively, the first operation data is matrix data, the number of matrix rows and the number of matrix columns of the matrix data are both 128, and the number of array rows of the systolic array is 64. Then, when the target data format is adopted, the number of matrix rows of the first operation data can be converted to 2×64, so that the number of matrix rows of the first operation data is equal to the number of array rows of the systolic array.

[0049] Optionally, the systolic array is in the form of a two-dimensional matrix, which can be a two-dimensional systolic, a one-dimensional systolic plus a one-dimensional broadcast, or a two-dimensional systolic and broadcast mixed structure.

[0050] The data transfer engine 120 is used to read the first operator data and the second operator data from the first memory 130 based on the bus bit width, and transfer the first operator data and the second operator data to the second memory 111 inside the matrix operation engine 110 through the bus. The bus bit width is greater than the data bit width corresponding to the array dimension.

[0051] In some embodiments, the bus width is an integer multiple of the data width corresponding to the array dimension. For example, the bus width is twice the data width corresponding to the array dimension. Of course, the bus width may not be an integer multiple of the data width corresponding to the array dimension. The following embodiments are merely illustrative of integer multiples, but are not intended to be limiting.

[0052] Optionally, the data transfer engine can be a Direct Memory Access (DMA) controller for transferring data between different memories, including an address bus, a data bus, and control registers. The data transfer engine can then read the first operand data and the second operand data from the first memory and transfer the first operand data and the second operand data to the second memory via the bus.

[0053] Due to the large amount of data being processed, the matrix operation engine cannot complete the operation in a single pass. Therefore, the operation data must be divided into sub-operation data and then operated on. This means that the operation process for the first and second operation data is split into multiple sub-operation data operations. Accordingly, the first and second operation data in the first memory must be moved to the second memory in segments.

[0054] Optionally, the first operator data belongs to the first operation data and is a part of the first operation data, and the second operator data belongs to the second operation data and is a part of the second operation data, that is, the AI ​​processor can realize the segmented transmission of the first operation data and the second operation data from the first memory to the second memory inside the matrix operation engine through the data transfer engine.

[0055] In the related art, when the data dimension represented by the data format used for the first operation data and the second operation data does not match the array dimension of the systolic array, in order not to affect the matrix operation process in the matrix operation unit, the AI ​​processor uses the data transfer engine to transfer data according to the array dimension, resulting in the sub-data transferred in adjacent time periods not having data continuity, which in turn leads to low bus bandwidth utilization when the bus bit width is greater than the data bit width corresponding to the array dimension.

[0056] In an embodiment of the present application, after storing the first operation data and the second operation data in the target data format, the AI ​​processor can use the data transfer engine to directly read the first operator data and the second operator data from the first memory according to the bus bit width, thereby fully utilizing the bus bit width and ensuring the continuity of data transfer.

[0057] Schematically, when the operation data does not adopt the target data format and the bus width is greater than the data width corresponding to the array dimension, in order to enable the pulse array of the matrix operation unit to directly perform matrix operations on the sub-operation data transported by the data transport engine, the data transport engine transports the data in sequence according to the data width corresponding to the array dimension of the pulse array and the number of data rows or columns. Since the data width of each data transported is less than the bus width, the bus width is not fully utilized. In the case of adopting the target data format, the operation data can be stored continuously in the physical storage in units of the array dimension of the pulse array. The data transport engine can directly transport the data according to the bus width, that is, the data width of the transported data is equal to the bus width, thereby fully utilizing the bus width.

[0058] In an illustrative example, when the bus width is twice the data width corresponding to the array dimension, if the target data format is not used to store the computational data, and the computational data is moved according to the data width corresponding to the array dimension, the bandwidth utilization rate is only 50%. However, after the computational data is stored in the target data format, the computational data can be stored continuously in the physical storage using the array dimension of the systolic array as a unit, so that the computational data can be moved according to the bus width, achieving a bandwidth utilization rate of 100%.

[0059] For example, when the bus width is 2048 bits, the array dimension of the systolic array is 64, and the data in each dimension are all fp16 (i.e., all are 16-bit floating-point numbers), the data width corresponding to the array dimension of the systolic array is 1024 bits. That is, the bus width (2048 bits) is greater than the data width corresponding to the array dimension of the systolic array (1024 bits). If the target data format is not used to store the computational data, the computational data is transferred according to the data width corresponding to the array dimension, occupying only 1024 bits of bus width each time, resulting in a bandwidth utilization rate of only 50%. However, if the target data format is used to store the computational data, the computational sub-data having a data width of 2048 bits equal to the bus width can be read from the first memory according to the bus width.

[0060] In some embodiments, a second memory 111 is provided in the matrix operation engine 110 , and the data transfer engine 120 transfers the first operator data and the second operator data to the second memory 111 inside the matrix operation engine 110 during data transfer through the bus.

[0061] Optionally, the data transport process can be understood as transmitting the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus.

[0062] The matrix operation engine 110 is used to perform matrix operations on the first operator data and the second operator data in the second memory 111 through the matrix operation unit 112 to obtain sub-operation results, and the data dimension of the sub-operation results matches the array dimension.

[0063] In some embodiments, the matrix operation engine also includes a matrix operation unit. After storing the first operator data and the second operator data in the second memory, the matrix operation engine can use the matrix operation unit to perform matrix operations on the first operator data and the second operator data using a systolic array to obtain a sub-operation result.

[0064] The data dimension of the sub-operation result matches the array dimension of the systolic array. For example, the number of data rows of the sub-operation result is equal to the number of array rows of the systolic array, and the number of data columns of the sub-operation result is equal to the number of array columns of the systolic array.

[0065] In some embodiments, the matrix operation engine may use the first operator data as the left input data of the systolic array and the second operator data as the right input data of the systolic array, and perform a matrix operation through the matrix operation unit to obtain a sub-operation result.

[0066] Optionally, the matrix operation unit processes the first operator data and the second operator data through a matrix multiplication operation, thereby obtaining a sub-operation result.

[0067] Schematically, as shown in Figure 2, the resultant data output by a matrix operation unit is in the form of a two-dimensional matrix [Matrix_M, MALU_N]. Taking a calculation based on a systolic array as an example, after determining the left input data 21 and right input data 22 of the systolic array, the matrix operation unit performs data operations from left to right and from top to bottom based on the properties of the systolic array, thereby outputting a set of data on the MALU_N side. The PEs (Process Elements) in the systolic array are the smallest units of the array and are used to implement one-dimensional multiplication and addition calculations.

[0068] The matrix operation engine 110 is further configured to accumulate the sub-operation results through the accumulator 113 to obtain a matrix operation result of the first operation data and the second operation data.

[0069] In some embodiments, the matrix operation engine also includes an accumulator (ACC). After performing matrix operations through the matrix operation unit to obtain a large number of sub-operation results, the matrix operation engine can also accumulate the sub-operation results through the accumulator to obtain the matrix operation results of the first operation data and the second operation data.

[0070] Optionally, the accumulator's accumulation process of sub-operation results refers to temporarily storing the sub-operation results output by the matrix operation unit, and after obtaining all the sub-operation results, performing data splicing on each sub-operation result according to the matrix element position of each sub-operation result in the operation result matrix, thereby obtaining the operation results of the first operation data and the second operation data.

[0071] Schematically, the first operation data and the second operation data are both 2×2 matrix data. By performing matrix operation through the matrix operation unit, the following can be obtained: a first sub-operation result corresponding to the first row matrix data of the first operation data and the first column matrix data of the second operation data, a second sub-operation result corresponding to the first row matrix data of the first operation data and the second column matrix data of the second operation data, a third sub-operation result corresponding to the second row matrix data of the first operation data and the first column matrix data of the second operation data, and a fourth sub-operation result corresponding to the second row matrix data of the first operation data and the second column matrix data of the second operation data. Therefore, the matrix element position of the first sub-operation data in the operation result matrix is ​​the first row and first column, the matrix element position of the second sub-operation data in the operation result matrix is ​​the first row and second column, the matrix element position of the third sub-operation data in the operation result matrix is ​​the second row and first column, and the matrix element position of the fourth sub-operation data in the operation result matrix is ​​the second row and second column. Then, the accumulator can splice the sub-operation data according to the matrix element positions corresponding to each sub-operation data, thereby obtaining the matrix operation results of the first operation data and the second operation data.

[0072] In some embodiments, after obtaining the calculation results of the first calculation data and the second calculation data, the AI ​​processor may further utilize a data transfer engine to transfer the calculation results to the first memory via a bus.

[0073] In summary, in an embodiment of the present application, by storing the first operation data and the second operation data in the target data format in the first memory, the data dimensions of the first operation data and the second operation data match the array dimensions of the systolic array in the matrix operation unit in the matrix operation engine, so that the data transfer engine can read the first operator data and the second operator data from the first memory according to the bus bit width, and transfer the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and then the matrix operation engine performs a matrix operation on the first operator data and the second operator data through the matrix operation unit to obtain a sub-operation result, and accumulates the sub-operation results through the accumulator to obtain a matrix operation result of the first operation data and the second operation data. Using the AI ​​processor provided by the embodiment of the present application, by storing the operation data in the target data format in the first memory, it is possible to read the operation data according to the bus bit width. Since the bus bit width is greater than the data bit width corresponding to the array dimension, compared with reading data according to the data bit width corresponding to the array dimension, the AI ​​processor provided by the embodiment of the present application can improve the bandwidth utilization of data transfer during the matrix operation process, thereby improving the efficiency of the matrix operation.

[0074] In some embodiments, a deep learning network may include a convolutional network for performing operations on convolutional data, or a matrix network for performing operations on matrix data. Therefore, after integrating an AI processor into the deep learning network framework, the AI ​​processor needs to process the convolutional data and matrix data separately to ensure that both conform to the target data format.

[0075] The following describes the target data formats used in convolution and matrix operation scenarios.

[0076] In a possible implementation, during the process of processing convolution data, the first operation data may be convolution input data, the second operation data may be convolution weight data, and the operation result of the first operation data and the second operation data may be convolution output data.

[0077] Optionally, the convolution input data can be an image or feature map to be convolved, and the convolution weight data can be a convolution kernel used to convolve the image or feature map. Typically, the size of the convolution input data is smaller than the size of the convolution weight data. For example, the convolution input data is a 100×100 matrix, and the convolution weight data is a 3×3 convolution kernel.

[0078] In some embodiments, in the expression of dimensional information of a convolutional network, H represents the height of an image or feature map, W represents the width of an image or feature map, C represents the feature dimension, and N represents a batch of images or feature maps.

[0079] Schematically, as shown in Figure 3, it shows the logical expression of convolutional network data and the physical storage forms of data in different formats in the related art. Among them, the logical expression forms corresponding to NCHW format data, NHWC format data, and CHWN format data are the same, that is, the data arrangement in the C direction, H direction, and W direction is the same. In the physical storage process, for NCHW format data, the W direction is taken first, followed by the H direction, then the C direction, and finally the N direction; for NHWC format data, the C direction is taken first, followed by the W direction, then the H direction, and finally the N direction; for CHWN format data, the N direction is taken first, followed by the W direction, then the H direction, and finally the C direction.

[0080] In related art, taking NHWC format data as an example, if the number of array rows in a systolic array, Ck0, is 3, data needs to be moved in groups of 000, 020, and 040, and in groups of 001, 021, and 041. Physically, these data values, 000, 020, and 040, and 001, 021, and 041, are stored discretely. Therefore, the data moving engine can only move one group of data at a time. Even if the bus bandwidth is greater than the data bit width corresponding to the number of array rows, the data moving engine can only read three data values ​​at a time.

[0081] In some embodiments, when the first operation data is convolution input data, in order to fully utilize bus bandwidth during data transfer of the convolution input data, the convolution input data can be stored in a first target data format, where the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array.

[0082] In one possible implementation, the number of array rows of the systolic array can be expressed as Ck0, and the convolution input data in the first target data format can be expressed as NCi1HiWiCk0, where N represents the number of groups of images or feature maps, Hi represents the height of the input image or feature map, Wi represents the width of the input image or feature map, Ci1 is the value of Ci / Ck0 rounded up, and Ci is the total number of input channels (as shown in FIG4 ).

[0083] Schematically, as shown in Figure 4, it shows the physical storage form of the first target data format provided by an exemplary embodiment of the present application. As can be seen from Figure 4, for the target data format NCi1HiWiCk0, the data is arranged in the following manner: first in the Ck0 direction, then in the Wi direction, then in the Hi direction, then in the Ci1 direction, and finally in the N direction.

[0084] For example, when the number of array rows of the systolic array Ck0 = 3, the three values ​​000, 020, and 040 are first taken in the Ck0 direction, and then 001, 021, and 041 are taken from the Wi direction. After the three values ​​in the Wi direction are taken (that is, 003, 023, and 043 are taken), the data is transferred from 043 to 004 on Hi, and the data is stored in this order, thereby ensuring that during the data transfer based on the number of array rows of the systolic array, the data transferred in adjacent times are in a continuous state.

[0085] In some embodiments, when the second operation data is convolution weight data, in order to fully utilize bus bandwidth during data transfer of the convolution weight data, the convolution weight data may be stored in a second target data format. The number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array.

[0086] In one possible implementation, the number of array rows of the systolic array can be represented as Ck0, the number of array columns of the systolic array can be represented as Cn0, and the convolution weight data in the second target data format can be represented as Co1Ci1KhKwCk0Cn0, where Kh represents the convolution kernel height, Kw represents the convolution kernel width, Ci1 is the value of Ci / Ck0 rounded up, Co1 represents the value of Co / Cn0 rounded up, Ci is the total number of input channels, and Co is the total number of output channels.

[0087] In some embodiments, when the operation result of the first operation data and the second operation data is convolution output data, a matrix operation unit may perform a matrix operation based on the convolution input data in the first target data format and the convolution weight data in the second target data format to obtain convolution output data in a third target data format, wherein the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.

[0088] In one possible implementation, the number of array columns of the systolic array may be represented by Cn0, and the convolution output data in the third target data format may be represented as NCo1HoWoCn0, where N represents the number of groups of images or feature maps, Ho represents the height of the output image or feature map, Wo represents the width of the output image or feature map, Co1 represents the value of Co / Cn0 rounded up, and Co is the total number of output channels.

[0089] In some embodiments, after the convolution input data in the first target data format and the convolution weight data in the second target data format are stored in the first memory of the AI ​​processor, the AI ​​processor can use the data transfer engine to transfer the convolution input sub-data and the convolution weight sub-data according to the bus bit width, and store them in the second memory inside the matrix operation engine.

[0090] In one possible implementation, the AI ​​processor utilizes a data transfer engine to sequentially read the first convolution input sub-data and the second convolution input sub-data from the first memory according to the bus bit width, wherein the first convolution input sub-data and the second convolution input sub-data are continuously stored data, and the data bit width of the first convolution input sub-data is equal to the bus bit width, and the data bit width of the second convolution input sub-data is equal to the bus bit width.

[0091] Schematically, as shown in Figure 4, when the systolic array has three rows and the bus width is twice the data width corresponding to the number of rows, the AI ​​processor, using the data transfer engine, can simultaneously transfer the six data items 000, 020, 040, 001, 021, and 041 as the first convolution input sub-data, based on the bus width. Furthermore, when continuing to transfer data based on the number of rows in the systolic array, the AI ​​processor can continue to transfer the six data items 002, 022, 042, 003, 023, and 043 as the second convolution input sub-data. This shows that the first and second convolution input sub-data are stored consecutively in the physical storage data arrangement.

[0092] In one possible implementation, the AI ​​processor utilizes a data transfer engine to sequentially read the first convolution weight sub-data and the second convolution weight sub-data from the first memory according to the bus bit width, wherein the first convolution weight sub-data and the second convolution weight sub-data are continuously stored data, and the data bit width of the first convolution weight sub-data is equal to the bus bit width, and the data bit width of the second convolution weight sub-data is equal to the bus bit width.

[0093] Schematically, Figure 5 shows a schematic diagram of data processing using a systolic array based on a general data format in the related art. Taking the use of a data handling engine to handle convolution input data as an example, when the AI ​​processor uses the data handling engine to handle data, in order to obtain a data value of length MALU_K from the convolution input data, it is necessary to read data from the convolution input data in discrete segments and sequentially handle the data through the bus. This results in low bandwidth utilization when the bus bandwidth is greater than MALU_K.

[0094] Schematically, taking the logical expression form corresponding to the convolution input data in Figure 5 as an example, when the data arrangement is in the NHiWiCi format, it includes the Wi direction 501, the Hi direction 502 and the Ci direction 503, and the data is first continuous in the Ci direction 503 during the physical storage process. When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, in the related art, the convolution input data is directly read in the Ci direction 503 according to the number of array rows 504 of MALU_K, so that a data block 505 (height is Cut_Hi, width is Cut_Wi, number of channels is MALU_K) can be obtained. Each segment of data in the MALU_K direction in the data block 505 is discretely distributed in the physical storage of the convolution input data, and when the bus bandwidth is greater than MALU_K, the data handling process cannot fully utilize the bus bandwidth.

[0095] Schematically, taking the logical expression form corresponding to the convolution weight data in Figure 5 as an example, when the data arrangement is in the KhKwCiCo format, including the Ci direction 506 and the Co direction 507, the data is first continuous in the Co direction 507 during the physical storage process. When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, in the related art, the convolution weight data is directly read in the Ci direction 506 according to the number of array rows of MALU_K and in the Co direction 507 according to the number of array columns of MALU_N, so that each segment of data in the MALU_N direction in each data block 508 (with a height of MALU_K and a width of MALU_N) is discretely distributed in the physical storage of the convolution weight data, and when the bus bandwidth is greater than MALU_N, the bus bandwidth cannot be fully utilized during data transportation.

[0096] Schematically, as shown in Figure 6, it shows a schematic diagram of data processing by a systolic array based on a target data format provided by an exemplary embodiment of the present application. Taking the use of a data handling engine to handle convolution input data as an example, during the process of using the data handling engine to handle data, in order to obtain a data value of length MALU_K from the convolution input data, since the data is stored continuously along the MALU_K direction and then along the Wi direction in the target data format, when the bus bandwidth is greater than MALU_K, the AI ​​processor can directly read the data values ​​greater than MALU_K along the MALU_K direction and then along the Wi direction, thereby fully utilizing the bus bandwidth.

[0097] Schematically, taking the logical expression corresponding to the convolution input data in Figure 6 as an example, when the data arrangement is the target data format NCi1HiWiCk0, it includes Wi direction 601, Hi direction 602 and Ci1 (Ci / Ck0) direction 603, and the data is first continuous in the Ci1 direction 603 during the physical storage process.

[0098] When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, since the data is first continuous in the Ci1 direction 603 during the physical storage process, the embodiment of the present application can directly read the data block 604 (height is Cut_Hi, width is Cut_Wi, number of channels is MALU_K) continuously according to the bus bit width, thereby improving bus bandwidth utilization.

[0099] Schematically, taking the logical expression form corresponding to the convolution weight data in Figure 6 as an example, when the data is arranged in the target data format of Co1Ci1KhKwCk0Cn0, including the Ci1 (Ci / Ck0) direction 605 and the Co1 (Co / Cn0) direction 606, the data is first continuous in the Co1 direction 606 during the physical storage process. When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, since the data is first continuous in the Co1 direction 606 during the physical storage process, the embodiment of the present application can directly read the data block 607 (height is MALU_K and width is MALU_N) continuously according to the bus bit width, thereby improving the bus bandwidth utilization.

[0100] In some embodiments, considering that the channel dimensions of convolution data in different deep learning networks may be different, and the total number of channels of the convolution data may be less than the number of array rows or array columns of the systolic array, for convolution input data with a total input channel number less than the number of array rows of the systolic array, or for convolution output data with a total output channel number less than the number of array columns of the systolic array, different target data formats need to be used to store the convolution data.

[0101] In one possible implementation, when the total number of input channels of the convolution input data is less than the number of array rows of the systolic array (i.e., Ci / Ck0 is less than 1), the convolution input data may be stored using a fourth target data format, where the number of input channels represented by the fourth target data format is the total number of input channels.

[0102] In one possible implementation, the number of array rows of the systolic array can be represented as Ck0, and the total number of input channels of the convolution input data is Ci, where Ci is less than Ck0. The convolution input data can be represented as NHiWiCi, where N represents the number of groups of images or feature maps, Hi represents the height of the input image or feature map, and Wi represents the width of the input image or feature map.

[0103] In one possible implementation, when the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, and the total number of output channels of the convolution output data is not less than the number of array columns of the systolic array, the convolution weight data may be stored using a fifth target data format, wherein the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array. When the fourth and fifth target data formats are used to store the convolution input data and convolution weight data, additional data padding can be avoided for the convolution input data and the convolution weight data.

[0104] In one possible implementation, the number of rows of the systolic array can be represented as Ck0, the number of columns of the systolic array can be represented as Cn0, Ci is the total number of input channels, Co is the total number of output channels, and Ci is less than Ck0, while Co is greater than Cn0. The convolution weight data can be represented as Co1KwCiKhCn0, where Kh represents the convolution kernel height, Kh represents the convolution kernel width, and Co1 represents the value of Co / Cn0 rounded up.

[0105] In one possible implementation, when the total number of output channels of the convolution output data is less than the number of array columns of the systolic array (i.e., Co / Cn0 is less than 1), the convolution output data may be stored using a sixth target data format, where the number of output channels represented by the sixth target data format is the total number of output channels.

[0106] In one possible implementation, the number of array columns of the systolic array can be represented as Cn0, and the total number of output channels of the convolution output data is Co, where Co is less than Cn0. The convolution output data can be represented as NHoWoCo, where N represents a set of images or feature maps, Ho represents the height of the output image or feature map, and Wo represents the width of the output image or feature map.

[0107] In one possible implementation, when the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, and the total number of input channels of the convolution input data is not less than the number of array rows of the systolic array, the convolution weight data may be stored using a seventh target data format, wherein the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is the total number of output channels.

[0108] In one possible implementation, the number of array rows of the systolic array can be represented as Ck0, the number of array columns of the systolic array can be represented as Cn0, Ci is the total number of input channels, Co is the total number of output channels, and Ci is greater than Ck0, Co is less than Cn0, and the convolution weight data can be expressed as Ci1KhKwCk0ECn0, where Kh represents the height of the convolution kernel, Kw represents the width of the convolution kernel, Ci1 is the value of Ci / Ck0 rounded up, and W is the expansion coefficient in the Co-specific optimization algorithm, which is jointly determined by the computing power, bandwidth, and convolution parameters.

[0109] When the sixth target data format and the seventh target data format are used to store the convolution output data and the convolution weight data, additional data padding (padding) of the convolution weight data can be avoided during the operation. In the above embodiment, for convolution data, by comparing the total number of input channels of the convolution input data with the number of array rows of the systolic array, and the total number of output channels of the convolution output data with the number of array columns of the systolic array, the target data formats corresponding to the convolution input data, the convolution weight data, and the convolution output data are determined, thereby achieving different target data formats according to different data situations. Since the bus bit width is greater than the data bit width corresponding to the array dimension, compared to reading data according to the data bit width corresponding to the array dimension, for convolution data, data can be read directly according to the bus bit width, so that the data bit widths of the input sub-data and the weight sub-data are equal to the bus bit width, thereby optimizing data handling efficiency.

[0110] In a possible implementation, during the process of processing matrix data, the first operation data may be left matrix input data, the second operation data may be right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data.

[0111] In some embodiments, the dimensionality of matrix networks is expressed as follows: the left matrix is ​​represented as MK (M rows and K columns), the right matrix is ​​represented as KN (K rows and N columns), and the result matrix is ​​represented as MN (M rows and N columns). Furthermore, in commonly used deep learning networks, batch data of other dimensions is also added on top of this.

[0112] In some embodiments, when the first operation data is left matrix input data, in order to fully utilize bus bandwidth during data transfer of the left matrix input data, the left matrix input data may be stored in an eighth target data format, where the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array.

[0113] In one possible implementation, the number of array rows of the systolic array is Wk0, and the left matrix input data in the eighth target data format can be expressed as B0B1W1HWk0, where B0 and B1 represent the number of batches in two dimensions, respectively, H represents the number of matrix rows of the left matrix, W1 is the value of K / Wk0 rounded up, and K is the number of matrix columns of the left matrix.

[0114] In some embodiments, when the second operation data is right matrix input data, in order to fully utilize bus bandwidth during data transfer of the right matrix input data, the right matrix input data may be stored in a ninth target data format, wherein the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array.

[0115] In one possible implementation, the number of columns in the systolic array is Wn0. In the ninth target data format, the right matrix input data can be represented as B0B1W1HWn0, where B0 and B1 represent the number of batches in two dimensions, respectively. H represents the number of rows in the right matrix, W1 is N / Wn0 rounded up, and N is the number of columns in the right matrix.

[0116] It should be noted that the batches in the left matrix input data and the right matrix input data may not be equal, but when they are not equal, one of them must be 1.

[0117] In some embodiments, when the operation result of the first operation data and the second operation data is matrix output data, matrix output data using the tenth target data format can be obtained by performing matrix operations through a matrix operation unit based on the left matrix input data using the eighth target data format and the right matrix input data using the ninth target data format, wherein the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.

[0118] In one possible implementation, under the tenth target data format, the matrix output data can be expressed as B0B1W1HWn0, where B0 in the matrix output data is the maximum value of B0 in the left matrix input data and B0 in the right matrix input data, B1 in the matrix output data is the maximum value of B1 in the left matrix input data and B1 in the right matrix input data, W1 and Wn0 are the same as the W1 and Wn0 values ​​of the right matrix input data, and H is the same as the H value of the left matrix input data.

[0119] In some embodiments, after storing the left matrix input data in the eighth target data format and the right matrix input data in the ninth target data format in the first memory of the AI ​​processor, the AI ​​processor can use the data transfer engine to transfer the left matrix input sub-data and the right matrix input sub-data according to the bus bit width, and store them in the second memory inside the matrix operation engine.

[0120] In one possible embodiment, the AI ​​processor uses a data transfer engine to read the first left matrix input sub-data and the second left matrix input sub-data from the first memory in sequence according to the bus bit width, wherein the first left matrix input sub-data and the second left matrix input sub-data are continuously stored data, and the data bit width of the first left matrix input sub-data is equal to the bus bit width, and the data bit width of the second left matrix input sub-data is equal to the bus bit width.

[0121] In one possible embodiment, the AI ​​processor uses a data transfer engine to read the first right matrix input sub-data and the second right matrix input sub-data from the first memory in sequence according to the bus bit width, wherein the first right matrix input sub-data and the second right matrix input sub-data are continuously stored data, and the data bit width of the first right matrix input sub-data is equal to the bus bit width, and the data bit width of the second right matrix input sub-data is equal to the bus bit width.

[0122] Schematically, Figure 7 illustrates data processing using a systolic array based on a general data format in the related art. Taking the data handling engine for left-matrix input data as an example, the AI ​​processor, in order to obtain a data value of length MALU_K from the left-matrix input data, needs to discretely read data from the left-matrix input data along the M direction and sequentially handle the data through the bus. This results in low bandwidth utilization when the bus bandwidth exceeds MALU_K.

[0123] Schematically, taking the left matrix input data in Figure 7 as an example, when the data arrangement is in MK format, it includes M direction 701 and K direction 702, and the data is first continuous in the K direction 702 during physical storage (i.e., first stored along the K direction, then stored along the M direction). When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, in the related art, data is read directly from the K direction 702 according to the number of array rows of MALU_K. Each segment of data in the data block 703 in the MALU_K direction is discretely distributed in the physical storage of the left matrix input data (the data block 703 in Figure 7 spans multiple rows in the M direction, so the data in the data block 703 is discretely distributed), and when the bus bandwidth is greater than MALU_K, the problem of low bandwidth utilization will also arise during data handling. The same is true for the right matrix input data in Figure 7, and the embodiments of the present application will not be described in detail here.

[0124] Schematically, as shown in FIG8 , a schematic diagram illustrating data processing via a systolic array based on a target data format, according to another exemplary embodiment of the present application, is shown. Taking the data handling engine for left-matrix input data as an example, during data handling by the AI ​​processor using the data handling engine, in order to retrieve a segment of data values ​​of length MALU_K from the left-matrix input data, since the target data format stores data continuously along the MALU_K and then along the H_Left direction, if the bus bandwidth is greater than MALU_K, the AI ​​processor can directly read data values ​​greater than MALU_K along the MALU_K and then along the H_Left direction, thereby fully utilizing the bus bandwidth.

[0125] Schematically, taking the left matrix input data in FIG8 as an example, when the data arrangement is the target data format of B0B1W1HWk0, including the H direction 801 (vertical) and the W1(K / Wk0) direction 802 (horizontal), the data is first continuous in the W1 direction 802 during the physical storage process (i.e., first stored along the W1 direction, then stored along the H direction). When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, since the data is first continuous in the W1 direction 802 during the physical storage process, the embodiment of the present application can directly read the data block 803 continuously according to the bus bit width (although the data block 803 in FIG8 spans multiple rows in the H direction, since the data in the target data format is stored continuously in the H direction, the data in the data block 803 is continuous), thereby improving bus bandwidth utilization. The same is true for the right matrix input data in FIG8, and the embodiment of the present application will not be described in detail here.

[0126] In the above embodiment, for matrix data, the specific target data formats used for the left matrix input data, right matrix input data, and matrix output data are determined based on the number of matrix columns of the left matrix input data and the right matrix input data, in combination with the number of array rows and array columns of the systolic array. Because the bus width is greater than the data width corresponding to the array dimensions, data can be read directly based on the bus width, rather than reading data based on the data width corresponding to the array dimensions. This ensures that the data width of the input sub-data is equal to the bus width, thereby optimizing the efficiency of data transfer for matrix data.

[0127] In some embodiments, considering that in deep learning networks there are often scenarios where convolution operators and matrix operators are used alternately, and that convolution data and matrix data have differences in data storage formats, in order to improve the data processing efficiency in the AI ​​processor, a data format converter may also be provided in the AI ​​processor.

[0128] Optionally, the data format converter is used to perform unidirectional conversion of the data storage formats of convolution data and matrix data (if mutual conversion is to be achieved, data format converters with different conversion directions need to be set), or bidirectional conversion.

[0129] In one possible scenario, when it is necessary to first use the convolution kernel to convolve the feature map (convolution input data), and then use the convolution result as the left matrix input data to perform matrix multiplication with the right matrix input data, it is necessary to convert the convolution output data into the data format used by the matrix input data.

[0130] In a possible implementation, when the operation result of the first operation data and the second operation data is convolution output data, and the convolution output data is used for subsequent matrix calculations, the data format converter can be used to convert the data format of the convolution output data based on the target data format corresponding to the matrix input data.

[0131] In one possible implementation, the data format of the convolution output data is NC1HWC0, and the data format of the matrix input data is B0B1W1HW0. Then, B0 can be set to 1, B1 to N, W1 to C1, H to H×W, and W0 to C0, thereby converting the convolution output data into equivalent matrix input data.

[0132] In one possible scenario, when it is necessary to first perform matrix multiplication on two matrices (left matrix input data and right matrix input data) and then use the convolution kernel to convolve the matrix multiplication result, it is necessary to convert the matrix data into the data format used by the convolution input data.

[0133] In a possible implementation, when the operation result of the first operation data and the second operation data is matrix output data, and the matrix output data is used for subsequent convolution calculations, the data format converter can be used to convert the data format of the matrix output data based on the target data format corresponding to the convolution input data.

[0134] In one possible implementation, the data format of the matrix output data is B0B1W1HW0, and the data format of the convolution input data is NC1HWC0. Then, N can be set to B0×B1, C1 is set to W1, H is set to 1, W is set to H, and C0 is set to W0, thereby making the matrix output data equivalent to the convolution input data.

[0135] In the above embodiment, when the convolution output data or the matrix output data adopts the target data format and there are subsequent operations with different data types, the data equivalence relationship between the convolution output data and the matrix input data, or the data equivalence relationship between the matrix output data and the convolution input data can be determined, that is, the data equivalence between the convolution output data and the matrix input data, or the data equivalence between the matrix output data and the convolution input data can be achieved, thereby improving the data processing efficiency in the AI ​​processor.

[0136] In some embodiments, considering that after the AI ​​processor is integrated into the deep learning network framework, data calculation through the matrix operation engine is only one of the links, that is, the calculation data obtained based on the previous process may not be in the target data format, so directly moving the data may affect the efficiency of data movement. Therefore, a data format converter can also be provided in the AI ​​processor to convert the data format of the calculation data into the target data format.

[0137] In one possible embodiment, a data format converter is used to read first operation data and second operation data from a first memory. The first operation data and the second operation data are in an original data format. The data dimension represented by the original data format does not match the array dimension of the systolic array in the matrix operation unit in the matrix operation engine. Then, the data format converter converts the data format of the first operation data and the second operation data from the original data format to the target data format by performing data format conversion on the first operation data and the second operation data, and writes the first operation data and the second operation data into the first memory.

[0138] In an illustrative example, the data format converter first reads the convolution input data in the original data format (NCHW) from the first memory, and performs data format conversion on the convolution input data according to the target data format (NCi1HiWiCk0), thereby storing the converted convolution input data in the target data format back into the first memory.

[0139] It should be noted that before performing data format conversion, the data needs to be converted to the target data format corresponding to the user based on the data usage. For example, the convolution input data is converted to the first target data format, the convolution weight data is converted to the second target data format, the left matrix input data is converted to the eighth target data format, the right matrix input data is converted to the ninth target data format, and so on.

[0140] In addition, for convolution data, it is also necessary to determine the target data format used for the convolution input data, convolution weight data, and convolution output data based on the relationship between the total number of input channels of the convolution input data and the number of array rows of the pulse array, and based on the relationship between the total number of output channels of the convolution output data and the number of array columns of the pulse array. This embodiment will not be elaborated here.

[0141] In one possible implementation, when an AI processor is integrated into a general deep learning framework as an operation acceleration device, in order to reduce the complexity of integrating the AI ​​processor into the deep learning framework, corresponding format processing can also be provided at the operator layer and the framework layer. For example, at the operator layer, full support for the original data format and the target data format is added; at the framework layer, format information about the original data format and the target data format is provided.

[0142] Schematically, as shown in Figure 9, Math Dim (Math Dimention) is set to correspond to the calculation data in the original data format (universal format) in the algorithm framework, and Data Dim (Data Dimention) corresponds to the calculation data in the target data format (dedicated format) on the AI ​​processor. During data calculations using the deep learning framework, calculation data in the target data format can be directly used for data processing. During debugging, when transferring data from the device to the host, calculation data in the target data format can be converted to calculation data in the original data format based on the format information of the original data format and the target data format (i.e., format conversion is completed during transmission) for viewing and analysis.

[0143] Please refer to FIG10 , which shows a flow chart of a data processing method provided by an exemplary embodiment of the present application. The method is used in the AI ​​processor in the above embodiment, and the method includes:

[0144] Step 1001: First operation data and second operation data are stored in a first memory. The first operation data and the second operation data are in a target data format. The data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine.

[0145] Step 1002: Based on the bus bit width, the first operator data and the second operator data are read from the first memory through the data transfer engine, and the first operator data and the second operator data are transferred to the second memory inside the matrix operation engine through the bus. The bus bit width is greater than the data bit width corresponding to the array dimension.

[0146] Step 1003: Perform matrix operations on the first operator data and the second operator data in the second memory through the matrix operation unit in the matrix operation engine to obtain sub-operation results, where the data dimension of the sub-operation results matches the array dimension.

[0147] Step 1004 , accumulating the sub-operation results through an accumulator in the matrix operation engine to obtain a matrix operation result of the first operation data and the second operation data.

[0148] In some embodiments, in the process of processing convolution data, the first operation data is the convolution input data, the second operation data is the convolution weight data, and the operation result of the first operation data and the second operation data is the convolution output data; in the process of processing matrix data, the first operation data is the left matrix input data, the second operation data is the right matrix input data, and the operation result of the first operation data and the second operation data is the matrix output data.

[0149] In some embodiments, when the first operation data is convolution input data, the second operation data is convolution weight data, and the operation results of the first operation data and the second operation data are convolution output data, the convolution input data adopts a first target data format, and the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array; the convolution weight data adopts a second target data format, and the number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array; and the convolution output data adopts a third target data format, and the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.

[0150] In some embodiments, when the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, the convolution input data adopts a fourth target data format, and the number of input channels represented by the fourth target data format is the total number of input channels; the convolution weight data adopts a fifth target data format, and the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array.

[0151] In some embodiments, when the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, the convolution output data adopts a sixth target data format, and the number of output channels represented by the sixth target data format is the total number of output channels; the convolution weight data adopts a seventh target data format, and the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is the total number of output channels.

[0152] In some embodiments, during the data transfer process, the first convolution input sub-data and the second convolution input sub-data are read from the first memory in sequence based on the bus bit width through the data transfer engine, the first convolution input sub-data and the second convolution input sub-data are continuously stored data, and the data bit width of the first convolution input sub-data and the data bit width of the second convolution input sub-data are both equal to the bus bit width; based on the bus bit width, the first convolution weight sub-data and the second convolution weight sub-data are read from the first memory in sequence, the first convolution weight sub-data and the second convolution weight sub-data are continuously stored data, and the data bit width of the first convolution weight sub-data and the data bit width of the second convolution weight sub-data are both equal to the bus bit width.

[0153] In some embodiments, when the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data, the left matrix input data adopts an eighth target data format, and the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array; the right matrix input data adopts a ninth target data format, and the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array; the matrix output data adopts a tenth target data format, and the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.

[0154] In some embodiments, during data transfer, the data transfer engine sequentially reads the first left matrix input sub-data and the second left matrix input sub-data from the first memory based on the bus bit width, the first left matrix input sub-data and the second left matrix input sub-data are continuously stored data, and the data bit width of the first left matrix input sub-data and the data bit width of the second left matrix input sub-data are both equal to the bus bit width; based on the bus bit width, the first right matrix input sub-data and the second right matrix input sub-data are sequentially read from the first memory, the first right matrix input sub-data and the second right matrix input sub-data are continuously stored data, and the data bit width of the first right matrix input sub-data and the data bit width of the second right matrix input sub-data are both equal to the bus bit width.

[0155] In some embodiments, the AI ​​processor further includes a data format converter;

[0156] The method also includes: when the operation results of the first operation data and the second operation data are convolution output data, and the convolution output data is used for subsequent matrix calculation, converting the data format of the convolution output data based on the target data format corresponding to the matrix input data through a data format converter; when the operation results of the first operation data and the second operation data are matrix output data, and the matrix output data is used for subsequent convolution calculation, converting the data format of the matrix output data based on the target data format corresponding to the convolution input data through a data format converter.

[0157] In some embodiments, the AI ​​processor further includes a data format converter;

[0158] The method further includes: reading first operation data and second operation data from a first memory through a data format converter, the first operation data and the second operation data being in an original data format, wherein a data dimension represented by the original data format does not match an array dimension of a systolic array in a matrix operation unit in a matrix operation engine;

[0159] The data formats of the first operation data and the second operation data are converted by a data format converter so that the data formats of the first operation data and the second operation data are converted from the original data format to the target data format, and the first operation data and the second operation data are written into the first memory.

[0160] In summary, in an embodiment of the present application, by storing the first operation data and the second operation data in the target data format in the first memory, the data dimensions of the first operation data and the second operation data match the array dimensions of the systolic array in the matrix operation unit in the matrix operation engine, so that the data transfer engine can read the first operator data and the second operator data from the first memory according to the bus bit width, and transfer the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and then the matrix operation engine performs a matrix operation on the first operator data and the second operator data through the matrix operation unit to obtain a sub-operation result, and accumulates the sub-operation results through the accumulator to obtain a matrix operation result of the first operation data and the second operation data. Using the AI ​​processor provided by the embodiment of the present application, by storing the operation data in the target data format in the first memory, it is possible to read the operation data according to the bus bit width. Since the bus bit width is greater than the data bit width corresponding to the array dimension, compared with reading data according to the data bit width corresponding to the array dimension, the AI ​​processor provided by the embodiment of the present application can improve the bandwidth utilization of data transfer during the matrix operation process, thereby improving the efficiency of the matrix operation.

[0161] Please refer to Figure 11, which shows a block diagram of a computer device 1100 according to an exemplary embodiment of the present application. Computer device 1100 may be a portable mobile terminal, such as a smartphone, a tablet computer, a Moving Picture Experts Group Audio Layer III (MP3) player, or a Moving Picture Experts Group Audio Layer IV (MP4) player. Computer device 1100 may also be referred to as a user device, a portable terminal, a workstation, a server, or other similar names.

[0162] Typically, the computer device 1100 includes an AI processor 1101 and a memory 1102 .

[0163] The AI ​​processor 1101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The AI ​​processor 1101 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The AI ​​processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the AI ​​processor 1101 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the AI ​​processor 1101 may also be used to process computing operations related to machine learning.

[0164] The memory 1102 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices and flash memory storage devices.

[0165] In some embodiments, the computer device 1100 may optionally further include a peripheral device interface 1103 and at least one peripheral device.

[0166] Those skilled in the art will appreciate that the structure shown in FIG11 does not limit the computer device 1100 and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0167] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0168] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An AI processor, comprising: A matrix operation engine, a data handling engine and a first memory, wherein the matrix operation engine and the data handling engine are connected via a bus; The first memory is used to store first operation data and second operation data, the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine; The data transport engine is configured to read the first operator data and the second operator data from the first memory based on a bus bit width, and transport the first operator data and the second operator data to a second memory inside the matrix operation engine through the bus, wherein the bus bit width is greater than a data bit width corresponding to the array dimension; The matrix operation engine is used to perform a matrix operation on the first operator data and the second operator data in the second memory through a matrix operation unit to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension; The matrix operation engine is further used to accumulate the sub-operation results through an accumulator to obtain a matrix operation result of the first operation data and the second operation data.

2. The AI ​​processor according to claim 1, wherein: In the process of processing convolution data, the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data; In the process of processing matrix data, the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data.

3. The AI ​​processor according to claim 2, wherein: In the case where the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data, The convolution input data adopts a first target data format, and the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array; The convolution weight data adopts a second target data format, the number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array; The convolution output data adopts a third target data format, and the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.

4. The AI ​​processor according to claim 3, wherein: In the case where the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, The convolution input data adopts a fourth target data format, and the number of input channels represented by the fourth target data format is the total number of input channels; The convolution weight data adopts a fifth target data format, the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array.

5. The AI ​​processor according to claim 3, wherein: In the case where the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, The convolution output data adopts a sixth target data format, and the number of output channels represented by the sixth target data format is the total number of output channels; The convolution weight data adopts the seventh target data format, the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is equal to the total output Number of channels.

6. The AI ​​processor according to any one of claims 3 to 5, wherein: The data handling engine is used to: Based on the bus bit width, first convolution input sub-data and second convolution input sub-data are sequentially read from the first memory, where the first convolution input sub-data and the second convolution input sub-data are continuously stored data, and the data bit width of the first convolution input sub-data and the data bit width of the second convolution input sub-data are both equal to the bus bit width; Based on the bus bit width, the first convolution weight sub-data and the second convolution weight sub-data are read from the first memory in sequence, the first convolution weight sub-data and the second convolution weight sub-data are continuously stored data, and the data bit width of the first convolution weight sub-data and the data bit width of the second convolution weight sub-data are both equal to the bus bit width.

7. The AI ​​processor according to any one of claims 2 to 6, wherein: In the case where the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data, The left matrix input data adopts an eighth target data format, and the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array; The right matrix input data adopts a ninth target data format, and the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array; The matrix output data adopts the tenth target data format, the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.

8. The AI ​​processor according to claim 7, wherein: The data handling engine is used to: Based on the bus bit width, first left matrix input sub-data and second left matrix input sub-data are sequentially read from the first memory, the first left matrix input sub-data and the second left matrix input sub-data are continuously stored data, and the data bit width of the first left matrix input sub-data and the data bit width of the second left matrix input sub-data are both equal to the bus bit width; Based on the bus bit width, the first right matrix input sub-data and the second right matrix input sub-data are read from the first memory in sequence, the first right matrix input sub-data and the second right matrix input sub-data are continuously stored data, and the data bit width of the first right matrix input sub-data and the data bit width of the second right matrix input sub-data are both equal to the bus bit width.

9. The AI ​​processor according to any one of claims 1 to 8, wherein: The AI ​​processor also includes a data format converter; The data format converter is used for: When the operation results of the first operation data and the second operation data are convolution output data, and the convolution output data is used for subsequent matrix calculation, converting the data format of the convolution output data based on the target data format corresponding to the matrix input data; When the operation results of the first operation data and the second operation data are matrix output data, and the matrix output data is used for subsequent convolution calculation, the data format of the matrix output data is converted based on the target data format corresponding to the convolution input data.

10. The AI ​​processor according to any one of claims 1 to 8, wherein: The AI ​​processor also includes a data format converter; The data format converter is used to read the first operation data and the second operation data from the first memory. calculation data, the first calculation data and the second calculation data are in original data format, and the data dimension represented by the original data format does not match the array dimension of the systolic array in the matrix calculation unit in the matrix calculation engine; The data format converter is also used to convert the data formats of the first operation data and the second operation data, so that the data formats of the first operation data and the second operation data are converted from the original data format to the target data format, and the first operation data and the second operation data are written into the first memory.

11. A data processing method, the method being executed by an AI processor, the AI ​​processor comprising a matrix operation engine, a data handling engine and a first memory, the matrix operation engine being connected to the data handling engine via a bus; The method comprises: storing first operation data and second operation data by the first memory, wherein the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine; Based on a bus bit width, reading first operator data and second operator data from the first memory through the data transfer engine, and transferring the first operator data and the second operator data to a second memory inside the matrix operation engine through the bus, wherein the bus bit width is greater than a data bit width corresponding to the array dimension; Performing a matrix operation on the first operator data and the second operator data in the second memory through a matrix operation unit in the matrix operation engine to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension; The sub-operation results are accumulated by an accumulator in the matrix operation engine to obtain a matrix operation result of the first operation data and the second operation data.

12. The method according to claim 11, wherein: In the process of processing convolution data, the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data; In the process of processing matrix data, the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data.

13. The method according to claim 12, wherein: In the case where the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data, The convolution input data adopts a first target data format, and the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array; The convolution weight data adopts a second target data format, the number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array; The convolution output data adopts a third target data format, and the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.

14. The method according to claim 13, wherein: In the case where the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, The convolution input data adopts a fourth target data format, and the number of input channels represented by the fourth target data format is the total number of input channels; The convolution weight data adopts a fifth target data format, the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array.

15. The method according to claim 13, wherein: In the case where the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, The convolution output data adopts a sixth target data format, and the number of output channels represented by the sixth target data format is the total number of output channels; The convolution weight data adopts a seventh target data format, the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is the total number of output channels.

16. The method according to claim 12, wherein: In the case where the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data, The left matrix input data adopts an eighth target data format, and the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array; The right matrix input data adopts a ninth target data format, and the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array; The matrix output data adopts the tenth target data format, the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.

17. The method according to any one of claims 11 to 16, wherein: The AI ​​processor also includes a data format converter; The method further comprises: When the operation results of the first operation data and the second operation data are convolution output data, and the convolution output data is used for subsequent matrix calculation, converting the data format of the convolution output data based on the target data format corresponding to the matrix input data by the data format converter; When the operation results of the first operation data and the second operation data are matrix output data, and the matrix output data is used for subsequent convolution calculation, the data format of the matrix output data is converted through the data format converter based on the target data format corresponding to the convolution input data.

18. The method according to any one of claims 11 to 16, wherein: The AI ​​processor also includes a data format converter; The method further comprises: Reading the first operation data and the second operation data from the first memory through the data format converter, wherein the first operation data and the second operation data are in original data format, and the data dimension represented by the original data format does not match the array dimension of the systolic array in the matrix operation unit in the matrix operation engine; The data formats of the first operation data and the second operation data are converted by the data format converter, so that the data formats of the first operation data and the second operation data are converted from the original data format to the target data format, and the first operation data and the second operation data are written into the first memory.

19. A computer device, characterized in that: The computer device comprises a memory and an AI processor as claimed in any one of claims 1 to 10, wherein the memory stores at least one instruction, and the at least one instruction is used to be executed by the AI ​​processor.

Citation Information

Patent Citations

  • Operation accelerator, processing method, and related device

    CN112840356A

  • Data processing method and device, electronic equipment and computer readable storage medium

    CN116911367A

  • Data processing method and device, computer equipment and storage medium

    CN116980277A

  • Variable-size problem solving with systolic arrays

    US20170161611A1

  • Neural processor

    US20220283984A1