AI processor, data processing method and computer equipment
By storing the operation data in the target data format in the memory and using the data transfer engine to carry data according to the bus bit width, the problem of low bus bandwidth utilization in the prior art is solved, and more efficient matrix computing and data transfer is achieved.
Patent Information
- Application Number
- CN202311579613.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
When data transfer is carried through the bus, the bus bandwidth utilization rate is low, mainly because the data format of the input data does not match the array dimensions of the pulsating array, resulting in low data handling efficiency.
By storing operation data in the target data format in the memory, the data dimensions match the array dimensions of the pulsating array in the matrix computing engine, and the data handling engine reads and transports data according to the bus bit width.
It improves bandwidth utilization during data handling, optimizes the efficiency of matrix computing, and ensures the continuity and efficiency of data handling.
Smart Images

Figure CN120031084A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of processor technology, and in particular to an AI processor, a data processing method, and a computer device. Background Art
[0002] The core calculations in deep learning algorithms mainly include two categories: convolution calculations and matrix calculations. Convolution calculations can be equivalent to matrix calculations. Therefore, by accelerating matrix calculations, deep learning algorithms can be accelerated.
[0003] In the related art, a matrix acceleration engine is set up to improve the speed of matrix calculation, and input data is moved based on the array dimension of a systolic array in the matrix acceleration engine. However, the data dimension represented by the data format adopted by different input data often does not match the array dimension of the systolic array. Therefore, in the process of moving data through a bus, directly moving data according to the systolic array dimension often leads to a problem of low bus bandwidth utilization. Summary of the invention
[0004] The embodiment of the present application provides an AI processor, a data processing method and a computer device, which improves the bandwidth utilization rate during data transfer. The technical solution is as follows:
[0005] On the one hand, an embodiment of the present application provides an AI processor, the AI processor comprising: a matrix operation engine, a data handling engine and a first memory, the matrix operation engine and the data handling engine being connected via a bus;
[0006] The first memory is used to store first operation data and second operation data, the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine;
[0007] The data transport engine is configured to read the first operator data and the second operator data from the first memory based on a bus bit width, and transport the first operator data and the second operator data to a second memory inside the matrix operation engine through the bus, wherein the bus bit width is greater than a data bit width corresponding to the array dimension;
[0008] The matrix operation engine is used to perform a matrix operation on the first operator data and the second operator data in the second memory through a matrix operation unit to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension;
[0009] The matrix operation engine is further used to accumulate the sub-operation results through an accumulator to obtain a matrix operation result of the first operation data and the second operation data.
[0010] On the other hand, an embodiment of the present application provides a data processing method, which is used for an AI processor, wherein the AI processor includes a matrix operation engine, a data handling engine, and a first memory, wherein the matrix operation engine is connected to the data handling engine via a bus; the method includes:
[0011] storing first operation data and second operation data by the first memory, wherein the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine;
[0012] Based on a bus bit width, reading first operator data and second operator data from the first memory through the data transfer engine, and transferring the first operator data and the second operator data to a second memory inside the matrix operation engine through the bus, wherein the bus bit width is greater than a data bit width corresponding to the array dimension;
[0013] Performing a matrix operation on the first operator data and the second operator data in the second memory by a matrix operation unit in the matrix operation engine to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension;
[0014] The sub-operation results are accumulated by an accumulator in the matrix operation engine to obtain a matrix operation result of the first operation data and the second operation data.
[0015] In some embodiments, in the process of processing convolution data, the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data; in the process of processing matrix data, the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation result of the first operation data and the second operation data is matrix output data.
[0016] In some embodiments, when the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data, the convolution input data adopts a first target data format, and the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array; the convolution weight data adopts a second target data format, the number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array; the convolution output data adopts a third target data format, and the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.
[0017] In some embodiments, when the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, the convolution input data adopts a fourth target data format, and the number of input channels represented by the fourth target data format is the total number of input channels; the convolution weight data adopts a fifth target data format, and the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array.
[0018] In some embodiments, when the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, the convolution output data adopts a sixth target data format, and the number of output channels represented by the sixth target data format is the total number of output channels; the convolution weight data adopts a seventh target data format, and the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is the total number of output channels.
[0019] In some embodiments, the data handling engine is used to read the first convolution input sub-data and the second convolution input sub-data from the first memory in sequence based on the bus bit width, the first convolution input sub-data and the second convolution input sub-data are continuously stored data, and the data bit width of the first convolution input sub-data and the data bit width of the second convolution input sub-data are both equal to the bus bit width; based on the bus bit width, read the first convolution weight sub-data and the second convolution weight sub-data from the first memory in sequence, the first convolution weight sub-data and the second convolution weight sub-data are continuously stored data, and the data bit width of the first convolution weight sub-data and the data bit width of the second convolution weight sub-data are both equal to the bus bit width.
[0020] In some embodiments, when the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data, the left matrix input data adopts an eighth target data format, and the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array; the right matrix input data adopts a ninth target data format, and the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array; the matrix output data adopts a tenth target data format, and the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.
[0021] In some embodiments, the data handling engine is used to read the first left matrix input sub-data and the second left matrix input sub-data from the first memory in sequence based on the bus bit width, the first left matrix input sub-data and the second left matrix input sub-data are continuously stored data, and the data bit width of the first left matrix input sub-data and the data bit width of the second left matrix input sub-data are both equal to the bus bit width; based on the bus bit width, read the first right matrix input sub-data and the second right matrix input sub-data from the first memory in sequence, the first right matrix input sub-data and the second right matrix input sub-data are continuously stored data, and the data bit width of the first right matrix input sub-data and the data bit width of the second right matrix input sub-data are both equal to the bus bit width.
[0022] In some embodiments, the AI processor further includes a data format converter;
[0023] The data format converter is used to convert the data format of the convolution output data based on the target data format corresponding to the matrix input data when the operation result of the first operation data and the second operation data is convolution output data and the convolution output data is used for subsequent matrix calculation; and to convert the data format of the matrix output data based on the target data format corresponding to the convolution input data when the operation result of the first operation data and the second operation data is matrix output data and the matrix output data is used for subsequent convolution calculation.
[0024] In some embodiments, the AI processor further includes a data format converter;
[0025] The data format converter is used to read the first operation data and the second operation data from the first memory, the first operation data and the second operation data are in original data format, and the data dimension represented by the original data format does not match the array dimension of the systolic array in the matrix operation unit in the matrix operation engine;
[0026] The data format converter is also used to convert the data formats of the first operation data and the second operation data, so that the data formats of the first operation data and the second operation data are converted from the original data format to the target data format, and the first operation data and the second operation data are written into the first memory.
[0027] On the other hand, an embodiment of the present application provides a computer device, comprising a memory and an AI processor as described in the above aspects, wherein the memory stores at least one instruction, and the at least one instruction is used to be executed by the AI processor.
[0028] In an embodiment of the present application, by storing the first operation data and the second operation data in the target data format in the first memory, the data dimension of the first operation data and the second operation data matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine, so that the data handling engine can read the first operator data and the second operator data from the first memory according to the bus width, and carry the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and then the matrix operation engine performs matrix operation on the first operator data and the second operator data through the matrix operation unit to obtain the sub-operation result, and accumulates the sub-operation result through the accumulator to obtain the matrix operation result of the first operation data and the second operation data. Using the AI processor provided in the embodiment of the present application, by storing the operation data in the target data format in the first memory, it is possible to realize the reading of operation data according to the bus width. Since the bus width is greater than the data width corresponding to the array dimension, compared with reading data according to the data width corresponding to the array dimension, the AI processor provided in the embodiment of the present application can improve the bandwidth utilization of data handling during matrix operation, thereby improving the efficiency of matrix operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0030] Figure 1 A schematic diagram of the structure of an AI processor provided by an exemplary embodiment of the present application is shown;
[0031] Figure 2 A schematic diagram of a systolic array operation provided by an exemplary embodiment of the present application is shown;
[0032] Figure 3 The logical expression of convolutional network data and the physical storage of data in different formats in related technologies are shown;
[0033] Figure 4 The physical storage form of the first target data format provided by an exemplary embodiment of the present application is shown;
[0034] Figure 5 A schematic diagram showing data processing by a systolic array based on a general data format in the related art is shown;
[0035] Figure 6 A schematic diagram of performing data processing through a systolic array based on a target data format provided by an exemplary embodiment of the present application is shown;
[0036] Figure 7 A schematic diagram showing data processing by a systolic array based on a general data format in another related art is shown;
[0037] Figure 8 A schematic diagram showing data processing by a systolic array based on a target data format provided by another exemplary embodiment of the present application is shown;
[0038] Fig. 9 A schematic diagram of data format conversion provided by an exemplary embodiment of the present application is shown;
[0039] Fig.10 A flowchart of a data processing method provided by an exemplary embodiment of the present application is shown;
[0040] Fig.11 A structural block diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0041] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0042] It should be understood that the "several" mentioned in this article refers to one or more, and "multiple" refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0043] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0044] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0045] Among them, deep learning methods are now widely used in various fields such as image processing, video processing, speech processing, content generation, etc., and the core characteristics of deep learning are large amount of calculation and large number of parameters. The existing mainstream deep learning networks can be divided into three categories: Convolutional Neural Networks (CNN), Transformer and ViT (Vision Transformer). The core calculations in these three categories of networks are mainly concentrated in convolution calculations and matrix calculations, and convolution calculations can be equivalent to matrix calculations. Therefore, the core of accelerating deep learning algorithms is to accelerate matrix operations, and the key issue of AI processors is to perfectly match the core matrix calculations with other commonly used vector calculations.
[0046] Optionally, the AI processor can be a processor based on a neural network algorithm and acceleration, such as a neural network processor (Neural network Processing Unit, NPU).
[0047] In the related art, an AI processor is integrated into a general deep learning framework, and matrix operations are performed through a matrix operation engine in the AI processor, wherein the structure of the matrix operation unit in the matrix operation engine is generally in the form of a two-dimensional systolic array. Therefore, in order to use the matrix operation unit to perform data operations on the operation data, the related art generally performs data transfer on the operation data according to the array dimension of the systolic array in the matrix operation unit. However, for a bus whose bus width is larger than the data bit width corresponding to the array dimension, the problem of insufficient bus bandwidth utilization often occurs during data transfer, resulting in low bus bandwidth utilization.
[0048] In an embodiment of the present application, by storing the first operation data and the second operation data in the target data format in the first memory, the first operator data and the second operator data can be directly read from the first memory according to the bus bit width through the data transfer engine, and the first operator data and the second operator data can be transferred to the second memory inside the matrix operation engine through the bus, thereby fully improving the utilization rate of the bus bit width and optimizing the data transfer process.
[0049] Please refer to Figure 1 , which shows a structural diagram of an AI processor provided by an exemplary embodiment of the present application. The AI processor 100 mainly includes a matrix operation engine 110, a data transfer engine 120 and a first memory 130, wherein the matrix operation engine 110 and the data transfer engine 120 are connected via a bus.
[0050] The first memory 130 is used to store the first operation data and the second operation data. The first operation data and the second operation data are in a target data format. The data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine.
[0051] Unlike the related art, the AI processor stores operation data in a dimensional expression unique to the deep learning network through the first memory, resulting in insufficient bandwidth utilization during data transfer through the data transfer engine. In the embodiment of the present application, the first operation data and the second operation data in the first memory are both stored in the target data format.
[0052] Optionally, the architecture of the matrix operation engine is a multi-dimensional systolic array composed of multiple physical matrices of multiply-accumulate (MAC) operations, which is used to perform computational processing on a series of matrix operations of a convolutional neural network.
[0053] Optionally, the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine, which may mean that the data dimension of each data channel represented by the target data format is equal to the array dimension of the systolic array, or that the data dimension of some data channels among the data channels represented by the target data format is equal to the array dimension of the systolic array.
[0054] Optionally, the array dimension of the systolic array may be the number of array rows or the number of array columns. When the operation data is convolution data, the data dimension represented by the target data format may be the number of feature channels; when the operation data is matrix data, the data dimension represented by the target data format may be the number of matrix rows or the number of matrix columns.
[0055] For example, the number of data channels represented by the target data format is equal to the array dimension of the systolic array, or the number of data rows or columns represented by the target data format is equal to the number of array rows or columns of the systolic array.
[0056] Illustratively, the first operation data is matrix data, the number of matrix rows and the number of matrix columns of the matrix data are both 128 dimensions, and the number of array rows of the systolic array is 64 dimensions. Then, when the target data format is adopted, the number of matrix rows of the first operation data can be converted into 2×64 dimensions, so that the number of matrix rows of the first operation data is equal to the number of array rows of the systolic array.
[0057] Optionally, the systolic array is generally in the form of a two-dimensional matrix, which can be a two-dimensional systolic, a one-dimensional systolic plus a one-dimensional broadcast, or a two-dimensional systolic and broadcast mixed structure.
[0058] The data transfer engine 120 is used to read the first operator data and the second operator data from the first memory based on the bus bit width, and transfer the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and the bus bit width is greater than the data bit width corresponding to the array dimension.
[0059] Optionally, the data transfer engine may be a direct memory access (DMA) controller for transferring data between memories, including an address bus, a data bus, and a control register. The data transfer engine may read the first operand data and the second operand data from the first memory in sequence, and transfer the first operand data and the second operand data to the second memory through the bus.
[0060] Optionally, the first operator data belongs to the first operation data and is a part of the first operation data, and the second operator data belongs to the second operation data and is a part of the second operation data. The AI processor can use the data transfer engine to realize the segmented transmission of the first operation data and the second operation data from the first memory to the second memory inside the matrix operation engine.
[0061] Different from the related art, the data dimension represented by the data format used for the first operation data and the second operation data does not match the array dimension of the systolic array. In order not to affect the matrix operation process in the matrix operation unit, the AI processor performs data transfer according to the array dimension through the data transfer engine, resulting in the sub-data transferred in adjacent time not having data continuity, and when the bus bit width is greater than the data bit width corresponding to the array dimension, the bus bandwidth utilization is low. In the embodiment of the present application, after storing the first operation data and the second operation data in the target data format, the AI processor can use the data transfer engine to directly read the first operation sub-data and the second operation sub-data from the first memory according to the bus bit width, while fully utilizing the bus bit width and ensuring the continuity of data transfer.
[0062] Indicatively, when the operation data does not adopt the target data format and the bus width is greater than the data width corresponding to the array dimension, in order to enable the matrix operation unit to directly perform matrix operations on the sub-operation data transported by the data transport engine, the data transport engine transports the data in sequence according to the data width corresponding to the array dimension and the number of data rows or columns. Since the data width of each data transport is less than the bus width, the bus width is not fully utilized. In the case of adopting the target data format, the operation data can be stored continuously in the physical storage in units of the array dimension of the systolic array, and the data transport engine can directly transport the data according to the bus width, that is, the data width of the transported data is equal to the bus width, thereby making full use of the bus width.
[0063] In an illustrative example, when the bus bit width is twice the data bit width corresponding to the array dimension, the bandwidth utilization rate is only 50% when the operation data is moved according to the data bit width corresponding to the array dimension. After the operation data is stored in the target data format, the operation data can be stored continuously in the physical storage in units of the array dimension of the systolic array, so that the operation data can be moved according to the bus bit width, so that the bandwidth utilization rate reaches 100%.
[0064] For example, when the bus width is 2048 bits, the array dimension of the systolic array is 64 dimensions, and the data in each dimension is fp16, the data width corresponding to the array dimension of the systolic array is 1024 bits, that is, the bus width (2048 bits) is greater than the data width (1024 bits) corresponding to the array dimension of the systolic array, so that the operation data is transported according to the data width corresponding to the array dimension, and only 1024 bits of bus width can be occupied each time, resulting in a bandwidth utilization rate of only 50%; and when the target data format is used to store the operation data, the operator data with a data width of 2048 bits equal to the bus width can be read from the first memory according to the bus width.
[0065] In some embodiments, a second memory is provided in the matrix operation engine, so that the data transfer engine can directly transfer the first operator data and the second operator data to the second memory inside the matrix operation engine during data transfer through the bus.
[0066] Optionally, the data transfer process may be understood as transmitting the first operator data and the second operator data to a second memory inside the matrix operation engine through a bus.
[0067] The matrix operation engine 110 is used to perform matrix operation on the first operator data and the second operator data in the second memory through a matrix operation unit to obtain a sub-operation result, and the data dimension of the sub-operation result matches the array dimension.
[0068] In some embodiments, the matrix operation engine also includes a matrix operation unit, so that after the first operator data and the second operator data are stored in the second memory, the matrix operation engine can perform matrix operations on the first operator data and the second operator data through the matrix operation unit using the systolic array to obtain a sub-operation result.
[0069] The data dimension of the sub-operation result matches the array dimension of the systolic array. For example, the number of data rows of the sub-operation result is equal to the number of array rows of the systolic array, and the number of data columns of the sub-operation result is equal to the number of array columns of the systolic array.
[0070] In some embodiments, the matrix operation engine may use the first operator data as the left input data of the systolic array and the second operator data as the right input data of the systolic array, and perform matrix operation through the matrix operation unit to obtain sub-operation results.
[0071] Optionally, the matrix operation unit processes the first operator data and the second operator data through a matrix multiplication operation, so as to obtain a sub-operation result.
[0072] Indicatively, Figure 2 As shown, the result data output by the matrix operation unit through the matrix operation unit are all in the two-dimensional matrix form of [Matrix_M, MALU_N]. Taking a calculation based on the systolic array as an example, after determining the left input data and the right input data of the systolic array, the matrix operation unit performs data operations from left to right and from top to bottom based on the properties of the systolic array, thereby outputting a set of data on the MALU_N side.
[0073] The matrix operation engine 110 is further used to accumulate the sub-operation results through an accumulator to obtain a matrix operation result of the first operation data and the second operation data.
[0074] In some embodiments, the matrix operation engine also includes an accumulator (ACC). After performing matrix operations through the matrix operation unit to obtain a large number of sub-operation results, the matrix operation engine can also accumulate the sub-operation results through the accumulator to obtain the matrix operation results of the first operation data and the second operation data.
[0075] Optionally, the accumulator's accumulation process of sub-operation results refers to temporarily storing the sub-operation results output by the matrix operation unit, and after obtaining all the sub-operation results, performing data splicing on each sub-operation result according to the matrix element position of each sub-operation result in the operation result matrix, thereby obtaining the operation results of the first operation data and the second operation data.
[0076] Schematically, the first operation data and the second operation data are both 2×2 matrix data, and matrix operation can be performed by the matrix operation unit to obtain: a first sub-operation result corresponding to the first row matrix data of the first operation data and the first column matrix data of the second operation data, a second sub-operation result corresponding to the first row matrix data of the first operation data and the second column matrix data of the second operation data, a third sub-operation result corresponding to the second row matrix data of the first operation data and the first column matrix data of the second operation data, and a fourth sub-operation result corresponding to the second row matrix data of the first operation data and the second column matrix data of the second operation data, so that the matrix element position of the first sub-operation data in the operation result matrix is the first row and the first column, the matrix element position of the second sub-operation data in the operation result matrix is the first row and the second column, the matrix element position of the third sub-operation data in the operation result matrix is the second row and the first column, and the matrix element position of the fourth sub-operation data in the operation result matrix is the second row and the second column, and then the accumulator can splice the sub-operation data according to the matrix element positions corresponding to each sub-operation data to obtain the matrix operation result.
[0077] In some embodiments, after obtaining the operation results of the first operation data and the second operation data, the AI processor may also use the data transfer engine to transfer the operation results to the first memory through the bus.
[0078] In summary, in the embodiment of the present application, by storing the first operation data and the second operation data in the target data format in the first memory, the data dimension of the first operation data and the second operation data matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine, so that the data handling engine can read the first operator data and the second operator data from the first memory according to the bus width, and carry the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and then the matrix operation engine performs matrix operation on the first operator data and the second operator data through the matrix operation unit to obtain the sub-operation result, and accumulates the sub-operation result through the accumulator to obtain the matrix operation result of the first operation data and the second operation data. Using the AI processor provided in the embodiment of the present application, by storing the operation data in the target data format in the first memory, it is possible to realize the reading of operation data according to the bus width. Since the bus width is greater than the data width corresponding to the array dimension, compared with reading data according to the data width corresponding to the array dimension, the AI processor provided in the embodiment of the present application can improve the bandwidth utilization of data handling during matrix operation, thereby improving the efficiency of matrix operation.
[0079] In some embodiments, the deep learning network can be a convolutional network that performs operations on convolution data; or it can be a matrix network that performs operations on matrix data. Therefore, after the AI processor is integrated into the deep learning network framework, the AI processor needs to process the convolution data and matrix data separately.
[0080] In a possible implementation, in the process of processing convolution data, the first operation data may be convolution input data, the second operation data may be convolution weight data, and the operation result of the first operation data and the second operation data may be convolution output data.
[0081] In some embodiments, in the expression of dimensional information of a convolutional network, H represents the height of an image or a feature map, W represents the width of an image or a feature map, C represents the feature dimension, and N represents a batch of images or feature maps.
[0082] Indicatively, Figure 3As shown, it shows the logical expression of convolutional network data and the physical storage form of data in different formats in related technologies. Among them, the logical expression forms corresponding to NCHW format data, NHWC format data and CHWN format data are the same, that is, the data arrangement methods in the C direction, H direction and W direction are the same, and in the physical storage process, for NCHW format data, the W direction is taken first, followed by the H direction, then the C direction, and finally the N direction; for NHWC format data, the C direction is taken first, followed by the W direction, then the H direction, and finally the N direction; for CHWN format data, the N direction is taken first, followed by the W direction, then the H direction, and finally the C direction.
[0083] In the related art, taking data arranged in NHWC format as an example, when the number of array rows of the systolic array Ck0=3, data needs to be moved in groups of 000, 020, 040 and 001, 021, 041. From the physical storage form, it can be seen that 000, 020, 040 and 001, 021, 041 are in discrete storage states. Therefore, the data moving engine can only move one group of data at a time. Even if the bus bandwidth is greater than the data bit width corresponding to the number of array rows, the data moving engine can only read 3 data values at a time.
[0084] In some embodiments, when the first operation data is convolution input data, in order to fully utilize the bus bandwidth in the process of data transfer of the convolution input data, the convolution input data can be stored in a first target data format, wherein the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array.
[0085] In one possible implementation, the number of array rows of the systolic array can be expressed as Ck0, so that the convolution input data can be expressed as NCi1HiWiCk0, where N represents a set of images or feature maps, Hi represents the height of the input image or feature map, Wi represents the width of the input image or feature map, Ci1 is the value of Ci / Ck0 rounded up, and Ci is the total number of input channels.
[0086] Indicatively, Figure 4 As shown, it shows the physical storage form of the first target data format provided by an exemplary embodiment of the present application, from Figure 4It can be seen that for the target data format NCi1HiWiCk0, the data is arranged in the Ck0 direction first, then the Wi direction, then the Hi direction, then the Ci1 direction, and finally the N direction. For example, when the number of array rows of the systolic array Ck0=3, first take the three values 000, 020, and 040 in the first Ck0 direction, and then from the Wi direction, take 001, 021, and 041. After taking the first Wi direction, transfer from 043 to 004 on the second Hi, and store the data in this order, so as to ensure that the data is in a continuous state during the data transfer process based on the number of array rows of the systolic array.
[0087] In some embodiments, when the second operation data is convolution weight data, in order to fully utilize the bus bandwidth in the process of data transfer of the convolution weight data, the convolution weight data can be stored in a second target data format, wherein the number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array.
[0088] In one possible implementation, the number of array rows of the systolic array can be represented as Ck0, and the number of array columns of the systolic array can be represented as Cn0, so that the convolution weight data can be represented as Co1Ci1KhKwCk0Cn0, wherein Kh represents the height of the convolution kernel, Kw represents the width of the convolution kernel, Ci1 is the value of Ci / Ck0 rounded up, Co1 represents the value of Co / Cn0 rounded up, Ci is the total number of input channels, and Co is the total number of output channels.
[0089] In some embodiments, when the operation result of the first operation data and the second operation data is convolution output data, convolution output data in a third target data format can be obtained by performing matrix operation through a matrix operation unit according to the convolution input data in the first target data format and the convolution weight data in the second target data format, wherein the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.
[0090] In one possible implementation, the number of array columns of the systolic array can be represented by Cn0, so that the convolution output data can be represented as NCo1HoWoCn0, where N represents a set of images or feature maps, Ho represents the height of the output image or feature map, Wo represents the width of the output image or feature map, Co1 represents the value of Co / Cn0 rounded up, and Co is the total number of output channels.
[0091] In some embodiments, after the convolution input data in the first target data format and the convolution output data in the second target data format are stored in the first memory of the AI processor, the AI processor can use the data transfer engine to transfer the convolution input sub-data and the convolution weight sub-data according to the bus bit width, and store them in the second memory inside the matrix operation engine.
[0092] In one possible implementation, the AI processor uses a data transfer engine to sequentially read the first convolution input sub-data and the second convolution input sub-data from the first memory according to the bus bit width, wherein the first convolution input sub-data and the second convolution input sub-data are continuously stored data, and the data bit width of the first convolution input sub-data is equal to the bus bit width, and the data bit width of the second convolution input sub-data is equal to the bus bit width.
[0093] Indicatively, Figure 4 As shown, when the bus width is 2048 bits, the number of array rows of the systolic array is 3, and the data width corresponding to the number of array rows is 1024 bits, the AI processor uses the data transfer engine to directly take six values 000, 020, 040, 001, 021, and 041 as the first convolution input sub-data according to the bus width, and transfer the six values at the same time, and when continuing to transfer data according to the number of array rows of the systolic array, it can continue to take 002, 022, 042, 003, 023, and 043 as the second convolution input sub-data, and the first convolution input sub-data and the second convolution input sub-data are continuously stored data in the physical storage data arrangement.
[0094] In one possible implementation, the AI processor uses a data transfer engine to sequentially read the first convolution weight sub-data and the second convolution weight sub-data from the first memory according to the bus bit width, wherein the first convolution weight sub-data and the second convolution weight sub-data are continuously stored data, and the data bit width of the first convolution weight sub-data is equal to the bus bit width, and the data bit width of the second convolution weight sub-data is equal to the bus bit width.
[0095] Indicatively, Figure 5 As shown, it shows a schematic diagram of data processing through a systolic array based on a general data format in the related art. Taking the use of a data handling engine to carry out data handling of convolution input data as an example, in the process of using the data handling engine to carry out data handling, in order to take a data value of a length of MALU_K from the convolution input data, the AI processor needs to discretely read data from the convolution input data in the Cut_Wi direction in sections, and carry out data handling through the bus in sequence, resulting in a problem of low bandwidth utilization when the bus bandwidth is greater than MALU_K.
[0096] Indicatively, Figure 5 Taking the logical expression form corresponding to the convolution input data in the example, when the data is arranged in NHWC format, it includes Wi direction 501, Hi direction 502 and Ci direction 503, and the data is first continuous in the Ci direction 503 during physical storage. When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, in the related technology, the convolution input data is directly read from the Ci direction 503 according to the array row number 504 of MALU_K, so as to obtain the data block 505, and each segment of data in the MALU_K direction in the data block 505 is discretely distributed in the physical storage of the convolution input data, and when the bus bandwidth is greater than MALU_K, the problem of low bandwidth utilization will also arise during data transportation.
[0097] Indicatively, Figure 5 Taking the logical expression form corresponding to the convolution weight data in as an example, when the data is arranged in the KhKwCiCo format, including the Ci direction 506 and the Co direction 507, the data is first continuous in the Co direction 507 during the physical storage process. When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, in the related technology, the convolution weight data is directly read from the Ci direction 506 according to the number of array rows of MALU_K and from the Co direction 507 according to the number of array columns of MALU_N, so that each segment of data in each data block 508 in the MALU_N direction is discretely distributed in the physical storage of the convolution weight data, and when the bus bandwidth is greater than MALU_N, the problem of low bandwidth utilization will also arise during data transportation.
[0098] Indicatively, Figure 6 As shown, it shows a schematic diagram of data processing through a systolic array based on a target data format provided by an exemplary embodiment of the present application. Taking the use of a data handling engine to carry out data handling of convolution input data as an example, in the process of using the data handling engine to carry out data handling, the AI processor, in order to take a data value of a length of MALU_K from the convolution input data, because the data in the target data format is continuously stored along the direction of MALU_K and then along the direction of Cut_Wi, when the bus bandwidth is greater than MALU_K, the AI processor can directly read the data values greater than MALU_K along the direction of MALU_K and then along the direction of Cut_Wi, thereby making full use of the bus bandwidth.
[0099] Indicatively, Figure 6Taking the logical expression form corresponding to the convolution input data as an example, when the data is arranged in the target data format of NCi1HiWiCk0, it includes Wi direction 601, Hi direction 602 and Ci1 (Ci / Ck0) direction 603, and the data is first continuous in the Ci1 direction 603 during the physical storage process. When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, since the data is first continuous in the Ci1 direction 603 during the physical storage process, the embodiment of the present application can directly read the data block 604 continuously according to the bus width, thereby improving the bus bandwidth utilization.
[0100] Indicatively, Figure 6 Taking the logical expression form corresponding to the convolution weight data in the example, when the data is arranged in the target data format of Co1Ci1KhKwCk0Cn0, including Ci1(Ci / Ck0) direction 605 and Co1(Co / Cn0) direction 606, the data is first continuous in the Co1 direction 606 during the physical storage process. When the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, since the data is first continuous in the Co1 direction 606 during the physical storage process, the embodiment of the present application can directly read the data block 607 according to the bus width continuity, thereby improving the bus bandwidth utilization.
[0101] In some embodiments, considering that the channel dimensions of convolution data in different deep learning networks may be different, and the total number of channels of the convolution data may be less than the number of array rows or array columns of the systolic array, for convolution input data with a total number of input channels less than the number of array rows of the systolic array, or convolution output data with a total number of output channels less than the number of array columns of the systolic array, different target data formats need to be used to store the convolution data.
[0102] In a possible implementation, when the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, the convolution input data may be stored in a fourth target data format, wherein the number of input channels represented by the fourth target data format is the total number of input channels.
[0103] In one possible implementation, the number of array rows of the systolic array can be represented as Ck0, and the total number of input channels of the convolution input data is Ci, Ci is less than Ck0, so that the convolution input data can be represented as NHiWiCi, where N represents a set of images or feature maps, Hi represents the height of the input image or feature map, and Wi represents the width of the input image or feature map.
[0104] In a possible implementation, when the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, and the total number of output channels of the convolution output data is not less than the number of array columns of the systolic array, the convolution weight data may be stored in a fifth target data format, wherein the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array.
[0105] In one possible implementation, the number of array rows of the systolic array can be represented as Ck0, the number of array columns of the systolic array can be represented as Cn0, Ci is the total number of input channels, Co is the total number of output channels, and Ci is less than Ck0, Co is greater than Cn0, so that the convolution weight data can be expressed as Co1KwCiKhCn0, wherein Kh represents the height of the convolution kernel, Kw represents the width of the convolution kernel, and Co1 represents the value of Co / Cn0 rounded up.
[0106] In a possible implementation, when the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, the convolution output data may be stored in a sixth target data format, wherein the number of output channels represented by the sixth target data format is the total number of output channels.
[0107] In one possible implementation, the number of array columns of the systolic array can be expressed as Cn0, and the total number of output channels of the convolution output data is Co, Co is less than Cn0, so that the convolution output data can be expressed as NHoWoCo, wherein N represents a set of images or feature maps, Ho represents the height of the output image or feature map, and Wo represents the width of the output image or feature map.
[0108] In a possible implementation, when the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, and the total number of input channels of the convolution input data is not less than the number of array rows of the systolic array, the convolution weight data may be stored in a seventh target data format, wherein the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is the total number of output channels.
[0109] In one possible implementation, the number of array rows of the systolic array can be represented as Ck0, the number of array columns of the systolic array can be represented as Cn0, Ci is the total number of input channels, Co is the total number of output channels, and Ci is greater than Ck0, and Co is less than Cn0, so that the convolution weight data can be expressed as Ci1KhKwCk0ECn0, wherein Kh represents the height of the convolution kernel, Kw represents the width of the convolution kernel, Ci1 is the value of Ci / Ck0 rounded up, and E is the expansion coefficient in the Co small special optimization algorithm, which is jointly determined by the computing power, bandwidth and convolution parameters.
[0110] In the above embodiment, for convolution data, by comparing the total number of input channels of convolution input data with the number of array rows of the systolic array, and the total number of output channels of convolution output data with the number of array columns of the systolic array, the target data formats corresponding to the convolution input data, convolution weight data, and convolution output data are determined, thereby achieving different target data formats according to different data situations. Since the bus width is greater than the data width corresponding to the array dimension, compared with data reading according to the data width corresponding to the array dimension, for convolution data, data reading can be achieved directly according to the bus width, so that the data widths of the input sub-data and the weight sub-data are equal to the bus width, thereby optimizing data handling efficiency.
[0111] In a possible implementation, during processing of matrix data, the first operation data may be left matrix input data, the second operation data may be right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data.
[0112] In some embodiments, in the expression of dimensional information of the matrix-type network, the left matrix is represented as MK, the right matrix is represented as KN, and the result matrix is represented as MN. In addition, in commonly used deep learning networks, batch data of other dimensions will be added on this basis.
[0113] In some embodiments, when the first operation data is the left matrix input data, in order to fully utilize the bus bandwidth in the process of data transfer of the left matrix input data, the left matrix input data may be stored in an eighth target data format, wherein the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array.
[0114] In one possible implementation, the number of array rows of the systolic array is Wk0, so that the left matrix input data can be expressed as B0B1W1HWk0, wherein B0 and B1 respectively represent batch data of two dimensions, H represents the number of matrix rows of the left matrix, W1 is the value rounded up of K / Wk0, and K is the number of matrix columns of the left matrix.
[0115] In some embodiments, when the second operation data is right matrix input data, in order to fully utilize the bus bandwidth during data transfer of the right matrix input data, the right matrix input data can be stored in a ninth target data format, wherein the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array.
[0116] In a possible implementation, the number of array columns of the systolic array is Wn0, so that the right matrix input data can be expressed as B0B1W1HWn0, wherein B0 and B1 respectively represent batch data of two dimensions, and the batches in the left matrix input data and the right matrix input data may not be equal, but when they are not equal, one of them must be 1, H represents the number of matrix rows of the right matrix, W1 is the value rounded up of N / Wn0, and N is the number of matrix columns of the right matrix.
[0117] In some embodiments, when the operation result of the first operation data and the second operation data is matrix output data, matrix output data using the tenth target data format can be obtained by performing matrix operations through a matrix operation unit based on left matrix input data using the eighth target data format and right matrix input data using the ninth target data format, wherein the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.
[0118] In one possible implementation, the matrix output data may be expressed as B0B1W1HWn0, wherein BO is the maximum value of B0 in the left matrix input data and the right matrix input data, B1 is the maximum value of B1 in the left matrix input data and the right matrix input data, W1 and Wn0 are the same as the W1 and Wn0 values of the right matrix input data, and H is the same as the H value of the left matrix input data.
[0119] In some embodiments, after the left matrix input data in the eighth target data format and the right matrix input data in the ninth target data format are stored in the first memory of the AI processor, the AI processor can use the data transfer engine to transfer the left matrix input sub-data and the right matrix input sub-data according to the bus bit width, and store them in the second memory inside the matrix operation engine.
[0120] In one possible implementation, the AI processor uses a data transfer engine to sequentially read the first left matrix input sub-data and the second left matrix input sub-data from the first memory according to the bus bit width, wherein the first left matrix input sub-data and the second left matrix input sub-data are continuously stored data, and the data bit width of the first left matrix input sub-data is equal to the bus bit width, and the data bit width of the second left matrix input sub-data is equal to the bus bit width.
[0121] In one possible implementation, the AI processor uses a data transfer engine to sequentially read the first right matrix input sub-data and the second right matrix input sub-data from the first memory according to the bus bit width, wherein the first right matrix input sub-data and the second right matrix input sub-data are continuously stored data, and the data bit width of the first right matrix input sub-data is equal to the bus bit width, and the data bit width of the second right matrix input sub-data is equal to the bus bit width.
[0122] Indicatively, Figure 7 As shown, it shows a schematic diagram of data processing through a systolic array based on a general data format in another related technology. Taking the use of a data handling engine to carry out data handling on the left matrix input data as an example, in the process of using the data handling engine to carry out data handling, in order to obtain a data value of a length of MALU_K from the left matrix input data, the AI processor needs to discretely read data from the convolution input data in the M direction in sequence, and carry out data handling through the bus in sequence, resulting in a problem of low bandwidth utilization when the bus bandwidth is greater than MALU_K.
[0123] Indicatively, Figure 7 Taking the input data of the left matrix as an example, when the data is arranged in the MK format, it includes the M direction 701 and the K direction 702, and the data is first continuous in the K direction 702 during the physical storage process, so that when the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, in the related art, data is read directly from the K direction 702 according to the number of array rows of MALU_K, so that each segment of data in the MALU_K direction in the data block 703 is discretely distributed in the physical storage of the left matrix input data, and when the bus bandwidth is greater than MALU_K, the problem of low bandwidth utilization will also arise during data transportation. Figure 7 The same is true for the input data of the middle and right matrices, and the embodiments of the present application will not be described here.
[0124] Indicatively, Figure 8 As shown, it shows a schematic diagram of data processing through a systolic array based on a target data format provided by another exemplary embodiment of the present application. Taking the use of a data handling engine to handle data on the left matrix input data as an example, in the process of using the data handling engine to handle data, the AI processor, in order to obtain a data value of a length of MALU_K from the left matrix input data, because the data is stored continuously along the MALU_K and then along the H_Left direction in the target data format, when the bus bandwidth is greater than MALU_K, the AI processor can directly read the data values greater than MALU_K along the MALU_K and then along the H_Left direction, thereby making full use of the bus bandwidth.
[0125] Indicatively, Figure 8 Taking the input data of the middle left matrix as an example, when the data is arranged in the target data format of B0B1W1HWk0, including the H direction 801 and the W1(K / Wk0) direction 802, the data is first continuous in the W1 direction 802 during the physical storage process. Therefore, when the number of array rows of the systolic array is MALU_K and the number of array columns is MALU_N, since the data is first continuous in the W1 direction 802 during the physical storage process, the embodiment of the present application can directly read the data block 803 continuously according to the bus width, thereby improving the bus bandwidth utilization. Figure 8 The same is true for the input data of the middle and right matrices, and the embodiments of the present application will not be described here.
[0126] In the above embodiment, for matrix data, the specific target data format adopted by the left matrix input data, the right matrix input data and the matrix output data is determined according to the matrix column number of the left matrix input data and the right matrix input data, and in combination with the array row number and the array column number of the systolic array. Since the bus width is larger than the data width corresponding to the array dimension, compared with data reading according to the data width corresponding to the array dimension, for matrix data, data reading can be directly performed according to the bus width, so that the data width of the input sub-data is equal to the bus width, thereby optimizing the efficiency of data transfer for matrix data.
[0127] In some embodiments, considering that in deep learning networks there are often scenarios where convolution operators and matrix operators are used alternately, in order to improve the data processing efficiency in the AI processor, a data format converter may also be provided in the AI processor.
[0128] In a possible implementation, when the operation result of the first operation data and the second operation data is convolution output data, and the convolution output data is used for subsequent matrix calculations, the data format converter can be used to convert the data format of the convolution output data based on the target data format corresponding to the matrix input data.
[0129] In one possible implementation, the data format of the convolution output data is NC1HWC0, and the data format of the matrix input data is B0B1W1HW0. Then, B0 can be set to 1, B1 can be set to N, W1 can be set to C1, H can be set to H×W, and W0 can be set to C0, thereby achieving the convolution output data being equivalent to the matrix input data.
[0130] In a possible implementation, when the operation result of the first operation data and the second operation data is matrix output data, and the matrix output data is used for subsequent convolution calculations, the data format converter can be used to convert the data format of the matrix output data based on the target data format corresponding to the convolution input data.
[0131] In one possible implementation, the data format of the matrix output data is B0B1W1HW0, and the data format of the convolution input data is NC1HWC0. Then, N can be set to B0×B1, C1 can be set to W1, H can be set to 1, W can be set to H, and C0 can be set to W0, thereby achieving the equivalent of the matrix output data to the convolution input data.
[0132] In the above embodiments, when the convolution output data or the matrix output data adopts the target data format and there are subsequent operations with different data types, the data equivalence relationship between the convolution output data and the matrix input data, or the data equivalence relationship between the matrix output data and the convolution input data can be determined, that is, the data equivalence between the convolution output data and the matrix input data, or the data equivalence between the matrix output data and the convolution input data can be achieved, thereby improving the data processing efficiency in the AI processor.
[0133] In some embodiments, considering that after the AI processor is integrated into the deep learning network framework, data calculation through the matrix operation engine is only one of the links, that is, the calculation data obtained based on the previous process may not be in the target data format, so directly moving the data may affect the efficiency of data movement. Therefore, a data format converter can also be provided in the AI processor to convert the data format of the calculation data into the target data format.
[0134] In a possible implementation, a data format converter is used to read first operation data and second operation data from a first memory, the first operation data and the second operation data are in an original data format, and the data dimension represented by the original data format does not match the array dimension of the systolic array in the matrix operation unit in the matrix operation engine, and then the data format converter converts the data format of the first operation data and the second operation data from the original data format to the target data format by performing data format conversion on the first operation data and the second operation data, and writes the first operation data and the second operation data into the first memory.
[0135] In an illustrative example, the data format converter first reads the convolution input data in the original data format (NCHW) from the first memory, and converts the convolution input data according to the target data format (NCi1HiWiCk0), thereby storing the converted convolution input data in the target data format back in the first memory.
[0136] In one possible implementation, when an AI processor is integrated into a general deep learning framework as a computing acceleration device, in order to reduce the complexity of integrating the AI processor into the deep learning framework, corresponding format processing can also be provided at the operator layer and the framework layer. For example, at the operator layer, full support for the original data format and the target data format is added; at the framework layer, format information about the original data format and the target data format is provided.
[0137] Indicatively, Fig. 9 As shown, Math Dim is set to correspond to the operation data in the original data format in the algorithm framework, and Data Dim corresponds to the operation data in the target data format on the AI processor, so that in the process of data calculation through the deep learning framework, the operation data in the target data format can be directly used for data processing. During the debugging process, when reading data, that is, when transmitting data from the device (Device) to the host (Host), the operation data in the target data format can be converted into the operation data in the original data format according to the format information of the original data format and the target data format for viewing and analysis.
[0138] Please refer to Fig.10 , which shows a flow chart of a data processing method provided by an exemplary embodiment of the present application, the method is used in the AI processor in the above embodiment, and the method includes:
[0139] Step 1001, storing first operation data and second operation data in a first memory, wherein the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine.
[0140] Step 1002, based on the bus bit width, read the first operator data and the second operator data from the first memory through the data transfer engine, and transfer the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and the bus bit width is greater than the data bit width corresponding to the array dimension.
[0141] Step 1003, performing matrix operation on the first operator data and the second operator data in the second memory through the matrix operation unit in the matrix operation engine to obtain a sub-operation result, and the data dimension of the sub-operation result matches the array dimension.
[0142] Step 1004, accumulating the sub-operation results through the accumulator in the matrix operation engine to obtain the matrix operation result of the first operation data and the second operation data.
[0143] In some embodiments, in the process of processing convolution data, the first operation data is the convolution input data, the second operation data is the convolution weight data, and the operation result of the first operation data and the second operation data is the convolution output data; in the process of processing matrix data, the first operation data is the left matrix input data, the second operation data is the right matrix input data, and the operation result of the first operation data and the second operation data is the matrix output data.
[0144] In some embodiments, when the first operation data is convolution input data, the second operation data is convolution weight data, and the operation results of the first operation data and the second operation data are convolution output data, the convolution input data adopts a first target data format, and the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array; the convolution weight data adopts a second target data format, and the number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array; the convolution output data adopts a third target data format, and the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.
[0145] In some embodiments, when the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, the convolution input data adopts a fourth target data format, and the number of input channels represented by the fourth target data format is the total number of input channels; the convolution weight data adopts a fifth target data format, and the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array.
[0146] In some embodiments, when the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, the convolution output data adopts a sixth target data format, and the number of output channels represented by the sixth target data format is the total number of output channels; the convolution weight data adopts a seventh target data format, and the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is the total number of output channels.
[0147] In some embodiments, the data handling engine is used to read the first convolution input sub-data and the second convolution input sub-data from the first memory in sequence based on the bus bit width, the first convolution input sub-data and the second convolution input sub-data are continuously stored data, and the data bit width of the first convolution input sub-data and the data bit width of the second convolution input sub-data are both equal to the bus bit width; based on the bus bit width, read the first convolution weight sub-data and the second convolution weight sub-data from the first memory in sequence, the first convolution weight sub-data and the second convolution weight sub-data are continuously stored data, and the data bit width of the first convolution weight sub-data and the data bit width of the second convolution weight sub-data are both equal to the bus bit width.
[0148] In some embodiments, when the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data, the left matrix input data adopts an eighth target data format, and the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array; the right matrix input data adopts a ninth target data format, and the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array; the matrix output data adopts a tenth target data format, and the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.
[0149] In some embodiments, a data handling engine is used to read first left matrix input sub-data and second left matrix input sub-data from a first memory in sequence based on a bus bit width, the first left matrix input sub-data and the second left matrix input sub-data are continuously stored data, and the data bit width of the first left matrix input sub-data and the data bit width of the second left matrix input sub-data are both equal to the bus bit width; based on the bus bit width, read first right matrix input sub-data and second right matrix input sub-data from a first memory in sequence, the first right matrix input sub-data and the second right matrix input sub-data are continuously stored data, and the data bit width of the first right matrix input sub-data and the data bit width of the second right matrix input sub-data are both equal to the bus bit width.
[0150] In some embodiments, the AI processor further includes a data format converter;
[0151] A data format converter is used to convert the data format of the convolution output data based on the target data format corresponding to the matrix input data when the operation result of the first operation data and the second operation data is the convolution output data and the convolution output data is used for the subsequent matrix calculation; and to convert the data format of the matrix output data based on the target data format corresponding to the convolution input data when the operation result of the first operation data and the second operation data is the matrix output data and the matrix output data is used for the subsequent convolution calculation.
[0152] In some embodiments, the AI processor further includes a data format converter;
[0153] A data format converter, used for reading first operation data and second operation data from the first memory, the first operation data and the second operation data are in original data format, and the data dimension represented by the original data format does not match the array dimension of the systolic array in the matrix operation unit in the matrix operation engine;
[0154] The data format converter is also used to convert the data formats of the first operation data and the second operation data, so that the data formats of the first operation data and the second operation data are converted from the original data format to the target data format, and the first operation data and the second operation data are written into the first memory.
[0155] In summary, in the embodiment of the present application, by storing the first operation data and the second operation data in the target data format in the first memory, the data dimension of the first operation data and the second operation data matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine, so that the data handling engine can read the first operator data and the second operator data from the first memory according to the bus width, and carry the first operator data and the second operator data to the second memory inside the matrix operation engine through the bus, and then the matrix operation engine performs matrix operation on the first operator data and the second operator data through the matrix operation unit to obtain the sub-operation result, and accumulates the sub-operation result through the accumulator to obtain the matrix operation result of the first operation data and the second operation data. Using the AI processor provided in the embodiment of the present application, by storing the operation data in the target data format in the first memory, it is possible to realize the reading of operation data according to the bus width. Since the bus width is greater than the data width corresponding to the array dimension, compared with reading data according to the data width corresponding to the array dimension, the AI processor provided in the embodiment of the present application can improve the bandwidth utilization of data handling during matrix operation, thereby improving the efficiency of matrix operation.
[0156] Please refer to Fig.11, which shows a block diagram of a computer device 1100 provided by an exemplary embodiment of the present application. The computer device 1100 may be a portable mobile terminal, such as a smart phone, a tablet computer, a Moving Picture Experts Group Audio Layer III (MP3) player, or a Moving Picture Experts Group Audio Layer IV (MP4) player. The computer device 1100 may also be called a user device, a portable terminal, a workstation, a server, or other names.
[0157] Typically, the computer device 1100 includes: an AI processor 1101 and a memory 1102 .
[0158] The AI processor 1101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The AI processor 1101 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The AI processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the AI processor 1101 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the AI processor 1101 may also be used to process computing operations related to machine learning.
[0159] The memory 1102 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash memory storage devices.
[0160] In some embodiments, the computer device 1100 may also optionally include a peripheral device interface 1103 and at least one peripheral device.
[0161] Those skilled in the art will understand that Fig.11The structure shown in the figure does not constitute a limitation on the computer device 1100, and the computer device 1100 may include more or less components than shown in the figure, or combine some components, or adopt a different component arrangement.
[0162] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0163] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. An AI processor, It is characterized in that The AI processor includes: a matrix operation engine, a data handling engine and a first memory, wherein the matrix operation engine is connected to the data handling engine via a bus; The first memory is used to store first operation data and second operation data, the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine; The data transport engine is configured to read the first operator data and the second operator data from the first memory based on a bus bit width, and transport the first operator data and the second operator data to a second memory inside the matrix operation engine through the bus, wherein the bus bit width is greater than a data bit width corresponding to the array dimension; The matrix operation engine is used to perform a matrix operation on the first operator data and the second operator data in the second memory through a matrix operation unit to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension; The matrix operation engine is further used to accumulate the sub-operation results through an accumulator to obtain a matrix operation result of the first operation data and the second operation data.
2. The AI processor according to claim 1, It is characterized in that In the process of processing convolution data, the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data; In the process of processing matrix data, the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data.
3. The AI processor according to claim 2, It is characterized in that In the case where the first operation data is convolution input data, the second operation data is convolution weight data, and the operation result of the first operation data and the second operation data is convolution output data, The convolution input data adopts a first target data format, and the number of sub-input channels represented by the first target data format is equal to the number of array rows of the systolic array; The convolution weight data adopts a second target data format, the number of sub-input channels represented by the second target data format is equal to the number of array rows of the systolic array, and the number of sub-output channels represented by the second target data format is equal to the number of array columns of the systolic array; The convolution output data adopts a third target data format, and the number of sub-output channels represented by the third target data format is equal to the number of array columns of the systolic array.
4. The AI processor according to claim 3, It is characterized in that In the case where the total number of input channels of the convolution input data is less than the number of array rows of the systolic array, The convolution input data adopts a fourth target data format, and the number of input channels represented by the fourth target data format is the total number of input channels; The convolution weight data adopts a fifth target data format, the number of input channels represented by the fifth target data format is the total number of input channels, and the number of sub-output channels represented by the fifth target data format is equal to the number of array columns of the systolic array.
5. The AI processor according to claim 3, It is characterized in that In the case where the total number of output channels of the convolution output data is less than the number of array columns of the systolic array, The convolution output data adopts a sixth target data format, and the number of output channels represented by the sixth target data format is the total number of output channels; The convolution weight data adopts a seventh target data format, the number of sub-input channels represented by the seventh target data format is equal to the number of array rows of the systolic array, and the number of output channels represented by the seventh target data format is the total number of output channels.
6. The AI processor according to any one of claims 3 to 5, It is characterized in that The data handling engine is used to: Based on the bus bit width, first convolution input sub-data and second convolution input sub-data are sequentially read from the first memory, where the first convolution input sub-data and the second convolution input sub-data are continuously stored data, and the data bit width of the first convolution input sub-data and the data bit width of the second convolution input sub-data are both equal to the bus bit width; Based on the bus bit width, the first convolution weight sub-data and the second convolution weight sub-data are read from the first memory in sequence, the first convolution weight sub-data and the second convolution weight sub-data are continuously stored data, and the data bit width of the first convolution weight sub-data and the data bit width of the second convolution weight sub-data are both equal to the bus bit width.
7. The AI processor according to any one of claims 2 to 6, It is characterized in that In the case where the first operation data is left matrix input data, the second operation data is right matrix input data, and the operation results of the first operation data and the second operation data are matrix output data, The left matrix input data adopts an eighth target data format, and the number of submatrix columns represented by the eighth target data format is equal to the number of array rows of the systolic array; The right matrix input data adopts a ninth target data format, and the number of submatrix columns represented by the ninth target data format is equal to the number of array columns of the systolic array; The matrix output data adopts the tenth target data format, the number of matrix rows represented by the tenth target data format is equal to the number of matrix rows represented by the eighth target data format, and the number of matrix columns represented by the tenth target data format is equal to the number of matrix columns represented by the ninth target data format.
8. The AI processor according to claim 7, It is characterized in that The data handling engine is used for: Based on the bus bit width, first left matrix input sub-data and second left matrix input sub-data are sequentially read from the first memory, the first left matrix input sub-data and the second left matrix input sub-data are continuously stored data, and the data bit width of the first left matrix input sub-data and the data bit width of the second left matrix input sub-data are both equal to the bus bit width; Based on the bus bit width, the first right matrix input sub-data and the second right matrix input sub-data are read from the first memory in sequence, the first right matrix input sub-data and the second right matrix input sub-data are continuously stored data, and the data bit width of the first right matrix input sub-data and the data bit width of the second right matrix input sub-data are both equal to the bus bit width.
9. The AI processor according to any one of claims 1 to 8, It is characterized in that The AI processor also includes a data format converter; The data format converter is used for: When the operation results of the first operation data and the second operation data are convolution output data, and the convolution output data is used for subsequent matrix calculation, performing data format conversion on the convolution output data based on the target data format corresponding to the matrix input data; When the operation results of the first operation data and the second operation data are matrix output data, and the matrix output data is used for subsequent convolution calculation, the data format of the matrix output data is converted based on the target data format corresponding to the convolution input data.
10. The AI processor according to any one of claims 1 to 8, It is characterized in that The AI processor also includes a data format converter; The data format converter is used to read the first operation data and the second operation data from the first memory, the first operation data and the second operation data are in original data format, and the data dimension represented by the original data format does not match the array dimension of the systolic array in the matrix operation unit in the matrix operation engine; The data format converter is also used to convert the data formats of the first operation data and the second operation data, so that the data formats of the first operation data and the second operation data are converted from the original data format to the target data format, and the first operation data and the second operation data are written into the first memory.
11. A data processing method, It is characterized in that The method is used for an AI processor, the AI processor includes a matrix operation engine, a data handling engine and a first memory, the matrix operation engine and the data handling engine are connected via a bus; The method comprises: storing first operation data and second operation data by the first memory, wherein the first operation data and the second operation data are in a target data format, and the data dimension represented by the target data format matches the array dimension of the systolic array in the matrix operation unit in the matrix operation engine; Based on a bus bit width, reading first operator data and second operator data from the first memory through the data transfer engine, and transferring the first operator data and the second operator data to a second memory inside the matrix operation engine through the bus, wherein the bus bit width is greater than a data bit width corresponding to the array dimension; Performing a matrix operation on the first operator data and the second operator data in the second memory by a matrix operation unit in the matrix operation engine to obtain a sub-operation result, wherein the data dimension of the sub-operation result matches the array dimension; The sub-operation results are accumulated by an accumulator in the matrix operation engine to obtain a matrix operation result of the first operation data and the second operation data.
12. A computer device, It is characterized in that The computer device comprises a memory and an AI processor as claimed in any one of claims 1 to 10, wherein the memory stores at least one instruction, and the at least one instruction is used to be executed by the AI processor.