Artificial intelligence chip and computing board, data processing method and electronic equipment
By introducing conversion circuits and pipeline modules of multiple buffers into the AI chip, the ping-pong cache mechanism is realized, which solves the problem of performance bottlenecks in the data preprocessing and computing process of artificial intelligence chips in the prior art, and improves processing efficiency and overall performance.
Patent Information
- Application Number
- CN202010877800.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-08-27
AI Technical Summary
There are performance bottlenecks in existing artificial intelligence chips in data preprocessing and computing, especially due to data format requirements and data continuity requirements, which leads to frequent calls of software-side operators, which increases execution overhead and affects overall performance.
The first pipeline module is introduced into the AI chip, including a conversion circuit and multiple buffers. By alternately writing and reading data, a ping-pong cache mechanism is realized, reducing the waiting time for the computing circuit, and simplifying the data format conversion and continuous conversion process.
By reducing the number of software-side operator calls, the overall performance is improved, the data processing process is simplified, and fine-grained pipeline operation is realized, which significantly improves the processing efficiency of the AI chip.
Smart Images

Figure CN114115995B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence chips, and in particular to artificial intelligence chips and computing boards, data processing methods and electronic devices. Background Art
[0002] At present, with the popularization of smart devices, artificial intelligence (AI) technology is developing rapidly. Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0003] As an important research direction in the field of artificial intelligence, deep learning (such as convolutional neural networks) is used to continuously train a large number of samples to obtain the characteristic information of the samples. These characteristic information can be used to identify a new test sample when it is exposed to the characteristic information in it. The model is applied to the machine, allowing the machine to obtain human-like recognition capabilities. The sample training and information extraction stages involve a very large amount of computing, and the AI chip needs to calculate a large amount of data during this operation.
[0004] like Figure 1 As shown, the AI chip includes a data preprocessing module (data preprocess), a cache (Cache) and an operation circuit. Among them, since the operation circuit has requirements for the data to be processed, for example, there are requirements for the format of the data (such as requiring the data to be in 5D format). Therefore, when the data preprocessing module performs data conversion, software intervention is required to call multiple operators, map multiple operators to the data preprocessing module, and the data processing module converts the data. On the software side, these operators are regarded as additional operator calls when called, increasing the number of operator calls on the software side, thereby increasing the execution overhead on the software side and affecting the overall performance. And by calling the operator to map to the data processing module, the data processing module needs to convert all the data before the operation circuit performs the calculation. There is waiting time during the execution of the operation circuit, and the chip operation efficiency is reduced. Summary of the invention
[0005] The embodiments of the present application provide an artificial intelligence chip and a computing board, a data processing method and an electronic device for improving the processing efficiency of the artificial intelligence chip.
[0006] In a first aspect, an embodiment of the present application provides an artificial intelligence chip, which includes: a first pipeline module and an operation circuit connected in sequence, the first pipeline module includes a conversion circuit and multiple buffers connected to the conversion circuit; wherein the conversion circuit is used to obtain data of a first feature to be converted, and convert the data of the first feature into data of a second feature, and the data of the second feature is data applicable to the operation circuit when performing operations; the conversion circuit writes the second feature data alternately into multiple buffers, wherein alternatingly writing into multiple buffers means writing the data into the first buffer, and when the buffer is full, continuing to write the data into the second buffer until the last buffer is full, and then switching the state, writing the data into the first buffer again, writing the data into the second buffer, and so on in a cycle; the operation circuit is used to read the data of the second feature from the full buffer when one of the multiple buffers is full, operate on the data of the second feature, obtain result data, and output the result data.
[0007] In this example, a first pipeline module is added to the AI chip, and the first pipeline module includes a conversion circuit and multiple buffers. For example, the conversion circuit is used to support continuous conversion and / or format conversion functions. In this application, a hardware-specific module is added to support the data conversion function. The conversion circuit writes the second feature data alternately into multiple buffers, and multiple buffers are used to implement a ping-pong cache mechanism. When one of the multiple buffers is full, the operation circuit can be started to operate on the read data. In this application, there is no need to call the operator multiple times to pre-process the data to meet the requirements of the operation circuit for the data, simplify the operator call on the software side, improve the overall performance and simplify the call overhead and process. In addition, the conversion circuit writes the converted data alternately into multiple buffers. When a buffer is full, the operation circuit directly reads the data from the full buffer, and implements a fine pipeline operation based on the ping-pong cache mechanism. The operation circuit and the conversion circuit process the data in parallel, and the operation circuit does not need to wait for time. Compared with the traditional method, all data need to be pre-processed before the operation is performed. In this application, the processing efficiency of the chip is greatly improved.
[0008] In an optional implementation, the conversion circuit is a first format conversion circuit, and the buffer is a conversion format buffer; the chip also includes a second pipeline module connected to the operation circuit, and the second pipeline module includes a second format conversion circuit and multiple target format buffers connected to the second format conversion circuit; wherein the first format conversion circuit is used to convert data in the first format into data in the second format, and write the data in the second format alternately into multiple conversion format buffers; the operation circuit is used to read the second format data from the conversion format buffer when one of the multiple conversion format buffers is full, process the second format data, and obtain result data in the second format; the second format conversion circuit is used to convert the result data in the second format into result data in the first format, and write the result data in the first format alternately into multiple target format buffers.
[0009] In this example, two pipeline modules are added. The first pipeline module includes a first format conversion circuit, and the second pipeline module includes a second format conversion circuit. A dedicated circuit supporting format conversion is added to realize the format conversion function. There is no need to call the format conversion operator to realize format conversion, which simplifies the operator call on the software side, improves the overall performance and simplifies the call overhead and process. Based on the ping-pong cache mechanism, fine pipeline operation is realized. The first format conversion circuit, the operation circuit and the second format conversion circuit process data in parallel. The operation circuit does not need to wait, which greatly improves the processing efficiency of the chip.
[0010] In an optional implementation, the first pipeline module also includes a continuous conversion circuit and multiple continuous data buffers connected to the continuous conversion circuit; wherein the continuous conversion circuit is used to convert discontinuous data into continuous data, and write the continuous data alternately into multiple continuous data buffers; the first format conversion circuit is also used to read continuous data from the full continuous data buffer when one of the multiple continuous data buffers is full of continuous data, the continuous data is data in a first format, and convert the data in the first format into data in a second format.
[0011] In this example, the first pipeline module also includes a continuous conversion circuit. When the data to be converted is discontinuous data and is in the first format, the discontinuous data is converted into continuous data through the continuous conversion circuit. There is no need to call the continuous conversion operator to achieve data continuity conversion, which simplifies the operator call on the software side, improves the overall performance and simplifies the call overhead and process. Based on the ping-pong cache mechanism, the fine pipeline operation is realized. The continuous conversion circuit and the first format conversion circuit process data in parallel. The operation circuit does not need to wait, which greatly improves the processing efficiency of the chip.
[0012] In an optional implementation, the conversion circuit is a continuous conversion circuit, and the buffer is a continuous data buffer; wherein the continuous conversion circuit is used to convert discontinuous data into continuous data, and write the continuous data alternately into multiple continuous data buffers; the operation circuit is used to read continuous data from the full continuous data buffer when one of the multiple continuous data buffers is full, perform operations on the continuous data, obtain result data, and output the result data.
[0013] In this example, when the data to be converted is discontinuous, the discontinuous data is converted into continuous data through the continuous conversion circuit, and there is no need to call the continuous conversion operator to achieve data continuity conversion, which simplifies the operator call on the software side, improves the overall performance and simplifies the call overhead and process. Based on the ping-pong cache mechanism, fine pipeline operation is implemented, the continuous conversion circuit and the operation circuit process data in parallel, and the operation circuit does not need to wait time, which greatly improves the processing efficiency of the chip.
[0014] In an optional implementation, the operation circuit includes a convolution calculation unit and / or a vector calculation unit.
[0015] In an optional implementation, the chip further includes a first buffer and a second buffer; wherein the first buffer is used to output the data to be converted to the first pipeline module; and the second buffer is used to cache the result data. In the present application, the first buffer and the second buffer are added, the first pipeline module reads data from the first buffer, and the second pipeline module outputs the result data to the second buffer. Compared with the traditional method, the present application reduces the number of interactions with the external memory, thereby improving the system performance.
[0016] In the second aspect, the present application provides a data processing method, which is applied to an artificial intelligence chip, the chip includes a first pipeline module and an operation circuit connected in sequence, the first pipeline module includes a conversion circuit and multiple buffers connected to the conversion circuit, the method also includes: the conversion circuit obtains data of a first feature to be converted, converts the data of the first feature into data of a second feature, the data of the second feature is data applicable to the operation circuit when performing operations; and writes the second feature data alternately into multiple buffers; when one of the multiple buffers is full, the operation circuit reads the data of the second feature from the full buffer, operates on the data of the second feature, and obtains result data; and outputs the result data.
[0017] In this example, a first pipeline module is added to the AI chip, and the first pipeline module includes a conversion circuit and multiple buffers. For example, the conversion circuit is used to support continuous conversion and / or format conversion functions. In this application, a hardware-specific module is added to support the data conversion function. The conversion circuit writes the second feature data alternately into multiple buffers, and the multiple buffers are used to implement the ping-pong cache mechanism. When one of the multiple buffers is full, the operation circuit can be started to operate on the read data. In this application, there is no need to call the operator multiple times to pre-process the data to meet the requirements of the operation circuit for the data, simplify the operator call on the software side, improve the overall performance and simplify the call overhead and process. In addition, the conversion circuit writes the converted data alternately into multiple buffers. When a buffer is full, the operation circuit directly reads the data from the full buffer, and implements a fine pipeline operation based on the ping-pong cache mechanism. The operation circuit and the conversion circuit process the data in parallel, and the operation circuit does not need to wait for time. Compared with the traditional method, all data need to be pre-processed before the operation is performed. In this application, the processing efficiency of the chip is greatly improved.
[0018] In an optional implementation, the conversion circuit is a first format conversion circuit, and the buffer is a conversion format buffer; the chip also includes a second pipeline module connected to the operation circuit, the second pipeline module includes a second format conversion circuit and multiple target format buffers connected to the second format conversion circuit; the first format conversion circuit converts data in the first format into data in the second format; and the second format data is alternately written into multiple conversion format buffers; the operation circuit reads the second format data from the conversion format buffer, processes the second format data, and obtains result data in the second format; the second format conversion circuit converts the result data in the second format into result data in the first format, and alternately writes the result data in the first format into multiple target format buffers.
[0019] In this example, two pipeline modules are added. The first pipeline module includes a first format conversion circuit, and the second pipeline module includes a second format conversion circuit, that is, a dedicated circuit supporting format conversion is added to realize the format conversion function. There is no need to call the format conversion operator to realize format conversion, which simplifies the operator call on the software side, improves the overall performance and simplifies the call overhead and process. Based on the ping-pong cache mechanism, fine pipeline operation is realized. The first format conversion circuit, the operation circuit and the second format conversion circuit process data in parallel. The operation circuit does not need to wait, which greatly improves the processing efficiency of the chip.
[0020] In an optional implementation, the first pipeline module also includes a continuous conversion circuit and multiple continuous data buffers connected to the continuous conversion circuit; the method may also include: the continuous conversion circuit converts discontinuous data into continuous data, and writes the continuous data alternately into multiple continuous data buffers; when one of the multiple continuous data buffers is full of continuous data, the first format conversion circuit reads continuous data from the full continuous data buffer, and the continuous data is data in the first format.
[0021] In this example, the first pipeline module also includes a continuous conversion circuit. When the data to be converted is discontinuous data and is in the first format, the discontinuous data is converted into continuous data through the continuous conversion circuit. There is no need to call the continuous conversion operator to achieve data continuity conversion, which simplifies the operator call on the software side, improves the overall performance and simplifies the call overhead and process. Based on the ping-pong cache mechanism, the fine pipeline operation is realized. The continuous conversion circuit and the first format conversion circuit process data in parallel. The operation circuit does not need to wait, which greatly improves the processing efficiency of the chip.
[0022] In an optional implementation, the conversion circuit is a continuous conversion circuit, and the buffer is a continuous data buffer; the continuous conversion circuit converts discontinuous data into continuous data, and writes the continuous data alternately into multiple continuous data buffers; when a continuous data buffer among the multiple continuous data buffers is full, the operation circuit reads continuous data from the full continuous data buffer, operates on the continuous data, obtains result data, and outputs the result data.
[0023] In this example, when the data to be converted is discontinuous, the discontinuous data is converted into continuous data through the continuous conversion circuit, and there is no need to call the continuous conversion operator to achieve data continuity conversion, which simplifies the operator call on the software side, improves the overall performance and simplifies the call overhead and process. Based on the ping-pong cache mechanism, fine pipeline operation is implemented, the continuous conversion circuit and the operation circuit process data in parallel, and the operation circuit does not need to wait time, which greatly improves the processing efficiency of the chip.
[0024] In a third aspect, an embodiment of the present application further provides an artificial intelligence computing board, characterized in that it comprises a communication interface and an artificial intelligence chip as described in any one of the first aspects above; wherein the communication interface is used to connect to a host.
[0025] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory coupled to the processor, and an artificial intelligence computing board according to the second aspect above, wherein the processor and the memory perform data transmission with the artificial intelligence computing board via a communication interface.
[0026] In a fifth aspect, the present application provides a chip system, which includes a processor and the artificial intelligence chip of the first aspect, and the processor transmits data with the artificial intelligence chip. In a possible design, the chip system also includes a memory, which is used to store data to be converted by the artificial intelligence chip and result data after the artificial intelligence chip calculates. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A schematic diagram of the structure of an example of an artificial intelligence chip in a traditional method;
[0028] Figure 2 This is a schematic diagram of the structure of an example of an artificial intelligence chip in an embodiment of the present application;
[0029] Figure 3 This is a structural diagram of another example of an artificial intelligence chip in an embodiment of the present application;
[0030] Figure 4 This is a schematic diagram of a step flow of an example of a data processing method in an embodiment of the present application;
[0031] Figure 5 A schematic diagram of another exemplary step flow of a data processing method in an embodiment of the present application;
[0032] Figure 6 A schematic diagram of another exemplary step flow of a data processing method in an embodiment of the present application;
[0033] Figure 7 This is a structural diagram of another example of an artificial intelligence chip in an embodiment of the present application;
[0034] Figure 8 A schematic diagram of another exemplary step flow of a data processing method in an embodiment of the present application;
[0035] Fig. 9 A schematic diagram of another exemplary step flow of a data processing method in an embodiment of the present application;
[0036] Fig.10 This is a structural diagram of another example of an artificial intelligence chip in an embodiment of the present application;
[0037] Fig.11 A schematic diagram of another exemplary step flow of a data processing method in an embodiment of the present application;
[0038] Fig.12 A schematic diagram of another exemplary step flow of a data processing method in an embodiment of the present application;
[0039] Fig.13This is a schematic diagram of the structure of an example of an artificial intelligence computing board in an embodiment of the present application;
[0040] Fig.14 This is a schematic diagram of the structure of an example of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0041] The technical solution in the embodiment of the present application will be described below in conjunction with the accompanying drawings in the embodiment of the present application. The term "and / or" appearing in the present application can be a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present application generally indicates that the associated objects before and after are in an "or" relationship.
[0042] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices.
[0043] In order to better understand the present application, the terms involved in the present application are first explained.
[0044] Data format: Multidimensional data is stored in a multidimensional array. Usually, the data format in the software is generally 4D format, which means that the data is saved in a four-dimensional array. For example, the feature map of a convolutional neural network is saved in a four-dimensional array, and the four dimensions are batch size (batch, N), feature map height (height, H), feature map width (width, W), and feature map channels (channels, C). Since data can only be stored linearly, because these four dimensions have a corresponding order. Different deep learning frameworks store feature map data in different orders. For example, the arrangement order is [batch, channels, height, width], that is, NCHW format (a 4D format). Or, the arrangement order is [batch, height, width, channels], that is, NHWC format (a 4D format). Due to the chip design, the size of its matrix calculation unit is required, for example, it can support 16x16 matrix operations. The operation circuit can operate on data in 5D format (such as 6D, or other formats, etc.), which can be understood as the 5D format carrying hardware information. The 5D format refers to a five-dimensional data format using NC1HWC0. The five-dimensional "C0" is strongly related to data alignment. For example, "C0, C1" is split from the "C" dimension of "NCHW". "C1 = C / C0". If it is not divisible, it needs to be padded with zeros for alignment. The general disassembly process is as follows: (1) Split the "NCHW" data in the C dimension to obtain C1 parts of NHWC0. (2) Arrange the C1 parts of "NHWC0" continuously in the memory to become "NC1HWC0".
[0045] Data continuity: When the operation circuit performs data operations, it has certain requirements for the data to be operated, requiring the data to be continuous. The continuity here refers to the continuous order of elements. It can be understood as whether the storage order of the elements of the underlying one-dimensional array of the tensor is consistent with the order of the elements of the tensor expanded in one dimension in row priority. If they are consistent, the data is continuous. If they are inconsistent, the data is considered to be discontinuous. The underlying implementation of the tensor multidimensional array is a one-dimensional array using a piece of continuous memory. The Tensor saves the shape of the multidimensional array in the metadata. When accessing elements, the corresponding data can be found by converting the multi-dimensional index into the offset (stride) of the one-dimensional array relative to the starting position of the array. This offset is called the stride. After some Tensor operations, the positions of adjacent elements change, that is, the data is discontinuous. For example: a two-dimensional array t is as follows:
[0046]
[0047] The above two-dimensional array t is converted into a one-dimensional array by row priority. The one-dimensional array is as follows:
[0048] [0,1,2,3,4,5,6,7,8], formula (2)
[0049] If the actual storage format of the above two-dimensional array is consistent with formula (2), the data is continuous. Accessing the next element in the matrix is achieved by offsetting one position (stride=1).
[0050] If the actual storage format of the above two-dimensional array is as follows:
[0051] [0,3,6,1,4,7,2,5,8], formula (3)
[0052] If the above formula (3) is inconsistent with formula (2), the data is discontinuous. Accessing the next element in the matrix is achieved by shifting 2 positions (stride=2).
[0053] Data of the first characteristic: "Characteristic" includes format and / or continuity. The data of the first characteristic may be data in a first format (such as 4D). Alternatively, the data of the first characteristic may be discontinuous data. Alternatively, the data of the first characteristic may be data in a first format and discontinuous data.
[0054] Data of the second characteristic: "Characteristic" includes format and / or continuity. The data of the first characteristic can be data in the second format (such as 5D). Alternatively, the data of the second characteristic can be continuous data. Alternatively, the data of the second characteristic can be data in the first format and continuous data.
[0055] The embodiment of the present application provides an AI chip, the AI chip includes a first pipeline module and an operation circuit connected in sequence, the first pipeline module includes a conversion circuit and a plurality of buffers connected to the conversion circuit. When the operation circuit performs data operation, it has certain requirements for the data to be operated. Exemplarily, there are requirements for the format of the data and / or there are requirements for the continuity of the data. The conversion circuit writes the second characteristic data alternately into multiple buffers. Among them, alternatingly writing multiple buffers means that the data is written into the first buffer, and when the buffer is full, the data is continuously written into the second buffer until the last buffer is full, and then the state is switched, and the data is written into the first buffer again, and the data is written into the second buffer, etc. For example, when the number of buffers is 2, when the first buffer is full, the data is written into the second buffer, and then the state is switched, and when the second buffer is full, the data is written into the first buffer. When one of the multiple buffers is full, the conversion circuit reads the data of the second characteristic from the full buffer, operates on the data of the second characteristic, obtains the result data, and outputs the result data. For example, when the first buffer is full, the operation circuit reads data from the first buffer, and when the second buffer is full, the operation circuit reads data from the second buffer.
[0056] In an embodiment of the present application, a first pipeline module is added to the AI chip, and the first pipeline module includes a conversion circuit and multiple buffers. For example, the conversion circuit is used to support the functions of continuous conversion and / or format conversion. In the present application, a hardware-specific module is added, a data conversion function, and the conversion circuit writes the second characteristic data alternately into multiple buffers, and multiple buffers are used to implement a ping-pong cache mechanism. When one of the multiple buffers is full, the operation circuit can be started to operate on the read data. In the present application, there is no need to call the operator multiple times to pre-process the data to meet the requirements of the operation circuit for the data, simplify the operator call on the software side, improve the overall performance and simplify the call overhead and process. In addition, the conversion circuit writes the converted data alternately into multiple buffers. When a buffer is full, the operation circuit directly reads the data from the full buffer, and implements a fine pipeline operation based on the ping-pong cache mechanism. The operation circuit and the conversion circuit process the data in parallel, and the operation circuit does not need to wait for time. Compared with the traditional method, all data need to be pre-processed before the operation is performed. In the present application, the processing efficiency of the chip is greatly improved.
[0057] In an embodiment of the present application, different implementation methods may be included according to the specific functions of the conversion circuit. Exemplarily, 1. When the data is non-continuous data and the data format is a first format (such as a 4D format), the conversion circuit may be a data format conversion circuit, and the first pipeline module also includes a continuous conversion circuit. The data format conversion circuit is used to convert the format of the data, and the continuous conversion circuit is used to convert the discontinuous data into continuous data. 2. When the data is continuous data, only the data format needs to be converted. In this implementation method, the conversion circuit is a format conversion circuit. 3. When the data is non-continuous data, only the data needs to be converted for continuity, and no format conversion is required. At this time, the conversion circuit is a continuous conversion circuit.
[0058] Example 1: When the data is non-continuous data and the format of the data is data in the first format, the first pipeline module includes a format conversion circuit and a continuous conversion circuit. Figure 2 As shown, the AI chip includes a first pipeline module 201, an operation circuit 202, and a second pipeline module 203 connected in sequence. Optionally, the AI chip may also include a first buffer 204 and a second buffer 205. One end of the first buffer 204 is connected to the memory 206, and the other end of the first buffer 204 is connected to the first pipeline module 201. One end of the second buffer 205 is connected to the second pipeline module 203, and the other end of the second buffer 205 is connected to the memory 206. Among them, the memory 206 can be an external memory or an internal memory of the AI chip. The memory can be a dynamic random access memory (DRAM). For example, the memory includes but is not limited to a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc.
[0059] See also Figure 3 As shown, the first pipeline module 201 includes a first format conversion circuit 2012 and a plurality of conversion format buffers connected thereto, a continuous conversion circuit 2011 and a plurality of continuous data buffers connected thereto. The second pipeline module 203 includes a second format conversion circuit 2031 and a plurality of target format buffers connected thereto. The storage format of the first format conversion circuit 2012 needs to be determined according to the storage format of the data by the operation circuit.
[0060] Exemplarily, the plurality of continuous data buffers include at least a first continuous data buffer (referred to as “1-buffer_0”) 20111 and a second continuous data buffer (referred to as “1-buffer_1”) 20112 .
[0061] The multiple conversion format buffers include at least a first conversion format buffer (referred to as “2-buffer_0”) 20121 and a second conversion format buffer (referred to as “2-buffer_1”) 20122 .
[0062] The multiple target format buffers include at least a first target format buffer (referred to as “3-buffer_0”) 20311 and a second target format buffer (referred to as “3-buffer_1”) 20312 .
[0063] In this example, the number of continuous data buffers, the number of conversion format buffers, and the number of target format buffers are only illustrated by taking two as examples. In actual applications, the number and size of continuous data buffers, the number and size of conversion format buffers, and the number and size of conversion format buffers are not limited. The size and number of each buffer can be determined according to the processing power of the operation circuit. In this example, multiple continuous data buffers constitute a ping-pong cache. Multiple conversion format buffers constitute a ping-pong cache. Multiple target format buffers also constitute a ping-pong cache. Through the ping-pong cache mechanism, the data processing of the AI chip realizes fine-grained pipeline processing, that is, there is no need to wait for a circuit (such as a continuous conversion circuit) to process all the data before continuing the processing process of the next circuit (such as a format conversion circuit). The first format conversion circuit, the operation circuit, and the second format conversion circuit maintain parallel computing for most of the time, which greatly improves the processing efficiency of the AI chip.
[0064] The continuous conversion circuit 2011 is used to read data from the first buffer, convert discontinuous data into continuous data, and write the continuous data into a plurality of continuous data buffers in an alternating manner.
[0065] The AI chip can receive instructions sent by the host computer and configure parameters according to the instructions. The configuration parameters are sent to the continuous conversion circuit 2011, the first format conversion circuit 2012 and the second format conversion circuit 2031. For example, the step size parameter of the continuous conversion circuit 2011 is configured (such as stride=2), and the continuous conversion circuit 2011 converts the discontinuous data into continuous data according to the step size. For example, please refer to the above formula (3). If the discontinuous data is: [0,3,6,1,4,7,2,5,8], the continuous conversion circuit 2011 needs to convert the discontinuous data into continuous data. The continuous conversion circuit 2011 reads the data according to the step size parameter (such as the step size is 2), and the read data is [0,1,2,3,4,5,6,7,8].
[0066] The first format conversion circuit 2012 is used to read data in the first format from the continuous data buffer, convert the data in the first format into data in the second format, and write the data in the second format alternately into a plurality of conversion format buffers. For example, the first format conversion circuit 2012 converts data in the first format (such as data in 4D format) into data in the second format (such as data in 5D format) according to format parameters (such as 5D parameters), and the storage format of the data in the operation circuit is data in the second format (such as data in 5D format). The purpose of converting the data in the first format into the data in the second format is that the data in the second format is suitable for the data storage format of the operation circuit, so that the operation circuit can operate on the data in the second format.
[0067] The operation circuit 202 reads the second format data from the conversion format buffer when one of the multiple conversion format buffers is full, processes the second format data, and obtains result data in the second format.
[0068] The operation circuit may include a plurality of processing units (process engines, PEs). Optionally, the processing unit may be a convolution calculation unit and / or a vector calculation unit.
[0069] The second format conversion circuit 2031 is used to receive the result data in the second format from the operation circuit, convert the result data in the second format into the result data in the first format, and write the result data in the first format alternately into a plurality of target format buffers. The purpose of the second format conversion circuit 2031 converting the result data in the second format into the result data in the first format is that the format of the output data can be suitable for the data format of the subsequent process.
[0070] The second buffer is used to read the result data of the first format from the target format buffer and output the result data of the first format. Since the data processing is in units of tensors, after conversion, the space occupied by the data is greater than the original data space, so it usually exceeds the capacity of the buffer. Therefore, in the traditional method, the data after data preprocessing usually needs to be written back to the memory, which increases the number of interactions with the memory data. In the present application, the continuous conversion circuit is connected to multiple continuous data buffers. After processing a part of the data (that is, when the first continuous data buffer is full), the next circuit (first format conversion circuit) can be started. Similarly, when a conversion format buffer is full, the operation circuit is started, the continuous conversion circuit, the first format conversion circuit, the operation circuit and the second format conversion circuit can all process data in parallel, so the amount of data that needs to be cached in the buffer will not be large. In the present application, a first buffer and a second buffer are added, the first pipeline module reads data from the first buffer, and the second pipeline module outputs the result data to the second buffer. Compared with the traditional method, the number of interactions with the memory is reduced, thereby improving the system performance.
[0071] It should be noted that, in order to distinguish the format conversion circuit in the first pipeline module (also referred to as the front pipeline module) from the format conversion circuit in the second pipeline module (also referred to as the rear pipeline module), the format conversion circuit in the first pipeline module is referred to as the "first format conversion circuit", and the format conversion circuit in the second pipeline module is referred to as the "second format conversion circuit". The first format conversion circuit is used to convert data in a first format (such as 4D) into data in a second format (such as 5D), and the second format conversion circuit is used to convert data in a second format (such as 5D) into data in a first format (such as 4D). In order to distinguish the buffer of the first format conversion circuit from the buffer of the first format conversion circuit, the buffer of the first format circuit is referred to as the "conversion format buffer", and the buffer of the second format circuit is referred to as the "target format buffer".
[0072] In this example, the AI chip includes a first pipeline module, an operation circuit and a second pipeline module. Among them, the first pipeline module includes a continuous conversion circuit and multiple continuous data buffers connected thereto, a first format conversion circuit and multiple conversion format buffers and an operation circuit connected thereto. The second pipeline module includes a second format conversion circuit and multiple target format buffers connected thereto. First, compared with the traditional method, for example, in the adaptation process of the AI framework (such as PyTorch) of the dynamic graph architecture, a large number of format conversion operators need to be inserted, and the software side needs to consider the storage format of the hardware implementation, which increases the difficulty and complexity of software development. In this application, the continuity conversion of data is realized by a continuous conversion circuit, and the format conversion of data is realized by a first format conversion circuit and a second format conversion circuit, which reduces the number of operator calls on the software side and improves the overall performance of the system. Secondly, in terms of hardware implementation, through the ping-pong cache mechanism, a fine-grained pipeline data processing mechanism can be realized, each circuit is processed in parallel, no waiting time is required, and the total execution time is basically consistent with the time of a simple operation (such as a convolution calculation), and each circuit maintains a parallel computing state, which improves the data processing efficiency of the chip. Finally, the ping-pong cache structure in the first pipeline module and the second pipeline module only buffers part of the data, and the area of each buffer is not large, so less area overhead can be used to ensure the implementation of the pipeline mechanism.
[0073] In this embodiment, please refer to Figure 4 As shown, a data processing method is applied to the above Figure 2 and Figure 3 For the corresponding AI chip, the steps executed by each circuit can be as follows:
[0074] Step 401: The continuous conversion circuit converts discontinuous data into continuous data, and writes the continuous data alternately into a plurality of continuous data buffers.
[0075] For details, please refer to Figure 5 shown.
[0076] S10, the continuous conversion circuit reads data from the first cache (cache_1), and writes continuous first format data into the first continuous data buffer (1-buffer_0). When the first continuous data buffer (1-buffer_0) is full, the first format conversion circuit is started.
[0077] S11, the continuous conversion circuit continues to write the continuous first format data into the second continuous data buffer (1-buffer_1). Then, S14 is executed.
[0078] Step 402: When one of the multiple continuous data buffers is full, the first format conversion circuit reads continuous data from the full continuous data buffer, where the continuous data is data in the first format.
[0079] S12, the first format conversion circuit reads the first format data from the first continuous data buffer (1-buffer_0), converts the first format data into the second format data, and writes the second format data into the first conversion format buffer (2-buffer_0). Continue to execute S13.
[0080] It should be noted that there is no sequential relationship between step S11 and step S12, and step S11 and step S12 are executed synchronously, that is, the continuous conversion circuit and the first format conversion circuit can perform data processing synchronously, and there is no need for the continuous conversion circuit to read all the data and then perform format conversion. The continuous conversion circuit reads part of the data, that is, when the first 1-buffer_0 is full, the operation circuit can be started, and the first format conversion circuit can process the data in "1-buffer_0" while the continuous conversion circuit continues to read data. The continuous conversion circuit and the first format conversion circuit can perform data processing synchronously, thereby improving the efficiency of data processing of the AI chip.
[0081] Step 403: The first format conversion circuit converts the data in the first format into data in the second format; and writes the data in the second format alternately into a plurality of conversion format buffers.
[0082] S13, when the second continuous data buffer (1-buffer_1) is full, the first format conversion circuit is started. The first format conversion circuit reads the first format data from the second continuous data buffer (1-buffer_1), converts the first format data into the second format data, and writes the second format data into the second conversion format buffer (2-buffer_1).
[0083] S14. The continuous conversion circuit continues to write the continuous first format data into the first continuous data buffer (1-buffer_0).
[0084] It should be noted that there is no sequential relationship between the above step S13 and step S14, and step S13 and step S14 are executed synchronously, that is, the continuous conversion circuit and the first format conversion circuit can perform data processing synchronously.
[0085] The above S11-S14 are repeatedly executed. The continuous conversion circuit writes the read continuous first format data into 1-buffer_0 and 1-buffer_1 alternately, and the first format conversion circuit reads the continuous first format data from 1-buffer_0 and 1-buffer_1 alternately, realizing a fine-grained pipeline mechanism.
[0086] S15. When the first conversion format buffer (2-buffer_0) is full, the operation circuit is started. The operation circuit reads the data in the second format from the first conversion format buffer (2-buffer_0), and then the operation circuit performs an operation operation (such as a convolution operation) on the data in the second format to obtain result data (the result data is the data in the second format). Continue to execute S17.
[0087] S16, the first format conversion circuit continues to write the data in the second format into the second conversion format buffer (2-buffer_1). Then, the process continues to execute S18.
[0088] It should be noted that there is no sequential relationship between step S15 and step S16, and step S15 and step S16 are executed synchronously, that is, the first format conversion circuit and the operation circuit can perform data processing synchronously, and the first format conversion circuit does not need to process all the data before performing the operation. The first format conversion circuit converts part of the data, that is, when the first conversion format buffer (2-buffer_0) is full, the operation circuit can be started, and the operation circuit can process the data in "2-buffer_0", while the first format conversion circuit continues to convert data operations. The first format conversion circuit and the operation circuit can perform data processing synchronously, thereby improving the efficiency of data processing of the AI chip.
[0089] Step 404: the operation circuit reads the second format data from the conversion format buffer, processes the second format data, and obtains result data in the second format;
[0090] S17, when the second conversion format buffer (2-buffer_1) is full, the operation circuit is started, the operation circuit reads the data in the second format from the second conversion format buffer (2-buffer_1), and then the operation circuit performs an operation operation (such as a convolution operation) on the data in the second format to obtain result data (the result data is the data in the second format). Continue to execute S20.
[0091] S18. The first format conversion circuit continues to write the data in the second format into the first conversion format buffer (2-buffer_0).
[0092] It should be noted that the above steps S17 and S18 are executed synchronously, and the order is not limited. In this example, the first pipeline module includes multiple conversion format buffers, and the reading and writing processes are performed alternately in multiple conversion format buffers. It is not necessary for the first format conversion circuit to convert all the data and then output the second format data to the operation circuit, and the operation circuit performs convolution operations. Instead, when the first conversion format buffer is full, the operation circuit can be started, and the operation circuit can read the data in the first format from the first conversion format buffer, and perform convolution operations on the data in the first format. Synchronously, the first format conversion circuit continues to perform conversion operations on the data in the first format. In other words, the first format conversion circuit and the operation circuit synchronously perform data operations, thereby realizing fine-grained pipeline operations, saving AI chip data processing time, and improving data processing efficiency.
[0093] Repeat the above steps S15 to S18 until the data processing is completed.
[0094] Then, the process of data processing by the second format conversion circuit in the second pipeline module is described:
[0095] Step 405: The second format conversion circuit converts the result data in the second format into result data in the first format, and writes the result data in the first format alternately into a plurality of target format buffers.
[0096] For details, please refer to Figure 6 shown.
[0097] S20, the operation circuit outputs the result data in the second format to the second format conversion circuit, and starts the second format conversion circuit.
[0098] S21. The second format conversion circuit converts the result data in the second format to obtain result data in the first format, and writes the result data in the first format into a first target format buffer (3-buffer_0).
[0099] S23, when the first target format buffer (3-buffer_0) is full, the second storage (cache_2) is started, and the second storage reads the result data of the first format from the first target format buffer (3-buffer_0). Then, S26 is executed.
[0100] S24, the second format conversion circuit writes the result data of the first format into the second target format buffer (3-buffer_1). Then, the process continues with S25.
[0101] Step S23 and step S24 are executed synchronously.
[0102] S25. When the second target format buffer (3-buffer_1) is full, the second format conversion circuit writes the result data of the first format into the first target format buffer (3-buffer_0).
[0103] S26. The second memory reads the result data in the first format from the second target format buffer (3-buffer_1).
[0104] Step S25 and step S26 are executed synchronously.
[0105] S21-S26 are repeatedly executed until the data processing is completed.
[0106] In this example, firstly, the continuity conversion of data is realized by a continuous conversion circuit, and the format conversion of data is realized by a first format conversion circuit and a second format conversion circuit, which reduces the number of operator calls on the software side and improves the overall performance of the system. Secondly, in terms of hardware implementation, through the ping-pong cache mechanism, a fine-grained pipeline data processing mechanism can be realized, and each circuit does not need waiting time, and the total execution time is basically consistent with the time of a simple operation (such as convolution calculation), and each circuit maintains a parallel computing state, which improves the data processing efficiency of the chip. Finally, for the ping-pong cache structure in the first pipeline module and the second pipeline module, only part of the data is buffered, and the area of each buffer is not large, that is, less area overhead can be used to ensure the implementation of the pipeline mechanism.
[0107] Example 2: When the data acquired by the AI chip is continuous data, only the data format needs to be converted. In this implementation, the conversion circuit is a format conversion circuit, which is used to convert the data format. The difference between this example and the first implementation is that the first pipeline module does not include a continuous conversion circuit, and the conversion circuit is a first format conversion circuit. Figure 2 and Figure 7 As shown, the first pipeline module 701 includes a first format conversion circuit 7012 and a conversion format buffer connected thereto. The number of multiple conversion format buffers includes at least a first conversion format buffer (2-buffer_0) 70121 and a second conversion format buffer (2-buffer_1) 70122, and the multiple conversion format buffers are used to implement a ping-pong cache mechanism. The second pipeline module 703 includes a second format conversion circuit 7031 and multiple target format buffers. The multiple target format buffers include at least a first target format buffer (3-buffer_0) 70311 and a second target format buffer ((3-buffer_1) 70312).
[0108] The first format conversion circuit 7012 is used to read data in a first format from a first cache (cache_1), convert the data in the first format into data in a second format, and write the data in the second format alternately into a plurality of conversion format buffers.
[0109] The operation circuit 702 is used for reading the second format data from the conversion format buffer when one of the multiple conversion format buffers is full, and processing the second format data to obtain result data in the second format.
[0110] The second format conversion circuit 7031 is used to receive the result data in the second format from the operation circuit 702, convert the result data in the second format into the result data in the first format, and write the result data in the first format into a plurality of target format buffers alternately. The purpose of the second format conversion circuit 7031 converting the result data in the second format into the result data in the first format is that the format of the output data can be suitable for the data format of the subsequent process.
[0111] The second buffer is used to read the result data in the first format from the full target format buffer and output the result data in the first format.
[0112] See also Figure 8 As shown, the data processing method in this example includes the following steps:
[0113] Step 801: A first format conversion circuit converts data in a first format into data in a second format; and writes the second format data alternately into a plurality of conversion format buffers.
[0114] For details, please refer to Fig. 9 shown.
[0115] S31 . The first format conversion circuit reads data in a first format from the first cache ( cache_1 ), converts the data in the first format into data in a second format, and writes the data in the second format into the first conversion format buffer ( 2 - buffer_0 ).
[0116] S32. When the first conversion format buffer (2-buffer_0) is full, the operation circuit is started. The operation circuit reads the data in the second format from the first conversion format buffer (2-buffer_0), and then the operation circuit performs an operation operation (such as a convolution operation) on the data in the second format to obtain result data (the result data is the data in the second format). Continue to execute S34.
[0117] S33, the first format conversion circuit continues to write the data in the second format into the second conversion format buffer (2-buffer_1). Then, the process continues to execute S35.
[0118] S31 and S32 are executed synchronously. Please refer to the relevant description of step S15 and step S16 in the above example 1, which will not be repeated here.
[0119] Step 802: The operation circuit reads the second format data from the conversion format buffer, processes the second format data, and obtains result data in the second format.
[0120] S34. When the second conversion format buffer (2-buffer_1) is full, the operation circuit is started, and the operation circuit reads the data in the second format from the second conversion format buffer (2-buffer_1). Then, the operation circuit performs an operation operation (such as a convolution operation) on the data in the second format to obtain result data (the result data is the data in the second format).
[0121] S35 , the first format conversion circuit continues to write the data in the second format into the first conversion format buffer ( 2 - buffer_0 ).
[0122] S34 and S35 are executed synchronously. Please refer to the relevant descriptions of step S17 and step S18 in the above example 1, which will not be repeated here.
[0123] Step 803: The second format conversion circuit converts the result data in the second format into result data in the first format, and writes the result data in the first format alternately into a plurality of target format buffers.
[0124] For this step, please refer to the relevant instructions of steps S20-S26 in the above example 1, which will not be repeated here.
[0125] Example 3: When the data acquired by the AI chip is non-continuous data, the format of the data meets the requirements of the operation circuit. It is only necessary to convert the continuity of the data, and it is not necessary to convert the format of the data. In this implementation, the conversion circuit is a continuous conversion circuit. The difference between this example and the first implementation is that the first format conversion circuit is not included in the first pipeline module, and the conversion circuit is a continuous conversion circuit. The AI chip does not include a second pipeline module.
[0126] See also Fig.10 As shown, the AI chip includes a first buffer 901, a first pipeline module 902, an operation circuit 903, and a second buffer memory 904 connected in sequence. The first pipeline module 902 includes a continuous conversion circuit 9021 and a plurality of continuous data buffers, for example, the plurality of continuous data buffers at least include a first continuous data buffer (1-buffer_0) 90211 and a second continuous data buffer (1-buffer_1) 90212.
[0127] The continuous conversion circuit 9021 is used for reading data from the first buffer 901, converting discontinuous data into continuous data, and writing the continuous data into a plurality of continuous data buffers in an alternating manner.
[0128] The operation circuit 903 is used to read continuous data from the full continuous data buffer when one of the multiple continuous data buffers is full, operate on the continuous data to obtain result data, and output the result data to the second buffer 904.
[0129] See also Fig.11 As shown in the figure, the data processing process in this example is explained:
[0130] Step 1101: The continuous conversion circuit converts discontinuous data into continuous data, and writes the continuous data alternately into a plurality of continuous data buffers.
[0131] For details, please refer to Fig.12 shown.
[0132] S40, the continuous conversion circuit reads data from the first cache (cache_1), and writes the continuous data into the first continuous data buffer (1-buffer_0). When the first continuous data buffer (1-buffer_0) is full, the operation circuit is started.
[0133] S41, the continuous conversion circuit continues to write the continuous data into the second continuous data buffer (1-buffer_1). Then, S42 is executed.
[0134] S42. When the second continuous data buffer (1-buffer_1) is full, the continuous conversion circuit continues to write the continuous data into the first continuous data buffer (1-buffer_0).
[0135] S41 and S42 are repeatedly executed.
[0136] Step 1102: When one of the multiple continuous data buffers is full, the operation circuit reads the continuous data from the full continuous data buffer, operates on the continuous data to obtain result data, and outputs the result data to the second cache (cache_2).
[0137] S43. When the first continuous data buffer (1-buffer_0) is full, the operation circuit reads continuous data from the first continuous data buffer (1-buffer_0), performs operation (such as convolution operation) on the continuous data, obtains result data, and outputs the result data to the second buffer.
[0138] S41 and S43 are executed synchronously.
[0139] S44. When the second continuous data buffer (1-buffer_1) is full, the operation circuit reads continuous data from the second continuous data buffer (1-buffer_1), performs operation (such as convolution operation) on the continuous data, obtains result data, and outputs the result data to the second buffer.
[0140] S42 and S44 are executed simultaneously. S43-S44 are executed repeatedly.
[0141] In this example, when the data acquired by the AI chip is non-continuous data, the format of the data meets the requirements of the operation circuit, and only the continuity of the data needs to be converted, and the format of the data does not need to be converted, which saves the data processing flow. In addition, the continuity conversion of the data is realized by a continuous conversion circuit, and the format conversion of the data is realized by a first format conversion circuit and a second format conversion circuit, which reduces the number of operator calls on the software side and improves the overall performance of the system. In terms of hardware implementation, a fine-grained pipeline data processing mechanism can be implemented through the ping-pong cache mechanism, and each circuit does not require waiting time, and the total execution time is basically consistent with the time for a simple operation (such as a convolution calculation), and each circuit maintains a parallel computing state, which improves the data processing efficiency of the chip.
[0142] The present application embodiment provides an artificial intelligence computing board. Fig.13 As shown, the artificial intelligence computing board 1300 includes a communication interface 1301 and the artificial intelligence chip 1302 in the above examples. The communication interface can be a high-speed serial computer expansion bus standard (Peripheral Component Interconnect Express, PCIE) interface, and the PCIE interface is used to connect to the host.
[0143] See also Fig.14 As shown, an electronic device 1400 is also provided in an embodiment of the present application. The electronic device may be a server, or the electronic device may also be a terminal device.
[0144] For example, the electronic device is described by taking a server as an example. The server may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1422 (for example, one or more processors) and memory 1432, one or more storage media 1430 (for example, one or more mass storage devices) storing application programs 1442 or data 1444, and artificial intelligence computing boards 1460. Among them, the memory 1432 and the storage medium 1430 may be short-term storage or permanent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1422 may be configured to communicate with the storage medium 1430 and execute a series of instruction operations in the storage medium 1430 on the server.
[0145] In the present application, the artificial intelligence computing board 1460 reads the data to be converted from the memory, or the artificial intelligence computing board 1460 outputs the result data after the operation to the memory. The AI chip is mounted on the CPU (also referred to as the main CPU) as a coprocessor. The CPU can assign data processing tasks to the artificial intelligence chip, and the CPU can send configuration parameters for continuity conversion and / or format conversion configuration parameters to the AI chip. The continuous conversion circuit and / or format conversion circuit (such as the first format conversion circuit and the second format conversion circuit) in the AI chip can convert the data according to the configuration parameters.
[0146] The server may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input and output interfaces 1458, and / or one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0147] Among them, the processor mentioned in any of the above places can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the program execution of the wireless communication method of the first aspect mentioned above.
[0148] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0149] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0150] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0151] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0152] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An artificial intelligence chip, characterized in that: include: A first pipeline module and an operation circuit connected in sequence, wherein the first pipeline module includes a conversion circuit and a plurality of buffers connected to the conversion circuit; The conversion circuit is used to obtain data of a first characteristic to be converted, convert the data of the first characteristic into data of a second characteristic, the data of the second characteristic being data applicable to the operation of the operation circuit; and alternately write the data of the second characteristic into the plurality of buffers; The operation circuit is used for reading the data of the second characteristic from the full buffer when one of the multiple buffers is full, operating the data of the second characteristic to obtain result data, and outputting the result data; The conversion circuit is a first format conversion circuit, and the buffer is a conversion format buffer; the chip further includes a second pipeline module connected to the operation circuit, and the second pipeline module includes a second format conversion circuit and a plurality of target format buffers connected to the second format conversion circuit; The first format conversion circuit is used to convert the data in the first format into data in the second format, and to write the data in the second format into the plurality of conversion format buffers alternately; The operation circuit is used for reading the second format data from the conversion format buffer when one of the multiple conversion format buffers is full, and processing the second format data to obtain result data in the second format; The second format conversion circuit is used for converting the result data in the second format into the result data in the first format, and alternately writing the result data in the first format into the plurality of target format buffers.
2. The chip according to claim 1, characterized in that: The first pipeline module also includes a continuous conversion circuit and a plurality of continuous data buffers connected to the continuous conversion circuit; The continuous conversion circuit is used to convert discontinuous data into continuous data, and write the continuous data into the plurality of continuous data buffers in an alternating manner; The first format conversion circuit is also used to read continuous data from the full continuous data buffer when one of the multiple continuous data buffers is full, the continuous data being data in the first format, and convert the data in the first format into data in the second format.
3. The chip according to claim 1 or 2, characterized in that: The operation circuit includes a convolution calculation unit and / or a vector calculation unit.
4. The chip according to claim 1 or 2, characterized in that: The chip also includes a first buffer and a second buffer; The first buffer is used to output the data to be converted to the first pipeline module; The second buffer is used to cache the result data.
5. A data processing method, characterized in that: The method is applied to an artificial intelligence chip, the chip comprising a first pipeline module and an operation circuit connected in sequence, the first pipeline module comprising a conversion circuit and a plurality of buffers connected to the conversion circuit, the method further comprising: The conversion circuit acquires data of a first characteristic to be converted, converts the data of the first characteristic into data of a second characteristic, the data of the second characteristic being data applicable to the operation circuit for operation; and alternately writes the data of the second characteristic into the plurality of buffers; When one of the plurality of buffers is full, the operation circuit reads the data of the second characteristic from the full buffer, operates on the data of the second characteristic to obtain result data; and outputs the result data; The conversion circuit is a first format conversion circuit, and the buffer is a conversion format buffer; the chip further includes a second pipeline module connected to the operation circuit, and the second pipeline module includes a second format conversion circuit and a plurality of target format buffers connected to the second format conversion circuit; The conversion circuit acquires data of a first feature to be converted, converts the data of the first feature into data of a second feature, and writes the second feature data alternately into the plurality of buffers, including: The first format conversion circuit converts the data in the first format into data in the second format; and the second format data is alternately written into the plurality of conversion format buffers; When one of the plurality of buffers is full, the operation circuit reads the data of the second feature from the full buffer, operates on the data of the second feature, and obtains result data, including: When one of the multiple conversion format buffers is full, the operation circuit reads the second format data from the full conversion format buffer, processes the second format data, and obtains result data in the second format; The method further comprises: The second format conversion circuit converts the result data in the second format into result data in the first format, and writes the result data in the first format alternately into the plurality of target format buffers.
6. The method according to claim 5, characterized in that The first pipeline module also includes a continuous conversion circuit and a plurality of continuous data buffers connected to the continuous conversion circuit; The method further comprises: The continuous conversion circuit converts discontinuous data into continuous data, and writes the continuous data alternately into the plurality of continuous data buffers; When one of the plurality of continuous data buffers is full of continuous data, the first format conversion circuit reads continuous data from the full continuous data buffer, where the continuous data is data in a first format.
7. An artificial intelligence computing board, characterized in that: It comprises a communication interface and an artificial intelligence chip as described in any one of claims 1 to 4; wherein the communication interface is used to connect to a host.
8. An electronic device, characterized in that: include: A processor, a memory coupled to the processor, and an artificial intelligence computing board as claimed in claim 7, wherein the processor and the memory perform data transmission with the artificial intelligence computing board via a communication interface.
Citation Information
Patent Citations
Hardware accelerator and method for realizing sparse GRU neural network based on FPGA
CN107229967A
Data transfer circuit and method of deep learning neural network
CN109858622A
Model-based prediction method and device
CN110033091A
Signal processing device and related product
CN110968285A
Winograd transform convolution operations for neural networks
US20200234124A1