Chip-based data processing method and apparatus, and related product

CN115203121BActive Publication Date: 2026-09-25CAMBRICON SINGGO (NANJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110743865.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2026-09-25
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

[0005]本发明实施例提供一种基于芯片的数据处理方法、装置及相关产品,用以解决现有技术中导致竞争片上存储器的端口带宽,从而导致IO操作带宽下降的技术问题

Benefits of technology

[0026]本发明实施例提供的基于芯片的数据处理方法、装置及相关产品,通过将所述片外存储器中的目标数据读取至所述待读取片上缓存存储器中;在将所述片外存储器中的目标数据读取至所述待读取片上缓存存储器中的同时,将任一其他片上缓存存储器中的目标数据移动到所述片上计算存储器中,以进行第一计算;将所述任一其他片上缓存存储器更新为所述待读取片上缓存存储器;以及重复执行所述读取、移动、更新步骤,直至所述片外存储器中的目标数据为空。在每次执行将片外存储器中的目标数据读取至片上缓存存储器中所针对的片上缓存存储器,与并行执行将缓存在片上缓存存储器中的目标数据移动到片上计算存储器中进行计算所针对的片上缓存存储器是不同的,这种并行处理方式为对称排流水设计,能够有效避免不同类型的操作竞争同一片上存储空间的端口带宽,进而提高IO操作带宽。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115203121B_ABST
    Figure CN115203121B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a chip-based data processing method and device and related products. The method comprises: reading target data in an off-chip memory into a to-be-read on-chip cache memory; while reading the target data in the off-chip memory into the to-be-read on-chip cache memory, moving target data in any other on-chip cache memory to an on-chip computing memory for first calculation; updating any other on-chip cache memory as the to-be-read on-chip cache memory; and repeating the reading, moving and updating steps until the target data in the off-chip memory is empty. This parallel processing method is a symmetric pipeline design, which can effectively avoid different types of operations competing for the port bandwidth of the same on-chip storage space, thereby improving the IO operation bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data processing technology, and in particular to a chip-based data processing method, apparatus and related products. Background Technology

[0002] Because on-chip storage offers high bandwidth and low latency, loading data onto the chip and performing on-chip computations has become a growing trend. This is especially true for operations requiring multiple reads of input data. When loading data onto the chip and performing on-chip computations, the data is first read into one on-chip memory for caching, and then moved to another on-chip memory for calculation. Alternatively, the data can be first read into one on-chip memory for calculation, and then moved to another on-chip memory for caching. Since different operations can be executed in parallel, the process of loading data onto the chip and performing on-chip computations can be piped.

[0003] Specifically, during each loop of data processing, the current data is read from off-chip memory and cached in on-chip memory, while the previous data is moved from that on-chip memory to another on-chip memory for computation. Alternatively, during each loop of data processing, the current data is read from off-chip memory and cached in on-chip memory, while the previous data is moved from that on-chip memory to another on-chip memory for computation.

[0004] Although different operations can be executed in parallel, they all require operation on the same on-chip memory, which leads to competition for the port bandwidth of the on-chip memory, resulting in a decrease in I / O operation bandwidth. Summary of the Invention

[0005] This invention provides a chip-based data processing method, apparatus, and related products to solve the technical problem in the prior art that leads to competition for on-chip memory port bandwidth, thereby causing a decrease in I / O operation bandwidth.

[0006] In a first aspect, embodiments of the present invention provide a chip-based data processing method, wherein the chip includes on-chip computing memory and multiple on-chip cache memories, wherein the on-chip cache memory includes an on-chip cache memory to be read, the on-chip cache memory communicates with off-chip memory, and the off-chip memory stores multiple target data, the method comprising:

[0007] Read the target data from the off-chip memory into the on-chip cache memory to be read;

[0008] While reading the target data from the off-chip memory into the on-chip cache memory to be read, the target data from any other on-chip cache memory is moved into the on-chip computing memory for the first calculation.

[0009] Update any other on-chip cache memory to the on-chip cache memory to be read; and

[0010] Repeat the read, move, and update steps until the target data in the off-chip memory is empty.

[0011] Secondly, embodiments of the present invention provide a chip-based data processing device, wherein the chip includes on-chip computing memory and multiple on-chip cache memories, wherein the on-chip cache memory includes an on-chip cache memory to be read, the on-chip cache memory communicates with off-chip memory, and the off-chip memory stores multiple target data, the device comprising:

[0012] The read operation module is used to read the target data from the off-chip memory into the on-chip cache memory to be read.

[0013] The moving operation module is used to simultaneously read target data from the off-chip memory into the on-chip cache memory to be read, and move target data from any other on-chip cache memory into the on-chip computing memory for performing a first calculation;

[0014] The update operation module is used to update any other on-chip cache memory to the on-chip cache memory to be read, and

[0015] The repeat execution module is used to repeatedly execute the read, move, and update steps until the target data in the off-chip memory is empty.

[0016] Thirdly, embodiments of the present invention provide a chip-based data processing device, comprising: a processor and a memory; wherein,

[0017] The memory is used to store program code;

[0018] The processor is configured to call the process code stored in the memory to execute the method described in the first aspect.

[0019] Fourthly, embodiments of the present invention provide an artificial intelligence chip, including: on-chip computing memory, multiple on-chip cache memories, and a chip-based data processing device as described in the second or third aspect.

[0020] Fifthly, embodiments of the present invention provide an electronic device, the electronic device including an off-chip memory and an artificial intelligence chip as described in the fourth aspect.

[0021] In a sixth aspect, embodiments of the present invention provide a board, the board comprising: a storage device, an interface device, a control device, and an artificial intelligence chip as described in the fourth aspect;

[0022] The artificial intelligence chip is connected to the storage device, the control device, and the interface device, respectively.

[0023] The storage device is used to store the target data;

[0024] The interface device is used to realize data transmission between the artificial intelligence chip and external devices;

[0025] The controller is used to monitor the state of the artificial intelligence chip.

[0026] The chip-based data processing method, apparatus, and related products provided in this invention involve: reading target data from the off-chip memory into the on-chip cache memory to be read; simultaneously moving target data from any other on-chip cache memory to the on-chip compute memory for a first calculation; updating the other on-chip cache memory to the on-chip cache memory to be read; and repeating the read, move, and update steps until the target data in the off-chip memory is empty. The on-chip cache memory targeted in each execution of reading target data from the off-chip memory into the on-chip cache memory is different from the on-chip cache memory targeted in the parallel execution of moving target data cached in the on-chip cache memory to the on-chip compute memory for calculation. This parallel processing method uses a symmetrical pipelining design, which can effectively avoid different types of operations competing for port bandwidth in the same on-chip memory space, thereby improving I / O operation bandwidth. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0028] Figure 1 A flowchart illustrating a chip-based data processing method provided in one embodiment of this application;

[0029] Figure 2 A network architecture diagram for a chip-based data processing method provided in this application embodiment;

[0030] Figure 3 This is a schematic diagram illustrating the working principle of a chip-based data processing method provided in one embodiment of this application.

[0031] Figure 4 A flowchart illustrating a chip-based data processing method provided in another embodiment of this application;

[0032] Figure 5 A schematic diagram illustrating the working principle of a chip-based data processing method provided in another embodiment of this application;

[0033] Figure 6 A flowchart illustrating a chip-based data processing method provided in another embodiment of this application;

[0034] Figure 7 This is a schematic diagram illustrating the working principle of a chip-based data processing method provided in another embodiment of this application.

[0035] Figure 8 A flowchart of a chip-based data processing method provided in another embodiment of this application;

[0036] Figure 9 This application also provides a schematic diagram of the working principle of a chip-based data processing method according to an embodiment;

[0037] Figure 10 This is a structural diagram of a board according to an embodiment of this application;

[0038] Figure 11 This is a structural diagram illustrating a combined processing apparatus according to an embodiment of this application;

[0039] Figure 12 This is a schematic diagram showing the internal structure of a computing device according to an embodiment of this application;

[0040] Figure 13 This is a schematic diagram illustrating the internal structure of a processor core according to an embodiment of the present disclosure.

[0041] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0042] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0043] To clearly understand the technical solution of this application, the solutions of the prior art will be described in detail first.

[0044] Because on-chip storage offers high bandwidth and low latency, loading data onto the chip and performing on-chip operations has become a growing trend. This is especially true for operations that require multiple reads of input data, such as normalization and batch normalization. To avoid repeatedly loading input data, residing it in on-chip storage leverages its high bandwidth and low latency, significantly reducing I / O time.

[0045] However, when loading data onto the chip and performing data operations on the chip, an additional on-chip input data move operation is required. This involves first reading the data into one on-chip memory for buffering, and then moving the data to another on-chip memory for computation. Alternatively, the data can be first read into one on-chip memory for computation, and then moved to another on-chip memory for buffering. Since these different operations can be executed in parallel, a pipelining design can be implemented for the process of loading data onto the chip and performing data operations on the chip.

[0046] Typically, the following two drainage methods are used to execute different operations in parallel.

[0047] The first drainage method involves reading the current data from off-chip memory into on-chip cache memory for caching during each processing cycle, while simultaneously moving the previous data from on-chip cache memory into on-chip compute memory for computation.

[0048] The second drainage method involves reading the current data directly from the off-chip memory into the on-chip compute memory in each processing cycle for computation, while simultaneously moving the previous data from the on-chip compute memory to the on-chip cache memory for caching.

[0049] Although both drainage methods allow for the parallel execution of two different types of instructions, they interfere with each other because both operate on the same on-chip memory. I / O operations and MOVE operations compete for port bandwidth within the same on-chip memory space. Specifically, in the first drainage method, MOVE operations and I / O operations compete for port bandwidth in the on-chip cache memory, resulting in a decrease in I / O operation bandwidth. In the second drainage method, MOVE operations and I / O operations compete for port bandwidth in the on-chip compute memory, also leading to a decrease in I / O operation bandwidth.

[0050] Therefore, when faced with the technical problems of existing technologies, the inventors, through creative research, discovered that to avoid different types of operations competing for port bandwidth in the same on-chip memory space, different operations can be assigned to different on-chip memory spaces during pipeline design, allowing different operations to be executed in parallel. Regarding the first type of pipeline design in existing technologies, to avoid MOVE operations and IO operations competing for port bandwidth in the on-chip cache memory, multiple on-chip cache memories and on-chip compute memories are set on the chip. The on-chip cache memory includes a target on-chip cache memory to be read, which communicates with external memory. The external memory stores multiple target data, including: reading the target data from the external memory into the target on-chip cache memory; simultaneously, while reading the target data from the external memory into the target on-chip cache memory, moving the target data from any other on-chip cache memory to the on-chip compute memory for a first calculation. After updating any other on-chip cache memory to the target on-chip cache memory, the read, move, and update steps are repeated until the target data in the external memory is empty. The ability to read target data from off-chip memory into the corresponding on-chip cache memory in each execution differs from the ability to move target data cached in the on-chip cache memory to the on-chip compute memory for the first computation in parallel execution. Furthermore, for each on-chip cache memory, read and move operations are performed alternately in a loop. By employing this method, when loading data onto the chip and performing data operations on the chip, the problem of different types of operations competing for port bandwidth in the same on-chip memory space can be avoided, thereby improving I / O operation bandwidth.

[0051] Therefore, based on the above-mentioned inventive discovery, the inventors proposed the technical solution of the embodiments of the present invention. The application scenarios of the chip-based data processing method provided by the embodiments of the present invention will be introduced below.

[0052] The chip-based data processing method provided in this invention can be applied to scenarios where, after determining multiple target data, a certain data operation is performed on the chip for these target data. The type of data operation is not limited; it can be an operation where each target data is subjected to one I / O operation on the chip, such as addition or subtraction of multiple target data. Alternatively, it can be an operation where each data is subjected to two I / O operations on the chip, such as normalization or batch normalization of the target data.

[0053] The technical solutions of the present invention and how they solve the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0054] Example 1

[0055] Figure 1 A flowchart of a chip-based data processing method provided in one embodiment of this application is shown below. Figure 1 As shown, the execution entity of the chip-based data processing method provided in this embodiment is a chip-based data processing device. Figure 2 As shown in the embodiment of this application, the network architecture of the chip-based data processing method further includes an off-chip memory 1 and a chip 3. Chip 3 includes an on-chip compute memory 31 and multiple on-chip cache memories. The on-chip compute memory 31 is communicatively connected to each on-chip cache memory. Figure 2 The chip includes N on-chip cache memories, namely the first on-chip cache memory 32, the second on-chip cache memory 33, ..., the Nth on-chip cache memory 36, where N is an integer greater than or equal to 2. The chip-based data processing device 2 is communicatively connected to the on-chip memory 1 and the chip 3. The chip-based data processing device 2 can also be located on the chip. The off-chip memory 1 is communicatively connected to each on-chip cache memory. Multiple target data are stored in the off-chip memory 1. The chip-based data processing method provided in this embodiment includes the following steps:

[0056] Step 101: Read the target data from the off-chip memory into the on-chip cache memory to be read.

[0057] The on-chip cache memory to be read is one of multiple on-chip cache memories.

[0058] In this embodiment, optionally, read / write instructions (also known as I / O instructions) can be used to read the target data from the off-chip memory into the on-chip cache memory to be read, so that the target data is cached in the on-chip cache memory to be read. Alternatively, the target data from the off-chip memory can be read into the on-chip cache memory to be read according to the address of the target data.

[0059] It should be noted that, in this embodiment, the method of reading the target data from the off-chip memory into the on-chip cache memory to be read can also be other, and this embodiment does not limit it.

[0060] In this embodiment, each time step 101 is executed, a target data or a batch of target data in the off-chip memory is read into the on-chip cache memory to be read for caching. The target data is image data, voice data, or video data, and this disclosure does not impose any restrictions on it.

[0061] Step 102: While reading the target data from the off-chip memory into the on-chip cache memory to be read, move the target data from any other on-chip cache memory into the on-chip computing memory to perform the first calculation.

[0062] Among them, any other on-chip cache memory is any on-chip cache memory other than the on-chip cache memory to be read.

[0063] In this embodiment, while reading the target data from the off-chip memory into the on-chip cache memory to be read, the target data from any other on-chip cache memory is moved to the on-chip compute memory in parallel. In the on-chip compute memory, the processor performs a first calculation on the target data.

[0064] For example, the on-chip cache memory to be read is Figure 2 The first on-chip cache memory is used, and any other on-chip cache memory is used as the second on-chip cache memory. Therefore, while reading the target data from the off-chip memory into the first on-chip cache memory, the target data from the second on-chip cache memory is moved into the on-chip compute memory in parallel, and the processor performs the first calculation on the target data.

[0065] Since reading the target data from off-chip memory to on-chip cache memory targets the first on-chip cache memory, and moving the target data to on-chip compute memory targets the second on-chip cache memory, and the two on-chip cache memories are different, the parallel execution of read and move operations occupies the port bandwidth of different on-chip cache memories, thereby effectively avoiding different types of operations competing for the port bandwidth of the same on-chip storage space.

[0066] Optionally, in this embodiment, a move instruction (also known as a MOVE instruction) can be used to move the target data in any other on-chip cache memory to the on-chip compute memory. Alternatively, other methods can be used to move the target data in any other on-chip cache memory to the on-chip compute memory; this embodiment does not limit the specific method used.

[0067] The first calculation involves reading each target data from off-chip memory into on-chip cache memory and moving the target data into on-chip compute memory for computation.

[0068] Step 103: Update any other on-chip cache memory to be read.

[0069] In this embodiment, any other on-chip cache memory is updated to be read from the on-chip cache memory to continue executing step 101, whereby the target data in the off-chip memory is read into the on-chip cache memory to be read. In step 102, the target data in any other on-chip cache memory is used to move to the on-chip compute memory. To ensure that each on-chip cache memory contains the target data the next time the target data is moved from the on-chip cache memory to the on-chip compute memory, this other on-chip cache memory is updated to be read from the on-chip cache memory, and the target data is read from the off-chip memory into this other on-chip cache memory.

[0070] Step 104: Repeat the read, move, and update steps until the target data in the off-chip memory is empty.

[0071] In this embodiment, steps 101-103 are executed repeatedly. After each execution, it is determined whether the target data in the off-chip memory is empty. If it is not empty, steps 101-103 are executed repeatedly until the target data in the off-chip memory is empty.

[0072] To better understand the chip-based data processing method provided in the embodiments of this application, the chip-based data processing method provided in the embodiments of this application will be described by way of example.

[0073] As an alternative implementation method, such as Figure 3 As shown, assume there are N on-chip cache memories, namely the first on-chip cache memory 32, the second on-chip cache memory 33, the third on-chip cache memory 34, the fourth on-chip cache memory 35, ..., the Nth on-chip cache memory 36. Where N is a positive integer greater than or equal to 2. During steps 101-103, the on-chip cache memory to be read is configured as the first on-chip cache memory 32. Firstly, the target data in the off-chip memory 1 is read into the first on-chip cache memory 32. Simultaneously, the target data cached in any other on-chip cache memory 33 is moved to the on-chip computing memory 31 for the first calculation. This other on-chip cache memory can be, for example, the third on-chip cache memory (...). Figure 3 The process is shown as step (1). If it is determined that the target data in the off-chip memory is not empty, the third on-chip cache memory 33 is updated to be read from the on-chip cache memory. While reading the target data from the off-chip memory 1 into the third on-chip cache memory 33, the target data cached in any other on-chip cache memory is moved to the on-chip computing memory 31 for computation. This other on-chip cache memory may be a fourth on-chip cache memory (…). Figure 3The process is shown as step (2). If it is determined that the target data in the off-chip memory is not empty, the target data in the off-chip memory 1 is read into the fourth on-chip cache memory 33, and at the same time, the target data cached in any other on-chip cache memory is moved to the on-chip computing memory 31 for calculation. This other on-chip cache memory may be the second on-chip cache memory ( Figure 3 As shown in step (3), the second on-chip cache memory 34 is updated to the on-chip cache memory to be read. This process continues until the target data in the off-chip memory is empty.

[0074] The chip-based data processing method provided in this embodiment reads target data from off-chip memory into an on-chip cache memory to be read; simultaneously, while reading the target data from off-chip memory into the on-chip cache memory to be read, target data from any other on-chip cache memory is moved to on-chip compute memory for a first calculation; the other on-chip cache memory is updated to the on-chip cache memory to be read; and the read, move, and update steps are repeated until the target data in off-chip memory is empty. The on-chip cache memory targeted in each execution of reading the target data from off-chip memory into the on-chip cache memory is different from the on-chip cache memory targeted in the parallel execution of moving the target data cached in the on-chip cache memory into the on-chip compute memory for calculation. This parallel processing method uses a symmetrical pipelining design, which can effectively avoid different types of operations competing for port bandwidth in the same on-chip memory space, thereby improving I / O operation bandwidth.

[0075] As an optional implementation, in this embodiment, step 101, reading the target data from the off-chip memory into the on-chip cache memory to be read, includes:

[0076] The target data in the off-chip memory is read into the on-chip cache memory to be read using I / O instructions.

[0077] In this embodiment, when reading target data from off-chip memory into on-chip cache memory, I / O instructions are specifically used, which can quickly read target data into on-chip cache memory, improve the efficiency of reading target data, and save the time of reading target data.

[0078] As an optional implementation, in this embodiment, step 102, moving the target data from any other on-chip cache memory to the on-chip compute memory, includes:

[0079] Move instructions are used to move target data from any other on-chip cache memory to on-chip compute memory.

[0080] In this embodiment, when moving target data from any other on-chip cache memory to on-chip compute memory, the MOVE instruction is specifically used, which can quickly move the target data to on-chip compute memory, improve the efficiency of moving target data, and save the time of moving target data.

[0081] Example 2

[0082] Figure 4 A flowchart of a chip-based data processing method provided in another embodiment of this application is shown below. Figure 4 As shown, the chip-based data processing method provided in this embodiment, compared to the chip-based data processing method provided in Embodiment 1, uses two on-chip cache memories: a first on-chip cache memory and an on-chip cache memory to be read. The data processing method employs a symmetrical drainage design. The chip-based data processing method provided in this embodiment is a further refinement of steps 101-104, specifically, it iteratively executing steps 201-202 until the target data in the off-chip memory is empty. Therefore, steps 201-202 are as follows:

[0083] Step 201: While reading the target data from the off-chip memory to the on-chip cache memory to be read, the target data cached in the first on-chip cache memory is moved to the on-chip computing memory for the first calculation.

[0084] Step 202: While reading the target data from the off-chip memory to the first on-chip cache memory, the target data cached in the on-chip cache memory to be read is moved to the on-chip computing memory for the first calculation.

[0085] In this embodiment, when there are only two on-chip cache memories on the chip, namely the on-chip cache memory to be read and the first on-chip cache memory, in order to minimize the storage space for cached data on the chip, when processing target data, when reading target data from external memory to the on-chip cache memory to be read, in order to ensure that the bandwidth of the first on-chip cache memory or the on-chip cache memory to be read is not affected when reading and moving data are performed in parallel, the data in the first on-chip cache memory is moved to the on-chip compute memory. The next time, when reading target data from external memory to the first on-chip cache memory, the data originally in the on-chip cache memory to be read is moved to the on-chip compute memory. This solution, in parallel processing, uses a symmetrical pipelining method to operate on different on-chip cache memories simultaneously, avoiding the decrease in I / O operation bandwidth caused by competition for the bandwidth of the same port of on-chip memory.

[0086] Optionally, the on-chip cache memory to be read is a shared random access memory (SRAM), and the first on-chip cache memory is a weighted random access memory (WRAM).

[0087] Optionally, the on-chip computing memory is Neuronal Random Access Memory (NRAM).

[0088] Optionally, the off-chip memory is Global Dynamic Random Access Memory (GDRAM).

[0089] like Figure 5 As shown, firstly, the target data in GDRAM (1a) is read into SRAM (32a) using an IO instruction, and simultaneously, the target data cached in WRAM (33a) is moved into NRAM (31a) using a MOVE instruction to perform a first calculation on the target data. The result of the calculation is the first calculation result data. In this parallel execution of the read and move operations, the read operation uses the port bandwidth of SRAM, and the move operation uses the port bandwidth of WRAM. The two operations do not compete for the port bandwidth of the same on-chip cache memory, achieving true independence from each other.

[0090] It should be noted that in step 201, step 202 is executed only after both parallel operations have been completed. That is, after the target data in the off-chip memory is read into the on-chip cache memory using I / O instructions, it is determined whether the first calculation has been completed by moving the target data cached in the first on-chip cache memory to the on-chip compute memory using move instructions. If it has been completed, then step 202 is executed.

[0091] like Figure 5 As shown, in step 202, while the target data in GDRAM (1a) is read into WRAM (33a) using IO instructions, the target data cached in SRAM (32a) is moved into NRAM (31a) using MOVE instructions to perform a first calculation on the target data, and the result is the first calculation result data. In this parallel execution of the read and move operations, the read operation uses the port bandwidth of WRAM, and the move operation uses the port bandwidth of SRAM. The two operations do not compete for the port bandwidth of the same on-chip cache memory, achieving true independence.

[0092] It should be noted that during the cyclic execution of steps 201-202, after each step, it is checked whether the target data in the off-chip memory is empty. If the target data is determined to be not empty, the next step is executed. If the target data is determined to be empty, it means that all target data has been processed on the chip, and the next step is stopped.

[0093] The process involves moving the target data to NRAM. The first calculation result data after the processor performs a first calculation on the target data is related to the operation type of the target data. When determining the operation type of the target data, a data processing instruction is received, which includes the data operation type. The data type is parsed to obtain the data operation type. If the data operation type is an operation type that performs one I / O operation for each target data, the first calculation result data is the final calculation result data corresponding to the target data. For example, if the operation type is an addition operation for multiple target data, the first calculation result data is the final addition result data after adding multiple target data. If the data operation type is an operation type that performs two I / O operations for each target data, the first calculation result data is the intermediate result data corresponding to the target data.

[0094] For example, if the operation is a batch normalization operator, the first I / O operation is to cache the target data stored in off-chip memory into on-chip cache memory and keep the data in on-chip cache memory. Then, the target data is moved from the on-chip cache memory to the computation cache memory so that the processor can calculate the mean and variance of the target data. Then, the target data is moved from the on-chip cache memory to the computation cache memory again. The processor uses the calculated mean and variance to normalize the target data to obtain the final result.

[0095] The chip-based data processing method provided in this embodiment uses two on-chip cache memories. While reading target data from external memory to the on-chip cache memory to be read, the target data cached in the first on-chip cache memory is moved to the on-chip compute memory for the first calculation. Similarly, while reading target data from external memory to the first on-chip cache memory, the target data cached in the on-chip cache memory to be read is moved to the on-chip compute memory for the first calculation. This allows for parallel execution of read and move operations without competition for the port bandwidth of the same on-chip cache memory, achieving true independence. This not only effectively improves on-chip data processing efficiency but also minimizes the on-chip storage space for cached data by using only two on-chip cache memories.

[0096] As an optional implementation, before reading the target data from the off-chip memory into the on-chip cache memory to be read in step 201, the method further includes receiving a data processing instruction, which includes: a data operation type; and determining whether the data operation type is an operation type that performs a second I / O operation for each target data according to the data processing instruction.

[0097] Specifically, in this embodiment, before reading the target data from the on-chip memory into the on-chip cache memory to be read, it is necessary to determine the data operation type performed in the on-chip compute memory. When determining the data operation type, data processing instructions sent by external devices are received, parsed, and the data type is determined. A mapping relationship between data operation types and the number of I / O operations performed for each target data can be pre-stored. Based on this mapping relationship, it is determined whether the data operation type is an operation type requiring two I / O operations for each target data. If it is not a data operation type requiring two I / O operations for each target data, it is determined whether it is a data operation type requiring one I / O operation for each target data. If it is determined to be a data operation type requiring one I / O operation for each target data, the first calculation result calculated in the on-chip compute memory needs to be written to the off-chip memory.

[0098] Therefore, in this embodiment, as an optional implementation, if it is determined that the target data in the off-chip memory is empty, the method further includes:

[0099] The first calculation result data in the on-chip computing memory is written to the off-chip memory using I / O instructions.

[0100] The first calculation result is the data obtained by the processor performing a first calculation on the target data.

[0101] Specifically, in this embodiment, if the data operation type is an operation type where each target data undergoes one I / O operation, then after the first calculation is completed in the arithmetic unit of the processor in the on-chip compute memory for multiple target data, the first calculation result data needs to be retrieved. Specifically, the off-chip memory communicates with the on-chip compute memory and uses I / O instructions to write the first calculation result data calculated in the arithmetic unit of the processor in the on-chip compute memory to the off-chip memory.

[0102] Example 3

[0103] Figure 6 A flowchart of a chip-based data processing method provided in another embodiment of this application is shown below. Figure 6As shown, the chip-based data processing method provided in this embodiment, based on the chip-based data processing method provided in Embodiment 1, if it is determined that the data operation type of the target data is an operation type that performs two I / O operations on each target data, then in step 102, the target data in any other on-chip cache memory is moved to the on-chip computing memory so that the target data resides in any other on-chip cache memory during the first calculation. Furthermore, if it is determined that the target data in the off-chip memory is empty, the method further includes the following steps:

[0104] Step 301: Move the target data residing in any other on-chip cache memory to the on-chip computing memory for the second calculation, and move the calculated second calculation result data to any other on-chip cache memory for caching; the second calculation is performed using the target data and the first calculation result data.

[0105] Among them, any other on-chip cache memory is any on-chip cache memory other than the on-chip cache memory to be read.

[0106] In this embodiment, optionally, a MOVE instruction can be used to move the target data residing in any other on-chip cache memory to the on-chip compute memory. The on-chip compute memory stores first computation result data. A second computation is performed based on the target data and the first computation result data, and the result of the second computation is the second computation result data. The second computation result data is the final computation result data of an operation type involving two I / O operations on each target data. After the on-chip compute memory has the second computation result data, the MOVE instruction is used to move the calculated second computation result data to the other on-chip cache memory for caching.

[0107] For example, if the data operation type is a normalization operation, the first calculation result data is the mean and variance of multiple target data. The second calculation result data is the result of subtracting the target data from the mean data and then dividing by the variance. The MOVE instruction is used to move the target data residing in any other on-chip cache memory to the on-chip compute memory. The on-chip compute memory stores the mean and variance. The difference between the target data and the mean is calculated, and then the difference is divided by the variance to obtain the second calculation result. After the on-chip compute memory has the second calculation result data, the MOVE instruction is used to move the calculated second calculation result data to that other on-chip cache memory for caching.

[0108] It should be noted that, in this embodiment, the method of moving the target data and the method of moving the second calculation result data are not limited.

[0109] Step 302: While moving the target data residing in any other on-chip cache memory to the on-chip computing memory for the second calculation, and moving the calculated second calculation result data to any other on-chip cache memory for caching, the second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory.

[0110] In this embodiment, each time the target data residing in any other on-chip cache memory is moved to the on-chip compute memory for the second calculation, and the calculated second calculation result data is moved to any other on-chip cache memory for caching, the second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory in parallel.

[0111] Optionally, in this embodiment, I / O instructions can be used to write the second calculation result data cached in the on-chip cache memory to the off-chip memory. Other methods can also be used to write the second calculation result data cached in the on-chip cache memory to the off-chip memory; this embodiment does not limit this method.

[0112] Step 303: Update the on-chip cache memory to be read to any other on-chip cache memory.

[0113] In this embodiment, any other on-chip cache memory is updated to the on-chip cache memory to be read in order to continue executing step 301.

[0114] Step 304: Repeat the move, write, and update steps until the target data in all on-chip cache memories is empty.

[0115] In this embodiment, steps 301-303 are executed repeatedly. After each execution, it is determined whether the target data in all on-chip cache memories is empty. If it is not empty, steps 301-303 are executed repeatedly until the target data in all on-chip cache memories is empty.

[0116] To better understand the chip-based data processing method provided in the embodiments of this application, the chip-based data processing method provided in the embodiments of this application will be described by way of example.

[0117] As an alternative implementation method, such as Figure 7As shown, assume there are N on-chip cache memories, namely the first on-chip cache memory 32, the second on-chip cache memory 33, ..., the Nth on-chip cache memory 36. Here, N is a positive integer greater than or equal to 2. Configure any other on-chip cache memory as the second on-chip cache memory. The cache memory to be read is the first on-chip cache memory. First, the target data residing in the second on-chip cache memory 33 is moved to the on-chip compute memory 31 for the second calculation. Simultaneously, the calculated second calculation result data is moved to the second on-chip cache memory 33 for caching, and the second calculation result data cached in the first on-chip cache memory 32 is written to the off-chip memory 1. Figure 7 The process is shown as step (1). If it is determined that the target data still exists in the on-chip cache memory, any other on-chip cache memory can be the first on-chip cache memory. Then, the target data residing in the first on-chip cache memory is moved to the on-chip compute memory for the second calculation. At the same time, the calculated second calculation result data is moved to the first on-chip cache memory 34 for caching. Simultaneously, the second calculation result data cached in the on-chip cache memory 33 to be read is written to the off-chip memory 1. The on-chip cache memory to be read can be the fourth on-chip cache memory (e.g., the fourth on-chip cache memory). Figure 7 The process is shown as step (2). Any on-chip cache memory can be a fourth on-chip cache memory. Then, the target data residing in the fourth on-chip cache memory is moved to the on-chip compute memory for the second calculation. Simultaneously, the calculated second calculation result data is moved to the fourth on-chip cache memory 34 for caching. At the same time, the second calculation result data cached in the on-chip cache memory 33 to be read is written to the off-chip memory 1. This on-chip cache memory to be read can be, for example, a second on-chip cache memory (…). Figure 7 (3) is shown in the figure. Continue in this column until the target data in the on-chip cache memory is empty.

[0118] The chip-based data processing method provided in this embodiment involves a data operation type where the target data is subjected to two I / O operations for each target data. If it is determined that the target data does not exist in the off-chip memory, the target data residing in any other on-chip cache memory is moved to the on-chip compute memory for a second calculation, and the calculated second calculation result data is moved to any other on-chip cache memory for caching. The second calculation uses the target data and the first calculation result data. Each time the target data residing in any other on-chip cache memory is moved to the on-chip compute memory for the second calculation, and the calculated second calculation result data is moved to any other on-chip cache memory for caching, the second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory, and any other on-chip cache memory is updated to the on-chip cache memory to be read. The move, write, and update steps are repeated until the target data in all on-chip cache memories is empty. The parallel processing method, which involves moving the target data residing in any other on-chip cache memory to the on-chip compute memory for a second computation and then moving the calculated result data to any other on-chip cache memory for caching, differs from the parallel execution method of writing the second computation result data cached in the on-chip cache memory to the on-chip cache memory in the off-chip memory. This parallel processing method is a symmetrical pipelining design, which can avoid different types of operations competing for the port bandwidth of the same on-chip memory space, thereby improving the I / O operation bandwidth and effectively improving the on-chip data processing efficiency.

[0119] As an optional implementation, in this embodiment, step 301, moving the target data residing in any other on-chip cache memory to the on-chip compute memory for the second calculation, and moving the calculated second calculation result data to any other on-chip cache memory for caching, includes:

[0120] The target data residing in any other on-chip cache memory is moved to the on-chip compute memory for the second computation using move instructions, and the calculated second computation result data is moved to any other on-chip cache memory for caching using move instructions.

[0121] In this embodiment, when moving the target data residing in any other on-chip cache memory to the on-chip compute memory for the second calculation, and when moving the calculated second calculation result data to any other on-chip cache memory for caching, the MOVE instruction is specifically used. This can quickly move the target data between any other on-chip cache memory and the on-chip compute memory, improving the efficiency of moving the target data and saving the time of moving the target data.

[0122] As an optional implementation, in this embodiment, step 302, writing the second calculation result data cached in the on-chip cache memory to be read to the off-chip memory, includes:

[0123] The second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory using I / O instructions.

[0124] In this embodiment, when writing the second calculation result data cached in the on-chip cache memory to the off-chip memory, I / O instructions are specifically used, which can quickly write the target data to the off-chip memory, improve the writing efficiency of the target data, and save the time of writing the target data.

[0125] Example 4

[0126] Figure 8 A flowchart of a chip-based data processing method provided in another embodiment of this application is shown below. Figure 8 As shown, the chip-based data processing method provided in this embodiment, compared to the chip-based data processing method provided in Embodiment 3, uses two on-chip cache memories: a first on-chip cache memory and an on-chip cache memory to be read. The data processing method employs a symmetrical drainage design. The chip-based data processing method provided in this embodiment is a further refinement of steps 301-304, specifically, it iteratively executing steps 401-402 until the target data in all on-chip cache memories is empty. Steps 401 and 402 are as follows:

[0127] Step 401: Move the target data residing in the first on-chip cache memory to the on-chip computing memory for the second calculation, and simultaneously move the calculated second calculation result data to the first on-chip cache memory for caching, while writing the second calculation result data cached in the on-chip cache memory to be read to the off-chip memory.

[0128] Step 402: Move the target data residing in the on-chip cache memory to be read to the on-chip computing memory for the second calculation, and simultaneously move the calculated second calculation result data to the on-chip cache memory to be read for caching, while writing the second calculation result data cached in the first on-chip cache memory to the off-chip memory.

[0129] In this embodiment, when there are only two on-chip cache memories on the chip—the on-chip cache memory to be read and the first on-chip cache memory—to minimize the storage space for cached data on the chip, when processing target data, when moving the target data residing in the on-chip cache memory to be read for a second calculation, and then moving the calculated second calculation result data back to the on-chip cache memory to be read for caching, to ensure that the bandwidth of the first on-chip cache memory or the on-chip cache memory to be read is not affected when moving and writing data in parallel, the second calculation result data cached in the first on-chip cache memory is written to external memory. Next time, the target data residing in the first on-chip cache memory is moved to the on-chip compute memory for a second calculation, and the calculated second calculation result data is moved back to the first on-chip cache memory for caching, while simultaneously writing the data originally in the on-chip cache memory to be read to external memory. This solution, during parallel processing, uses a symmetrical pipelining method to operate on different on-chip cache memories simultaneously, avoiding the decrease in I / O operation bandwidth caused by competition for the bandwidth of the same port of on-chip memory.

[0130] Optionally, the on-chip cache memory to be read is a shared random access memory (SRAM), and the first on-chip cache memory is a weighted random access memory (WRAM).

[0131] Optionally, the on-chip computing memory is Neuronal Random Access Memory (NRAM).

[0132] Optionally, the off-chip memory is Global Dynamic Random Access Memory (GDRAM).

[0133] like Figure 9As shown, firstly, a move instruction is used to move the target data residing in WRAM (33a) to NRAM (31a) for the second calculation. Simultaneously, the calculated second calculation result is moved back to WRAM (33a) for caching. At the same time, an I / O instruction is used to write the second calculation result cached in SRAM (32a) to GDRAM (1a). In this parallel execution of the move and write operations, the move operation utilizes the port bandwidth of WRAM (33a), and the write operation utilizes the port bandwidth of SRAM (32a). The two operations do not compete for the port bandwidth of the same on-chip cache memory, achieving true independence.

[0134] It should be noted that in step 401, after both parallel operations are completed, step 402 is executed. That is, the target data residing in WRAM is moved to the on-chip computing memory for the second calculation, and the calculated second calculation result data is moved to WRAM for caching. After the execution is completed, it is determined whether the writing of the second calculation result data cached in SRAM to the off-chip memory is completed. If the execution is completed, step 402 is executed.

[0135] like Figure 9 As shown, in step 402, a move instruction is used to move the target data residing in SRAM (32a) to NRAM (31a) for the second calculation, and the calculated second calculation result data is moved to SRAM (32a) for caching. Simultaneously, an I / O instruction is used to write the second calculation result data cached in WRAM (33a) to GDRAM (1a). In this parallel execution of the move and write operations, the move operation uses the port bandwidth of SRAM (32a), and the write operation uses the port bandwidth of WRAM (33a). The two operations do not compete for the port bandwidth of the same on-chip cache memory, achieving true independence from each other.

[0136] It should be noted that during the cyclic execution of steps 401-402, after each step, it is determined whether the target data in the on-chip cache memory is empty. If the target data is determined to be not empty, the next step is executed. If the target data is determined to be empty, it means that all target data has been calculated on the chip, and the next step is stopped.

[0137] The chip-based data processing method provided in this embodiment uses two on-chip cache memories. Target data residing in the first on-chip cache memory is moved to the on-chip compute memory for a second calculation, and the calculated second calculation result is moved back to the first on-chip cache memory for caching. Simultaneously, the second calculation result data cached in the on-chip cache memory to be read is written to external memory. Similarly, when target data residing in the on-chip cache memory to be read is moved to the on-chip compute memory for a second calculation, and the calculated second calculation result data is moved back to the on-chip cache memory to be read for caching, the second calculation result data cached in the first on-chip cache memory is written to external memory. This ensures that when each move and write operation is executed in parallel, the two operations do not compete for the port bandwidth of the same on-chip cache memory, achieving true independence. This not only effectively improves on-chip data processing efficiency but also minimizes the on-chip storage space for cached data by using only two on-chip cache memories.

[0138] In one possible implementation, an artificial intelligence chip is also disclosed, which includes on-chip computing memory, multiple on-chip cache memories, and the aforementioned chip-based data processing device.

[0139] In one possible implementation, a board is also disclosed. Figure 10 This is a structural diagram illustrating a board according to an embodiment of this application. For example... Figure 10 As shown, board 100 includes an artificial intelligence chip 1001, which is a system-on-chip (SoC) that integrates one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 100 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and substantial computing power.

[0140] The artificial intelligence chip 1001 is connected to an external device 1003 via an external interface device 1002. The external device 1003 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from the external device 1003 to the artificial intelligence chip 1001 via the external interface device 1002. The calculation results of the artificial intelligence chip 1001 can be transmitted back to the external device 1003 via the external interface device 1002. Depending on the application scenario, the external interface device 1002 may have different interface forms, such as a PCIe interface.

[0141] The board 100 also includes a storage device 1004 for storing data, which includes one or more memory cells 1005. The storage device 1004 is connected to and transmits data with the controller 1006 and the artificial intelligence chip 1001 via a bus. The controller 1006 in the board 100 is configured to regulate the state of the artificial intelligence chip 1001. Therefore, in one application scenario, the controller 1006 may include a microcontroller (MCU).

[0142] In one possible implementation, a combined processing device is also provided. Figure 11 This is a structural diagram illustrating the combined processing apparatus in chip 1001 of this embodiment. (As shown...) Figure 11 As shown, the combined processing device 110 includes a computing device 1101, an interface device 1102, a processing device 1103, and a GDRAM 1104.

[0143] The computing device 1101 is configured to perform user-specified operations. It is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 1103 through the interface device 1102 to jointly complete the user-specified operations.

[0144] Interface device 1102 is used to transmit data and control commands between computing device 1101 and processing device 1103. For example, computing device 1101 can obtain input data from processing device 1103 via interface device 1102 and write it to on-chip storage device of computing device 1101. Further, computing device 1101 can obtain control commands from processing device 1103 via interface device 1102 and write them to on-chip control cache of computing device 1101. Alternatively or optionally, interface device 1102 can also read data from storage device of computing device 1101 and transmit it to processing device 1103.

[0145] The processing device 1103, as a general-purpose processing device, performs basic controls including but not limited to data transfer, and starting and / or stopping the computing device 1101. Depending on the implementation, the processing device 1103 may be one or more types of processors, including but not limited to digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, the computing device 1101 of this application can be considered as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 1101 and the processing device 1103 are considered together, they are considered to form a heterogeneous multi-core structure.

[0146] GDRAM 1104 is used to store data to be processed. It is DDR memory, typically 16G or larger, and is used to store data in computing device 1101 and / or processing device 1103.

[0147] In one possible implementation, a computing device is also provided. Figure 12 This is a schematic diagram illustrating the internal structure of a computing device according to an embodiment of this application. Figure 12 As shown, computing device 1101 is used to process input data such as computer vision, speech, natural language, and data mining. The computing device 1101 in the figure adopts a multi-core hierarchical structure design. As a system-on-a-chip, computing device 1101 includes multiple clusters, and each cluster includes multiple processor cores. In other words, computing device 1101 is composed of a system-on-a-chip-cluster-processor core hierarchy.

[0148] From the perspective of system-on-a-chip hierarchy, such as Figure 12 As shown, the computing device 1101 includes an external storage controller 1201, a peripheral communication module 1202, an on-chip interconnect module 1203, a synchronization module 1204, and multiple clusters 1205.

[0149] There can be multiple external storage controllers 1201; two are shown as an example in the figure. These controllers are used to respond to access requests from the processor core to access external storage devices, such as… Figure 2The GDRAM 1104 in the chip allows data to be read from or written to external devices. The peripheral communication module 1202 receives control signals from the processing device 1103 via the interface device 1102, initiating the computing device 1101 to execute tasks. The on-chip interconnect module 1203 connects the external memory controller 1201, the peripheral communication module 1202, and multiple clusters 1205 to transmit data and control signals between modules. The synchronization module 1204 is a global barrier controller (GBC) used to coordinate the working progress of each cluster and ensure information synchronization. The multiple clusters 1205 are the computing core of the computing device 1101. Four are shown exemplary in the figure; however, with hardware development, the computing device 1101 of this application may also include 8, 16, 64, or even more clusters 1205. The clusters 1205 are used to efficiently execute deep learning algorithms.

[0150] From the perspective of cluster hierarchy, such as Figure 12 As shown, each cluster 1205 includes multiple processor cores (IPU cores) 1206 and one memory core (MEM core) 1207.

[0151] Four processor cores 1206 are shown in the figure as an example; this application does not limit the number of processor cores 1206. Its internal architecture is as follows: Figure 13 As shown. Each processor core 1206 includes a control module 131 and a storage module 133.

[0152] The control module 131 controls the operation of the storage module 133 to complete the deep learning task. It includes an instruction fetch unit (IFU) 1311 and an instruction decode unit (IDU) 1312. The instruction fetch unit 1311 fetches instructions from the processing device 1103, and the instruction decode unit 1312 decodes the fetched instructions and sends the decoding result as control information to the storage module 43.

[0153] The storage module 133 is used to store or move relevant data, including neuron RAM (NRAM) 1331, weight RAM (WRAM) 1332, and NRAM 1331 is used to store input, output data and intermediate results for the processor core 1206 to calculate.

[0154] Storage core 1207 is primarily used for storage and communication, namely storing shared data or intermediate results among processor cores 1206, and performing communication between cluster 1205 and GDRAM 1104, communication between clusters 1205, and communication between processor cores 1206. In other embodiments, storage core 1207 has scalar operation capabilities and is used to perform scalar operations.

[0155] Storage core 1207 includes a shared memory unit (SRAM) 1208, a broadcast bus 1209, a cluster direct memory access (CDMA) module 3110, and a global direct memory access (GDMA) module 311. SRAM 1208 acts as a high-performance data relay station. Data multiplexed between different processor cores 1206 within the same cluster 1205 does not need to be obtained from GDRAM 1104 by each processor core 1206 individually. Instead, it is relayed between processor cores 1206 via SRAM 1208. Storage core 1207 only needs to quickly distribute the multiplexed data from SRAM 1208 to multiple processor cores 1206, thereby improving inter-core communication efficiency and significantly reducing on-chip and off-chip I / O access. Each processor core 1206 includes a Neuronal Random Access Memory (NRAM).

[0156] like Figure 13 As shown in this embodiment, the control module 131 acquires IO and MOVE instructions from the processing device 1103. The instruction decoding unit 1312 decodes the acquired instructions and sends the decoding result as control information to the storage module 133. The storage module 133 also sends the control information to the off-chip GDRAM 1104. While using IO instructions to read the target data from GDRAM 1104 into SRAM 1208, the MOVE instructions move the target data from WRAM 1332 into NRAM 1331 for the first calculation. The SRAM 1208 is updated to WRAM 1332, and the read, move, and update steps are repeated until the target data in GDRAM 1104 is empty.

[0157] If a secondary I / O operation is performed on each target data on the chip, the control module 131 obtains I / O instructions and MOVE instructions from the processing device 1103. The instruction decoding unit 1312 decodes the obtained instructions and sends the decoding result as control information to the storage module 133. The MOVE instruction moves the target data residing in WRAM 1332 to NRAM for a second calculation, and moves the calculated second calculation result data to WRAM 1332 for caching. At the same time, the I / O instruction writes the second calculation result data cached in SRAM 1208 to GDRAM 1104; updates SRAM 1208 to WRAM 1332; and repeats the move, write, and update steps until the target data in WRAM 1332 and SRAM are empty.

[0158] Broadcast bus 1209, CDMA 3110, and GDMA 311 are used to perform communication between processor cores 1206, communication between clusters 1205, and data transfer between cluster 1205 and GDRAM 1104, respectively. These will be explained below.

[0159] The broadcast bus 1209 is used to complete high-speed communication between the processor cores 1206 within the cluster 1205. In this embodiment, the broadcast bus 1209 supports inter-core communication methods including unicast, multicast, and broadcast. Unicast refers to point-to-point (i.e., single processor core to single processor core) data transmission. Multicast is a communication method that transmits a piece of data from SRAM 1208 to several specific processor cores 1206. Broadcast is a communication method that transmits a piece of data from SRAM 1208 to all processor cores 1206, and is a special case of multicast.

[0160] CDMA 3110 is used to control SRAM 1208 access between different clusters 1205 within the same computing device 1101.

[0161] The GDMA 311 works in conjunction with the external memory controller 1201 to control memory access from SRAM 1208 to GDRAM 1104 in cluster 1205, or to read data from GDRAM 1104 into SRAM 1208. For example... Figure 13 As shown, GDRAM 1104 communicates with NRAM 1331 or WRAM 1332. SRAM 1208 communicates with NRAM 1331.

[0162] This application also provides a chip-based data processing device, which includes a processor and a memory.

[0163] This includes a memory that is communicatively connected to at least one processor.

[0164] The memory stores program code, and the processor calls the program code stored in the memory to execute it. Figure 1 , Figure 4 , Figure 6 or Figure 8 The chip-based data processing method provided in any of the embodiments shown.

[0165] For relevant instructions, please refer to the corresponding text. Figure 1 , Figure 4 , Figure 6 or Figure 8 The relevant descriptions and effects corresponding to the steps will be understood, and will not be elaborated on here.

[0166] In this embodiment, the memory and the processor are connected via a bus.

[0167] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform actions to achieve... Figure 1 , Figure 4 , Figure 6 or Figure 8 The chip-based data processing method provided in any of the embodiments shown.

[0168] In one possible implementation, an electronic device is also provided, which includes the aforementioned off-chip memory and artificial intelligence chip. The electronic device can be a data processing device, robot, computer, printer, scanner, tablet computer, smart terminal, mobile phone, dashcam, navigator, sensor, camera, server, cloud server, camera, camcorder, projector, watch, earphone, mobile storage, wearable device, vehicle, home appliance, and / or medical device.

[0169] Transportation includes airplanes, ships and / or vehicles; household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; medical equipment includes MRI machines, ultrasound machines and / or electrocardiographs.

[0170] The foregoing may be better understood in view of the following clauses:

[0171] Clause A1, a chip-based data processing method, the chip including on-chip computing memory and multiple on-chip cache memories, wherein the on-chip cache memories include an on-chip cache memory to be read, the on-chip cache memory communicating with off-chip memory, the off-chip memory storing multiple target data, the method comprising:

[0172] Read the target data from the off-chip memory into the on-chip cache memory to be read;

[0173] While reading the target data from the off-chip memory into the on-chip cache memory to be read, the target data from any other on-chip cache memory is moved into the on-chip computing memory for the first calculation.

[0174] Update any other on-chip cache memory to the on-chip cache memory to be read; and

[0175] Repeat the read, move, and update steps until the target data in the off-chip memory is empty.

[0176] Clause 2. According to the method described in Clause 1, when there are two on-chip cache memories, the on-chip cache memory further includes a first on-chip cache memory, the data processing method adopts a symmetrical drainage design, and the method includes:

[0177] Read the target data from the off-chip memory into the on-chip cache memory to be read;

[0178] While reading the target data from the off-chip memory to the on-chip cache memory to be read, the target data cached in the first on-chip cache memory is moved to the on-chip computing memory for the first calculation.

[0179] While reading the target data from the off-chip memory into the first on-chip cache memory, the target data cached in the on-chip cache memory to be read is moved into the on-chip computing memory for the first calculation.

[0180] Clause 3. According to the method described in Clause 1, reading the target data from the off-chip memory into the on-chip cache memory to be read includes:

[0181] The target data in the off-chip memory is read into the on-chip cache memory to be read using I / O instructions.

[0182] Clause 4. The method described in Clause 1, wherein moving the target data from any other on-chip cache memory to the on-chip compute memory comprises:

[0183] The target data in any other on-chip cache memory is moved to the on-chip compute memory using a move instruction.

[0184] Clause 5. The method described in any one of Clauses 1-4, wherein the off-chip memory communicates with the on-chip computing memory;

[0185] If it is determined that the target data in the off-chip memory is empty, the method further includes:

[0186] The first calculation result data in the on-chip computing memory is written to the off-chip memory using I / O instructions. The first calculation result is the data obtained by the on-chip computing memory performing a first calculation on the target data.

[0187] Clause 6. According to the method described in Clause 1, if a second I / O operation is performed on the chip for each target data, then when moving the target data from any other on-chip cache memory to the on-chip compute memory for the first computation, the method further includes:

[0188] The target data resides in any of the other on-chip cache memories.

[0189] Clause 7. The method described in Clause 6, before reading the target data from the off-chip memory into the on-chip cache memory to be read, further includes:

[0190] Receive data processing instructions, wherein the data processing instructions include: data operation type;

[0191] Based on the data processing instructions, determine whether the data operation type is an operation type that performs a secondary I / O operation on each target data.

[0192] Clause 8. According to the method described in Clause 6, if it is determined that the target data in the off-chip memory is empty, the method further includes:

[0193] The target data residing in any other on-chip cache memory is moved to the on-chip compute memory for a second calculation, and the calculated second calculation result data is moved to the other on-chip cache memory for caching; the second calculation is performed using the target data and the first calculation result data;

[0194] Each time the target data residing in any other on-chip cache memory is moved to the on-chip compute memory for the second calculation, and the calculated second calculation result data is moved to any other on-chip cache memory for caching, the second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory.

[0195] Update the on-chip cache memory to be read to any other on-chip cache memory; and

[0196] Repeat the move, write, update steps until the target data in all on-chip cache memories is empty.

[0197] Clause 9. According to the method described in Clause 8, when there are two on-chip cache memories, the on-chip cache memory further includes a first on-chip cache memory; the data processing method employs a symmetrical drainage design, and the method includes:

[0198] The target data residing in the first on-chip cache memory is moved to the on-chip compute memory for the second computation, and the calculated second computation result data is moved to the first on-chip cache memory for caching;

[0199] The target data residing in the first on-chip cache memory is moved to the on-chip computing memory for the second calculation, and the calculated second calculation result data is moved to the first on-chip cache memory for caching. At the same time, the second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory.

[0200] The target data residing in the on-chip cache memory to be read is moved to the on-chip computing memory for the second calculation, and the calculated second calculation result data is moved to the on-chip cache memory to be read for caching. At the same time, the second calculation result data cached in the first on-chip cache memory is written out to the off-chip memory.

[0201] Clause 10. The method described in Clause 8, wherein moving the target data residing in any other on-chip cache memory to the on-chip compute memory for a second computation, and moving the calculated second computation result data to the any other on-chip cache memory for caching, comprises:

[0202] The target data residing in any other on-chip cache memory is moved to the on-chip compute memory for a second computation using a move instruction, and the calculated second computation result data is moved to the any other on-chip cache memory for caching using a move instruction.

[0203] Clause 11. The method described in Clause 8, wherein writing the second calculation result data cached in the on-chip cache memory to be read to the off-chip memory, includes:

[0204] The second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory using I / O instructions.

[0205] Clause 12. In accordance with the method described in Clause 2 or 9, the on-chip cache memory to be read is a shared random access memory (SRAM), the first on-chip cache memory is a weighted random access memory (WRAM), the on-chip compute memory is a neuron random access memory (NRAM), and the off-chip memory is a global dynamic random access memory (GDRAM).

[0206] Clause 13. A chip-based data processing device, comprising: a processor and a memory; wherein,

[0207] The memory is used to store program code;

[0208] The processor is configured to invoke the level code stored in the memory to execute the method described in any one of clauses 1 to 12.

[0209] Clause 14. An artificial intelligence chip, comprising: on-chip computing memory, a plurality of on-chip cache memories, and a chip-based data processing apparatus as described in Clause 13.

[0210] Clause 15. An electronic device comprising: off-chip memory and an artificial intelligence chip as described in Clause 14.

[0211] Clause 16. A board, the board comprising: a storage device, an external interface device, and a control device, as well as an artificial intelligence chip as described in Clause 14;

[0212] The artificial intelligence chip is connected to the storage device, the control device, and the external interface device, respectively.

[0213] The storage device is used to store the target data;

[0214] The external interface device is used to realize data transmission between the artificial intelligence chip and external devices;

[0215] The controller is used to monitor the state of the artificial intelligence chip.

[0216] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0217] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0218] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0219] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0220] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, an AI processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, storage units can be any suitable magnetic or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0221] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0222] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

Claims

1. A chip-based data processing method, characterized in that, The chip includes on-chip computing memory and multiple on-chip cache memories, wherein the on-chip cache memory includes an on-chip cache memory to be read, the on-chip cache memory communicates with off-chip memory, and the off-chip memory stores multiple target data, the method includes: Read the target data from the off-chip memory into the on-chip cache memory to be read; While reading the target data from the off-chip memory into the on-chip cache memory to be read, the target data from any other on-chip cache memory is moved into the on-chip computing memory for the first calculation. Update any other on-chip cache memory to the on-chip cache memory to be read; and Repeat the read, move, and update steps until the target data in the off-chip memory is empty; When the data operation type of the target data is determined to be an operation type that involves two I / O operations for each target data item, During the process of moving target data from any other on-chip cache memory to on-chip compute memory for performing a first computation, after the target data resides in any other on-chip cache memory and the target data in off-chip memory is empty, the method further includes: The target data residing in any other on-chip cache memory is moved to the on-chip compute memory for a second calculation, and the calculated second calculation result data is moved to the other on-chip cache memory for caching; the second calculation is performed using the target data and the first calculation result data; Each time the target data residing in any other on-chip cache memory is moved to the on-chip compute memory for the second calculation, and the calculated second calculation result data is moved to any other on-chip cache memory for caching, the second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory. Update the on-chip cache memory to be read to any other on-chip cache memory; and Repeat the move, write, update steps until the target data in all on-chip cache memories is empty.

2. The method according to claim 1, characterized in that, When there are two on-chip cache memories, the on-chip cache memory further includes a first on-chip cache memory. The data processing method adopts a symmetrical drainage design, and the method includes: Read the target data from the off-chip memory into the on-chip cache memory to be read; While reading the target data from the off-chip memory to the on-chip cache memory to be read, the target data cached in the first on-chip cache memory is moved to the on-chip computing memory for the first calculation. While reading the target data from the off-chip memory into the first on-chip cache memory, the target data cached in the on-chip cache memory to be read is moved into the on-chip computing memory for the first calculation.

3. The method according to claim 1, characterized in that, The step of reading the target data from the off-chip memory into the on-chip cache memory to be read includes: The target data in the off-chip memory is read into the on-chip cache memory to be read using I / O instructions.

4. The method according to claim 1, characterized in that, Moving the target data from any other on-chip cache memory to the on-chip compute memory includes: The target data in any other on-chip cache memory is moved to the on-chip compute memory using a move instruction.

5. The method according to any one of claims 1-4, characterized in that, The off-chip memory communicates with the on-chip computing memory; If it is determined that the target data in the off-chip memory is empty, the method further includes: The first calculation result data in the on-chip computing memory is written to the off-chip memory using I / O instructions. The first calculation result is the data obtained by performing a first calculation on the target data.

6. The method according to claim 1, characterized in that, Before reading the target data from the off-chip memory into the on-chip cache memory to be read, the method further includes: Receive data processing instructions, wherein the data processing instructions include: data operation type; Based on the data processing instructions, determine whether the data operation type is an operation type that performs a secondary I / O operation on each target data.

7. The method according to claim 1, characterized in that, When performing the second calculation, if there are two on-chip cache memories, the on-chip cache memory further includes a first on-chip cache memory; the data processing method adopts a symmetrical drainage design, and the method includes: The target data residing in the first on-chip cache memory is moved to the on-chip compute memory for the second computation, and the calculated second computation result data is moved to the first on-chip cache memory for caching; The target data residing in the first on-chip cache memory is moved to the on-chip computing memory for the second calculation, and the calculated second calculation result data is moved to the first on-chip cache memory for caching. At the same time, the second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory. The target data residing in the on-chip cache memory to be read is moved to the on-chip computing memory for the second calculation, and the calculated second calculation result data is moved to the on-chip cache memory to be read for caching. At the same time, the second calculation result data cached in the first on-chip cache memory is written out to the off-chip memory.

8. The method according to claim 1, characterized in that, When performing the second calculation, the step of moving the target data residing in any other on-chip cache memory to the on-chip compute memory for the second calculation, and moving the calculated second calculation result data to the other other on-chip cache memory for caching, includes: The target data residing in any other on-chip cache memory is moved to the on-chip compute memory for a second computation using a move instruction, and the calculated second computation result data is moved to the any other on-chip cache memory for caching using a move instruction.

9. The method according to claim 1, characterized in that, When performing the second calculation, writing the second calculation result data cached in the on-chip cache memory to the off-chip memory includes: The second calculation result data cached in the on-chip cache memory to be read is written out to the off-chip memory using I / O instructions.

10. The method according to claim 2 or 7, characterized in that, The on-chip cache memory to be read is a shared random access memory (SRAM), the first on-chip cache memory is a weighted random access memory (WRAM), the on-chip computation memory is a neuron random access memory (NRAM), and the off-chip memory is a global dynamic random access memory (GDRAM).

11. A chip-based data processing device, characterized in that, include: Processor and memory; among which, The memory is used to store program code; The processor is configured to invoke the level code stored in the memory to execute the method according to any one of claims 1 to 10.

12. An artificial intelligence chip, characterized in that, include: On-chip computing memory, multiple on-chip cache memories, and the chip-based data processing device as described in claim 11.

13. An electronic device, characterized in that, The electronic device includes: off-chip memory and the artificial intelligence chip as described in claim 12.

14. A circuit board, characterized in that, The board includes: a storage device, an external interface device, a control device, and an artificial intelligence chip as described in claim 12; The artificial intelligence chip is connected to the storage device, the control device, and the external interface device, respectively. The storage device is used to store the target data; The external interface device is used to realize data transmission between the artificial intelligence chip and external devices; The controller is used to monitor the state of the artificial intelligence chip.

Citation Information

Patent Citations

  • Data double-layer caching method suitable for special neural network accelerator

    CN111783967A