A neural network computing system and method based on data flow architecture

By setting up conversion units in the artificial intelligence acceleration chip of the data flow architecture, efficient data transmission between on-chip and off-chip memory is achieved, the bottlenecks in data transmission efficiency and cost of traditional chips are solved, and the performance of neural network computing is improved.

CN111860788BActive Publication Date: 2025-05-06SHENZHEN CORERAIN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010733604.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2025-05-06
Estimated Expiration
2040-07-27

AI Technical Summary

Technical Problem

The existing data flow architecture's artificial intelligence acceleration chip has a performance bottleneck between data transmission efficiency and chip cost. The slow access speed of traditional buses and large on-chip RAM lead to large chip area and high cost.

Method used

By setting a conversion unit between the on-chip memory and the off-chip memory, data transmission between different bus protocols is realized, and data flow selection is optimized to reduce transmission time.

Benefits of technology

The speed and efficiency of neural network computing are improved, and the cost and area problems caused by slow access speed of traditional buses and large capacity RAM on chip are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111860788B_ABST
    Figure CN111860788B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a neural network computing system and method based on a data flow architecture. The neural network computing system based on the data flow architecture includes: an off-chip memory and a neural network acceleration module, the neural network acceleration module includes a conversion unit, an on-chip memory and a computing unit; the off-chip memory is used to store data; the conversion unit is connected between the on-chip memory and the off-chip memory, and is used to realize the conversion between the first communication bus of the off-chip memory and the second communication bus of the on-chip memory, so as to store the data in the on-chip memory; the computing unit is directly connected to the on-chip memory through the second communication bus, and is used to perform calculations based on the data received from the on-chip memory. By arranging a conversion unit between the on-chip memory and the off-chip memory, the speed and efficiency of the neural network computing of the data flow architecture are improved while ensuring the cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of neural network technology, for example, a neural network computing system and method based on a data flow architecture. Background Art

[0002] With the rapid development of computers, the calculation of neural network data is becoming more and more important. The calculation of neural networks requires a large amount of data. The data flow architecture completes the entire calculation process based on the continuous flow of data, without the participation of the instruction set. The traditional instruction set architecture has to go through several stages to complete a complete operation, including the instruction fetch stage, instruction decoding stage, instruction execution stage, memory access stage, and result write-back stage. The entire process is extremely inefficient. Compared with the instruction set architecture, the data flow architecture can maximize the performance and efficiency of the chip, but because the data flow architecture requires data to flow at all times, it places strict requirements on storage, bus, and bandwidth.

[0003] At present, for AI acceleration chips with data flow architecture, the data for these calculations mainly comes from two ways. Method 1: Access external storage through the bus, such as accessing the off-chip double data rate (DDR) memory through the AXI (Advanced eXtensible Interface) bus; Method 2: By placing a relatively large random access memory (RAM) on the chip, the output of each layer of the neural network is stored on the chip.

[0004] However, although the cost of method 1 is relatively low, the transmission rate is limited by the bus rate and bandwidth. Because often a chip will have multiple devices accessing external storage at the same time, and the external storage can only be read or written at the same time, and reading and writing cannot be simultaneous. This seriously affects the efficiency of data transmission, and the neural network must receive data before it can start calculating, so the data transmission time and data processing time of method 1 are difficult to overlap, causing a performance bottleneck. Although method 2 saves the transmission time of a copy of data through local storage, the capacity of the local storage must be greater than the layer with the largest amount of data in the neural network to achieve full functionality, which will cause the chip area to be very large, resulting in a sharp increase in chip costs. Summary of the invention

[0005] The embodiments of the present invention provide a neural network computing system and method based on a data flow architecture, so as to improve the speed and efficiency of the neural network computing of the data flow architecture while ensuring the cost.

[0006] In a first aspect, an embodiment of the present invention provides a flow-based neural network computing system, comprising:

[0007] An off-chip memory and a neural network acceleration module, wherein the neural network acceleration module includes a conversion unit, an on-chip memory and a computing unit;

[0008] The off-chip memory is used to store data;

[0009] The conversion unit is connected between the on-chip memory and the off-chip memory, and is used to implement conversion between the first communication bus of the off-chip memory and the second communication bus of the on-chip memory, so as to store the data in the on-chip memory;

[0010] The computing unit is directly connected to the on-chip memory via a second communication bus, and is configured to perform computing based on the data received from the on-chip memory.

[0011] Optionally, the computing unit is directly connected to the off-chip memory via a first communication bus, and is configured to perform computing based on the data directly received from the off-chip memory.

[0012] Optionally, the computing unit includes:

[0013] The data flow direction selection subunit is used to control the flow direction of the data.

[0014] Optionally, the first communication bus is an AXI bus, and the conversion unit and the off-chip memory interact with the data through an AXI bus protocol, or the computing unit and the off-chip memory interact with the data through an AXI bus protocol.

[0015] Optionally, the second communication bus is a Sif bus, and the conversion unit and the on-chip memory interact with each other on the data through a Sif bus protocol, and the on-chip memory and the computing unit interact with each other on the data through the Sif bus protocol.

[0016] Optionally, the Sif bus has at least one channel.

[0017] Optionally, the computing unit includes an operator.

[0018] In a second aspect, an embodiment of the present invention provides a neural network computing method based on a data flow architecture, which is applied to a neural network computing system as described in any embodiment of the present invention, wherein the neural network is multi-layered, and the method includes:

[0019] Get the data that needs to be calculated for the current layer of the neural network;

[0020] estimating the calculation time and transmission time of the current layer based on the data;

[0021] A transmission path of the data to be calculated is determined according to the calculation time and the transmission time.

[0022] Optionally, the transmission time includes a first transmission time and a second transmission time, the first transmission time is the data transmission time between the computing unit and the off-chip memory, the second transmission time is the data transmission time between the computing unit and the on-chip memory, and the first transmission time is greater than the second transmission time;

[0023] The determining, according to the calculation time and the transmission time, a transmission path of the data to be calculated includes:

[0024] Determining whether the calculation time is greater than the first transmission time;

[0025] When the computing time is greater than the first transmission time, controlling the computing unit to interact with an off-chip memory for data, or controlling the computing unit to interact with an on-chip memory for data;

[0026] When the calculation time is less than or equal to the first transmission time, the calculation unit interacts with the on-chip memory for data.

[0027] Optionally, controlling the computing unit to interact with an off-chip memory for data, or controlling the computing unit to interact with an on-chip memory for data, includes:

[0028] determining whether the data is in order;

[0029] When the data is arranged in order, controlling the computing unit to interact with the off-chip memory for the data;

[0030] In the case that the data are not arranged in order, the computing unit is controlled to interact with the on-chip memory for the data.

[0031] The neural network computing system based on the data flow architecture of the embodiment of the present invention includes: an off-chip memory and a neural network acceleration module, the neural network acceleration module includes a conversion unit, an on-chip memory and a computing unit; the off-chip memory is used to store data; the conversion unit is connected between the on-chip memory and the off-chip memory, and is used to realize the conversion between the first communication bus of the off-chip memory and the second communication bus of the on-chip memory, so as to store the data in the on-chip memory; the computing unit is directly connected to the on-chip memory through the second communication bus, and is used to perform calculations based on the data received from the on-chip memory. By arranging a conversion unit between the on-chip memory and the off-chip memory, the problem of too slow transmission rate caused by accessing the off-chip memory or placing a relatively large RAM on the chip causing a large chip area is solved, and the speed and efficiency of the neural network computing of the data flow architecture are improved while ensuring the cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a structural diagram of a neural network computing system based on a data flow architecture provided in one embodiment of the present application;

[0033] Figure 2 It is a structural diagram of another neural network computing system based on a data flow architecture provided in one embodiment of the present application;

[0034] Figure 3 It is a flowchart of a neural network calculation method based on a data flow architecture provided by an embodiment of the present application;

[0035] Figure 4 It is a flowchart of another neural network calculation method based on data flow architecture provided in one embodiment of the present application;

[0036] Figure 5 It is a flowchart of another neural network calculation method based on data flow architecture provided in one embodiment of the present application. DETAILED DESCRIPTION

[0037] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0038] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0039] In addition, the terms "first", "second", etc. can be used in this article to describe various directions, actions, steps or elements, but these directions, actions, steps or elements are not limited by these terms. These terms are only used to distinguish a first direction, action, step or element from another direction, action, step or element. For example, without departing from the scope of the present application, the first data can be referred to as the second data, and similarly, the second data can be referred to as the first data. The first data and the second data are both data, but they are not the same data. The terms "first", "second", etc. cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first" and "second" can expressly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0040] Example

[0041] Figure 1 It is a schematic diagram of a neural network computing system based on a data flow architecture provided in Example 1 of the present application. This embodiment can be applied to scenarios in which the data of the neural network is calculated through this structure.

[0042] The neural network computing system based on the data flow architecture provided in the embodiment of the present application includes an off-chip memory 110 and a neural network acceleration module 120. Among them, the neural network acceleration module 120 includes a conversion unit 121, an on-chip memory 122 and a computing unit 123. In this embodiment, the neural network acceleration module 120 includes a conversion unit 121, an on-chip memory 122 and a computing unit 123 integrated in the same chip.

[0043] The off-chip memory 110 is used to store data.

[0044] The conversion unit 121 is connected between the on-chip memory 122 and the off-chip memory 110, and is used to implement conversion between the first communication bus 130 of the off-chip memory 110 and the second communication bus 140 of the on-chip memory 122, so as to sequentially transmit data through the first communication bus 130, the conversion unit 121, and the second communication bus 140 and store the data in the on-chip memory 122. In one embodiment, the conversion unit 121 receives the data from the off-chip memory 110 through the first communication bus 130, and stores the data in the on-chip memory 122 through the second communication bus 140.

[0045] The computing unit 123 is directly connected to the on-chip memory 122 via the second communication bus 140 , and is configured to perform computing based on data received from the on-chip memory 122 .

[0046] Among them, the main function of the off-chip memory 110 is to store various data, and to complete data access at high speed and automatically during operation on the computer or chip. The off-chip memory 110 is a device with a "memory" function, which uses a physical device with two stable states to store information. The storage capacity of the off-chip memory 110 should be large to meet the needs of neural network data calculation. Compared with simply using the on-chip memory 122, the chip cost can be greatly reduced. Exemplarily, the off-chip memory 110 can be a dynamic random access memory (DRAM) or a double data rate (DDR) synchronous dynamic random access memory. For example, the off-chip memory 110 is a DDR memory to meet higher data transmission efficiency.

[0047] The on-chip memory 122 is in the neural network acceleration module 120, mainly to improve the efficiency of chip data transmission, and is the on-chip storage core of the data flow architecture. The size of the on-chip memory 122 needs to be determined considering the comprehensive cost, and the size can be set according to the configuration of the commonly used network. It does not need to be too large, as long as it can support the performance bottleneck of the neural network. In one embodiment, the size of the on-chip memory 122 needs to be larger than the amount of data that needs to be calculated in the current layer of the neural network. Compared with simply using the off-chip memory 110, the on-chip memory 122 has a large bandwidth and can greatly increase the speed and efficiency of data reading. The on-chip memory 122 can be a high-speed cache memory Cache, or it can be a high-speed cache memory Buffer. There is no restriction on the type of the on-chip memory 122 here. Exemplarily, the on-chip memory 122 is a high-speed cache memory Buffer.

[0048] Since the communication bus protocols of the off-chip memory 110 and the on-chip memory 122 are inconsistent, data cannot be directly transmitted between the off-chip memory 110 and the on-chip memory 122. It is necessary to add a conversion unit 121 between the off-chip memory 110 and the on-chip memory 122. Through bus conversion, data can be transmitted between memories with different bus protocols. Because the amount of data in each layer of the neural network is different, by adding a conversion unit 121 and combining the two types of storage, performance and cost can be balanced, greatly improving product competitiveness.

[0049] The computing unit 123 represents the neural network accelerator core computing engine (ENGINE), which contains various operators required for neural network computing and is the computing core of the data flow architecture. The conversion unit 121 is used to implement the conversion of the bus communication protocol and is the core of the data flow architecture responsible for data exchange.

[0050] In one embodiment, the calculation of the neural network has multiple layers, and the calculation result of the data of the previous layer needs to be output as the data input of the current layer. Exemplarily, the neural network has multiple layers A, B, and C, B is the next layer of A, and C is the next layer of B, which is also the last layer. The data before the neural network starts to calculate is stored in the off-chip memory 110. In the case where the amount of data calculated in the A layer can be stored in the on-chip memory 122, all the data required to be calculated in the A layer can be transmitted to the on-chip memory 122 in the off-chip memory 110 through the first communication bus 130, the conversion unit 121 and the second communication bus 140 before the neural network calculation is performed, so that the calculation unit 123 can pull data from the on-chip memory 122 when starting to calculate the A layer. When the calculation unit 123 completes the calculation of the data of the neural network of the A layer, the calculation result is transmitted to the on-chip memory 122 for simple sorting, and then transmitted to the calculation unit 123, so as to continue the data calculation of the B layer. When the computing unit 123 completes the calculation of the data of the C layer, the final calculation result can be transmitted to the on-chip memory 122, and the on-chip memory 122 transmits the final calculation result to the off-chip memory 110 through the conversion unit 121. After completing the calculation of the neural network once, the final calculation result is stored in the off-chip memory 110. In an alternative embodiment, the final calculation result can also be directly stored from the computing unit 123 to the on-chip memory 122 through the second communication bus 140.

[0051] In one embodiment, after the computing unit 123 executes the current computing task, if a small amount of data in the off-chip memory 110 is needed to perform the next task calculation, during the execution of the current computing task, the data corresponding to the next task can be transferred to the on-chip memory 122 through the conversion unit 121 for pre-sorting, and when the computing unit 123 completes the current task calculation, it can immediately obtain the data corresponding to the next task from the on-chip memory 122 through the faster second bus and start the calculation. This embodiment can avoid the need to obtain the data corresponding to the next task from the off-chip memory 110 through the slower first communication bus 130 after the current computing task is executed. In addition, since the data corresponding to the next task in the on-chip memory 122 has been pre-sorted, the computing unit 123 can directly perform the calculation by pulling the data corresponding to the next task, thereby improving the calculation efficiency of the neural network.

[0052] In one embodiment, sorting data includes sorting different types of data and sorting the same type of data. Taking the sorting of different types of data as an example, the data type may include data data, weight weight and bias, etc. Different computing units 123 have different requirements for the order of transmitted data types. Data data may be transmitted first, then weight bias, or bias may be transmitted first and then weight weight. Therefore, it needs to be determined according to the order of the calculated data types required by the computing unit 123. Pre-sorting the order of the data types required by the computing unit 123 can improve the computing power of the computing unit 123.

[0053] The same type of data also needs to be sorted. Taking the sorting of data as an example, pictures are also a type of data. Suppose the size of a picture is 10*8 pixels, including three colors, red, blue and yellow. The total data volume of this picture is 10*8*3. You can sort by each color, for example, sort the pixels by red, then by yellow, and finally by blue. You can also sort the red, blue and yellow colors of the first pixel, and then sort the three colors of the second pixel. Sorting of the same type of data needs to be determined according to the needs of the computing unit 123 and the data storage type.

[0054] The technical solution of the embodiment of the present application arranges a conversion unit 121 between the on-chip memory 122 and the off-chip memory 110, so that data can be transmitted between the on-chip memory 122 and the off-chip memory 110. During the calculation process, data conversion through the on-chip memory 122 can reduce the time of data transmission and thus improve the calculation efficiency. After the calculation is completed, the final calculation result is output to the off-chip memory 110, avoiding the situation where the off-chip DDR memory is accessed through the AXI bus, the transmission rate is limited by the bus rate and bandwidth, or the chip area is large by placing a relatively large RAM on the chip, thereby improving the speed and efficiency of neural network calculation.

[0055] Based on the above embodiment, this embodiment refines some structures. This embodiment is applicable to the scenario of calculating the data of the neural network through this structure, as follows:

[0056] like Figure 2 As shown, the computing unit 123 is directly connected to the off-chip memory 110 via the first communication bus 130 for performing computing based on data.

[0057] Among them, the computing unit 123 is directly connected to the off-chip memory 110, and data can be directly pulled from the off-chip memory 110 for calculation. The computing unit 123 can pull data from the on-chip memory 122 for neural network calculation, or directly pull data from the off-chip memory 110, which improves the flexibility of data acquisition. The path through which the computing unit 123 directly pulls data from the off-chip memory 110 is recorded as path A, and the path through which the computing unit 123 pulls data from the on-chip memory 122 is recorded as path B. The data transmission time of path A is recorded as the first transmission time, and the data transmission time of path B is recorded as the second transmission time, and the second transmission time is less than the first transmission time. Similarly, the final calculation result of the neural network can also be directly transmitted to the off-chip memory 110 through the first communication bus 130, which also improves the efficiency of data transmission. The computing unit 123 includes an operator 1232 required for neural network calculation. Operator 1232 refers to the smallest computing unit that can be executed by the CPU / GPU.

[0058] In one embodiment, the computing unit 123 includes a data flow direction selection subunit 1231, which is used to control the flow direction of data and is the data flow gating control core of the data flow architecture. Exemplarily, the data flow direction selection subunit 1231 can control the computing unit 123 to receive data stored in the off-chip memory 110 through the first communication bus 130, and can also control the computing unit 123 to receive data from the on-chip memory 122 through the second communication bus 140. Similarly, the data flow direction selection subunit 1231 can also control whether the final data result is stored in the on-chip memory 122 or in the off-chip memory 110.

[0059] In one embodiment, when all the data required to be calculated in the current layer of the neural network can be stored in the on-chip memory 122, and the calculation time of the calculation unit 123 of the current layer of the neural network is greater than the first transmission time, the calculation unit 123 can receive the data stored in the off-chip memory 110 through the first communication bus 130, and can also receive the data of the on-chip memory 122 through the second communication bus 140, that is, the data required to be calculated in the current layer of the neural network can be obtained through path A or path B. At this time, when the data in the off-chip memory 110 is in order, the calculation unit 123 can read the data from the off-chip memory 110. When the data is not in order, the data can be sent to the on-chip memory 122 through the conversion unit 121 for simple processing and then sent to the calculation unit 123 for calculation, thereby reducing the preparation time before the calculation unit 123 calculates.

[0060] When all the data to be calculated in the current layer of the neural network can be stored in the on-chip memory 122, and the calculation time of the calculation unit 123 of the current layer of the neural network is less than or equal to the first transmission time, the calculation unit 123 receives data through the on-chip memory 122, that is, the data to be calculated in the current layer of the neural network is obtained through path B. In one embodiment, when the calculation unit 123 calculates the data of the previous layer, the output result is sent to the on-chip memory 122 for sorting and then sent to the calculation unit 123. When the calculation result of the current layer is used in a subsequent layer, the data is first saved in the on-chip memory 122 and waits for a subsequent layer to calculate.

[0061] When the capacity of the on-chip memory can store the calculation results of the previous layer, the calculation results of the previous layer are stored in the on-chip memory. When the capacity of the on-chip memory cannot store the calculation results of the previous layer, the calculation results of the previous layer are stored in the off-chip memory.

[0062] When the data that need to be calculated in the current layer of the neural network cannot be stored in the on-chip memory 122, and the calculation time of the calculation unit 123 of the current layer of the neural network is greater than the first transmission time, the calculation unit 123 receives the data stored in the off-chip memory 110 through the first communication bus 130, that is, the data that need to be calculated in the current layer of the neural network is obtained through path A.

[0063] When the current layer is the last layer, the final calculation result of the current layer can be directly output to the off-chip memory 110 without being transmitted to the off-chip memory 110 through the on-chip memory 122, thereby improving the efficiency of neural network calculations.

[0064] In one embodiment, the first communication bus 130 may be an AXI bus, and the conversion unit 121 and the off-chip memory 110 interact with the data via the AXI bus protocol, or the computing unit 123 and the off-chip memory 110 interact with the data via the AXI bus protocol.

[0065] The second communication bus 140 may be a Sif bus, and the conversion unit 121 and the on-chip memory 122 interact with the data through the Sif bus protocol, and the on-chip memory 122 and the computing unit 123 interact with the data through the Sif bus protocol. Sif (Streaming Interface) is a simple, practical, and compatible bus protocol. Based on the handshake protocol of stream control, there is no requirement for the data format, data transmission can be full-duplex, fully adapted to the AXI bus, and can efficiently realize data transmission between the source and destination ends.

[0066] Because neural network calculations require switching back and forth between multiple data types, the Sif bus protocol with low resource overhead is used and multiple channels are formed to meet multiple data transmission requirements. Therefore, the conversion unit 121 is required to convert the communication bus protocol. If the chip efficiency needs to be maximized and the two memories are fully utilized, the data flow selection subunit 1231 needs to flexibly switch the storage mode. The data calculated by the calculation unit 123 can be received through the AXI bus of the first communication bus 130, and the data stored in the off-chip memory 110 can also be received through the Sif bus of the second communication bus 140, and the data cached in the on-chip memory 122 can be received.

[0067] Among them, there is at least one channel of the Sif bus. In one embodiment, the number of Sif channels is determined by the type of data and the amount of data transmitted. A Sif channel can only transmit one type of data, but one type of data can be transmitted through multiple Sif channels. Exemplarily, there are three types of data, A, B, and C, so there are at least three Sif channels. However, when the amount of data of type A is relatively large relative to the amount of data of types B and C, the data of type A can be transmitted simultaneously through two Sif channels, thereby reducing the transmission time of the data and preventing the computing efficiency of the neural network from being affected by too slow data transmission.

[0068] The technical solution of the embodiment of the present application arranges a conversion unit between the on-chip memory and the off-chip memory, so that data can be transmitted between the on-chip memory and the off-chip memory. During the calculation process, data conversion through the on-chip memory can reduce the time of data transmission and thus improve the calculation efficiency. After the calculation is completed, the final calculation result is output to the off-chip memory, avoiding the situation where the off-chip DDR memory is accessed through the AXI bus, the transmission rate is limited by the bus rate and bandwidth, or the chip area is large by placing a relatively large RAM on the chip, thereby improving the speed and efficiency of neural network calculation.

[0069] Figure 3 The figure is a flow chart of a neural network computing method based on a data flow architecture provided in one embodiment of the present application. This embodiment is applicable to scenarios where calculations are performed on neural network data. The method can be executed by the above-mentioned neural network computing system based on the data flow architecture, including steps S310 to S330.

[0070] In step S310, data that needs to be calculated in the current layer of the neural network is obtained.

[0071] Among them, the current layer refers to a layer that currently needs to perform data calculation. In one embodiment, the neural network has multiple layers, and the calculation results of the previous layer need to be output to the next layer after the data calculation is completed, as the input data that needs to be calculated in the next layer. Exemplarily, the neural network includes four layers A, B, C, and D. When the data calculation of layer A is completed, the calculation results of layer A are output to layer B. Layer B has not yet calculated and is ready to start calculating, then layer B is the current layer.

[0072] In step S320, the calculation time and the transmission time are estimated according to the data.

[0073] The calculation time refers to the time required for the current layer to calculate the data. In one embodiment, since the calculation configuration of each layer of the neural network acceleration module is fixed, the calculation time can be estimated based on the amount of data and the calculation configuration of the current layer. The transmission time refers to the time required to transmit the data to the calculation unit in the neural network acceleration module.

[0074] In step S330, a transmission path of the data to be calculated is determined according to the calculation time and the transmission time.

[0075] The transmission path refers to the way in which data is transmitted to the computing unit. In this embodiment, there are two transmission paths, path A and path B. For example, the computing unit can transmit data to the on-chip memory (path B) or to the off-chip memory (path A).

[0076] refer to Figure 4, step S330 includes step S331 to step S333.

[0077] In step S331, it is determined whether the calculation time is greater than the first transmission time.

[0078] Among them, since the calculation time and the transmission time can be estimated, the size of the calculation time and the transmission time can be determined. In one embodiment, since there are two transmission paths, the data transmission time of the two paths needs to be calculated respectively. Exemplarily, there are two transmission paths A and B, then the transmission time of path A and the transmission time of path B need to be calculated, the transmission time of path A is recorded as the first transmission time, and the transmission time of path B is recorded as the second transmission time, and the second transmission time is less than the first transmission time.

[0079] In step S332, when the calculation time is greater than the first transmission time, the calculation unit is controlled to interact with the off-chip memory for data or the calculation unit is controlled to interact with the on-chip memory for data.

[0080] Among them, all the data that need to be calculated in the current layer of the neural network can be stored in the on-chip memory. In one embodiment, path A can be used to control the data interaction between the computing unit and the off-chip memory. Path B can be used to control the data interaction between the computing unit and the on-chip memory. Among them, interaction refers to sending and receiving data. When the calculation time is greater than the first transmission time, data can be transmitted along any path. Exemplarily, the path with the shortest transmission time is selected for data transmission. For example, the first data is first transferred from the off-chip memory to the on-chip memory, thereby realizing data interaction between the on-chip memory and the computing unit.

[0081] In step S333, when the calculation time is less than the first transmission time, the calculation unit interacts with the on-chip memory for the data.

[0082] refer to Figure 5 , step S332 includes step S3321 to step S3323.

[0083] In step S3321, determine whether the data is in order.

[0084] The data refers to data stored in an off-chip memory. The order refers to being arranged according to certain rules. In this embodiment, the order means that the data to be calculated at the current layer is classified according to the type of data, that is, the data of the same type are arranged together as the order.

[0085] In step S3322, when the data is arranged in order, the computing unit is controlled to interact with the off-chip memory for the data.

[0086] In step S3323, when the data is not sorted, the computing unit is controlled to interact with the on-chip memory for the data.

[0087] In one embodiment, the off-chip memory sends the data to the on-chip memory for simple sorting, and then transmits the data to the computing unit.

[0088] In an alternative embodiment, it can also be determined whether the calculation result of the current layer needs to be used by a subsequent layer. In the case where the calculation result of the current layer needs to be used by a subsequent layer, the data of the current layer can be temporarily stored in the on-chip memory and sorted until it is used by a certain layer. Exemplarily, the neural network is sequentially A layer, B layer, and C layer. When the data of the A layer is needed for calculation in the C layer, the data of the A layer is temporarily stored in the on-chip memory, and the on-chip memory transfers the data of the A layer to the computing unit for calculation while waiting for the C layer to calculate. In the case where the calculation result of the current layer does not need to be used by a subsequent layer, the result is directly output to the off-chip memory to free up the storage space of the on-chip memory.

[0089] In one embodiment, sorting data includes sorting different types of data and sorting the same type of data. Taking the sorting of different types of data as an example, the data type may include data, weight, bias, etc. Different computing units 123 have different requirements for the order of transmitted data types. Data may be transmitted first, then weight bias, or bias may be transmitted first and then weight weight. Therefore, it needs to be determined according to the order of the data types required by the computing unit 123. Pre-sorting the order of data types required by the computing unit can improve the computing power of the computing unit.

[0090] The same type of data also needs to be sorted. Taking the sorting of data as an example, pictures are also a type of data. Generally, the first layer of a neural network only calculates three colors, red, blue and yellow, when calculating pictures. Suppose the size of a picture is 10*8 pixels, including three colors, red, blue and yellow. Then the total data volume of this picture is 10*8*3. You can sort by each color, for example, sort the pixels by red, then by yellow, and finally by blue. You can also sort the red, blue and yellow colors of the first pixel, and then sort the three colors of the second pixel. Sorting of the same type of data needs to be determined according to the needs of the computing unit and the type of data storage.

[0091] The technical solution of the embodiment of the present application obtains the amount of data that needs to be calculated in the current layer of the neural network, estimates the transmission time and calculation time based on the data amount, determines the transmission path based on the calculation time and transmission time, and thus selects an optimal path to meet the efficiency of the neural network calculation, avoids the situation where the calculation efficiency is too low or the chip area is too large, and improves the speed and efficiency of the neural network calculation.

[0092] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A neural network computing system based on data flow architecture, characterized in that: include: An off-chip memory and a neural network acceleration module, wherein the neural network acceleration module includes a conversion unit, an on-chip memory and a computing unit; wherein the conversion unit is used to implement the conversion of the bus communication protocol and is the core of the data flow architecture responsible for data exchange; the size of the on-chip memory is set based on the performance bottleneck of the neural network; The off-chip memory is used to store data; The conversion unit is connected between the on-chip memory and the off-chip memory, and is used to realize the conversion between the first communication bus of the off-chip memory and the second communication bus of the on-chip memory, so as to store the data in the on-chip memory; wherein the first communication bus is an AXI bus, and the second communication bus is a Sif bus; The computing unit is directly connected to the on-chip memory via a second communication bus, and is used for performing computing based on the data received from the on-chip memory; The computing unit is directly connected to the off-chip memory via a first communication bus, and is used for performing computing based on the data directly received from the off-chip memory; The computing unit comprises: The data flow direction selection subunit is used to control the flow direction of the data.

2. The neural network computing system according to claim 1, characterized in that: The conversion unit and the off-chip memory interact with each other on the data through an AXI bus protocol, or the calculation unit and the off-chip memory interact with each other on the data through an AXI bus protocol.

3. The neural network computing system according to claim 1, characterized in that The conversion unit and the on-chip memory interact with each other on the data through the Sif bus protocol, and the on-chip memory and the computing unit interact with each other on the data through the Sif bus protocol.

4. The neural network computing system according to claim 3, characterized in that: The Sif bus has at least one channel.

5. The neural network computing system according to claim 1, wherein: The calculation unit includes an operator.

6. A neural network computing method based on data flow architecture, characterized in that: Applied to the neural network computing system as claimed in any one of claims 1 to 5, the neural network is multi-layered, and the method comprises: Get the data that needs to be calculated for the current layer of the neural network; estimating the calculation time and transmission time of the current layer based on the data; A transmission path of the data to be calculated is determined according to the calculation time and the transmission time.

7. The method according to claim 6, characterized in that The transmission time includes a first transmission time and a second transmission time, the first transmission time is the data transmission time between the computing unit and the off-chip memory, the second transmission time is the data transmission time between the computing unit and the on-chip memory, and the first transmission time is greater than the second transmission time; The determining, according to the calculation time and the transmission time, a transmission path of the data to be calculated includes: Determining whether the calculation time is greater than the first transmission time; When the computing time is greater than the first transmission time, controlling the computing unit to interact with an off-chip memory for data, or controlling the computing unit to interact with an on-chip memory for data; When the calculation time is less than or equal to the first transmission time, the calculation unit interacts with the on-chip memory for data.

8. The method according to claim 7, characterized in that The controlling the computing unit to interact with the off-chip memory for data, or the controlling the computing unit to interact with the on-chip memory for data, comprises: determining whether the data is in order; When the data is arranged in order, controlling the computing unit to interact with the off-chip memory for the data; In the case that the data are not arranged in order, the computing unit is controlled to interact with the on-chip memory for the data.

Citation Information

Patent Citations

  • Caching data writing system and method, caching data reading system and method

    CN101246460A

  • Training device and method

    CN110909870A

  • Data processing method and device

    JP4345217B2

  • Artificial neural network computation acceleration apparatus for distributed processing, artificial neural network acceleration system using same, and artificial neural network acceleration method therefor

    WO2020075957A1

  • CN11118161A