A hybrid system based on the FPGA+NPU architecture
Through the hybrid system of FPGA+NPU architecture, the problem of poor flexibility of the existing AI computing platform is solved, efficient algorithm processing and computing performance improvement is achieved, and different algorithm needs are adapted to the needs of different algorithms.
Patent Information
- Application Number
- CN202211219965.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-30
AI Technical Summary
Existing AI computing platforms or systems have poor flexibility when facing emerging algorithms, cannot adapt quickly, and have low computing frequency, resulting in insufficient computing power.
A hybrid system using FPGA+NPU architecture is used to optimize bus utilization and realize algorithm processing of soft core NPU algorithm processing by constructing network layer parameter analysis module, parameter package calculation module, reading module, row data conversion module, algorithm and algorithm selection module and write back module.
It improves the flexibility and computing performance of the AI algorithm platform, reduces the need for chip redesign, and improves processing efficiency and bus utilization.
Smart Images

Figure CN115470164B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a hybrid system based on the FPGA+NPU architecture Background Art
[0002] With the upsurge of artificial intelligence and the wide application of AI algorithms, deep learning has become the focus of current AI research, and is involved in fields such as environmental perception and sensor fusion in the field of autonomous driving
[0003] To achieve high efficiency and reliability while facing parallel operations of massive data, it means that the computing platform or system carrying the AI algorithm needs to provide excellent computing performance. The algorithms implemented by NPU chips are changing with each passing day. Some algorithms are relatively fixed, such as convolution operations, etc. However, when new breakthrough algorithms appear, the existing platforms or systems cannot adapt to the cutting-edge technologies accordingly, and can only redesign the compatible algorithms of the chips. This process is time-consuming and costly, resulting in poor flexibility of the platform or system. In addition, the computing platform or system carrying the AI algorithm is restricted by its own architecture and has a low frequency when facing specific operations, so the computing power far fails to meet the requirements of intelligent algorithms Summary of the Invention
[0004] The present invention provides a hybrid system based on the FPGA+NPU architecture. By constructing a system architecture in which each functional module cooperates with each other, based on the concept of integration of storage and processing, the bus utilization rate is optimized, so that users can add new algorithms to the algorithm module in the system for application, thereby improving the flexibility and computing performance of the AI algorithm platform
[0005] To solve the above technical problems, an embodiment of the present invention provides a hybrid system based on the FPGA+NPU architecture. The PL part of the FPGA in the hybrid system is designed to implement the algorithm of the soft-core NPU. The hybrid system includes
[0006] A network layer parameter parsing module, configured to configure the configuration parameters of the current NPU network layer register, where the configuration parameters at least include the source address, total length, algorithm type, and write-back address of the stored feature map FM
[0007] A parameter packet calculation module, configured to calculate several groups of parameter packets containing the start address and length of a single AXI4-MM packet corresponding to each specific algorithm according to the source address of the feature map FM and the total length
[0008] A parameter packet selection module, configured to select and output a group from each group of the parameter packets according to the enable signal of the corresponding algorithm and the number of the current feature map FM
[0009] A read module, which is used to respond to a DMA read request and read out the feature map FM data under the guidance of the output parameter packet;
[0010] A row data conversion module, which is used to convert the feature map FM data into row data;
[0011] An algorithm and algorithm selection module, which includes several groups of addable algorithm types and is used to select a group of the algorithm types according to the enable signal of the corresponding algorithm, process the row data and output it to obtain the data to be written back;
[0012] A write-back module, which is used to respond to a DMA write request, process the data to be written back, and write it into the memory under the guidance of the write-back address until the processing of the network layer of the algorithm for implementing the soft-core NPU ends.
[0013] As one of the preferred solutions, the network layer parameter parsing module takes effect through configuration by the CPU via the AXI4-Lite protocol.
[0014] As one of the preferred solutions, the parameter packet calculation module includes a first parameter packet calculation module, a second parameter packet calculation unit, and a third parameter packet calculation unit;
[0015] The first parameter packet calculation module includes a Route parameter selection unit and a first parameter packet calculation unit. The Route parameter selection unit is used to select and output under the configuration parameters corresponding to the Route algorithm, and send the output result to the first parameter packet calculation unit, so that the first parameter packet calculation unit calculates a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Route algorithm;
[0016] The second parameter packet calculation unit is used to calculate a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Shortcut algorithm;
[0017] The third parameter packet calculation unit is used to calculate a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Activate algorithm according to the Activate table in the memory.
[0018] As one of the preferred solutions, the parameter packet calculation module is configured as:
[0019] Taking the merged start address as the start address of the single AXI4-MM packet and the merged length as the length of the single AXI4-MM packet.
[0020] As one of the preferred solutions, the configuration parameters further include a layer start signal.
[0021] As one of the preferred solutions, the row data conversion module includes a read request signal generation unit and a returned data row conversion unit;
[0022] The read request signal generation unit is configured to generate a read request signal according to the layer start signal, the AXI4-MM bus initiating a DMA read request but not yet returning data, and the change in the number of outstanding addresses under the DMA read request;
[0023] The returned data row conversion unit is configured to shift the data returned by the feature map FM data according to the read request signal to form row data, where the row data includes the complete W×H×C feature map FM data, where W is the number of Bytes in a row, H is the number of rows of one layer of FM, and C is the number of layers of FM.
[0024] As one of the preferred solutions, when writing the data to be written back into the data buffer at the write-back address, the write-back module is further configured to remove the inter-line bubbles in the data to be written back to make it have the maximum burst length.
[0025] As one of the preferred solutions, the write-back module is configured as follows:
[0026] Determine the previous data of the current data in the data to be written back;
[0027] Concatenate the current data and the previous data, and write the concatenated valid data into the data buffer;
[0028] When the number of valid data in the data buffer reaches a preset value, read out the data, form an AXI4-MM write packet, and initiate a DMA write operation.
[0029] As one of the preferred solutions, the write-back module is further configured to send an interrupt request signal after the response of the last DMA write request, so that the CPU sends a clear interrupt signal through AXI4-Lite to clear the interrupt request signal.
[0030] Compared with the prior art, the beneficial effects of the embodiments of the present invention are at least one of the following:
[0031] (1) In the traditional von Neumann architecture, storage and processing are separated. The present invention breaks through the limitations of the classical von Neumann architecture, reads the feature map FM from the memory, then processes it through the algorithm of the soft-core NPU, and writes it back to the memory, thereby improving the processing efficiency of the artificial intelligence chip;
[0032] (2) Construct a system architecture in which the network layer parameter parsing module, parameter packet calculation module, parameter packet selection module, reading module, row data conversion module, algorithm and algorithm selection module, and write-back module cooperate with each other, improving the entire deep learning algorithm process. Among them, various algorithm types can be added to the algorithm and algorithm selection module, the algorithm interface is simple, and the flexibility is relatively high. Only the most cutting-edge target algorithm needs to be added correspondingly, and the AI algorithm platform can be effectively applied without re-designing the chip.
[0033] (3) Fully optimize the read and fetch processes of the feature map FM, adopt multi-stage pipelining and parallel technologies, with short latency. At the same time, the read and write of AXI4-MM use the maximum burst length as much as possible, improving the bus utilization rate and thus enhancing the computing performance of the AI algorithm platform. Description of the Drawings
[0034] Figure 1 is a brief framework of the hybrid system based on the FPGA+NPU architecture in one embodiment of the present invention;
[0035] Figure 2 is a schematic structural diagram of the hybrid system based on the FPGA+NPU architecture in one embodiment of the present invention;
[0036] Figure 3 is a schematic flowchart of calculating the start address and length of a single AXI4-MM packet in one embodiment of the present invention;
[0037] Figure 4 is a schematic flowchart of selecting a group of parameter packets for output in one embodiment of the present invention;
[0038] Figure 5 is a schematic flowchart of the generation process of the DMA read request in one embodiment of the present invention;
[0039] Figure 6 is a timing diagram of the FM conversion to row data output in one embodiment of the present invention;
[0040] Reference Numerals:
[0041] Among them, 1, network layer parameter parsing module; 211, Route parameter selection unit; 212, first parameter packet calculation unit; 22, second parameter packet calculation unit; 23, third parameter packet calculation unit; 3, parameter packet selection module; 4, reading module; 51, first row data conversion module; 52, second row data conversion module; 61, Activate algorithm module; 62, Pooling algorithm module; 63, Upsample algorithm module; 64, Route algorithm module; 65, Shortcut algorithm module; 66, algorithm selection module; 7, write-back module. Detailed implementation mode
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The purpose of providing these embodiments is to make the public content of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0043] In the description of this application, the terms "first", "second", "third", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0044] In the description of this application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are only for the purpose of illustration, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0045] In the description of this application, it should be noted that unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which this belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0046] Explanation of related terms:
[0047] FPGA: Field Programmable Gate Array.
[0048] NPU: Embedded Neural Network Processor.
[0049] Soft-NPU: Soft-core NPU, which refers to the NPU algorithm implemented in the PL of the FPGA here.
[0050] DMA: Direct Memory Access, which allows hardware devices with different speeds to communicate without relying on a large number of interrupt loads of the CPU.
[0051] AXI: An on-chip bus protocol proposed by ARM for high performance, high bandwidth, and low latency.
[0052] HP: High-Performance Address Mapping Interface.
[0053] GP: General-Performance Address Mapping Interface.
[0054] IRQ: Interrupt Request.
[0055] FM: Abbreviation of the English name Feature Map, Feature Map.
[0056] PS: Programmable System.
[0057] PL: Programmable Logic.
[0058] FIFO: First In First Out Memory.
[0059] RAM: Random Access Memory.
[0060] In the field of artificial intelligence, if one wants to imitate human neurons with a circuit, each neuron has to be abstracted as an activation function, and the input of this function is jointly determined by the outputs of the neurons connected to it and the synapses connecting the neurons. To express specific knowledge, the user usually needs to (through certain specific algorithms) adjust the values of the synapses in the artificial neural network, the topological structure of the network, etc. This process is called "learning". After learning, the artificial neural network can solve specific problems with the acquired knowledge.
[0061] In the von Neumann architecture, storage and processing are separated and are implemented by a memory and an arithmetic unit respectively, and there are huge differences between the two. When using existing classical computers based on the von Neumann architecture (such as X86 processors and NVIDIA GPUs) to run neural network applications, it is inevitably restricted by the separated structure of storage and processing, thus affecting efficiency.
[0062] In the present invention, considering that the basic operations of deep learning are the processing of neurons and synapses, while traditional processor instruction sets (including x86 and ARM, etc.) are developed for general-purpose computing, and their basic operations are arithmetic operations (addition, subtraction, multiplication, and division) and logical operations (AND, OR, NOT), often hundreds or even thousands of instructions are required to complete the processing of a single neuron, resulting in low processing efficiency for deep learning. Therefore, the present invention adopts a system structure combining an NPU and an FPGA chip, which can not only meet the requirements of artificial intelligence algorithms, high bandwidth, and powerful computing capabilities, but also meet the requirements of different algorithm scenarios.
[0063] The NPU (Neural network Processing Unit) is a dedicated chip for neural networks with small size, low power consumption, high computing performance, and high computing efficiency. Compared with CPUs and GPUs, the NPU improves the operating efficiency by highlighting the integration of the storage and calculation of neural network algorithm weights and is specifically used to implement AI algorithm applications. Typical representatives of the NPU include Cambricon chips in China and IBM's TrueNorth. Taking Cambricon in China as an example, in March 2016, the research group of Chen Yunji and Chen Tianshi from the Institute of Computing Technology, Chinese Academy of Sciences proposed the world's first deep learning processor instruction set, DianNaoYu. The DianNaoYu instruction directly faces the processing of a large number of neurons and synapses. One instruction can complete the processing of a group of neurons and provides a series of special supports for the transmission of neuron and synapse data on the chip.
[0064] The FPGA (Field Programmable Gate Array) is a hardware programmable chip with powerful parallel computing capabilities. It can design and optimize hardware circuits according to different algorithms. In the case of the same computing power, its power consumption is approximately one-tenth of that of a GPU. Its main feature is that the hardware circuit can be reset and it can support the acceleration of different algorithm models. With the continuous development of modern deep learning algorithms towards more difficult and complex directions, the FPGA begins to demonstrate its unique acceleration capabilities. The FPGA is composed of programmable logic units, and almost all hardware resources can be dynamically programmed to achieve the desired logical functions. This ability makes the FPGA highly adaptable and widely applicable. If optimized properly, the FPGA can have characteristics such as high performance and low power consumption, and has incomparable advantages compared with other hardware with firmware capabilities, such as the NPU.
[0065] It should be noted that the inventor found in the research that due to its own architecture limitations, the FPGA cannot run at a very high frequency. Under the same process, the main frequency of the FPGA can generally only reach less than half of that of the hardcore NPU. That is to say, the FPGA is suitable for processing large quantities of parallel but simple operations. Therefore, in the hybrid system of the present invention, please refer to Figure 1 , Figure 1It shows a brief framework of the hybrid system based on the FPGA+NPU architecture provided by the present invention, which is similar to a system-on-chip (SOC). The FPGA includes two parts: the PS and the PL. In the present invention, the Hard-NPU in the PS is designed to complete algorithms with high performance requirements such as convolution, which are not the core improvement points of the present invention and will not be introduced in detail here; the Soft-NPU (soft-core NPU) in the PL part completes algorithms that are simple but variable in operation, such as the Shortcut algorithm, the Activate algorithm, etc., which will be described in detail below.
[0066] Specifically, an embodiment of the present invention provides a hybrid system based on the FPGA+NPU architecture. Specifically, please refer to Figure 2 , Figure 2 It shows a schematic structural diagram of the hybrid system based on the FPGA+NPU architecture in one embodiment of the present invention.
[0067] In order to implement the algorithms that are simple but variable in operation of the Soft-NPU (soft-core NPU), the hybrid system in this embodiment designs a variety of functional modules, and each functional module cooperates with each other to jointly improve the flexibility and computing performance of the AI algorithm platform.
[0068] Among them, the description of each functional module is as follows:
[0069] The network layer parameter parsing module 1 is used to configure the configuration parameters of the current NPU network layer register, where the configuration parameters at least include the source address, total length, algorithm type, and write-back address of the stored feature map FM;
[0070] The parameter packet calculation module is used to calculate several groups of parameter packets containing the start address and length of a single AXI4-MM packet corresponding to each specific algorithm according to the source address and the total length of the feature map FM;
[0071] The parameter packet selection module 3 is used to select a group of outputs from each group of the parameter packets according to the enable signal of the corresponding algorithm and the number of the current feature map FM;
[0072] The reading module 4 is used to respond to the DMA read request and read out the feature map FM data under the guidance of the output parameter packet;
[0073] The row data conversion module is used to convert the feature map FM data into row data;
[0074] The algorithm and algorithm selection module includes several groups of addable algorithm types and is used to select a group of the algorithm types according to the enable signal of the corresponding algorithm, process the row data and output it to obtain the data to be written back;
[0075] The write-back module 7 is used to respond to the DMA write request, process the data to be written back, and write it into the memory under the guidance of the write-back address until the processing of the network layer of the algorithm for implementing the soft-core NPU ends.
[0076] In short, an NPU network needs many layers to be implemented. In this embodiment, the process of implementing a certain network layer in Soft-NPU mainly includes: initially, parsing the configuration parameters sent by the CPU through AXI4-Lite; then reading the FM from the DDR through AXI4-MM and executing the corresponding algorithm, and while executing, writing the new FM back to the DDR through AXI4-MM; finally, after the write-back is completed, applying for an interruption to the CPU through the Int bus.
[0077] The principles of the above-mentioned functional modules will be described in detail below.
[0078] Network layer parameter parsing module 1: The network layer parameter parsing module 1 mainly uses the AXI protocol provided by ARM. Here, the AXI4-Lite bus is used to implement the configuration of a certain network layer register. The configuration parameters include the source address and write-back address for storing the FM, the total length (length * width * number of layers, unit: byte), the size (length * width * number of layers), the layer start signal, the algorithm type, and the corresponding configuration parameters, etc. Specifically, before the start of each layer, the network layer parameter parsing module 1 is configured and takes effect by the CPU through the AXI4-Lite protocol.
[0079] Parameter packet calculation module: The function of the parameter packet calculation module is to output the start address and length of a single AXI4-MM packet as a guide. However, for different algorithms, the actual operation processes are different. Below, three algorithms are used as examples for illustration. Of course, the following three specific algorithms are only for more clearly explaining the function of the parameter packet calculation module. Those skilled in the art can select other algorithm types according to needs. When it is other algorithms, since they are all single FMs, they are unified to the following first parameter packet calculation unit for implementation, which will not be elaborated here.
[0080] Preferably, the parameter packet calculation module includes a first parameter packet calculation module, a second parameter packet calculation unit 22, and a third parameter packet calculation unit 23;
[0081] The first parameter packet calculation module includes a Route parameter selection unit 211 and a first parameter packet calculation unit 212. The Route parameter selection unit 211 is used to select and output under the configuration parameters corresponding to the Route algorithm, and send the output result to the first parameter packet calculation unit 212, so that the first parameter packet calculation unit 212 calculates a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Route algorithm; The Route algorithm may need to read 1 or 2 or 3 or 4 FMs from the DDR. The source addresses, total lengths, and sizes of these FMs need to take effect alternately according to the Route algorithm configuration parameter selection, and one output is selected. The input of the first parameter packet calculation unit 212 is the output of the Route parameter selection unit, and when the Route enable signal is invalid, it is the parameter of the first parameter packet calculation unit 212.
[0082] The second parameter packet calculation unit 22 is used to calculate a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Shortcut algorithm. That is, the second parameter packet calculation unit 22 is only used when the Shortcut algorithm is enabled.
[0083] The third parameter packet calculation unit 23 is used to calculate a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Activate algorithm according to the Activate table in the memory. The input of the third parameter packet calculation unit 23 is the source address and total length of the Activate table (the Activate table stores some non-linear function lookup tables, such as the prelu activation function, etc.), and it is also only used when the Activate algorithm is enabled.
[0084] The inputs of the above first parameter packet calculation module, second parameter packet calculation unit 22, and third parameter packet calculation unit 23 are all FM source addresses and total lengths. The difference lies in the input attributes of the three. After the above input signals enter their respective parameter packet calculation modules or units, the start address and length of a single AXI4-MM packet need to be calculated before a DMA read can be initiated. Specifically, please refer to Figure 3 , Figure 3 which shows a schematic flowchart of calculating the start address and length of a single AXI4-MM packet in one embodiment of the present invention ( Figure 3On the left is length calculation, and on the right is address calculation. The principle of the calculation is as follows: The starting address after merging is used as the starting address of the single AXI4-MM packet, and the length after merging is used as the length of the single AXI4-MM packet. Specifically, the above calculation mainly includes two parts: calculating the packet starting address and the packet length. There is a step in calculating the packet length to estimate the length of the first packet. The lower bits of the base address involved are related to the burst packet length, and the value is the lower (N - 1) bits of the burst length valid bit (assuming N bits); in another process, "output packet lengths are merged together", and the output is the length of the AXI4-MM packet. Similarly, "output packet starting addresses are merged together", and the output is the starting address of the AXI4-MM packet.
[0085] Parameter packet selection module 3: Taking the above three algorithm types as an example, after obtaining the above three groups of parameter packets (the parameter packet of the starting address and length of the single AXI4-MM packet corresponding to the Route algorithm, the parameter packet of the starting address and length of the single AXI4-MM packet corresponding to the Shortcut algorithm, the parameter packet of the starting address and length of the single AXI4-MM packet corresponding to the Activate algorithm), select a group of outputs according to the algorithm enable signal and the current FM number. Specifically, please refer to Figure 4 , Figure 4 which is shown as a schematic flow diagram of selecting a group of parameter packets for output ( Figure 4 still taking the above three algorithm types as an example), Figure 4 "Ready" in it means that the packet starting address and length have been calculated; Figure 4 The prerequisite for "the Activate packet parameter calculation module (in the figure, the Activate packet parameter calculation module refers to the above-mentioned third parameter packet calculation unit 23, the FM packet parameter calculation module refers to the above-mentioned second parameter packet calculation unit 22, and the FM packet parameter calculation module 1 refers to the above-mentioned first parameter packet calculation unit 212) being ready" is that Activate has been enabled; when the Activate algorithm is executed, the lookup table needs to be read from the DDR first, and then the FM is read, so the calculation of the packet parameters is divided into two steps; the Shortcut algorithm needs to read two FMs alternately, so two packet parameter calculation modules are used; the input of the Route algorithm may be multiple FMs, and the output is serial, so the packet parameters of multiple FMs are calculated sequentially; for other algorithms, they are all single FMs, so the parameters are uniformly calculated by the FM packet parameter calculation module 0;
[0086] Reading module 4: The selected parameter packet enters the reading module to initiate a DMA read to read out the corresponding FM or Activate table, and the protocol also refers to the AXI protocol mentioned above. Since the above read is completed in stages, the data only needs to be sent to different private data buses according to the state.
[0087] It should be noted that the continuous data read out by the above DMA is not suitable for some row processing algorithms, such as Pooling, Upsampling, etc. Considering generality at this time, all the data of FM is converted into row-by-row data and then sent to the algorithm module (that is, it needs to go through the row data conversion module).
[0088] Row data conversion module: It should be noted in advance that this embodiment includes a first row data conversion module 51 and a second row data conversion module 52. The enable, data, etc. of the FM packet's first-layer private bus returned by the reading module 4 are sent to the first row data conversion module 51 (due to the Shortcut enable mentioned above, two-way cross-reading of FM generates two-way data, and the second-way data is sent to the second row data conversion module 52), and the activated lookup table data is sent to the subsequent Activate algorithm module 61; then the first row data conversion module 51 converts the incoming private bus into another set of private buses (fmc_frm, fmc_lnvld, fmc_lnlstp, fmc_dt, fmc_dten) that are convenient for subsequent algorithm implementation and sends them to all subsequent algorithm modules. Since the second row data conversion module 52 is only used by the Shortcut algorithm module for the time being, it is only sent to this module.
[0089] Regarding how to convert into row data, specifically, the row data conversion module includes a read request signal generation unit and a returned data row conversion unit. The read request signal generation unit is used to generate a read request signal corresponding to the layer start signal, the AXI4-MM bus initiating a DMA read request but not yet returning data, and the change in the outstanding address quantity under the DMA read request;
[0090] The returned data row conversion unit is used to shift the data returned by the feature map FM data according to the read request signal to form row data, and the row data includes the complete W×H×C feature map FM data, where W is the number of Bytes in a row, H is the number of rows in one layer of FM, and C is the number of layers of FM.
[0091] As can be seen from the above, the operation of converting from FM to rows is completed in two steps: generating a read request signal and converting the data returned by the read request into rows. First, a read request is generated. Specifically, please refer to Figure 5 , Figure 5 which shows the schematic diagram of the generation process of the DMA read request in one embodiment of the present invention. In order to reduce the bubbles in the data flow path and improve the bandwidth utilization rate of AXI4-MM, the read request scheduling method as Figure 5 is adopted. Figure 5The "data in transit" here refers to the situation where an AXI4-MM bus initiates a DMA read request but the data has not been returned yet. Additionally, since the base address of the first AXI4-MM packet is a random value, that is, the data is not valid starting from Byte0, it is necessary to shift the data according to the lower bits of the base address to form consecutive data and then write it into the FIFO. The FIFO write enable is effectively decremented by 1, and the DMA read request is effectively incremented by the burst length. Then, the "data in transit" is "merged", that is, the two are added together. The sum of the "remaining FIFO space" and the "merged data in transit" is the data volume that the FIFO space will reach without initiating a new read request. By comparing the value of "FIFO depth minus burst length" with the value that the FIFO will reach, it can be determined whether the FIFO can hold another packet of the burst length size. If the difference is positive, then check whether Outstanding (indicating the packets in progress, incremented by 1 for each request and decremented by 1 for each return) is less than the preset value. If it is less, it means a new DMA read request can be initiated; otherwise, wait until both of the above conditions are met. The initiation of the next read request follows the same control method as the current read request and will not be introduced here.
[0092] The DMA read requests generated according to the above method have the returned data shifted and written into the FIFO. When the input backpressure signal is invalid, the data is read out row by row, and a total of H*C rows are formed. Here, H is the number of rows in one layer of FM, and C is the number of layers of FM. At this time, it should be noted that since the row width is not necessarily an integer multiple of the AXI4-MM data bit width, the extra data from the previous row needs to be shifted to the next row, and so on to form a complete FM (W*H*C), where W is the number of bytes in one row (in this article, the data type of FM is all int8 or uint8). In addition, to further optimize the latency of subsequent algorithms, such as shortening the row spacing and reducing the time to detect the end of a row, the following timing waveforms are output, as Figure 6 shown, Figure 6 which shows the timing diagram of FM row conversion data output in one embodiment of the present invention, Figure 6 where sys_clk is the module clock signal; fmc_frm is the layer synchronization signal, which encloses all the rows inside the FM; fmc_lnvld is the row synchronization signal, enclosing the data of one row. When the number of remaining bytes after multiple previous rows is greater than the number of bytes required at the fmc_lnlstp position of the current row, there is no gap between rows; fmc_dten is the data enable signal, which is consistent with fmc_lnvld when the FM row conversion data module outputs, and may be inconsistent after subsequent other algorithms, such as when the pooling algorithm step size is not 1; fmc_lnlstp is the signal indicating the last data of the row, which is valid as a single pulse; fmc_dt is the FM data, and the data bit width is the same as that of AXI4-MM.
[0093] Algorithm and algorithm selection module: It includes an algorithm module and an algorithm selection module 66. In practical applications, the two can be combined into the same module or split into separate modules. Figure 2 The split structure in this does not constitute a limitation to the present invention and will not be elaborated here.
[0094] Algorithm module: The algorithms in this embodiment are as Figure 2 shown, including an Activate algorithm module 61, a Pooling algorithm module 62, an Upsample algorithm module 63, a Route algorithm module 64, and a Shortcut algorithm module 65. Taking the Figure 6 above timing diagram as an example, the above timing waveforms enter the Figure 2 algorithm module in parallel. Figure 2 Only the algorithms used in the YOLO object detection model are shown in this, such as the Activate algorithm module 61, the Pooling algorithm module 62, etc. The timing of other algorithm models has not been deeply studied. Of course, since the algorithms are addable, only the most advanced target algorithms need to be correspondingly added, and there is no need to redesign the chip to achieve the effective application of the AI algorithm platform. The implementation of these algorithms uses a multi-stage pipeline technology with short delays. Taking the Activate algorithm module 61 as an example, the data entering this module passes through operations such as look-up table and scale transformation in multiple steps and then outputs the first data of the FM. Subsequent data is continuously output. The overall FM timing is just delayed in multiple steps of output, and the timing has not changed.
[0095] Algorithm selection module 66: The above algorithm outputs are sent to the algorithm selection module 66, and one is selected according to the algorithm enable signal, and it can be output only after a one-beat delay. Algorithms without enable will not work, thus reducing power consumption. The selected algorithm output enters the following write-back module 7.
[0096] Write-back module 7: It should be noted that when writing the data to be written back into the data buffer under the write-back address, the write-back module 7 is also configured to remove the inter-line bubbles in the data to be written back to make it have the maximum burst length.
[0097] Specifically, the write-back module 7 is configured to:
[0098] Determine the previous data of the current data in the data to be written back;
[0099] Concatenate the current data and the previous data, and write the concatenated valid data into the data buffer;
[0100] When the number of valid data in the data buffer reaches the preset value, read out the data, form an AXI4-MM write packet, and initiate a DMA write operation.
[0101] As can be seen from the above, since the number of Bytes per line is not necessarily an integer multiple of the AXI4-MM data width, but to improve the efficiency of DMA writes, the maximum burst length is taken for each packet as much as possible. Therefore, it is considered to squeeze out the inter-line bubbles. This process is divided into two steps:
[0102] 1. Determine the previous data of the current data, that is, fmc_dt_d1;
[0103] 2. Concatenate the current data with its previous data to obtain continuous data.
[0104] Assume that the AXI4-MM data width is 64bit. Select the lower 3 bits of the row width W and the lower 3 bits of the remaining byte count RM after concatenating with the previous row data, and then use them as variables for selective output. If fmc_lnlstp is low, all are complete 64bit data and only need to be pipelined once. If fmc_lnlstp is high, it means it is the last data of a row and may need to be concatenated. Since {W[2:0], RM[2:0]} is a total of 6 bits and 2^6 = 64 cases, the concatenation idea is that "the previous data is concatenated with the remaining Bytes and the valid Bytes at the fmc_lnlstp position and then placed in the high valid bits". The specific concatenation method is shown in the following table:
[0105]
[0106]
[0107] After the fmc_dt_d1 obtained at the fmc_lnlstp position and the fmc_dt_d1 at other positions are merged together, the final fmc_dt_d1 is obtained. The current data fmc_dt and the final fmc_dt_d1 are concatenated according to the RM[2:0] variable to form the final data fifo_di and written into the FIFO. The concatenation method at this time is shown on the right side of the above table. After the concatenation is completed, the valid data is written into the FIFO. When the number of valid data in the FIFO is greater than the packet length, the data is read out, grouped into an AXI4-MM write packet, and a DMA write is initiated. At this time, the method of calculating the AXI4-MM write packet length and the initial address is similar to the process of calculating the AXI4-MM read packet length and the initial address, which will not be introduced here. However, it should be noted that the length of the last packet should be based on Figure 6It is determined by the number of remaining data in the FIFO after the falling edge of fmc_frm. When the amount of data in the FIFO is greater than the maximum burst length, packets are sent according to the maximum burst length. When it is less than or equal to the maximum burst length, the number of data in the FIFO at that time is the burst length. At the same time, Outstanding also needs to be considered during AXI4-MM write requests, and the counting method is similar to that of AXI4-MM read requests. There is a pointer counter in the FM write-back module. When a write request is sent, the counter is incremented by 1. When a write response is returned, the counter is decremented by 1. Only when the value of Outstanding is less than the preset value is a new request allowed to be initiated, otherwise wait.
[0108] Further, after the last DMA write request is sent and responded to, an interrupt request signal is sent. When the CPU clears the interrupt through AXI4-Lite, the interrupt request is cleared to 0, otherwise the interrupt request signal remains valid. Thus, one layer of the network is completed, waiting for the start signal of the next layer of the network, or all layers are completed and ended.
[0109] Compared with the prior art, the beneficial effects of the embodiments of the present invention are at least one of the following:
[0110] (1) In the traditional von Neumann architecture, storage and processing are separated. The present invention breaks through the limitations of the classical von Neumann architecture, reads the feature map FM from the memory, and then writes it back to the memory after being processed by the algorithm of the soft-core NPU, thereby improving the processing efficiency of the artificial intelligence chip;
[0111] (2) A system architecture is constructed with the cooperation of a network layer parameter parsing module, a parameter packet calculation module, a parameter packet selection module, a reading module, a row data conversion module, an algorithm and algorithm selection module, and a write-back module, which improves the entire deep learning algorithm process. Among them, various algorithm types can be added to the algorithm and algorithm selection module, the algorithm interface is simple, and the flexibility is relatively high. Only the most advanced target algorithm needs to be added correspondingly, and the effective application of the AI algorithm platform can be realized without re-designing the chip;
[0112] (3) The read and fetch processes of the feature map FM are fully optimized, using multi-stage pipelining and parallel technologies, with short delays. At the same time, the read and write of AXI4-MM use the maximum burst length as much as possible, improving the utilization rate of the bus, and thus enhancing the computing performance of the AI algorithm platform.
[0113] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. A hybrid system based on the FPGA + NPU architecture, characterized in that, In the hybrid system, the PL part of the FPGA is designed to implement the algorithm of the soft-core NPU. The hybrid system includes: A network layer parameter parsing module, configured to configure the configuration parameters of the current NPU network layer register, where the configuration parameters at least include the source address, total length, algorithm type, and write-back address of the stored feature map FM; A parameter packet calculation module, configured to calculate several groups of parameter packets containing the start address and length of a single AXI4-MM packet corresponding to each specific algorithm according to the source address and the total length of the feature map FM; A parameter packet selection module, configured to select and output one group from each group of the parameter packets according to the enable signal of the corresponding algorithm and the number of the current feature map FM; A reading module, configured to respond to a DMA read request and read the feature map FM data under the guidance of the output parameter packet; A row data conversion module, configured to convert the feature map FM data into row data; An algorithm and algorithm selection module, which includes several groups of addable algorithm types, and is configured to select one group of the algorithm types according to the enable signal of the corresponding algorithm, process and output the row data to obtain data to be written back; A write-back module, configured to respond to a DMA write request, process the data to be written back, and write it into the memory under the guidance of the write-back address until the processing of the network layer implementing the algorithm of the soft-core NPU ends.
2. The hybrid system based on the FPGA + NPU architecture according to claim 1, wherein The network layer parameter parsing module becomes effective through configuration by the CPU via the AXI4-Lite protocol.
3. The hybrid system based on the FPGA+NPU architecture according to claim 1, characterized in that, The parameter packet calculation module includes a first parameter packet calculation module, a second parameter packet calculation unit, and a third parameter packet calculation unit; The first parameter packet calculation module includes a Route parameter selection unit and a first parameter packet calculation unit. The Route parameter selection unit is configured to select and output under the configuration parameters corresponding to the Route algorithm, and send the output result to the first parameter packet calculation unit, so that the first parameter packet calculation unit calculates a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Route algorithm; The second parameter packet calculation unit is configured to calculate a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Shortcut algorithm; The third parameter packet calculation unit is configured to calculate a parameter packet containing the start address and length of a single AXI4-MM packet corresponding to the Activate algorithm according to the Activate table in the memory.
4. The hybrid system based on the FPGA+NPU architecture according to claim 1, characterized in that, The parameter packet calculation module is configured as: Using the merged start address as the start address of the single AXI4-MM packet and the merged length as the length of the single AXI4-MM packet.
5. The hybrid system based on the FPGA + NPU architecture according to claim 1, characterized in that The configuration parameters further include a layer start signal.
6. The hybrid system based on the FPGA+NPU architecture according to claim 5, wherein, The row data conversion module includes a read request signal generation unit and a return data conversion unit; The read request signal generation unit is configured to correspondingly generate a read request signal according to the layer start signal, the AXI4-MM bus initiating a DMA read request but not yet returning data, and the change in the number of outstanding addresses under the DMA read request; The returned data line conversion unit is configured to shift the data returned from the feature map FM data according to the read request signal to form line data, where the line data includes complete W×H×C feature map FM data, where W is the number of bytes in a row, H is the number of rows in one layer of FM, and C is the number of layers of FM.
7. The hybrid system based on the FPGA+NPU architecture according to claim 1, characterized in that When writing the data to be written back into the data buffer at the write-back address, the write-back module is further configured to remove the inter-line bubbles in the data to be written back to have the maximum burst length.
8. The hybrid system based on the FPGA+NPU architecture according to claim 7, characterized in that, The write-back module is configured as follows: Determine the previous data of the current data in the data to be written back; Concatenate the current data with the previous data, and write the concatenated valid data into the data buffer; When the number of valid data in the data buffer reaches a preset value, read out the data, form an AXI4-MM write packet, and initiate a DMA write operation.
9. The hybrid system based on the FPGA+NPU architecture according to claim 1, characterized in that, The write-back module is further configured to issue an interrupt request signal after the response of the last DMA write request, so that the CPU issues a clear interrupt signal through AXI4-Lite to clear the interrupt request signal.
Citation Information
Patent Citations
Method and system for realizing YOLOv2 detection network based on FPGA
CN110175670A
Configurable parallel universal convolutional neural network accelerator based on BNRP
CN110390385A