FPGA (Field Programmable Gate Array) type system on chip, heterogeneous computing acceleration platform and computing acceleration method

By designing an FPGA-like system on-chip that integrates ARM processor core, programmable logic zone and artificial intelligence engine, and deploying deep learning processors and algorithm acceleration cores, the problem of the inability to accelerate the entire process of AI inference algorithms in the existing technology is solved, and efficient computing acceleration and system scalability are achieved.

CN120216451APending Publication Date: 2025-06-27SHANDONG MARS ENERGY ENGINEERING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510251484.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing artificial intelligence edge computing system on chip cannot accelerate the calculation of the entire process of AI inference algorithms, and the ASIC's AI-specific integrated circuit design and chipping are costly, making it unable to compatible with future design needs.

Method used

Design an FPGA-type system on chip, integrate the ARM processor core, programmable logic area and artificial intelligence engine, deploy deep learning processors and algorithm acceleration cores, and realize efficient communication between modules and the construction of data processing pipelines through on-chip networks.

Benefits of technology

The hardware-level acceleration of the entire process of data preprocessing and postprocessing of artificial intelligence inference has been achieved, reducing the computing delay and data interactions, and improving the overall computing speed and system scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216451A_ABST
    Figure CN120216451A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of chip design, and provides a system-on-chip, a heterogeneous computing acceleration platform and a data processing method in order to improve the computing speed problem of an existing artificial intelligence edge computing system-on-chip. Wherein the system on chip comprises an ARM processor core, a programmable logic area and an artificial intelligence engine; a deep learning processor is deployed on the artificial intelligence engine; at least one target deep learning model is loaded in the deep learning processor; the ARM processor core controls a target deep learning model in the deep learning processor to run by setting a driving program; at least one algorithm acceleration kernel is integrated in the programmable logic area; each algorithm acceleration kernel is independently designed according to a set function and then transplanted to the programmable logic area, and a data processing assembly line with independent functions is formed according to a preset requirement. Data pre-processing and post-processing of artificial intelligence reasoning can be deployed on the same chip to achieve hardware-level full-process acceleration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of chip design, and particularly relates to an FPGA-based system-on-chip, a heterogeneous computing acceleration platform, and a computing acceleration method. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] An artificial intelligence edge computing system-on-chip can execute artificial intelligence computing tasks at a location close to the data source (i.e., the edge) without transmitting data to a central server or the cloud for processing. Existing artificial intelligence edge computing systems-on-chip usually integrate a general-purpose processor (such as an ARM CPU), a dedicated computing accelerator (such as a Cuda GPU, an AI application-specific integrated circuit, a tensor computing core, etc.), and an on-chip interconnection system inside. Although they can process and analyze data in real time, such as image recognition, digital signal processing, and sensor data fusion, there are still the following problems:

[0004] (1) It can only perform AI inference operations based on convolution operations and cannot accelerate the entire process of AI inference algorithms; the pre-processing and post-processing of data in the artificial intelligence edge computing system-on-chip, such as video transcoding and compression storage, are all performed by the ARM CPU, which reduces the overall computing speed.

[0005] (2) The design and tape-out costs of AI application-specific integrated circuits and tensor processors as ASICs are very high. Once put into production, they cannot be modified and cannot be compatible with future design requirements; all other hardware architectures cannot complete the full-process hardware acceleration of AI inference on the same chip. Summary of the Invention

[0006] To solve the technical problems existing in the above background art, the present invention provides an FPGA-based system-on-chip, a heterogeneous computing acceleration platform, and a computing acceleration method, which can deploy the pre-processing and post-processing of artificial intelligence inference on the same chip to achieve full-process acceleration at the hardware level and are also applicable to the acceleration calculation of other algorithms in specific application scenarios.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention provides an FPGA-based system-on-chip.

[0009] In one or more embodiments, an FPGA-based system-on-chip includes: an ARM processor core, a programmable logic area, and an artificial intelligence engine;

[0010] A deep learning processor is deployed on the artificial intelligence engine; at least one target deep learning model is loaded in the deep learning processor;

[0011] The ARM processor core controls the operation of the target deep learning model in the deep learning processor through a set driver program;

[0012] At least one algorithm acceleration core is integrated in the programmable logic area; each algorithm acceleration core is independently designed according to a set function and then transplanted into the programmable logic area, and forms an independent data processing pipeline according to preset requirements;

[0013] Any two of the ARM processor core, the programmable logic area and the artificial intelligence engine communicate with each other through the on-chip network; the on-chip network is also connected to the DDR memory.

[0014] The ARM processor core here is the on-chip CPU of the ARM architecture.

[0015] The advantage of the above solution is that in this way, the ARM processor core, the programmable logic area and the artificial intelligence engine are all connected to the DDR memory through the on-chip network, so that the ARM processor core does not participate in data processing and data interaction, and only is responsible for the control of each module and the data processing pipeline; the DDR memory is used to temporarily store the intermediate results of data processing, avoiding jams in the data processing process and ensuring the accuracy of data processing; moreover, the pipeline design can minimize the interaction with the DDR memory to reduce the calculation latency.

[0016] As an implementation, the AXI-stream interface is used between the data processing pipeline cores in the programmable logic area to design the data processing pipeline; the ARM processor core controls the state of the data processing pipeline.

[0017] The advantage of the above solution is that the pipeline is controlled by using the automatic data flow method (without handshake mechanism) or handshake control signals or AXILite bus. By utilizing the simple and lightweight characteristics of the AXI-stream interface, data is transmitted between the cores on the pipeline in the programmable logic area in a way without addresses, which is suitable for high-throughput data stream transmission scenarios and improves the computing speed of the on-chip system.

[0018] As an implementation, at least one programmable logic area core is deployed with an Ethernet controller, and the Ethernet controller is connected to a fiber optic interface.

[0019] The advantage of the above solution is that in this way, the on-chip system can communicate with other external communication devices through the Ethernet controller and the fiber optic interface, expanding the application field of the on-chip system.

[0020] As an implementation manner, a peripheral controller is further deployed in the programmable logic area, and the peripheral controller is connected to peripheral electronic devices; the programmable logic area is further connected to an FMC interface, and the FMC interface is connected to FMC extended peripherals.

[0021] In some other embodiments, an FPGA-based system-on-chip is provided, which includes: an ARM processor core and a programmable logic area;

[0022] A deep learning processor is deployed on the programmable logic area; at least one target deep learning model is loaded in the deep learning processor; at least one algorithm acceleration core is further integrated in the programmable logic area; each of the algorithm acceleration cores is independently designed according to a set function and then transplanted to the programmable logic area;

[0023] The ARM processor core communicates with the programmable logic area through a communication interface; the ARM processor core controls the operation of the target deep learning model in the deep learning processor through a set driver program.

[0024] As an implementation manner, an AXI-stream interface is adopted between the data processing pipeline cores in the programmable logic area to design a data processing pipeline.

[0025] As an implementation manner, at least one of the programmable logic area cores is deployed with an Ethernet controller, the Ethernet controller is connected to an optical fiber interface, and the Ethernet controller uses a DMA controller to perform data interaction with a DDR memory.

[0026] As an implementation manner, a peripheral controller is further deployed in the programmable logic area, and the peripheral controller is connected to peripheral electronic devices; the programmable logic area is further connected to an FMC interface, and the FMC interface is connected to FMC extended peripherals.

[0027] The advantages of the above solution are that the peripheral controller is used to obtain the data transmitted by the peripheral electronic devices or control the data acquisition instructions of the peripheral electronic devices, so as to realize the multi-source acquisition of data processing; the FMC interface is used to connect the FMC extended peripherals, which improves the scalability of the system-on-chip.

[0028] The second aspect of the present invention provides a heterogeneous computing acceleration platform.

[0029] A heterogeneous computing acceleration platform includes: the FPGA-based system-on-chip as described above and its peripheral circuits.

[0030] The third aspect of the present invention provides a computing acceleration method for a heterogeneous computing acceleration platform.

[0031] A computing acceleration method for a heterogeneous computing acceleration platform includes:

[0032] Using the bus interface controlled by the driver on the ARM processor core to access each starting programmable logic area core on the predefined data processing pipeline, or using the enable signal issued by the ARM processor core to control the start and stop of each starting programmable logic area core on the data processing pipeline;

[0033] Each programmable logic area core located on the data processing pipeline receives the data processing pipeline opening and closing instruction issued by the ARM processor core, and each programmable logic area automatically processes the corresponding data in sequence; wherein, the data processing pipeline is pre-configured according to the attributes of the data to be processed; the data processed by the data processing pipeline enters the corresponding address of the DDR memory and waits for the instruction of the programmable logic controller for the next step of processing;

[0034] Using the ARM processor core to control the deep learning processor to fetch data from the corresponding address of the DDR memory and perform artificial intelligence operations based on the preset target deep learning model, and then the deep learning processor stores the artificial intelligence operation result in the corresponding address of the DDR memory;

[0035] Using the ARM processor core to control the peripheral controller of the programmable logic area to read out the artificial intelligence operation result from the corresponding address of the DDR memory and send it to the set peripheral.

[0036] As an implementation manner, the deep learning processor only fetches one deep learning model each time, and the fetched deep learning model matches the scale and architecture of the current deep learning processor;

[0037] When the number of deep learning models is at least two, the deep learning models are stored on the SD card in the form of compiled binary files and are loaded by the control program under set conditions, and only one deep learning model is loaded on the deep learning processor each time.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] (1) The present invention integrates an ARM processor core, a programmable logic area, and an artificial intelligence engine to construct an FPGA-like system-on-chip. A deep learning processor is deployed on the artificial intelligence engine, and a target deep learning model trained according to design requirements is also deployed on the deep learning processor. The corresponding target deep learning model can be designed according to a general design process and then transplanted to the artificial intelligence engine according to the quantization and recompilation processes of the AI engine; the algorithm acceleration kernel is independently designed according to the set functions and then transplanted to the programmable logic area to form a functionally independent data processing pipeline according to preset requirements; finally, the ARM processor core is used to schedule the operation of the target deep learning model in the deep learning processor, realizing the deployment of the entire data processing process of artificial intelligence inference on the same chip to achieve the purpose of full-process hardware acceleration at the hardware level.

[0040] (2) The present invention uses a network-on-chip to improve the information exchange speed and communication quality between any two of the ARM processor core, the programmable logic area, and the artificial intelligence engine, enabling the data processed by the data processing pipeline to directly enter the corresponding address in the DDR memory, avoiding congestion during the data processing process and ensuring the accuracy of data processing; the ARM processor core controls the orderly memory access of each module to achieve hardware-level data processing.

[0041] (3) The present invention uses a communication interface to directly communicate between the ARM processor core and the programmable logic area, improving the communication speed and communication quality between the two and enhancing the data processing and computing speed of the entire FPGA-like system-on-chip.

[0042] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0044] Figure 1 is a schematic structural diagram of an FPGA-like system-on-chip according to Embodiment 1 of the present invention;

[0045] Figure 2 is a schematic structural diagram of an FPGA-like system-on-chip according to Embodiment 2 of the present invention;

[0046] Figure 3 is a schematic structural diagram of a heterogeneous computing acceleration platform according to an embodiment of the present invention.

[0047] Figure 4 is a flowchart of a computing acceleration method according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0048] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0049] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0050] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0051] Term Explanation:

[0052] AI (Artificial Intelligence): Artificial intelligence, referring to the intelligent behavior exhibited by computer systems. It includes multiple fields such as machine learning, natural language processing, and visual recognition.

[0053] ARM: (Advanced RISC Machines): A widely used 32-bit and 64-bit reduced instruction set microprocessor architecture, suitable for various application scenarios such as mobile devices and embedded systems, and is the central processing unit CPU in a system-on-chip.

[0054] ASIC (Application-Specific Integrated Circuit): Application-specific integrated circuit, an integrated circuit designed for a specific purpose, which has higher performance and lower cost compared to general-purpose chips.

[0055] AXI4 Bus (Advanced eXtensible Interface 4 Bus): A high-performance, high-bandwidth, low-latency on-chip interconnection standard proposed by ARM, widely used in SoC designs, and supports multiple data widths and transfer modes.

[0056] The AXI Lite (Advanced eXtensible Interface Lite) bus is a simplified version of the AXI4 protocol. AXI Lite is suitable for simple, low-throughput address-mapped communications, especially for read and write operations of individual registers. In this case, the CPU acts as the master device of the AXI Lite and is connected to multiple AXI Lite slave devices through the AXI bus arbiter. Each slave device has its own address. This data communication is two-way.

[0057] CPU (Central Processing Unit): The central processing unit, which is the main computing component in a computer and is responsible for executing instruction sets to perform various computing tasks.

[0058] CUDA (Compute Unified Device Architecture): A parallel computing platform and programming model developed by NVIDIA that allows developers to write programs in high-level languages such as C / C++ and directly utilize the parallel computing capabilities of the GPU.

[0059] DDR memory (Double Data Rate Memory): Double Data Rate Synchronous Dynamic Random Access Memory, a type of RAM that can transfer data twice within one clock cycle, improving the data transfer rate.

[0060] DEBUG: Debugging, which refers to the process of finding and fixing errors or defects during software development. It is also used in hardware development to detect and solve hardware problems.

[0061] DMA (Direct Memory Access): Direct Memory Access, a data transfer mechanism and a hardware module in a system-on-chip that allows direct data transfer between peripherals and memory without the CPU directly operating on the data from memory for peripherals. One only needs to control the status of the DMA controller to enable peripherals to directly operate on memory, thereby improving data transfer efficiency and system performance.

[0062] FMC interface (FPGA Mezzanine Card Interface): The FPGA Mezzanine Card Interface, a standard interface for connecting an FPGA to peripheral devices, facilitating the expansion of FPGA functions.

[0063] GPU (Graphics Processing Unit): The Graphics Processing Unit, initially used to accelerate graphics rendering and now also widely applied to parallel computing tasks such as deep learning and scientific computing.

[0064] HDL (Hardware Description Language): A programming language used for circuit design, such as VHDL and Verilog, which can describe the functions, structures, and behaviors of digital and mixed-signal systems.

[0065] HDMI Interface (High Definition Multimedia Interface): A fully digital audio-video interface technology used to transmit uncompressed audio and video data.

[0066] JTAG Interface (Joint Test Action Group Interface): An international standard test protocol mainly used for chip-level testing, debugging, and programming, especially in embedded systems.

[0067] NOC (Network on Chip): A communication architecture implemented within a single chip, used to connect different IP cores or modules to improve data transmission efficiency and reduce power consumption.

[0068] DPU (Deep learning Processing Unit): A hardware unit designed specifically to accelerate deep learning algorithms, capable of efficiently processing large amounts of data and complex mathematical operations.

[0069] SFP Fiber (Small Form-factor Pluggable Fiber Optic Module): A hot-pluggable input / output device used for the conversion between electrical and optical signals, commonly used in high-speed network communications.

[0070] SD Card (Secure Digital Card): A small flash memory card used for data storage, widely used in various electronic devices such as mobile phones and cameras.

[0071] SoC (System on a Chip): An integrated circuit that integrates multiple components of a computer or other electronic system onto a single chip. These components usually include a processor, memory, input / output ports, and other necessary functional modules, aiming to improve performance, reduce power consumption, and lower costs.

[0072] Example 1

[0073] According to Figure 1, an embodiment of the present invention provides an FPGA-based system-on-chip, including: an ARM processor core, a programmable logic area, and an artificial intelligence engine;

[0074] A deep learning processor (DPU) is deployed on the artificial intelligence engine; at least one target deep learning model is loaded in the deep learning processor;

[0075] The ARM processor core controls the operation of the target deep learning model in the deep learning processor by setting a driver program;

[0076] At least one algorithm acceleration core is integrated in the programmable logic area; each algorithm acceleration core is independently designed according to a set function and then transplanted into the programmable logic area, and forms an independent data processing pipeline according to preset requirements; among them, the algorithm acceleration core in the programmable logic area is a computing unit responsible for at least one data processing function; one or more algorithm acceleration cores form a complete functional data processing pipeline.

[0077] Any two of the ARM processor core, the programmable logic area, and the artificial intelligence engine communicate with each other through an on-chip network; the on-chip network is also connected to the DDR memory.

[0078] In this way, the ARM processor core, the programmable logic area, and the artificial intelligence engine are all connected to the DDR memory through the on-chip network, so that the ARM processor core does not participate in data processing and data interaction, and only is responsible for the control of each module; the DDR memory is used to temporarily store the intermediate results of data processing, avoiding congestion during the data processing process and ensuring the accuracy of data processing; moreover, the pipeline design can minimize the interaction with the DDR memory to reduce the computing latency.

[0079] It should be noted here that the scale of the deep learning processor is variable to meet different design requirements.

[0080] It should be noted here that the ARM processor core can include MPSoC series and Versal series chips. These series of chips are the continuation of programmable logic gate array (FPGA) chips and can be regarded as a system-on-chip (SoC) that adds on-chip CPU, signal processing, deep learning engine, high-speed serial deserializer and other resources on traditional FPGA chips. In this way, hardware-level algorithm acceleration of any algorithm can be achieved to obtain performance similar to that of ASIC, and the resource waste of data flow can be minimized to the greatest extent. Moreover, the internal structure can be flexibly changed to support the upgrade of design requirements.

[0081] Among them, the target deep learning model can be a target deep learning or artificial intelligence model designed and trained in standard AI development processes such as PyTorch, TensorFlow, and Caffe. It needs to go through processes such as quantization, compilation, and pruning to generate a compiled binary file that can be deployed on the SoC of this design, and then it can be written into the deep learning processor by a program at any time.

[0082] The target deep learning model pre-cleans and organizes data according to the input requirements of the deep learning processor. For example, if it is for image operations, operations such as transcoding, compression, and color gamut conversion are required first. Another example is that in the field of signal processing, operations such as converting between digital signals and analog signals, designing filters, Fourier transforms, and signal processing are required. These data processing tasks need to be completed by corresponding programs, which can be written in C / C++ language and then compiled into hardware language (HDL) by a high-level synthesizer or directly written in hardware language.

[0083] The artificial intelligence engine (proprietary to AMD Versal series chips) of this embodiment is computing hardware resources optimized specifically for deep learning tasks. The artificial intelligence engine is the hardware foundation, and on top of it, a hardware-based logic design, that is, the deep learning processor DPU (deep learning processing unit), needs to be deployed. The DPU needs to have corresponding control programs / driver programs in the operating system. The deep learning processor is responsible for controlling the computing tasks of the AI model.

[0084] The control of the deep learning processor and the transplantation of the AI model are completed in AMD's tool Vitis AI. The Vitis AI tool needs to perform a series of processes on the model, such as quantization, compilation, pruning, transplantation, and deployment. The computing acceleration methods for the AI model can be design frameworks such as widely used PyTorch, Caffe, and TensorFlow2.

[0085] In this embodiment, each algorithm acceleration kernel needs to be independently designed and transplanted to the programmable logic area according to its function. The compiled algorithm acceleration kernel and the compiled deep learning model file are saved in an external storage device (SD card) in file form and can be loaded and replaced when necessary.

[0086] In this embodiment, the Network on Chip (NoC) can be implemented using the AXI NoC (proprietary to AMD Versal series chips). Each data processing unit can be connected in a pipeline to save the results to the DDR memory via the AXI NoC or directly interact with other modules. The programmable logic area is also responsible for deploying peripheral controllers, such as the physical connection of Ethernet, to connect the input and output of data to the data processing pipeline. The AXI NoC is a communication mechanism within the system on a chip. The AXI NoC can use the AXI4 interface as the input / output interface to allow other modules to access the NoC network.

[0087] In the specific implementation process, the AXI-stream interface is used between the programmable logic area cores to design the data processing pipeline. The pipeline is controlled using the automatic data flow method (without a handshake mechanism), or a handshake control signal, or the AXI Lite bus. Utilizing the simple and lightweight characteristics of the AXI-stream interface, data is transferred between the cores on the pipeline in the programmable logic area without an address, which is suitable for high-throughput data flow transmission scenarios and improves the computing speed of the system on a chip.

[0088] Generally, the design of the data pipeline changes according to different requirements. Specifically, there are two possible cases for the control of each module. In the first control method, an AXI Lite bus interface controlled by a driver on the CPU accesses its internal registers to change its state, thereby controlling the states of modules such as on / off / wait / obtain debug on the data pipeline. The completion of the module operation is notified to the CPU through its unique interrupt number. The CPU also resets each module by controlling a certain bit of the internal GPIO. In the second control method, the CPU only uses a certain bit of the internal GPIO to control the enable signal of the module, thereby controlling the start and stop of the module.

[0089] Specifically, the control methods of the data pipeline include the following two methods:

[0090] In the first control method, an AXI Lite bus interface controlled by a driver on the ARM processor core accesses its internal registers to change its state, thereby controlling the states of modules such as on / off / wait / obtain debug on the data pipeline. The completion of each module operation is notified to the ARM processor core through its unique interrupt number. The ARM processor core also resets each module by controlling a certain bit of the internal GPIO. Some of the drivers deployed on the ARM processor core are provided or predefined by Xilinx. In the second control method, the ARM processor core only uses a certain bit of the internal GPIO to control the enable signal of the controlled module, thereby controlling the start and stop of the module.

[0091] Data transfer between various modules on the data pipeline is completed through AXI-Stream interfaces. The data flow of these interfaces does not require the intervention of the CPU and automatically completes the flow from top to bottom of the pipeline through the handshaking mechanism between modules. Finally, it accesses a preset address in the memory through a memory access module. The memory access module will tell the ARM processor core through the AXI lite bus and generate an interrupt signal when the memory access is completed. To sum up, the ARM processor core only controls each module and does not intervene in data processing. Data will flow between various modules on the pipeline and finally access the memory. The programs of the ARM processor core are executed sequentially. After issuing control instructions to each module in the first control mode above, it will wait for the module to return an interrupt to know that the module has completed running, thus avoiding data congestion. After the pipeline in the second control mode is enabled, it will run automatically and only return an interrupt after memory access to know that the entire pipeline has completed running, thus avoiding data congestion.

[0092] In this embodiment, at least one of the programmable logic region cores is deployed with an Ethernet controller, and the Ethernet controller is connected to a fiber optic interface. The Ethernet controller uses a DMA controller to interact with the DDR memory for data. This enables the on-chip system to communicate with other external communication devices through the Ethernet controller and the fiber optic interface, expanding the application field of the on-chip system.

[0093] Each design in this embodiment is only targeted at a specific application scenario; if the application scenario changes, it is necessary to re-match the data processing core of the design data processing pipeline to ensure compliance with the needs of the new application scenario.

[0094] In one or more embodiments, a peripheral controller is further deployed in the programmable logic region, and the peripheral controller is connected to peripheral electronic devices. In this way, the peripheral controller is used to obtain data transmitted by the peripheral electronic devices or control the data acquisition instructions of the peripheral electronic devices to achieve multi-source acquisition of processed data.

[0095] In one or more embodiments, examples of tasks of the data processing pipeline can include operations such as demosaicing, gray balance, data compression, video format transcoding, video size adjustment, frame buffering, and writing to memory on the original data obtained from a data sensor.

[0096] In one or more embodiments, examples of tasks of the data processing pipeline can be obtaining data from a fiber optic, MAC frame recognition, data frame recognition, frame verification, initial data processing, and DMA writing to memory.

[0097] In one or more embodiments, examples of data processing pipeline tasks may include acquiring signal samples from an ADC sensor, filtering, demodulating, performing Fourier transform, feature extraction, and then writing to memory via DMA. The data written to memory from the data processing pipeline may be data for inference provided to a deep learning model on a deep learning processor, may be processed video data, may be network packets for accessing an operating system, or may be processed data for other purposes.

[0098] In some specific embodiments, the programmable logic area is also connected to an FMC interface, and the FMC interface is connected to an FMC expansion peripheral. The FMC interface can more flexibly expand various peripheral communication protocols and high-speed physical interfaces. Connecting an FMC expansion peripheral using the FMC interface improves the scalability of the system-on-chip.

[0099] In some specific embodiments, the ARM processor is also connected to a debug Ethernet interface inside the ARM processor, and the debug Ethernet interface is connected to a debug host.

[0100] In other embodiments, the ARM processor core is also connected to at least one of a USB interface, a serial port interface, and an SD card interface.

[0101] The programmable logic area of this embodiment can deploy valuable algorithms with high time complexity in the programmable logic area and perform computing acceleration at the hardware level to achieve a computing speed close to that of an application-specific integrated circuit; it can also deploy the pre-processing and post-processing of AI inference data in the same chip in the field of AI computing to achieve full-process acceleration at the hardware level.

[0102] Embodiment 2

[0103] According to Figure 2 , this embodiment provides an FPGA-based system-on-chip, which includes: an ARM processor core and a programmable logic area;

[0104] A deep learning processor is deployed on the programmable logic area; at least one target deep learning model is loaded in the deep learning processor; at least one algorithm acceleration kernel is also integrated in the programmable logic area; each algorithm acceleration kernel is independently designed according to a set function and then transplanted to the programmable logic area;

[0105] The ARM processor core communicates with the programmable logic area through a communication interface; the ARM processor core controls the operation of the target deep learning model in the deep learning processor through a set driver program.

[0106] The FPGA-based system-on-chip in this embodiment is implemented based on AMD Ultrascale+ MPSoC series chips. Among them, the deep learning processor is designed and deployed in the programmable logic area in the form of a hardware acceleration kernel in the programmable logic area. The scale of the deep learning processor is variable to meet different design requirements;

[0107] In this embodiment, the ARM processor core is connected to the DDR memory and its controller through the AXI4 bus, and the programmable logic areas are all connected to the DDR memory controller through the AXI4 bus.

[0108] In the specific implementation process, the AXI-stream interface is adopted between the programmable logic area cores to design a data processing pipeline. The pipeline is controlled using the automatic data flow method (without a handshake mechanism), or a handshake control signal, or the AXI Lite bus. Utilizing the simple and lightweight characteristics of the AXI-stream interface, data is transferred between the cores on the pipeline in the programmable logic area in a way without addresses, which is suitable for high-throughput data stream transmission scenarios and improves the computing speed of the system-on-chip.

[0109] In this embodiment, at least one of the programmable logic area cores is deployed with an Ethernet controller, and the Ethernet controller is connected to a fiber optic interface. In this way, through the Ethernet controller and the fiber optic interface, the system-on-chip can communicate with other external communication devices, expanding the application field of the system-on-chip.

[0110] In one or more embodiments, a peripheral controller is also deployed in the programmable logic area, and the peripheral controller is connected to peripheral electronic devices. In this way, the peripheral controller is used to obtain the data transmitted by the peripheral electronic devices or control the data acquisition instructions of the peripheral electronic devices to achieve multi-source acquisition of processed data.

[0111] In some specific embodiments, the programmable logic area is also connected to an FMC interface, and the FMC interface is connected to an FMC expansion peripheral. The FMC interface can more flexibly expand various peripheral communication protocols and high-speed physical interfaces. Connecting the FMC expansion peripheral using the FMC interface improves the scalability of the system-on-chip.

[0112] In some specific embodiments, the ARM processor is also connected to a debug Ethernet interface, and the debug Ethernet interface is connected to a debug host.

[0113] In some other embodiments, the ARM processor core is also connected to at least one of a USB interface, a serial port interface, and an SD card interface.

[0114] Embodiment III

[0115] In one or more embodiments, a heterogeneous computing acceleration platform is further provided, which includes: an FPGA-based system-on-chip and its peripheral circuits as described above, as Figure 3 shown.

[0116] It should be noted here that the heterogeneous computing acceleration platform of the present invention can be applied to fields such as vehicle-person recognition, SSD vehicle recognition, 4K resolution quad-screen recognition tasks, image sensor real-time task recognition fields based on the YOLO algorithm, and Mobilenet image sensor real-time task recognition fields, etc.

[0117] Embodiment 4

[0118] Figure 4 is a flowchart of a computing acceleration method based on a heterogeneous computing acceleration platform according to an embodiment of the present invention; referring to Figure 4 , the computing acceleration method based on a heterogeneous computing acceleration platform includes:

[0119] Step 1: Use the bus interface controlled by the driver on the ARM processor core to access each starting programmable logic area core on the predefined data processing pipeline, or use the enable signal issued by the ARM processor core to control the start and stop of each starting programmable logic area core on the data processing pipeline;

[0120] Step 2: Each programmable logic area core located on the data processing pipeline receives the data processing pipeline opening and closing instruction issued by the ARM processor core, and each programmable logic area automatically processes the corresponding data in sequence; among them, the data processing pipeline is pre-configured according to the attributes of the data to be processed; the data processed by the data processing pipeline enters the corresponding address of the DDR memory and waits for the instruction of the programmable logic controller for the next step of processing;

[0121] Step 3: Use the ARM processor core to control the deep learning processor to fetch data from the corresponding address of the DDR memory and perform artificial intelligence operations based on a preset target deep learning model, and then the deep learning processor stores the artificial intelligence operation results in the corresponding address of the DDR memory;

[0122] Step 4: Use the ARM processor core to control the peripheral controller of the programmable logic area to read out the artificial intelligence operation results from the corresponding address of the DDR memory and send them to the set peripherals.

[0123] It should be noted here that the design of the data pipeline changes due to different requirements, and the specific control methods include the following two methods:

[0124] The first control method accesses its internal registers through an AXI Lite bus interface controlled by a driver on the ARM processor core to change its state, thereby controlling the states of modules such as on / off / wait / obtain debug on the data pipeline. The completion of the operation of each module notifies the ARM processor core through its unique interrupt number. The ARM processor core also resets each module by controlling a certain bit of the internal GPIO. Some of the drivers deployed on the ARM processor core are provided or predefined by Xilinx.

[0125] The second control method is that the ARM processor core only uses the enable signal of a certain bit of the internal GPIO of the controlled module to control the start and stop of the module.

[0126] The data transfer between each module on the data pipeline is completed through the AXI-Stream interface. The data flow of these interfaces does not require the intervention of the CPU and automatically completes the flow from top to bottom of the pipeline through the handshaking mechanism between modules and finally accesses the preset address of the memory through a memory access module. The memory access module will tell the ARM processor core through the AXI lite bus and generate an interrupt signal when the memory access is completed. To sum up, the ARM processor core only controls each module and does not intervene in data processing. The data will flow between each module on the pipeline and finally access the memory. The program of the ARM processor core is executed sequentially. After it issues a control instruction to each module in the first control mode, it will wait for the module to return an interrupt to know that the module has completed its operation, thereby avoiding data congestion. After the pipeline in the second control mode is started, it will run automatically and only return an interrupt after memory access to know that the entire pipeline has completed its operation, thereby avoiding data congestion.

[0127] In the specific implementation process, it is necessary to match the appropriate number of programmable logic area cores according to the design requirements or the data throughput of the input / output device or the available hardware resources in the programmable logic area. The number of programmable logic area cores determines its data processing ability for specific tasks. Among them, the same programmable logic area cores are executed in parallel. They are started simultaneously and report to the ARM CPU through their respective interrupt numbers or buses when they complete their operations. The control program of the CPU uniformly controls all programmable logic area cores.

[0128] As an implementation method, the deep learning processor only retrieves one deep learning model each time, and the retrieved deep learning model matches the scale and architecture of the current deep learning processor;

[0129] When the number of deep learning models is at least two, the deep learning models are stored on the SD card in the form of compiled binary files and are loaded by the control program under set conditions. Only one deep learning model is loaded onto the deep learning processor each time.

[0130] The following presents a performance comparison between the existing NVIDIA Jetson series of embedded accelerators and the FPGA-based system-on-chip of the present invention. The system-on-chip of the existing NVIDIA Jetson series of embedded accelerators consists of an ARM processor, a CUDA GPU, and a tensor engine.

[0131] Among them, the FPGA-based system-on-chip in this embodiment is implemented using AMD UltraScale+ MPSoC ZU5EV; the AMD UltraScale+ MPSoC ZU5EV in the FPGA-based system-on-chip of this embodiment and the NVIDIA Jetson Nano 4GB in the NVIDIA Jetson series are similar in performance, both having a performance of about 1 TOPS, and both being cost-sensitive hardware for edge computing accelerators; moreover, the power consumption of both hardware is between 5 - 10W. In Table 1, two experiments were conducted using MPSoC 5EV and Jetson Nano respectively. The AI models in the experiments were the YOLO4 model, trained on the COCO2014 dataset and tested for inference using the validation set of COCO2014; and the SSD-Mobilenet-v2 model, trained on the COCO2014 dataset and tested for inference using the validation set of COCO2014. The batch processing in the test = 4, the resolution = 320 * 320, the above models were quantized to INT8, and the test input was the original photo saved on the Micro SD card. In ZU5EV, the input-output pipeline was not deployed, only the video processing pipeline and the deep learning processor were deployed, in order to make the scale of the deep learning processor the largest by leaving as many FPGA resources as possible. The test environment used on Jetson Nano was the Nvidia official original Jetson Inference Docker environment. The specific comparison performance is measured in "frames per second". It represents how many pictures can be processed per second, as shown in Table 1.

[0132] Table 1 Performance comparison between the AMD UltraScale+ MPSoC ZU5EV of the present invention and the existing NVIDIA Jetson series of embedded accelerators

[0133] Hardware AI Model Frames per Second ZU5EV yolov4@coco dataset 14.7 Jetson Nano yolov4@coco dataset 2 ZU5EV SSD-Mobilenet-v2@coco dataset 103.7 Jetson Nano SSD-Mobilenet-v2@coco dataset 23

[0134] As can be seen from Table 1, during the processing of the same AI model, the frame per second processing speed of the FPGA-based system-on-chip in this embodiment is higher than that of the existing NVIDIA Jetson series embedded accelerators. Therefore, the computing acceleration performance of the FPGA-based system-on-chip in this embodiment is better than that of the existing NVIDIA Jetson series embedded accelerators.

[0135] Similar to the above example, in another embodiment, the test model of the Versal chip is the AMD Versal VE2802 chip; in contrast, the NVIDIA Jetson AGX Orin with similar TOPS (number of operations per second) performance has a performance of about 400 TOPS. They are both top-level hardware for edge computing accelerators. The power consumption of the two hardware is also similar, both around 50W. Two experiments were conducted in Table 2 using the AMD Versal VE2802 and the NVIDIA Jetson AGX Orin respectively. And a larger deep learning model than the above example was used. Yolo7 was trained on the COCO2017 data and the validation set of COCO2017 was used for inference testing; using the Resnet50 model, it was trained on the ILSVRC2012 (ImageNet Large Scale Visual Recognition Challenge 2012) dataset and the validation set of ILSVRC2012 was used for inference testing. The batch processing in the test = 10, the resolution = 416*416, Vitis AI quatize was used for INT8 quantization for the VE2802, and TensorRT was used for INT8 quantization for the Jetson AGX Orin. The test images were saved on the eMMC 5.1 flash memory of the NVIDIA Jetson AGX Orin and the Micro SD of the VE2802. No input-output pipeline was deployed in the VE2802, and only the largest scale of deep learning processors located in the artificial intelligence engine was deployed. The specific comparison performance is measured by "frames per second". It represents how many pictures can be processed per second. As shown in Table 2.

[0136] Table 2 Performance comparison of the present invention taking the Versal chip as an example with existing embedded accelerators

[0137]

[0138]

[0139] As can be seen from Table 2, during the processing of the same AI model, the frame per second processing speed of the FPGA-based system-on-chip in this embodiment is higher than that of the existing NVIDIA Jetson AGX Orin accelerator. Therefore, the computing acceleration performance of the FPGA-based system-on-chip in this embodiment is superior to that of the existing NVIDIA Jetson series of embedded accelerators.

[0140] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An FPGA-type system on chip, characterized in that: include: ARM processor core, programmable logic area and artificial intelligence engine; A deep learning processor is deployed on the artificial intelligence engine; At least one target deep learning model is loaded in the deep learning processor; The ARM processor core controls the operation of the target deep learning model in the deep learning processor by setting a driver program; At least one algorithm acceleration core is integrated in the programmable logic area; each algorithm acceleration core is independently designed according to a set function and then transplanted to the programmable logic area, and forms a functionally independent data processing pipeline according to preset requirements; Any two of the ARM processor core, programmable logic area and artificial intelligence engine communicate with each other through the on-chip network; the on-chip network is also connected to the DDR memory.

2. The FPGA-based system on chip as claimed in claim 1, characterized in that: The data processing pipeline cores of the programmable logic area are designed using an AXI-stream interface; the ARM processor core controls the state of the data processing pipeline.

3. The FPGA-based system on chip as claimed in claim 1, characterized in that: At least one of the programmable logic area cores is deployed with an Ethernet controller, and the Ethernet controller is connected to the optical fiber interface.

4. The FPGA-based system on chip as claimed in claim 1, characterized in that: A peripheral controller is also deployed in the programmable logic area, and the peripheral controller is connected to the peripheral electronic device; the programmable logic area is also connected to the FMC interface, and the FMC interface is connected to the FMC expansion peripheral.

5. An FPGA-type system on chip, characterized in that: include: ARM processor core and programmable logic area; A deep learning processor is deployed on the programmable logic area; At least one target deep learning model is loaded in the deep learning processor; at least one algorithm acceleration core is also integrated in the programmable logic area; each algorithm acceleration core is independently designed according to the set function and then transplanted to the programmable logic area; The ARM processor core communicates with the programmable logic area through a communication interface; the ARM processor core controls the operation of the target deep learning model in the deep learning processor by setting a driver.

6. The FPGA-based system on chip as claimed in claim 5, characterized in that: The data processing pipeline is designed using an AXI-stream interface between the programmable logic area cores; or at least one of the programmable logic area cores is deployed with an Ethernet controller, and the Ethernet controller is connected to the optical fiber interface.

7. The FPGA-based system on chip as claimed in claim 5, characterized in that: A peripheral controller is also deployed in the programmable logic area, and the peripheral controller is connected to the peripheral electronic device; the programmable logic area is also connected to the FMC interface, and the FMC interface is connected to the FMC expansion peripheral.

8. A heterogeneous computing acceleration platform, characterized in that: include: The FPGA-type system on chip and its peripheral circuit as described in any one of claims 1 to 4; Or an FPGA-type system-on-chip and its peripheral circuits as described in any one of claims 5-7.

9. A computing acceleration method based on the heterogeneous computing acceleration platform as claimed in claim 8, characterized in that: The steps include: Using the bus interface controlled by the driver on the ARM processor core to access each starting programmable logic area core on the predefined data processing pipeline, or using the enable signal sent by the ARM processor core to control the start and stop of each starting programmable logic area core on the data processing pipeline; Each programmable logic area core located on the data processing pipeline receives the data processing pipeline start and stop instructions issued by the ARM processor core, and each programmable logic area automatically processes the corresponding data in turn; wherein, the data processing pipeline is pre-configured according to the attributes of the data to be processed; the data processed by the data processing pipeline enters the corresponding address of the DDR memory, waiting for the instruction of the programmable logic controller for the next step of processing; The ARM processor core is used to control the deep learning processor to fetch data from the corresponding address of the DDR memory and perform artificial intelligence calculations based on the preset target deep learning model. The deep learning processor then stores the artificial intelligence calculation results into the corresponding address of the DDR memory. The ARM processor core is used to control the peripheral controller in the programmable logic area to read the artificial intelligence calculation results from the corresponding address of the DDR memory and send them to the set peripherals.

10. The computing acceleration method according to claim 9, characterized in that: The deep learning processor only calls one deep learning model at a time, and the called deep learning model matches the scale and architecture of the current deep learning processor; When the number of deep learning models is at least two, the deep learning models are stored on the SD card in the form of compiled binary files and loaded by the control program when conditions are set, and only one deep learning model is loaded on the deep learning processor at a time.