A VPX-type heterogeneous acceleration module based on NPU+FPGA architecture
By designing a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture, the problems of insufficient CPU computing power and high GPU power consumption are solved, and efficient algorithm acceleration and data transmission in special environments are achieved, which is suitable for scenarios such as industry and aerospace.
Patent Information
- Application Number
- CN202210194756.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-01
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-03-01
AI Technical Summary
In the existing technology, the CPU computing power is insufficient to meet the requirements of intelligent algorithms, GPU chips consume large amounts of power and are expensive, and the PCIE interface cannot meet the requirements of efficient and stable data transmission under special circumstances.
A VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture is designed. It uses domestic FPGA and NPU chips, and provides PCIE interface and Ethernet interface through the VPX connector. The NPU module can operate in RC or EP mode, and the FPGA module is reconfigurable to support acceleration in different intelligent application scenarios.
It achieves efficient and stable data transmission and algorithm acceleration in special environments, uses NPU acceleration to meet scenarios with high computing power requirements, and uses FPGA acceleration to meet scenarios where algorithm hardware is programmable. It has good mechanical structure and high bandwidth characteristics.
Smart Images

Figure CN114610483B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of heterogeneous acceleration technology, and specifically relates to a VPX-type heterogeneous acceleration module based on NPU+FPGA architecture. Background Art
[0002] Real-world scenarios such as object detection, speech recognition, and image processing require massive amounts of data processing, placing increasingly high demands on processor performance. While the CPU's primary strength lies in task scheduling and resource management, its computing power falls far short of the requirements of intelligent algorithms. To meet these demands, specialized processor chips for intelligent algorithms, such as GPUs, NPUs, and FPGAs, have emerged. These chips boast powerful parallel computing capabilities and, when combined with the CPU, form a heterogeneous acceleration platform that meets the needs of real-world scenarios.
[0003] Heterogeneous computing platforms can greatly improve the computing performance of the system. Data transmission between the host and the acceleration module is generally achieved through the PCIE interface. PCIE can provide a standard interface to achieve high-speed and stable data transmission, and has different specifications and different transmission rates to meet different application requirements. However, there are also the following problems: the physical form of the PCIE interface is relatively simple and is only suitable for general office environments. For special environments such as industry and aerospace, it cannot meet the requirements of efficient and stable work. As a new generation of industrial bus standard, VPX introduces high-speed buses such as Rapid IO, PCIE, and high-speed Ethernet based on its mechanical structure and cooling and anti-seismic advantages, solving bandwidth and high-speed interconnection problems, and greatly increasing the communication rate between internal components of the system. The VPX system architecture is suitable for special environments such as industry and aerospace due to its modularity, universality, high reliability and high bandwidth.
[0004] GPU is an image processor, a single-instruction, multi-data processing processor with high main frequency, large bandwidth, a large number of computing units, and powerful parallel computing capabilities. It is mostly used in application scenarios such as image processing and target recognition, but GPUs are expensive and power-hungry.
[0005] NPU neural network processing unit, NPU is a special chip for neural networks with small size, low power consumption, high computing performance and high computing efficiency. Compared with CPU and GPU, NPU improves operating efficiency by highlighting the integration of storage and calculation of neural network algorithm weights, and is specially designed to realize the application of AI algorithms. FPGA programmable gate array is a hardware programmable chip with powerful parallel computing capabilities. It can design and optimize hardware circuits according to different algorithms. With the same computing power, its power consumption is about one-tenth of that of GPU. Its main feature is that the hardware circuit is reset and can support the acceleration of different algorithm models. The acceleration module using the combination of NPU + FPGA chips can not only meet the requirements of artificial intelligence algorithms, high bandwidth and powerful computing power, but also meet different application scenarios. Summary of the Invention
[0006] (1) Technical issues to be solved
[0007] The technical problem to be solved by the present invention is how to provide a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture to solve the problems of high computing power requirements of intelligent algorithms, high power consumption and high price of GPU chips.
[0008] (2) Technical solution
[0009] In order to solve the above technical problems, the present invention proposes a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture. The acceleration module mainly includes an FPGA module, an NPU module, a PCIE switching module, a VPX connector, a storage module and a BMC management module.
[0010] The FPGA module is connected to the VPX connector through the PCIE switch module, providing a PCIE interface to the outside. The FPGA module is connected to the storage module through the memory interface. The FPGA module is connected to the ETH PHY chip, which is connected to the VPX connector. The FPGA module brings out a Gigabit network through the ETH PHY chip.
[0011] The NPU module is connected to the VPX connector via the PCIE switch module, providing a PCIE interface to the outside world. The NPU module is connected to the FPGA module via the GPIO interface. When the NPU module operates in RC mode, the FPGA algorithm is reconfigured through the NPU module. The NPU module is connected to the ETH PHY chip, which is connected to the VPX connector. The NPU module provides a Gigabit network to the outside world through the ETH PHY chip.
[0012] The BMC management module is connected to the VPX connector via the I2C interface to accelerate the transmission of module temperature, current, and voltage signal acquisition data;
[0013] The VPX connector of the heterogeneous acceleration module provides two interface modes: PCIE interface and Ethernet interface for host connection.
[0014] Furthermore, the FPGA module uses the V7690T chip, the NPU module uses the Atlas200 module, the storage module uses the JM3D512 chip, and the BMC management module uses the JS32F103 chip.
[0015] Furthermore, the VPX connector is a 6U VPX connector.
[0016] Furthermore, the NPU module can work in both RC mode and EP mode. In EP mode, the host accesses the FPGA module and NPU module through the PCIE interface to accelerate the algorithm; in RC mode, the NPU module realizes data transmission and algorithm acceleration through the Ethernet interface, and at the same time reconfigures the algorithm of FPGA through NPU.
[0017] Furthermore, the VPX connector of the acceleration module provides two interface modes: PCIE interface and Ethernet interface. The external host controls the bus simulation switch and chooses to use the PCIE interface or Ethernet interface on the VPX connector to access the FPGA or NPU module.
[0018] Furthermore, the NPU module supports PCIE3.0 X4, the FPGA module supports PCIE3.0 X8, and the VPX connector supports PCIE3.0 X8.
[0019] Furthermore, the NPU module and FPGA module bring out the Gigabit network through the ETH PHY chip and VPX connector, and support access to the PCIE interface provided by the VPX connector; the ETH PHY chip also supports Ethernet interface access; the Ethernet interface outputs a 1000BaseT network data stream, and the VPX connector outputs a 1000BaseX network data stream.
[0020] Furthermore, the BMC management module implements sampling and monitoring of the input 12V voltage through voltage and current sampling circuits; collects the temperature of the power supply and FPGA chip respectively through temperature sensors; the BMC module is connected to the external unit management module through the IPMB bus, and the IPMB bus is composed of two independent hardware I2C buses in hardware.
[0021] Furthermore, the FPGA module uses Veri log or VHDL language to implement the graphical design of the FPGA chip IP core, encapsulating commonly used data processing algorithms into IP cores. The circuit in the FPGA module adopts a data pipeline architecture, and data flows and calculates according to a pre-designed process.
[0022] Furthermore, the VPX connector provides 12V voltage to other power-consuming modules and a separate 3.3V voltage to the BMC to ensure that the BMC can operate normally even if problems occur in other circuits; the 12V voltage is converted to 5V through voltage conversion 1, and the 5V voltage is converted to 0.9V and 1.8V through voltage conversion 2 to power the PCIE switch module; the 5V voltage is converted to 1.2V and 1.5V through voltage conversion 3 to power the DDR3 module; the 5V voltage is converted to 3.8V through voltage conversion 4, and the 3.8V voltage is converted to 3.3V through voltage conversion 5, and then the 3.8V and 3.3V voltages are converted to 1.0V and 1.8V through voltage conversion 6 to power the NPU module; the 5V voltage is converted to 1.0V, 1.8V and 2.0V through voltage conversion 7, and the 12V and 5V are converted to 3.3V through voltage conversion 8, and then the 3.3V voltage is converted to 0.75V through voltage conversion 9 to power the FPGA module.
[0023] (3) Beneficial effects
[0024] This paper designs a VPX-type heterogeneous acceleration module based on an NPU+FPGA architecture. This independently designed deep learning accelerator module utilizes a domestically produced FPGA and a domestically produced NPU as its core. This module supports and accelerates mainstream AI algorithms, with all core components being domestically produced. The module can accelerate applications in various intelligent applications, using the NPU for scenarios requiring high computing power and the FPGA for scenarios requiring hardware-programmable algorithms. The NPU and FPGA share a PCIE interface, connected to the VPX connector via a PCIE switch. The designed NPU module can operate in both RC and EP modes. In RC mode, the NPU can reconfigure the FPGA algorithm.
[0025] The present invention proposes a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture. The present invention designs a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture, which can realize acceleration for different intelligent application scenarios; NPU is used to accelerate computing power for scenarios with high computing power requirements, and FPGA is used to accelerate computing power for scenarios with programmable algorithm hardware. The VPX-type connector can provide a PCIE communication interface on the basis of a good mechanical structure to ensure efficient and stable data transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a schematic diagram of the connection of the VPX heterogeneous acceleration module based on the NPU+FPGA architecture of the present invention;
[0027] Figure 2 This is a schematic diagram of voltage conversion for the VPX heterogeneous acceleration module based on the NPU+FPGA architecture of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below with reference to the accompanying drawings and examples.
[0029] This invention designs a VPX-type heterogeneous acceleration module based on an NPU+FPGA architecture, belonging to the field of heterogeneous acceleration modules. It can achieve acceleration for different intelligent application scenarios. The NPU is used for computing acceleration in scenarios requiring high computing power, while the FPGA is used for computing acceleration in scenarios requiring hardware-programmable algorithms. The NPU and FPGA modules have PCIE interfaces, which are connected to the VPX connector via a PCIE switch. The designed NPU module can operate in both RC mode and EP mode. In RC mode, the NPU can reconfigure the FPGA algorithm. A VPX-type heterogeneous acceleration module with an NPU+FPGA architecture primarily includes an FPGA module, an NPU module, a PCIE switch module, a VPX connector, a storage module, and peripheral electronic components. The FPGA module uses the V7690T chip produced by Fudan Microelectronics, and the NPU module uses the Atlas200 module produced by Huawei. The heterogeneous acceleration module provides PCIE and Ethernet interfaces, allowing external hosts to access the FPGA and NPU modules through these interfaces. The acceleration module accelerates the inference of intelligent algorithms and returns the calculation results to the host.
[0030] This invention designs a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture. It uses large-capacity, high-performance domestically produced FPGA and NPU chips as the main chips, and is designed with a standard 6U VPX connector, which can provide a standard PCIE interface to support the acceleration of intelligent algorithms. The specific implementation plan is as follows:
[0031] The present invention proposes a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture. The acceleration module mainly includes an FPGA module, an NPU module, a PCIE switching module, a VPX connector, a storage module, and a BMC management module.
[0032] The FPGA module is connected to the VPX connector through the PCIE switch module, providing a PCIE interface to the outside. The FPGA module is connected to the storage module through the memory interface. The FPGA module is connected to the ETH PHY chip, and the ETH PHY chip is connected to the VPX connector. The FPGA module leads to the Gigabit network through the RJ45 connector or VPX connector.
[0033] The NPU module is connected to the VPX connector through the PCIE switch module, providing a PCIE interface to the outside world. The NPU module is connected to the FPGA module through the GPIO interface. When the NPU module works in RC mode, the FPGA algorithm is reconfigured through the NPU module. The NPU module is connected to the ETH PHY chip, which is connected to the VPX connector. The NPU module brings out the Gigabit network through the RJ45 connector or VPX connector.
[0034] The BMC management module is connected to the VPX connector via the I2C interface to accelerate the transmission of module temperature, current, and voltage signal acquisition data;
[0035] The VPX connector of the heterogeneous acceleration module provides two interface modes: PCIE interface and Ethernet interface for host connection.
[0036] A VPX heterogeneous acceleration module connection diagram based on NPU+FPGA architecture is shown in the figure. Figure 1 As shown, it includes FPGA module, NPU module, PCIE switching module, VPX connector, storage module, BMC management module and peripheral electronic components. The FPGA module uses the V7690T chip produced by Fudan Micro, the NPU module uses the Atlas200 module produced by Huawei, the storage module uses the JM3D512 chip produced by China Electronics 58 Institute, and the BMC management module uses the JS32F103 chip produced by China Electronics 58 Institute.
[0037] The FPGA module is connected to the VPX connector, the NPU module is connected to the VPX connector, and the BMC management module is connected to the VPX connector. Specifically: the FPGA module is connected to the VPX connector via the PCIE switch module, providing a PCIE interface to the outside world. The FPGA module is connected to the storage module via the memory interface. The storage module stores the voice, image, video, and other data to be processed and the results of accelerated processing. The NPU module is connected to the VPX connector via the PCIE switch module, providing a PCIE interface to the outside world. The host can achieve high-speed and stable data transmission with the acceleration module through the PCIE interface of the VPX connector. The NPU module is connected to the FPGA module via the GPIO interface. When the NPU is working in RC mode, the FPGA algorithm can be reconfigured through the NPU. The BMC management module is connected to the VPX connector via the I2C interface to achieve the transmission of temperature, current, and voltage signal acquisition data from the acceleration module.
[0038] A VPX-based heterogeneous acceleration module based on an NPU+FPGA architecture can accelerate different intelligent application scenarios. The NPU is used for computing acceleration in scenarios requiring high computing power, while the FPGA is used for computing acceleration in scenarios requiring hardware-programmable algorithms. The NPU and FPGA have PCIe interfaces, which are connected to the VPX connector via a PCIe switch, providing an external PCIe interface. The designed NPU module can operate in both RC mode and EP mode. In EP mode, the host accesses the NPU module via the PCIe interface for algorithm acceleration. In RC mode, the NPU module uses an Ethernet interface for external data transmission and algorithm acceleration, and can also reconfigure the FPGA through the NPU. The heterogeneous acceleration module provides both PCIe and Ethernet interfaces. An external host can use a bus analog switch to select either the PCIe or Ethernet interface on the VPX connector to access the FPGA or NPU module, providing different ways to accelerate intelligent algorithm inference. An analog switch is a manual switch in the hardware circuit that switches between different modes by flipping the switch.
[0039] A VPX-type heterogeneous acceleration module based on an NPU+FPGA architecture supports NPU acceleration or FPGA acceleration, and supports independent operation of the NPU module. The NPU module supports PCIE 3.0 x4, the FPGA module supports PCIE 3.0 x8, and the VPX connector supports PCIE 3.0 x8. The NPU and FPGA modules are connected to an ETH PHY chip, which is connected to a VPX connector. The NPU and FPGA modules connect to the Gigabit network through the ETH PHY chip and VPX connector, supporting access via the PCIE interface provided by the VPX connector. The ETH PHY chip also supports access via the Ethernet interface, which outputs a 1000BaseT network data stream, while the VPX connector outputs a 1000BaseX network data stream.
[0040] A power conversion diagram of a VPX heterogeneous acceleration module based on NPU+FPGA architecture is shown in the figure. Figure 2 As shown, the VPX connector provides 12V voltage to other power-consuming modules and a separate 3.3V voltage to the BMC, ensuring normal operation of the BMC even if other circuits fail. Voltage converter 1 converts 12V to 5V, and voltage converter 2 converts 5V to 0.9V and 1.8V to power the PCIE switch module. Voltage converter 3 converts 5V to 1.2V and 1.5V to power the DDR3 module. Voltage converter 4 converts 5V to 3.8V, and voltage converter 5 converts 3.8V to 3.3V. Voltage converter 6 converts 3.8V and 3.3V to 1.0V and 1.8V to power the NPU module. Voltage converter 7 converts 5V to 1.0V, 1.8V, and 2.0V. Voltage converter 8 converts 12V and 5V to 3.3V. Voltage converter 9 converts 3.3V to 0.75V to power the FPGA module.
[0041] A BMC health management system in a VPX heterogeneous acceleration module based on an NPU+FPGA architecture supports time synchronization in the standard IPMI 2.0 format and OEM IPMI format get / set module data commands, responding to unit module inquiries and enabling monitoring and management of the module. This system primarily includes sampling and monitoring the 12V input voltage through voltage and current sampling circuits, and collecting temperatures at the power supply and FPGA chip using temperature sensors. The BMC module connects to an external unit management module via the IPMB bus, which consists of two independent I2C buses.
[0042] A VPX-type heterogeneous acceleration module based on an NPU+FPGA architecture contains a large number of arithmetic and logic units within the FPGA chip, which can fully leverage the inherent parallelism of algorithms. FPGAs are programmable gate arrays, and Verilog or VHDL languages can be used to graphically design FPGA chip IP cores. Common data processing algorithms can be encapsulated into IP cores, thereby accelerating the inference of various intelligent algorithms. The circuits within FPGA chips generally utilize a data pipeline architecture, where data flows and is calculated according to pre-designed processes. This eliminates the need for complex instruction control and avoids the time-consuming process of fetching and decoding instructions. Compared to CPUs, if properly designed, programs containing a large number of computationally intensive functions, such as intelligent algorithms, can run on FPGA chips with lower latency and higher performance.
[0043] This paper designs a VPX-type heterogeneous acceleration module based on an NPU+FPGA architecture. This independently designed deep learning accelerator module utilizes a domestically produced FPGA and a domestically produced NPU as its core. This module supports and accelerates mainstream AI algorithms, with all core components being domestically produced. The module can accelerate applications in various intelligent applications, using the NPU for scenarios requiring high computing power and the FPGA for scenarios requiring hardware-programmable algorithms. The NPU and FPGA share a PCIE interface, connected to the VPX connector via a PCIE switch. The designed NPU module can operate in both RC and EP modes. In RC mode, the NPU can reconfigure the FPGA algorithm.
[0044] The present invention proposes a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture. The present invention designs a VPX-type heterogeneous acceleration module based on the NPU+FPGA architecture, which can realize acceleration for different intelligent application scenarios; NPU is used to accelerate computing power for scenarios with high computing power requirements, and FPGA is used to accelerate computing power for scenarios with programmable algorithm hardware. The VPX-type connector can provide a PCIE communication interface on the basis of a good mechanical structure to ensure efficient and stable data transmission.
[0045] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A VPX heterogeneous acceleration module based on NPU+FPGA architecture, characterized by: The acceleration module mainly includes FPGA module, NPU module, PCIE switch module, VPX connector, storage module and BMC management module; The FPGA module is connected to the VPX connector through the PCIE switch module, providing a PCIE interface to the outside. The FPGA module is connected to the storage module through the memory interface. The FPGA module is connected to the ETH PHY chip, which is connected to the VPX connector. The FPGA module brings out a Gigabit network through the ETH PHY chip. The NPU module is connected to the VPX connector via the PCIE switch module, providing a PCIE interface to the outside world. The NPU module is connected to the FPGA module via the GPIO interface. When the NPU module operates in RC mode, the FPGA algorithm is reconfigured through the NPU module. The NPU module is connected to the ETH PHY chip, which is connected to the VPX connector. The NPU module provides a Gigabit network to the outside world through the ETH PHY chip. The BMC management module is connected to the VPX connector via the I2C interface to accelerate the transmission of module temperature, current, and voltage signal acquisition data; The VPX connector of the heterogeneous acceleration module provides two interface modes: PCIE interface and Ethernet interface for host connection; in, The NPU module operates in RC mode or EP mode. In EP mode, the host accesses the FPGA module and NPU module through the PCIE interface to accelerate the algorithm. In RC mode, the NPU module implements external data transmission and algorithm acceleration through the Ethernet interface, and at the same time implements algorithm reconfiguration of the FPGA through the NPU.
2. The VPX heterogeneous acceleration module based on the NPU+FPGA architecture according to claim 1, characterized in that: The FPGA module uses the V7690T chip, the NPU module uses the Atlas200 module, the storage module uses the JM3D512 chip, and the BMC management module uses the JS32F103 chip.
3. The VPX heterogeneous acceleration module based on NPU+FPGA architecture according to claim 1, characterized in that: The VPX connector is a 6U VPX connector.
4. The VPX heterogeneous acceleration module based on NPU+FPGA architecture according to claim 1, characterized in that: The VPX connector of the acceleration module provides two external interfaces: PCIE interface and Ethernet interface. The external host controls the bus simulation switch and chooses to use the PCIE interface or Ethernet interface on the VPX connector to access the FPGA or NPU module.
5. The VPX heterogeneous acceleration module based on NPU+FPGA architecture according to claim 1, characterized in that: The NPU module supports PCIE 3.0 x4, the FPGA module supports PCIE 3.0 x8, and the VPX connector supports PCIE 3.0 x8.
6. The VPX heterogeneous acceleration module based on NPU+FPGA architecture according to claim 1, characterized in that: The NPU module and FPGA module bring out the Gigabit network through the ETH PHY chip and VPX connector, and support access through the PCIE interface provided by the VPX connector; the ETH PHY chip also supports Ethernet interface access; the Ethernet interface outputs a 1000BaseT network data stream, and the VPX connector outputs a 1000BaseX network data stream.
7. The VPX heterogeneous acceleration module based on NPU+FPGA architecture according to claim 1, characterized in that: The BMC management module samples and monitors the input 12V voltage through voltage and current sampling circuits; collects the temperature of the power supply and FPGA chip through temperature sensors; the BMC module is connected to the external unit management module through the IPMB bus, and the IPMB bus is composed of two independent hardware I2C buses in hardware.
8. The VPX heterogeneous acceleration module based on NPU+FPGA architecture according to claim 1, characterized in that: The FPGA module uses Verilog or VHDL language to implement the graphical design of FPGA chip IP cores, encapsulating commonly used data processing algorithms into IP cores. The circuits in the FPGA module use a data pipeline architecture, and data flows and calculates according to a pre-designed process.
9. The VPX heterogeneous acceleration module based on the NPU+FPGA architecture according to any one of claims 4 to 8, characterized in that: The VPX connector provides 12V voltage to other power-consuming modules and a separate 3.3V voltage to the BMC to ensure that the BMC can operate normally even if other circuits fail. Voltage conversion 1 converts 12V to 5V, and voltage conversion 2 converts 5V to 0.9V and 1.8V to power the PCIE switch module. Voltage conversion 3 converts 5V to 1.2V and 1.5V to power the DDR3 module. Voltage conversion 4 converts 5V to 3.8V, and voltage conversion 5 converts 3.8V to 3.3V. Voltage conversion 6 converts 3.8V and 3.3V to 1.0V and 1.8V to power the NPU module. Voltage conversion 7 converts 5V to 1.0V, 1.8V, and 2.0V. Voltage conversion 8 converts 12V and 5V to 3.3V. Voltage conversion 9 converts 3.3V to 0.75V to power the FPGA module.
Citation Information
Patent Citations
Heterogeneous processor platform based management architecture and management method therefor
CN105630723A
FPGA heterogeneous computing accelerating system and method
CN107346170A