A heterogeneous multi-core processor architecture based on NIC-400 crossbar
By using a heterogeneous multi-core processor architecture based on the NIC-400 cross matrix, the problems of inefficient parallel transaction processing and high power consumption of system-level heterogeneous multi-core processors are solved. High-performance parallel communication and low-latency heterogeneous core interconnection are achieved, optimizing the power consumption and performance of mobile terminal multimedia devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2026-03-31
AI Technical Summary
Existing system-level heterogeneous multi-core processors suffer from problems such as low efficiency, high power consumption in global synchronization transmission, and low communication bandwidth due to bus arbitration in parallel computing, making it difficult to meet the requirements of real-time high-performance parallel computing.
It adopts a heterogeneous multi-core processor architecture based on NIC-400 cross matrix. Through unified addressing, asynchronous communication, physical layer implementation and multi-node parallel computing, the heterogeneous cores communicate with the interconnection network through a splittable asynchronous bridge. The global ID bit width and frequency are configured to meet the performance requirements of the heterogeneous cores. QoS-400 and QVN-400 virtual networks are used to prevent arbitration blocking and reduce power consumption.
It improves parallel transmission rate and interaction latency, maximizes the performance of heterogeneous cores, reduces power consumption, optimizes the power area product (PPA) performance of system-level multimedia SoC, and enhances the communication quality of mobile terminal multimedia devices.
Smart Images

Figure CN115658594B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of heterogeneous multi-core processor technology, and in particular to a heterogeneous multi-core processor architecture based on the NIC-400 cross matrix. Background Technology
[0002] With the rapid development of emerging mobile terminal multimedia-related industries such as multimedia applications, image processing, virtual reality, and computer vision, more stringent requirements are being placed on system-level multimedia solutions. The parallelism in transaction processing makes the interconnection of heterogeneous multi-core processors more complex, and real-time high-performance parallel computing is becoming increasingly sensitive to communication bandwidth and latency. Timing convergence of heterogeneous multi-core processors at advanced process nodes is becoming increasingly difficult. Therefore, the optimization of system-level multimedia architecture, parallel computing, and physical implementation have become current research hotspots.
[0003] System-on-a-multimedia (SoC) chips integrate various dedicated processor cores, such as CPUs for general-purpose complex logic operations, GPUs for massively parallel computing rendering architectures, video codec processors (VPUs), display format processors (DPUs) for various display formats, data-driven neural network accelerators (NPUs), and dedicated mass data processing units (DPUs), etc. Appropriate bus architectures can unlock the performance of these heterogeneous processor cores to address issues such as parallel computing, communication bandwidth, system power consumption, and access latency. Summary of the Invention
[0004] Therefore, the technical problem to be solved by this invention is to overcome the problems of low efficiency in parallel transaction processing, high power consumption in global synchronous transmission, and low communication bandwidth caused by bus arbitration in current system-level heterogeneous multi-core processors. It proposes a parallel transmission interconnection architecture based on asynchronous cross matrix to release the maximum performance of each heterogeneous core and improve the parallel transmission rate and interaction latency.
[0005] To address the aforementioned technical problems, this invention provides a heterogeneous multi-core processor architecture based on a NIC-400 cross-matrix. This architecture includes unified addressing of all heterogeneous processor cores, asynchronous communication between the heterogeneous cores and the NIC-400 to meet bandwidth requirements, physical layer implementation, and multi-node parallel computing. Furthermore, the heterogeneous processor comprises multiple heterogeneous cores, including a CPU core for complex computation, a high-throughput parallel computing unit (GPU), a video decoding unit (VPU) supporting multiple video formats, and a display engine (DPU) supporting multiple formats. These heterogeneous cores have different pipeline structures, each acting as a host and capable of independent computation. Each heterogeneous core communicates with the internet via a detachable asynchronous bridge. Moreover, its fully configurable, non-blocking, low-latency, and low-power characteristics can address the PPA requirements of mobile terminal multimedia SoCs at the architectural level.
[0006] In one embodiment of the present invention, the NIC-400 includes routing nodes corresponding one-to-one with network interfaces, and adjacent nodes are connected by connection lines that cut off data feedback. The bit width and frequency of each heterogeneous core interface based on the NIC-400 cross matrix can be arbitrarily configured to meet the target host's requirement for high-performance parallel communication computing with lower power consumption and area.
[0007] In one embodiment of the present invention, the heterogeneous core interfaces in the NIC-400 cross-connect matrix adopt a fully asynchronous clock design. The cross-connect matrix has multiple configurable parameters to meet the system's PPA requirements. According to the OT requirements that each host can perform, the global ID bit width can be arbitrarily adjusted to reduce the average memory access latency. According to the application scenario requirements, the interface type, clock domain, and bit width of each heterogeneous core can be arbitrarily configured. In addition, each interface channel can be independently inserted into a register chip to meet the timing and latency requirements.
[0008] The NIC-400 maximizes the performance of high-throughput applications, minimizes power consumption of mobile devices, and ensures the quality of service for system communications. The advanced QoS-400 can dynamically adjust the communication transactions of the entire network according to different configurations, the QVN-400 QoS virtual network effectively prevents arbitration congestion, and the TLX-400 reduces routing congestion and easily achieves timing closure for long paths.
[0009] In one embodiment of the present invention, the asynchronous bridge is based on a detachable structure integration in the form of a handshake. At the same time, both the host side and the slave side of the asynchronous bridge have FIFOs for each clock domain. The five channels of the AXI interface each have independent FIFOs, and the width of each FIFO is the sum of the signal bit widths of all channels in the current channel. The FIFOs can be used to complete the reliable transmission of large amounts of data to meet the high bandwidth parallel computing requirements of heterogeneous cores.
[0010] In one embodiment of the present invention, the CPU and GPU in the heterogeneous core adopt a shared memory approach, which can effectively reduce the communication overhead between the heterogeneous cores. Frequent memory accesses of the two cores will cause blocking at the system memory controller. Based on the differences in latency and bandwidth requirements of the two cores, the NIC-400 with QVN-400 and QoS-400 functions can effectively solve these problems.
[0011] In one embodiment of the present invention, the heterogeneous processor has an NIC-400 with scalable interface bandwidth; the GPU performs parallel computations on a large number of similar features, and the CPU performs high-performance computations on more general transactions and more complex logic. The scalable NIC-400 improves the parallel computing performance and concurrent transaction processing capability of the heterogeneous multi-core system and can be flexibly configured into a multi-level bus architecture to reduce memory access latency.
[0012] In one embodiment of the present invention, the NIC-400 cross-connect matrix employs switches and bridges with different clock domains and configurable data widths to meet the performance requirements of the hosts. The use of hierarchical clock gating technology can effectively reduce the power consumption of hosts in idle clock domains. In addition, switches connecting multiple hosts can be split into smaller switches to increase frequency and reduce critical path latency.
[0013] In one embodiment of the present invention, the NIC-400 has system-level characteristics and can automatically optimize the system architecture, and the parameters of the NIC-400 can be dynamically configured through GPV to meet actual bandwidth or performance requirements. Thin-Link can solve the cabling congestion problem in physical implementation.
[0014] In one embodiment of the invention, the NIC-400 also has an IP suite that can evaluate and optimize the area of a subsystem or submodule and the latency of host access to a slave device through built-in algorithms.
[0015] Compared with the prior art, the above-mentioned technical solution of the present invention has the following advantages: The heterogeneous multi-core processor architecture based on NIC-400 cross matrix described in the present invention releases the performance of each heterogeneous processor core through a suitable bus architecture to optimize efficiency, improve parallel transmission rate and interaction latency, and makes the interface between heterogeneous cores and cross matrix simpler. The simple and efficient wiring is more conducive to achieving physical timing convergence. Attached Figure Description
[0016] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0017] Figure 1 This is a schematic diagram of the heterogeneous multi-core processor architecture based on the NIC architecture of the present invention;
[0018] Figure 2 This is a schematic diagram of the processor bus topology based on the NIC-400 cross-matrix of the present invention;
[0019] Figure 3 This is a schematic diagram of the QoS policies for various types of hosts described in this invention;
[0020] Figure 4 This is a schematic diagram of the QVN strategy for various types of hosts described in this invention. Detailed Implementation
[0021] like Figure 1As shown, this embodiment provides a heterogeneous multi-core processor architecture based on a NIC-400 cross-matrix. The architecture includes unified addressing of each heterogeneous processor core, asynchronous communication between the heterogeneous cores and the NIC-400 to meet bandwidth requirements, physical layer implementation, and multi-node parallel computing. In addition, the heterogeneous processor is composed of multiple heterogeneous cores, which are a CPU core for complex calculations, a high-throughput parallel computing unit GPU, a VPU supporting multiple video decoding units, and a DPU supporting multiple formats. The heterogeneous cores have different pipeline structures, each heterogeneous core can independently complete the operation as a host, and each heterogeneous core communicates with the Internet through a detachable asynchronous bridge.
[0022] Based on the asynchronous CoreLink-NIC400 cross-matrix system-level heterogeneous multi-core processor architecture, the parameters of each host interface of the cross-matrix are configured completely independently and asynchronous communication is used for physical implementation; in order to meet the bandwidth and memory access latency of each heterogeneous core, the CPU occupies the processor bus, the GPU and VPU occupy the multimedia bus, and the real-time host DPU occupies the display bus.
[0023] Furthermore, such as Figure 1 As shown, the CPU uses a 64AXI3@667MHz interface. The ROM on the processor bus is used for CPU startup level 1 boot, and the SRAM serves as a variable cache or stack space for CPU applications to improve the CPU startup speed. Both the ROM and SRAM use the AXI3 interface to support multiple OTs in order to improve instruction and data parallelism.
[0024] The GPU and VPU have interfaces of 128AXI4@500MHz and 250MHz respectively. This greedy host occupies the audio and video bus at the same time, while the slave side has only one DDR. This strategy meets the needs of the GPU and VPU for communication bandwidth and low memory access latency.
[0025] The DPU uses a 64AXI4@250MHz interface. To meet the real-time requirements of the DPU, the DPU occupies a dedicated display bus, while the slave side has only one DDR to meet the DPU's real-time and bandwidth requirements.
[0026] like Figure 2 The diagram shows the topology of a heterogeneous multi-core processor bus architecture, where each AXI interface is configurable. The processor bus uses a synchronous design with only one global clock. The host interface is configured with 10 read / write transmit transactions and 22 write transactions, respectively. The QoS priority type is host-controlled and unlocked transfer, and the QoS policy uses an arbitration strategy of TRR and LR. Register slices can be inserted to resolve timing issues; this configuration is optional. Furthermore, the host interface accessing DDR uses ThinLinks, and the switch interface type is AXI.
[0027] The multimedia bus and display bus in a heterogeneous multi-core processor can be configured similarly to the processor bus. The QoS priority type of the multimedia bus host interface is host-controlled and unlocked transmission, and the QoS policy adopts an OTT strategy. The Switch structure is referenced from... Figure 1 The host interface of the display bus has a QoS priority type of host control and unlocked transmission, and the QoS policy adopts the LR arbitration strategy.
[0028] Configure the global ID and optimization options to obtain the optimal bus architecture. Start RTLValidation to validate the RTL code and ensure the accuracy of the configuration parameters.
[0029] Each heterogeneous processor core is interconnected with the host interface of the cross-matrix bus via an asynchronous bridge with adjustable FIFO depth. The FIFO reassembles all signals of the channel and sends them to the host interface simultaneously. The FIFO at the host interface then disperses the signals to ensure reliable transmission of channel information. Clock gating, dynamic voltage, and frequency adjustment can effectively reduce power consumption as needed.
[0030] This system employs an interlocked bidirectional asynchronous communication bridge. The width of the handshake signal automatically adjusts based on transmission conditions. There is no common clock reference between the master and slave sides, eliminating the need for strict timing relationships between heterogeneous cores and cross-matrix systems. This asynchronous bridge effectively matches the bandwidth of heterogeneous cores and cross-matrix buses. Its detachable nature simplifies the interface between heterogeneous cores and cross-matrix systems, and its simple and efficient wiring facilitates physical timing convergence.
[0031] The NIC-400 cross-connect matrix bus is compatible with various AMBA protocol bus interfaces for both master and slave devices. The master interface is connected to the slave interface through a multi-level switch, which effectively reduces the communication lines between the master and slave devices. This reduction in interconnects effectively increases bus frequency and solves bandwidth issues, resulting in overall performance better suited to the needs of heterogeneous multi-core applications.
[0032] like Figure 3 As shown in the table, the priority policies and QoS values for each type of host are listed. Based on this table, the arbitration policy for each host interface is configured to meet the bandwidth, latency, and other requirements of each host.
[0033] like Figure 4 As shown, configure the QVN feature for each host interface, assigning a virtual network to each interface and configuring a pre-allocated token policy to ensure that each transaction is received by the slave. Configure QVN to prevent transaction blocking at the shared switch.
[0034] Advanced Quality of Service (QoS) and the QVN-400 virtual network are configured to efficiently manage data transmission and meet the acceptable bandwidth and latency constraints of the hosts. The CPU is a latency-sensitive host, the GPU is a transactional host, and the DPU (Display Processing Unit) is a real-time host. Therefore, the QoS modulation strategy first ensures minimal latency for real-time hosts, then minimizes latency for latency-sensitive hosts, and finally allocates bandwidth to transactional hosts.
[0035] The CPU interface employs TRR and LR strategies to ensure the quantity of OT transfers and reduce transaction latency. The GPU interface uses an OTT strategy to ensure the quantity of synchronous OT transfers and prevent transaction blocking. The real-time DPU interface uses an LR strategy to ensure the real-time performance of transactions. QoS priority values are derived from host transaction auxiliary information; typically, DPUs and other real-time display-related devices have the highest QoS values, GPUs and other hosts with numerous OT transactions have the lowest QoS values, and the CPU's QoS is in the middle range.
[0036] QVN virtual networks prevent interface congestion to ensure that every transaction can be received. Each transaction sent by a host must obtain a token from the corresponding slave. Each host is configured with pre-allocated tokens.
[0037] Furthermore, the general-purpose processor employs a symmetric multi-core architecture with a multimedia acceleration unit. This architecture improves transactional parallel processing capabilities and performance while effectively reducing power consumption. To maximize the performance of the general-purpose processor, the separate L1 caches all use a 32KB set-associative mapping structure, along with a 1MB L2 cache for multi-core shared caching. An accelerator coherence port is configured to maintain cache coherence across cores. Moreover, this port allows sharing cache contents with other modules and supports all standard read and write operations without requiring additional coherence modules, thus improving the overall performance of the multi-core processor without increasing power consumption.
[0038] The NEON multimedia acceleration unit of the general-purpose processor can assist the VPU in completing fast decoding operations of various encoding standards. NEON improves the performance of complex video codecs by 60-150%.
[0039] NEON's clock and power supply can be gated independently, resulting in significant savings in both dynamic and static power consumption.
[0040] Deep task partitioning and scheduling are employed to hide the communication overhead between the general-purpose processor and heterogeneous cores. Sub-threads are further split, with some sub-tasks executed by the general-purpose CPU and others by other heterogeneous cores. The CPU is responsible for task scheduling, while other heterogeneous cores are responsible for accelerating computation. Therefore, GPUs, VPUs, and data acquisition devices with high data correlation can be placed on the same bus to reduce inter-access latency.
[0041] The increase in master-slave interfaces leads to an increase in fan-out, capacitors, and wiring, necessitating the addition of registers to increase the bus frequency. Without reducing the number of master over-the-air (OT) accesses, access latency increases dramatically. Therefore, the NIC-400 cross-connect matrix employs a 2x4 cross-connect matrix configuration with single-stage switches in series.
[0042] The NIC-400 features a cross-matrix clock synchronization system, with asynchronous bridges decoupling the connections between Master and Slave interfaces. The DDR controller is synchronized with the cross-matrix design to improve communication bandwidth and reduce memory access latency. The interface has a 128-bit width, and the CPU core bus frequency is 667MHz to meet the DDR frequency requirements.
[0043] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A heterogeneous multi-core processor architecture based on NIC-400 crossbar, characterized in that, The application relates to a heterogeneous processor architecture. The heterogeneous processor cores include CPU cores for complex computation, GPU for high-throughput parallel computation, VPU for supporting multiple video decoding, and DPU for supporting multiple format display; The heterogeneous processor cores are connected with a NIC-400 cross matrix through a detachable asynchronous bridge to realize unified addressing and asynchronous communication; the master side and the slave side of the asynchronous bridge each have FIFOs of different clock domains, and five channels of the AXI interface each have independent FIFOs; The NIC-400 cross matrix adopts a full-asynchronous clock design and supports independent configuration of the bit width and frequency of the interfaces of the heterogeneous cores; the NIC-400 cross matrix comprises routing nodes corresponding to network interfaces one by one, and a cutoff data feedback connection line is arranged between adjacent nodes; the cross matrix has multiple configurable parameters to meet the PPA demand of a system, and the global ID bit width is adjusted according to the OT demand of each host to reduce the average memory access delay; The architecture supports multi-node parallel computation and is designed for physical layer timing convergence and wiring optimization.
2. The NIC-400 crossbar based heterogeneous multi-core processor architecture according to claim 1, wherein: The bit width and frequency of the interfaces of the heterogeneous cores based on the NIC-400 cross matrix can be arbitrarily configured.
3. The NIC-400 crossbar based heterogeneous multi-core processor architecture according to claim 1, wherein: The asynchronous bridge is based on a handshake form of detachable structure integration, and the width of each FIFO is the sum of the bit widths of all signals of the current channel.
4. The NIC-400 crossbar based heterogeneous multi-core processor architecture of claim 1, wherein: The CPU and the GPU in the heterogeneous cores adopt a shared memory mode, which can effectively reduce the communication overhead between the heterogeneous cores.
5. The NIC-400 crossbar based heterogeneous multi-core processor architecture according to claim 1, wherein: The NIC-400 has an expandable interface bandwidth.
6. The NIC-400 crossbar based heterogeneous multi-core processor architecture according to claim 3, wherein: The Switch and the Bridge in the NIC-400 cross matrix have different clock domains and configurable data widths, and adopt a hierarchical clock gating technology.
7. The NIC-400 crossbar based heterogeneous multi-core processor architecture according to claim 1, wherein: The NIC-400 has system-level characteristics and can automatically optimize the system architecture and dynamically realize parameter configuration of the NIC-400 through a GPV.
8. The NIC-400 crossbar based heterogeneous multi-core processor architecture according to claim 1, wherein: The NIC-400 also has an IP suite which can evaluate and optimize the area of a subsystem or a sub-module and the delay of host access to a slave through a built-in algorithm.
Citation Information
Patent Citations
Device and method for running MULTI-CORE SYSTEM and multi-core system
CN107918557A