Heterogeneous multi-core CPU with discrete GPU super-chip processor system for accelerating high-performance computers
The integration of a heterogeneous multi-core CPU with a discrete GPU via a high-speed interface addresses inefficiencies in traditional architectures by optimizing CPU core clusters and workload distribution, achieving efficient power use and improved performance metrics.
Patent Information
- Application Number
- DE202025103050
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-06-02
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2035-06-30
AI Technical Summary
Traditional homogeneous computing architectures based on symmetric multi-core CPUs struggle to balance high throughput rates and low power consumption, while loosely coupled CPU-GPU systems face inefficiencies in communication latency, synchronization, and workload distribution, with a persistent bottleneck in harmonizing control flow and data parallel workloads.
A novel super-chip processor architecture integrates a heterogeneous multi-core CPU with discrete GPU via a high-speed PCI Express interface, organizing CPU cores into power-based clusters and utilizing a unified cache structure, with a thread director module for intelligent workload distribution, ensuring optimal resource allocation and minimal latency.
The architecture achieves efficient power consumption and scalability, significantly improving execution time, throughput, and instructions per clock cycle, overcoming inefficiencies of traditional systems and enabling seamless task sharing between CPU and GPU.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of the Invention:The invention relates to an HPC processor architecture having a heterogeneous multi-core central processing unit (CPU) integrated with a discrete graphics processing unit (GPU) via a multichip platform.BACKGROUND OF THE INVENTION;The exponential growth of data intensive applications and computational workloads in recent years has placed non-exemplary demands on modern processor systems. High performance computing (HPC) areas such as scientific simulations, artificial intelligence (AI), machine learning, weather forecasts, financial modelling and real time analyses require not only immense computing power but also intelligent resource allocation and optimal energy efficiency. In parallel, the variety of computing tasks - from highly parallelizable processes to sequential operations - requires architectures capable of executing different classes of instructions in the most efficient manner. Traditional homogeneous computing architectures based on symmetric multi-core CPUs with uniform instruction processing capabilities are increasingly constrained in such heterogeneous requirements.As semiconductor technologies progress, multi-core CPUs have become the standard, but have had difficulty in accommodating high throughput rates and low power consumption. These CPUs have been optimized for either peak power or energy efficiency, but rarely for both simultaneously. Simultaneously, discrete graphics processors (GPUs) began to develop from graphics-specific processors to powerful, universal parallel computing engines. Their architecture, consisting of thousands of light cores capable of executing thousands of threads in parallel, made them ideal for data parallel computations. However, their integration with CPUs in computing systems has often been loosely coupled, which introduced inefficiencies in communication latency, synchronization of memory access, and workload distribution.Efforts to integrate CPUs and GPUs into heterogeneous computing platforms have shown promising results, particularly in mobile and embedded computing environments where system-on-a-chip (SoCs) with integrated graphics cores may balance performance and efficiency. However, such closely integrated systems often suffer from thermal throttling, limited memory bandwidth, and suboptimal scaling due to shared resources in desktop, server, and HPC systems. Moreover, the challenge of harmonizing control flow intensive workloads (suitable for CPUs) and data parallel workloads (ideal for GPUs) within the same processing pipeline remains a persistent bottleneck.Another complication is the discrepancy in architectural philosophy between CPUs and GPUs. CPUs are distinguished by low-latency execution of single or easily parallelized instructions, thanks to deep pipelines, large caches and advanced branch prediction. GPUs, on the other hand, are throughput-oriented devices that are optimized for executing the same instruction over large sets of data with high arithmetic intensity. While offloading workloads from the CPU to the GPU is conceptually simple, efficient interworking between the two remains difficult to grasp due to interconnect speed, memory coherency models, and scheduling policies constraints.Recognizing these challenges, researchers and system architects have begun to explore the ability to develop "super-chip" solutions that combine heterogeneous CPU and GPU architectures in a synergistic and scalable manner. The idea is to coexist both architectures in a tightly integrated system in which the CPU can plan and orchestrate tasks while swapping parallel processing to the GPU seamlessly and with minimal overhead. This approach requires innovations in architectural integration, high speed interconnections, cache coherency, workload distribution algorithms, and memory bandwidth utilization.In response to these systemic limitations and the increasing need for more efficient heterogeneous computing solutions, the present invention introduces a novel super-chip processor architecture. This design includes a hybrid multi-core CPU that includes both high power and low power cores that operate in conjunction with a discrete high power GPU that is connected via a high speed PCI express interface. This architectural approach aims to achieve a fine-grained balance between performance, scalability and energy consumption. Through smart core cluster management and parallel execution capabilities, the invention removes many years of inefficiencies in heterogeneous computing systems and opens new ways to execute modern HPC workloads with higher efficiency and lower latency.Summary of the Invention:The invention described herein relates to a sophisticated and highly efficient computer architecture specifically tailored to the complex requirements of high performance computer applications. In the core, the system consists of a heterogeneous multi-core CPU with eight processing cores, which are clearly grouped into performance-oriented and efficiency-oriented clusters. This targeted organization reflects an architectural philosophy that accommodates processor capabilities with the various features of modern computational loads. The high performance cluster, often referred to as a "big" core cluster, includes modern cores such as cortex A-75, cortex A-65, cortex X3, and cortex X4. These cores are designed to operate at high frequencies and support simultaneous multi-threading, making them ideal for instruction-intensive, sequential tasks that require fast execution.In contrast, the cluster of "small" cores consists of more energy efficient cores such as cortex A-55, cortex A-15, cortex A-32, and cortex A-34. These cores consume less power and are optimized for processing background tasks, low-priority threads, and event driven processes. Each core in the CPU is equipped with its own dedicated level 1 and level 2 caches that allow low latency access to frequently used data and instructions. At a higher level, all eight cores share a unified last level cache (LLC) that facilitates inter-core communication and memory coherency, minimizing contention and improving throughput.This heterogeneous CPU is integrated over a PCIe x16 interface with a discrete, high performance GPU, which ensures a fast and efficient communication channel between the two processing units. The GPU itself is a parallel processing wounding facility and has a scalable array of processing clusters that consists of graphics processing cores (GPCs), texture processor clusters (TPCs), and a series of streaming multiprocessors (SMs). These components operate in unison to execute thousands of threads simultaneously, whereby the GPU is particularly well to handle data-parallel tasks such as image rendering, machine learning, and matrix operations. Moreover, a specialized thread director module within the GPU dynamically allocates workloads to different cores and clusters based on current workload and priority to ensure optimal resource allocation and minimize processing delays.Support of the parallel execution capabilities of the GPU is provided by a high bandwidth memory block (HBM) that enables fast data sampling and efficient data streaming during intensive computing tasks. The presence of a discrete GPU instead of an integrated also means that the system can avoid thermal and power bottle-outs, which are typically associated with on-chip integration. This modular design allows both the CPU and GPU components to operate at their respective optimal performance levels without causing resource conflicts or thermal disturbances.The novel cooperation between the CPU and the GPU in this architecture enables task sharing based on the execution features. The CPU, acting as a central scheduler, determines the type of incoming tasks and allocates them to either the corresponding CPU core cluster or the GPU. Sequential and control intensive tasks remain within the area of the CPU while highly parallelizable workloads are transmitted to the GPU for acceleration. This sharing not only provides for faster task completion, but also maximizes energy efficiency by only activating the components that are best suited to the particular task.To validate the performance of the proposed system, the entire architecture was modeled and simulated using the GPGPUSIM simulation environment. The results of this simulation confirmed that the invention significantly outfits the current commercial heterogeneous computational models of manufacturers such as Intel and NVIDIA. Important performance metrics such as execution time, throughput, and per clock cycle (IPC) all exhibited significant improvement, demonstrating the ability of the invention to effectively speed up HPC workloads.In summary, the invention presents a coherent and future-oriented approach to computer architecture by seamlessly integrating a heterogeneous CPU with a powerful discrete GPU over a high-speed link. This design not only eliminates many years of inefficiencies in heterogeneous systems, but also bases future scalable and energy efficient computing platforms, particularly in the context of high performance and data intensive applications.Brief Description of the DrawingsFigure 1 shows a block diagram of the system of the invention.DETAILED DESCRIPTION OF THE INVENTIONThe present invention relates to a novel processor system architecture that has been developed to meet the increasing demands on computational and parallel processing capabilities of modern high performance computational loads (HPC). It introduces synergistic integration of a heterogeneous multi-core central processing unit (CPU) with a discrete graphics processing unit (GPU) interconnected via a high speed communication interface and forming a unitary processor-multichip platform. This architecture is designed not only to increase raw performance, but also to provide efficient power consumption, dynamic workload handling, and scalability in data intensive environments.The focus of the invention is on the central processing unit (CPU) which comprises a total of eight cores divided into two different power classes based on its intended application: four high power cores and four power saving cores. These are collectively orchestrated to manage the computational diversity ranging from low latency, sequential task processing to energy efficient background processing. The high performance cores, also referred to as "big" core clusters, consist of cortex A-75, cortex A-65, cortex X3, and cortex X4 cores. These cores are equipped with wide instruction pipelines, out-of-order execution engines, and aggressive branch prediction algorithms that allow them to process sophisticated workloads at high clock frequencies. Moreover, these cores are configured to support simultaneous multi-threading (SMT), thereby enabling execution of two threads per core simultaneously. This design increases instruction throughput rate and significantly conceals latency, which is particularly advantageous for compute-intensive applications.The supplemental cluster, referred to as the "small" core cluster, consists of cortex A-55, cortex A-15, cortex A-32, and cortex A-34 cores. These cores are optimized for energy efficiency rather than peak power. They are mainly responsible for executing background processes, handling event-driven interrupts and managing the workload of the peripheral interfaces. By assigning less demanding operations to these clusters, the system saves energy without compromising responsiveness or throughput. This sharing of processing responsibilities is fundamental to the energy efficient scheduling system of the invention, in which the computational loads are dynamically allocated to the most suitable processing unit based on the availability of resources and the execution characteristics.Each CPU core in the architecture has its own level 1 (L1) instruction and data caches and an independent level 2 (L2) cache to reduce latency associated with memory accesses. These private caches are designed to have low latency and high speed, thereby ensuring that each core operates with minimal dependency on external memory hierarchies for frequently used data. Above these private cache layers is a shared last level cache (LLC) accessible to all cores within the CPU complex. The LLC acts as a unified cache level that not only improves data locality, but also facilitates inter-core communication and data exchange, which is critically important for workloads spanning multiple threads or cores. Communication between the cores and the shared cache is orchestrated over a high speed ring network that provides low latency, high band data transfer across the processor chip.The CPU is connected to system memory via a dedicated memory bus. The memory subsystem uses double data rate random access memory (DDRAM) which serves as main memory for the execution of processes and temporary data storage. Efficient memory access mechanisms, cache prefetch algorithms, and parallel memory access strategies are incorporated into the system to ensure that memory bottle bottle bottle necks are minimized even under peak load conditions. The system bus also connects the CPU to peripherals and other hardware components required for complete system integration.The true novelty of the invention is evidenced by the seamless integration of a discrete GPU into the processor architecture. Unlike conventional systems that loosely couple GPUs to CPUs and often result in significant latency and underutilization, the present invention proposes a highly coordinated approach. The discrete GPU is connected to the CPU via a high-speed PCI Express (PCIe)-x16 interface that supports low-latency communication and high-speed data exchange. This interface acts as a communication backbone that allows the CPU to delegate data-parallel and graphics-intensive workloads to the GPU in real-time, thereby enabling dynamic load balancing and task parallelism.The discrete GPU itself is designed as a scalable computing unit having multiple layers of processing units. It includes a grid of graphics processing cores (GPCs) divided into texture processor clusters (TPCs) and further into streaming multiprocessors (SMs). Each SM accommodates multiple integer and floating point execution units, register files, load memory units, and warp schedulers. This hierarchical structure allows enormous thread level parallelism, in which thousands of concurrent threads can be efficiently executed across multiple SMs. Each GPC also includes shared memory and level 2 (L2) cache modules that reduce access to the global memory and improve kernel performance. The inclusion of multiple GPCs ensures that the GPU can scale horizontally to handle larger, more complex workloads.These processing units are supported by a comprehensive storage system based on high speed storage modules (HBM). A series of HBM stacks, each independently accessible, serve as the main storage pool for the GPU. This memory architecture has been chosen due to its superior throughput, minimal access delay, and ability to process large data sets essential for HPC workloads such as simulations, neural network training, or scientific computations. In addition, the layout of the GPU includes a thread director unit, a scheduler responsible for distributing incoming work to GPCs and SMs based on thread priority, resource availability, and thermal constraints. This component provides a balanced distribution of workload and optimum use of GPU resources.In operation, the CPU functions as a central control unit of the system. When a program or application is executed, the CPU analyzes the type of workload and evaluates whether it should best be executed on the CPU itself or on the GPU. For tasks that include sequential logic, branches, or low parallelism execution paths, the large core cluster of the CPU handles processing. On the other hand, workloads characterized by large area data parallelism, such as rendering, data mining, image processing, or neural network inference are packaged and forwarded to the GPU via the PCIe channel.Once the GPU receives these workloads, the thread director evaluates its properties and distributes them to the available SMs within the corresponding GPCs. Each SM executes its assigned threads simultaneously with warp-based scheduling and single instruction, multiple data (SIMD) paradigms. Intermediate results are stored in shared or global memory depending on the data access patterns. Upon completion, the GPU returns the results to the CPU, which may then perform post processing or coordinate further steps in the execution pipeline. This sharing of responsibilities allows the overall system to operate with high efficiency, minimum idle time, and reduced latency.One of the prominent strengths of the invention lies in its validation by simulation with GPGPUSIM, an established simulation tool for evaluating the GPU performance and the CPU-GPU integration. Through extensive benchmarking tests, the simulated model of the invention showed superior results in the most important performance metrics compared to leading commercial systems of Intel and NVIDIA. In particular, it has outful heterogeneous architectures with respect to instructions per clock cycle (IPC), throughput, and execution time. These improvements confirm the practical advantages of the proposed system in real HPC applications.Moreover, the modularity of the design allows scalable implementations for different market segments. For example, server-based implementations may integrate larger GPU rasters and extended HBM banks, while power efficient mobile or embedded applications may benefit from reduced versions of the CPU GPU architecture. Design inherent flexibility allows for future adaptations as newer generations of processor cores, interconnect standards, and memory technologies become available.In summary, the present invention introduces a deeply integrated heterogeneous processor system that combines the strengths of the multi-core CPU architecture with the massive parallelism of discrete GPUs. By organizing CPU cores into power level-based clusters, enabling high speed interconnection with a scalable GPU subsystem, and intelligent task sharing, the invention achieves an optimal balance between power, efficiency, and adaptability. The architecture overcomes many of the disadvantages observed with traditional CPU-GPU combinations and provides a new paradigm for powerful, latency-free, and scalable computing that is suitable for the most demanding applications in scientific, commercial, and industrial fields.List of reference characters100 System 101 Central Processing Unit (CPU) 102 Discrete Graphics Unit (GPU) 103 High-speed link 104 Thread Director
Claims
A heterogeneous processor system for accelerating high performance computational loads, comprising: o a central processing unit (CPU) (101) comprised of a plurality of processing cores, the cores being grouped into a high performance core cluster and a low performance core cluster, each core having private level 1 (L1) and level 2 (L2) caches, and all cores sharing a common load level cache (LLC); ◯ a discrete graphics unit (GPU) (102) comprised of a plurality of graphics processing cores (GPCs), texture processor clusters (TPCs), streaming multiprocessors (SMs), and high speed memory (HBM); a high speed link (103) configured as a PCIe x16 interface operatively connecting the CPU and the GPU; o a thread director (104) configured to dynamically schedule workloads across streaming multiprocessors in the GPU; wherein the CPU is configured to delegate parallelizable workloads to the GPU and maintain control flow intensive tasks to enable collaborative execution that improves performance, energy efficiency, and throughput for high performance computing applications.The heterogeneous processor system of claim 1, wherein the energy efficient core cluster comprises at least the cortex A-55, cortex A-15, cortex A-32, and cortex A-34 cores optimized for performing background tasks and low energy consumption.The heterogeneous processor system of claim 1, wherein the GPU comprises at least four graphics processing cores (GPCs), each including a plurality of streaming multiprocessors (SMs) capable of executing thousands of threads simultaneously.The heterogeneous processor system of claim 1, wherein the shared last level cache (LLC) is configured to support cache coherency across all CPU cores and facilitate data exchange between the cores.The heterogeneous processor system of claim 1, wherein the thread director monitors GPU core usage and dynamically allocates workloads based on execution priority, data dependencies, and available computing resources.The heterogeneous processor system of claim 1, wherein the simulation with GPGPUSIM shows an improved number of instructions per clock cycle (IPC), reduced execution time, and increased throughput compared to existing heterogeneous CPU GPU architectures.