Integrated Circuit Star Bus Architecture for Multi-Core Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current CPU architectures require increasing die surface area for linear clock speed improvements, achieving only 10% to 20% of their theoretical maximum performance under typical conditions.
Innovation Solution
A processing system with multiple processing cores coupled via star buses and dedicated random access memories, utilizing a hierarchical structure with unidirectional star buses and interleaved shared RAMs to enhance memory access speed and reduce die area usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If a single CPU architecture is used to increase clock speed, then processing speed is improved, but die surface area increases linearly without corresponding increases in processed instructions
Solution Approach 1:
The patent divides the processing system into multiple independent processing cores (first group and second group of processing cores) that can operate simultaneously. Each core handles a portion of the instruction stream, enabling parallel processing. This segmentation allows the system to achieve higher total instruction throughput without requiring a single core to operate at impractically high clock speeds, thereby improving processing capacity without linearly increasing die area for a single high-speed core.
Solution Approach 2:
The patent transitions from a single-dimensional sequential processing model to a multi-dimensional parallel processing architecture. By organizing cores into multiple groups that can execute instructions concurrently, the system adds a temporal parallelism dimension. The instruction stream is distributed across multiple cores in different groups, allowing simultaneous execution across multiple processing pipelines, thus achieving higher throughput without proportionally increasing the clock speed of individual cores.
2Device complexity
If a single CPU architecture is used, then design is simplified, but theoretical maximum performance is not achieved and only 10% to 20% utilization is obtained under typical conditions
Solution Approach 1:
The processing system is segmented into multiple independent core groups that can be independently configured and activated. Each group of processing cores can handle specific instruction streams, allowing the system to scale processing capacity by activating additional groups as needed. This segmentation enables better utilization of processing resources under typical operating conditions by distributing workloads across multiple cores rather than relying on a single overloaded core.
Solution Approach 2:
The patent creates a universal processing architecture where multiple groups of processing cores can handle various types of instruction streams. The system is designed to be multi-functional, capable of executing different instruction sets and handling diverse computational tasks across multiple core groups. This universality allows the system to achieve higher theoretical maximum performance by efficiently utilizing all available cores for different types of processing tasks simultaneously.
3Speed
If dedicated random access memories are assigned to each processing core, then memory access speed is improved, but die area increases
Solution Approach 1:
The memory system is segmented into dedicated random access memories assigned to specific processing core groups. Each group of processing cores has its own dedicated memory resources, allowing simultaneous independent memory access without contention. This segmentation of memory resources ensures that memory access speed is maintained at high levels for each core group while the overall memory capacity is distributed across the die area rather than concentrated in a single large memory block.
Solution Approach 2:
The patent implements a hierarchical memory architecture that adds a spatial dimension to memory organization. By distributing dedicated memories across multiple processing core groups rather than using a single centralized memory, the system creates a multi-dimensional memory access structure. This allows parallel memory access operations to occur simultaneously in different core groups, effectively increasing total memory bandwidth without requiring a proportional increase in total memory capacity or die area.
4Productivity
If star buses are used to couple processing cores, then communication efficiency is improved, but the system complexity increases with hierarchical bus structures
Solution Approach 1:
The bus system is segmented into multiple star-bus structures, each serving a specific group of processing cores. Rather than implementing a single complex centralized bus for all cores, the system divides the communication infrastructure into separate star buses for different core groups. Each star bus provides efficient communication within its group, and higher-level buses interconnect the groups. This segmentation reduces the complexity of individual bus structures while maintaining communication efficiency within each segment.
Solution Approach 2:
The patent implements a hierarchical bus architecture that adds a vertical dimension to the communication structure. Multiple star buses operate at different hierarchical levels, with lower-level buses serving individual core groups and higher-level buses interconnecting these groups. This multi-dimensional bus hierarchy allows efficient communication within groups at the lower level while providing inter-group communication at higher levels, thereby maintaining overall communication efficiency without requiring a single overly complex centralized bus structure.
Data Source
AI summary
A processing system on an integrated circuit includes a group of processing cores. A group of dedicated random access memories are severally coupled to one of the group of processing cores or shared among the group. A star bus couples the group of processing cores and random access memories. Additional layer(s) of star bus may couple many such clusters to each other and to an off-chip environment.


