NoC Processing Clusters for Fast Microcode Distribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer systems face challenges in efficiently processing massive data due to limitations in processor architecture and microcode distribution, leading to inefficiencies in power consumption, area usage, throughput, resource utilization, and processing speed.
Innovation Solution
A Network on Chip (NoC) processing system with microcode-programmable Processing Elements (PEs) organized in clusters, each with a Cluster Controller and Cluster Memory, enabling efficient microcode distribution and request-based transfer, allowing swift changes in programmable functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If microcode is distributed to all processing elements centrally, then system control is simplified, but data transfer time and power consumption increase
Solution Approach 1:
The patent divides the processing elements into multiple clusters, each with its own cluster controller that can independently receive and distribute microcode. This segmentation allows parallel microcode distribution to different clusters, reducing overall distribution time while maintaining simplified centralized control through the root node.
Solution Approach 2:
The patent implements a two-phase microcode distribution approach where microcode is first pre-distributed to cluster memories during an initialization phase, and then processing elements can quickly access it during runtime. This preliminary action eliminates the need for repeated microcode transmission during operation.
2Productivity
If more processing elements are added to increase processing capacity, then data processing throughput improves, but chip area and power consumption increase
Solution Approach 1:
The patent organizes processing elements into clusters with shared cluster memories and controllers. This segmentation allows multiple PEs to share common resources (cluster memory, controller interfaces), increasing processing throughput without proportionally increasing chip area, as resources are reused across the cluster.
Solution Approach 2:
The patent merges multiple processing elements into clusters that share common infrastructure including cluster memories, controllers, and interconnect resources. This combining approach achieves high processing capacity while reducing overall chip area compared to fully distributed architectures where each PE has dedicated resources.
3Speed
If processing elements have dedicated microcode memory, then microcode access speed improves, but chip area and cost increase
Solution Approach 1:
The patent combines multiple processing elements into clusters that share common cluster memories for microcode storage. This merging approach maintains fast microcode access speeds through localized shared memory while reducing total chip area compared to giving each PE its own dedicated microcode memory.
Solution Approach 2:
The patent pre-loads microcode into cluster memories during system initialization, allowing processing elements to quickly access pre-positioned microcode without real-time transmission delays. This preliminary action optimizes access speed while avoiding the need for extensive dedicated storage at each PE.
4Ease of manufacture
If fixed processor architecture is used, then manufacturing simplicity is maintained, but adaptability to different applications decreases
Solution Approach 1:
The patent implements microcode-programmable processing elements that can dynamically change their functionality by loading different microcode from cluster memories. This dynamic reconfigurability allows the same hardware architecture to adapt to different applications and processing tasks while maintaining manufacturing simplicity of a fixed physical structure.
Solution Approach 2:
The patent enables functional changes in processing elements by modifying microcode parameters and instructions stored in cluster memories. This allows the same physical processor architecture to perform different operations by changing software-defined parameters, achieving application adaptability without hardware reconfiguration.
Data Source
AI summary
A Network on Chip, NoC, processing system configured to perform data processing. The NoC processing system is configured for interconnection with a control processor connectable to said NoC processing system. The NoC processing system comprises a plurality of microcode-programmable Processing Elements, PEs, organized in multiple clusters, each cluster comprising a multitude of said programmable PEs, the functionality of each microcode-programmable PE being defined by internal microcode in a microprogram memory associated with the PE. The clusters of programmable PE are arranged on a Network on Chip, NoC, the NoC having a root and a plurality of peripheral nodes, wherein the clusters of PE are arranged at peripheral nodes of the NoC, and the NoC is connectable to the control processor. Each cluster further includes a Cluster Controller, CC, and an associated Cluster Memory, CM, shared by the multitude of programmable PE within the cluster.


