Hierarchical Register Files for Scalable VLIW Processors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Wide-issue VLIW processors face scalability issues due to limitations in inter-cluster communication and register file structure, particularly in deep submicron silicon processes, where implicit operand transfer hampers performance across a large number of functional unit clusters.
Innovation Solution
A VLIW processor with a hierarchical structure that uses explicit control in the instruction stream for inter-cluster communication, featuring a cluster-level switch network for data transfer between sub-clusters, allowing data permutations and operand broadcasting between sub-clusters, global register files, and memory.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a large centralized register file is used to store variables for all functional units, then all functional units can access operands, but the structure does not scale as the issue width grows
Solution Approach 1:
The patent divides the centralized register file into multiple distributed register files, with each cluster having its own local register file. This segmentation allows the system to scale by adding more clusters without requiring a single large centralized register file, thus resolving the scalability issue while maintaining manageable complexity through modular organization.
Solution Approach 2:
The patent transitions from a single-dimensional centralized register file structure to a multi-dimensional hierarchical structure where register files are distributed across multiple clusters arranged in a hierarchical manner. This dimensional change enables scalable expansion by adding clusters in additional dimensions rather than expanding a single centralized structure.
2Productivity
If implicit cross-cluster path connection is used for inter-cluster communication, then operand transfer between clusters is enabled, but performance is impeded by interconnect delays in deep submicron silicon processes
Solution Approach 1:
The patent implements preliminary action by performing data permutation operations in the switch network before data transfer, and by using explicit control instructions that are issued in dedicated slots parallel to computation instructions. This preliminary organization of data paths and parallel execution of transfer operations reduces the effective delay impact on performance.
Solution Approach 2:
The patent introduces a switch network as an intermediary between clusters, which actively manages and optimizes data transfer paths. This intermediary enables efficient routing and permutation operations, reducing the negative impact of interconnect delays by providing intelligent data movement management rather than relying on simple implicit connections.
3Speed
If implicit operand transfer with short latency functional unit operation is used, then fast computation is achieved, but inter-cluster communication is limited for a large number of functional unit clusters
Solution Approach 1:
The patent merges data transfer operations with computation operations by allowing transfer instructions to issue in parallel with computation instructions through dedicated instruction issue slots. This combining of transfer and computation capabilities enables both fast computation and scalable inter-cluster communication to occur simultaneously without one limiting the other.
Solution Approach 2:
The patent implements dynamic control of data transfer through explicit instructions in the instruction stream, allowing the system to adaptively manage inter-cluster communication based on computational needs. This dynamic approach enables the system to maintain fast computation speeds while providing flexible, scalable communication capabilities across a large number of clusters.
Data Source
AI summary
A VLIW processor has a hierarchy of functional unit clusters that communicate through explicit control in the instruction stream and store data in register files at each level of the hierarchy. Explicit instructions transfer values between sub-clusters through a cluster level switch network. Transfer instructions issue in dedicated instruction issue slots in parallel with instructions that perform computation in functional units. The switch network can perform permutations on the data being moved. The switch network enables for operands to be broadcast between the sub-clusters, global register file and memory.


