Dynamic Graph and Communication Scheduling for Reconfigurable Processors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional instruction set architectures struggle to efficiently execute complex applications like machine learning and artificial intelligence due to limitations in performance and efficiency, especially in multicore processors and GPUs, necessitating a more flexible and adaptable reconfigurable processor architecture.
Innovation Solution
A system comprising a reconfigurable processor, a runtime execution engine, a graph scheduler, and a communication scheduler that generates new schedules based on user-defined and static schedules to optimize the execution of dataflow graphs on an array of reconfigurable units, enabling dynamic changes and improved resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional instruction set architectures are used, then hardware simplicity is maintained, but computational efficiency and adaptability for complex applications deteriorate
Solution Approach 1:
The patent implements dynamic reconfiguration of the processor architecture at runtime, allowing the hardware structure to adapt to different computational workloads. The reconfigurable units can be dynamically connected and disconnected, and their configurations can be changed without requiring physical hardware modifications, thus achieving high computational efficiency for complex applications while maintaining manageable system complexity through software-based control.
Solution Approach 2:
The patent changes the operational parameters of the reconfigurable units dynamically based on the execution requirements. By modifying connection topologies, data flow paths, and unit configurations through the runtime execution engine, the system optimizes computational efficiency for specific workloads without requiring permanent architectural changes, effectively resolving the contradiction between performance and complexity.
2Adaptability or versatility
If fixed processor architecture is used, then device complexity is low, but adaptability to different workloads deteriorates
Solution Approach 1:
The processor architecture employs dynamic reconfiguration mechanisms where the topological connections between reconfigurable units can be changed at runtime. The runtime execution engine manages these changes by loading new configurations and updating the physical connectivity, enabling the same hardware to adapt to different workloads (machine learning, high-performance computing, etc.) without requiring multiple fixed architectures, thus achieving high versatility with controlled complexity.
Solution Approach 2:
The reconfigurable units are designed with universal interfaces and standardized connection protocols that allow them to serve multiple functions across different applications. By combining these universal units in various topologies under the guidance of the runtime execution engine, the system achieves multi-functionality and workload adaptability without requiring specialized hardware for each application type, effectively balancing versatility with complexity.
3Adaptability or versatility
If dynamic reconfiguration is enabled, then workload adaptability improves, but execution time for configuration changes deteriorates
Solution Approach 1:
The system performs preliminary actions by pre-loading configuration data and pre-computing reconfiguration paths into the runtime execution engine before actual workload changes are needed. When a workload change occurs, the engine can quickly retrieve and apply pre-prepared configurations rather than computing and loading configurations from scratch, significantly reducing the time loss associated with dynamic reconfiguration while maintaining full adaptability.
Solution Approach 2:
The runtime execution engine incorporates feedback mechanisms that monitor workload characteristics and automatically select optimal reconfiguration paths based on historical performance data. By learning from previous configuration changes and workload patterns, the system can predict the best reconfiguration strategies and minimize execution time delays, effectively balancing adaptability with speed through intelligent feedback-driven decision making.
Data Source
AI summary
A system including a reconfigurable processor, a runtime execution engine, a graph scheduler, and a communication scheduler is presented. The graph scheduler and the communication scheduler receive a dataflow graph and static schedules of graph and communication operations from a compiler. The graph scheduler and the communication scheduler generate new schedules of graph and communication operations based on user-defined schedules of graph and communication operations and the static schedules of graph and communication operations. The runtime execution engine uses the dataflow graph and the new schedules of graph and communication operations to configure an array of reconfigurable units in the reconfigurable processor for execution of the dataflow graph. The present technology also relates to a method of operating such a system, and to a non-transitory computer-readable storage medium including instructions that, when executed by a processing unit, cause the processing unit to operate such a system.


