Dynamic Graph and Communication Scheduling for Reconfigurable Processors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional instruction set architectures struggle to efficiently execute complex applications like machine learning and artificial intelligence due to limitations in performance and efficiency, especially in multicore processors and GPUs, necessitating a more flexible and adaptable reconfigurable processor architecture.

Innovation Solution

A system comprising a reconfigurable processor, a runtime execution engine, a graph scheduler, and a communication scheduler that generates new schedules based on user-defined and static schedules to optimize the execution of dataflow graphs on an array of reconfigurable units, enabling dynamic changes and improved resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional instruction set architectures are used, then hardware simplicity is maintained, but computational efficiency and adaptability for complex applications deteriorate

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidprocessor architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements dynamic reconfiguration of the processor architecture at runtime, allowing the hardware structure to adapt to different computational workloads. The reconfigurable units can be dynamically connected and disconnected, and their configurations can be changed without requiring physical hardware modifications, thus achieving high computational efficiency for complex applications while maintaining manageable system complexity through software-based control.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the operational parameters of the reconfigurable units dynamically based on the execution requirements. By modifying connection topologies, data flow paths, and unit configurations through the runtime execution engine, the system optimizes computational efficiency for specific workloads without requiring permanent architectural changes, effectively resolving the contradiction between performance and complexity.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If fixed processor architecture is used, then device complexity is low, but adaptability to different workloads deteriorates

Engineering Contradiction:
Improveworkload adaptabilityVSAvoidprocessor configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The processor architecture employs dynamic reconfiguration mechanisms where the topological connections between reconfigurable units can be changed at runtime. The runtime execution engine manages these changes by loading new configurations and updating the physical connectivity, enabling the same hardware to adapt to different workloads (machine learning, high-performance computing, etc.) without requiring multiple fixed architectures, thus achieving high versatility with controlled complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The reconfigurable units are designed with universal interfaces and standardized connection protocols that allow them to serve multiple functions across different applications. By combining these universal units in various topologies under the guidance of the runtime execution engine, the system achieves multi-functionality and workload adaptability without requiring specialized hardware for each application type, effectively balancing versatility with complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If dynamic reconfiguration is enabled, then workload adaptability improves, but execution time for configuration changes deteriorates

Engineering Contradiction:
Improveruntime reconfiguration capabilityVSAvoidconfiguration loading time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-loading configuration data and pre-computing reconfiguration paths into the runtime execution engine before actual workload changes are needed. When a workload change occurs, the engine can quickly retrieve and apply pre-prepared configurations rather than computing and loading configurations from scratch, significantly reducing the time loss associated with dynamic reconfiguration while maintaining full adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The runtime execution engine incorporates feedback mechanisms that monitor workload characteristics and automatically select optimal reconfiguration paths based on historical performance data. By learning from previous configuration changes and workload patterns, the system can predict the best reconfiguration strategies and minimize execution time delays, effectively balancing adaptability with speed through intelligent feedback-driven decision making.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250258678A1Flexible Runtime Execution and Communication Scheduler for Reconfigurable Processors
Publication Date: 2025.08.14 SAMBANOVA SYSTEMS INC
  • US20250258678A1 patent drawing
  • US20250258678A1 patent drawing
  • US20250258678A1 patent drawing

AI summary

A system including a reconfigurable processor, a runtime execution engine, a graph scheduler, and a communication scheduler is presented. The graph scheduler and the communication scheduler receive a dataflow graph and static schedules of graph and communication operations from a compiler. The graph scheduler and the communication scheduler generate new schedules of graph and communication operations based on user-defined schedules of graph and communication operations and the static schedules of graph and communication operations. The runtime execution engine uses the dataflow graph and the new schedules of graph and communication operations to configure an array of reconfigurable units in the reconfigurable processor for execution of the dataflow graph. The present technology also relates to a method of operating such a system, and to a non-transitory computer-readable storage medium including instructions that, when executed by a processing unit, cause the processing unit to operate such a system.