On-Chip Interface Unit for Flexible Accelerator Instruction Routing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current processing systems with integrated accelerators face challenges in efficiently managing compute-intensive tasks for complex machine learning and AI models, as existing configurations lack flexibility in accommodating varying computing demands.
Innovation Solution
A processing system with an interface unit that converts processor instructions into accelerator instructions and a bus network that routes these instructions to the appropriate accelerators, allowing for flexible configuration and management of multiple accelerators and processors, enabling efficient processing of various machine learning and AI tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If accelerators are connected through external bus (PCIe), then device complexity is reduced, but communication speed and processing efficiency deteriorate
Solution Approach 1:
The patent merges the accelerator and processor onto the same chip, integrating previously separate components into a unified system. This integration eliminates the need for external bus connections while providing high-speed internal communication pathways, thereby resolving the contradiction between reduced device complexity and improved communication speed.
Solution Approach 2:
The patent transitions from external bus communication (spatial separation) to on-chip integration (temporal/signal dimension optimization). By changing the dimensional relationship between components from externally connected to internally integrated, the system achieves both simplified architecture and enhanced communication performance.
2Speed
If accelerators are integrated on the same chip, then communication speed improves, but device complexity increases
Solution Approach 1:
The patent implements a unified instruction conversion unit that handles multiple instruction types and formats, and a flexible bus network that supports various communication patterns. This multi-functional design allows the system to manage complex on-chip integration through standardized, versatile components that reduce overall system complexity despite the integrated architecture.
3Productivity
If processor instructions are directly executed by accelerators, then processing efficiency improves, but instruction compatibility and adaptability deteriorate
Solution Approach 1:
The patent introduces an instruction conversion unit as an intermediary between the processor and accelerator. This mediator translates processor instructions into accelerator-specific instructions, enabling efficient execution while maintaining compatibility with various instruction sets. The conversion unit acts as a buffer that preserves processing efficiency without sacrificing adaptability.
Solution Approach 2:
The patent transforms instruction parameters and formats through the conversion unit, adapting processor instructions to match accelerator requirements. By dynamically changing instruction parameters during translation, the system achieves both high processing efficiency and broad instruction compatibility across different accelerator types.
Data Source
AI summary
Embodiments of this disclosure provide a processing system and an instruction transmission method. The instruction transmission method includes: receiving a processor instruction from a main processor communicatively coupled to an interface unit; generating an accelerator instruction corresponding to the processor instruction; determining a target accelerator corresponding to the accelerator instruction from a plurality of accelerators communicatively coupled to a first bus network; and transmitting the accelerator instruction to the target accelerator through the first bus network.


