Training Accelerator System Decoder for CPU-Free Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for deep learning training in AI systems face substantial overhead and resource waste due to cumbersome communication between accelerators, which is exacerbated by software-mediated coordination and limited scalability, especially in edge cloud computing environments.
Innovation Solution
The system decoder extends the accelerator inter-chip link (ICL) with programmable logic to provide transparent communication between AI appliances, abstracting physical and network channels, enabling seamless coordination across distributed devices without traversing memory tiers or requiring CPU intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If software-mediated coordination is used for accelerator communication, then ease of operation is improved, but device complexity and resource overhead increase
Solution Approach 1:
The patent introduces a system decoder as an intermediary component that mediates between accelerators and the CPU. The system decoder handles communication protocol translation, data routing, and coordination tasks, allowing accelerators to communicate directly without requiring complex software mediation. This resolves the contradiction by providing simplified operation through the intermediary while reducing software overhead.
Solution Approach 2:
The patent replaces software-mediated coordination with a hardware-based system decoder that operates at the hardware level. By substituting the software coordination mechanism with a dedicated hardware decoder, the system achieves faster communication with lower overhead, resolving the contradiction between ease of operation and device complexity.
2Adaptability or versatility
If data traverses memory tiers for accelerator communication, then compatibility is improved, but speed and latency performance worsen
Solution Approach 1:
The system decoder acts as an intermediary that enables direct accelerator-to-accelerator communication by translating different memory addressing schemes. This allows accelerators with different memory architectures to communicate efficiently without traversing CPU memory tiers, thus maintaining compatibility while improving speed.
Solution Approach 2:
The patent segments the communication path into dedicated accelerator interconnect channels that bypass the general-purpose memory hierarchy. By creating separate communication pathways for accelerator data exchange, the system achieves both compatibility across different accelerator types and high-speed data transfer.
3Ease of operation
If CPU intervention is required for accelerator coordination, then ease of operation is improved, but productivity and resource utilization worsen
Solution Approach 1:
The system decoder serves as an autonomous intermediary that handles accelerator coordination independently of the CPU. It manages data routing, protocol translation, and synchronization tasks without requiring CPU intervention, thereby maintaining operational simplicity while significantly improving training throughput and productivity.
Solution Approach 2:
The system decoder enables accelerators to self-coordinate and self-manage their communication needs. Each accelerator can independently initiate data exchange operations, and the system decoder automatically handles the coordination details, eliminating the need for CPU intervention and maximizing productivity.
4Speed
If existing high-speed links are used for accelerator interconnection, then speed is improved, but adaptability and scalability worsen
Solution Approach 1:
The system decoder provides universal functionality that can adapt to different accelerator types, communication protocols, and system configurations. It serves as a multi-functional interface that maintains high-speed communication while providing the flexibility and scalability needed for diverse accelerator interconnection scenarios.
Data Source
AI summary
There is disclosed an example of an artificial intelligence (AI) system, including: a first hardware platform; a fabric interface configured to communicatively couple the first hardware platform to a second hardware platform; a processor hosted on the first hardware platform and programmed to operate on an AI problem; and a first training accelerator, including: an accelerator hardware; a platform inter-chip link (ICL) configured to communicatively couple the first training accelerator to a second training accelerator on the first hardware platform without aid of the processor; a fabric ICL to communicatively couple the first training accelerator to a third training accelerator on a second hardware platform without aid of the processor; and a system decoder configured to operate the fabric ICL and platform ICL to share data of the accelerator hardware between the first training accelerator and second and third training accelerators without aid of the processor.


