Training Accelerator System Decoder for CPU-Free Data Sharing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for deep learning training in AI systems face substantial overhead and resource waste due to cumbersome communication between accelerators, which is exacerbated by software-mediated coordination and limited scalability, especially in edge cloud computing environments.

Innovation Solution

The system decoder extends the accelerator inter-chip link (ICL) with programmable logic to provide transparent communication between AI appliances, abstracting physical and network channels, enabling seamless coordination across distributed devices without traversing memory tiers or requiring CPU intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If software-mediated coordination is used for accelerator communication, then ease of operation is improved, but device complexity and resource overhead increase

Engineering Contradiction:
Improveease of coordinationVSAvoidsoftware mediation overhead
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces a system decoder as an intermediary component that mediates between accelerators and the CPU. The system decoder handles communication protocol translation, data routing, and coordination tasks, allowing accelerators to communicate directly without requiring complex software mediation. This resolves the contradiction by providing simplified operation through the intermediary while reducing software overhead.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces software-mediated coordination with a hardware-based system decoder that operates at the hardware level. By substituting the software coordination mechanism with a dedicated hardware decoder, the system achieves faster communication with lower overhead, resolving the contradiction between ease of operation and device complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If data traverses memory tiers for accelerator communication, then compatibility is improved, but speed and latency performance worsen

Engineering Contradiction:
Improvememory compatibilityVSAvoiddata exchange speed
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The system decoder acts as an intermediary that enables direct accelerator-to-accelerator communication by translating different memory addressing schemes. This allows accelerators with different memory architectures to communicate efficiently without traversing CPU memory tiers, thus maintaining compatibility while improving speed.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the communication path into dedicated accelerator interconnect channels that bypass the general-purpose memory hierarchy. By creating separate communication pathways for accelerator data exchange, the system achieves both compatibility across different accelerator types and high-speed data transfer.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If CPU intervention is required for accelerator coordination, then ease of operation is improved, but productivity and resource utilization worsen

Engineering Contradiction:
Improvecoordination simplicityVSAvoidtraining throughput
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system decoder serves as an autonomous intermediary that handles accelerator coordination independently of the CPU. It manages data routing, protocol translation, and synchronization tasks without requiring CPU intervention, thereby maintaining operational simplicity while significantly improving training throughput and productivity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system decoder enables accelerators to self-coordinate and self-manage their communication needs. Each accelerator can independently initiate data exchange operations, and the system decoder automatically handles the coordination details, eliminating the need for CPU intervention and maximizing productivity.

Inventive Principle:
Principle #25Self-service

4Speed

If existing high-speed links are used for accelerator interconnection, then speed is improved, but adaptability and scalability worsen

Engineering Contradiction:
Improvelink speedVSAvoidsystem scalability
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system decoder provides universal functionality that can adapt to different accelerator types, communication protocols, and system configurations. It serves as a multi-functional interface that maintains high-speed communication while providing the flexibility and scalability needed for diverse accelerator interconnection scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250272261A1System decoder for training accelerators
Publication Date: 2025.08.28 INTEL CORP
  • US20250272261A1 patent drawing
  • US20250272261A1 patent drawing
  • US20250272261A1 patent drawing

AI summary

There is disclosed an example of an artificial intelligence (AI) system, including: a first hardware platform; a fabric interface configured to communicatively couple the first hardware platform to a second hardware platform; a processor hosted on the first hardware platform and programmed to operate on an AI problem; and a first training accelerator, including: an accelerator hardware; a platform inter-chip link (ICL) configured to communicatively couple the first training accelerator to a second training accelerator on the first hardware platform without aid of the processor; a fabric ICL to communicatively couple the first training accelerator to a third training accelerator on a second hardware platform without aid of the processor; and a system decoder configured to operate the fabric ICL and platform ICL to share data of the accelerator hardware between the first training accelerator and second and third training accelerators without aid of the processor.