Neural Processing Device Distributed L3 Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional neural processing units experience high latency and increased overhead due to centralized control processors managing synchronization signals, especially as the number of processing units and cores increases, limiting the efficiency of deep-learning tasks.
Innovation Solution
A neural processing device with multiple neural processors that generate and manage L3 sync targets, using shared memory and semaphore memories for synchronization, and a global interconnection for transmitting synchronization signals, allowing processors to perform synchronization independently without centralized control, thereby reducing latency and scheduling overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a centralized control processor manages synchronization signals, then synchronization control is simplified, but synchronization latency increases and control processor overhead increases
Solution Approach 1:
The patent divides the centralized synchronization control function into distributed synchronization units, where each processing unit has its own synchronization controller. This segmentation eliminates the single-point bottleneck of centralized control, allowing parallel synchronization operations that reduce latency while maintaining organized control through modular synchronization units.
Solution Approach 2:
The patent introduces synchronization signals as intermediaries that coordinate between processing units without requiring direct centralized control. These signals act as mediators that enable autonomous synchronization at each processing unit, reducing the control overhead on the main processor while maintaining coordination.
2Device complexity
If a centralized control processor manages synchronization signals, then synchronization coordination is centralized, but control processor overhead increases
Solution Approach 1:
Each processing unit is equipped with its own synchronization controller that autonomously manages synchronization without requiring continuous intervention from the main control processor. This self-service approach distributes the computational burden, reducing the energy consumption and overhead of the centralized control processor while maintaining synchronization coordination.
3Productivity
If more processing units and cores are included, then processing capability increases, but synchronization latency and overhead increase
Solution Approach 1:
The patent segments the synchronization control function across multiple processing units, allowing each unit to handle its own synchronization independently. This enables linear scaling of processing capability without proportional increases in synchronization latency, as each segment operates autonomously rather than waiting for centralized coordination.
Solution Approach 2:
The patent transitions from a single-dimensional centralized control model to a multi-dimensional distributed control architecture. By adding the dimension of local synchronization control at each processing unit, the system achieves scalable performance where processing capability can increase without proportionally increasing synchronization overhead.
Data Source
AI summary
A neural processing device is provided. The neural processing device comprises a plurality of neural processors, a shared memory shared by the plurality of neural processors, a plurality of semaphore memories, and global interconnection. The plurality of neural processors generates a plurality of L3 sync targets, respectively. Each semaphore memory is associated with a respective one of the plurality of neural processors, and the plurality of semaphore memories receive and store the plurality of L3 sync targets, respectively. Synchronization of the plurality of neural processors is performed according to the plurality of L3 sync targets. The global interconnection connects the plurality of neural processors with the shared memory, and comprises an L3 sync channel through which an L3 synchronization signal corresponding to at least one L3 sync target is transmitted.


