Bus Device Computing Function Unit for Gradient Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed processing systems face communication congestion and bottlenecks during group communication in distributed deep learning, leading to reduced processing efficiency and increased latency.
Innovation Solution
The system incorporates a computing function unit in a bus device connected to computing devices and an interconnect device, with a DMA controller for managing data transfer, and a control unit for allocating learning jobs, enabling high-speed processing of gradient data without congestion by integrating computing functions within the bus device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the number of weights and sample data are increased to improve inference accuracy, then inference accuracy is improved, but communication time and deep learning time increase
Solution Approach 1:
The patent segments the aggregation process by introducing intermediate aggregation nodes that collect gradients from multiple computing devices before forwarding to the central aggregation processing node. This segmentation reduces the communication burden on the central node and enables parallel processing of gradient aggregation across multiple intermediate nodes, thereby reducing overall communication time while maintaining the ability to process large numbers of weights and sample data
Solution Approach 2:
The patent introduces a hierarchical aggregation structure that adds a spatial dimension to the gradient aggregation process. By organizing intermediate aggregation nodes in a hierarchical manner between computing devices and the central aggregation node, the system creates multiple communication pathways and reduces the dimensionality of communication bottlenecks, enabling efficient handling of increased data volumes
2Productivity
If multiple computing devices are mounted in distributed processing nodes at high density to improve processing performance, then processing performance is improved, but communication congestion occurs at the interconnect device
Solution Approach 1:
The patent introduces intermediate aggregation nodes as mediators between multiple computing devices and the central aggregation processing node. These intermediate nodes collect gradients from multiple computing devices locally before forwarding aggregated results to the central node, thereby reducing the communication load on the interconnect device and eliminating congestion while maintaining high processing performance across multiple computing devices
Solution Approach 2:
The patent segments the aggregation function across multiple intermediate aggregation nodes distributed among the processing nodes. Each intermediate node handles a subset of computing devices, dividing the total communication load into manageable segments that can be processed in parallel, thereby preventing congestion at any single interconnect device while maintaining high overall processing performance
3Adaptability or versatility
If integrated communication and distributed communication are performed frequently for aggregation processing, then deep learning processing is enabled, but processing time increases due to communication overhead
Solution Approach 1:
The patent performs preliminary gradient aggregation at intermediate aggregation nodes before the final aggregation at the central node. By pre-aggregating gradients from multiple computing devices at intermediate nodes, the system reduces the volume of data requiring communication in subsequent steps, thereby enabling distributed deep learning capability while significantly reducing the time consumed by communication overhead
Solution Approach 2:
The patent establishes a continuous hierarchical aggregation pipeline where intermediate nodes continuously collect and aggregate gradients from computing devices, and the central node continuously receives aggregated results. This continuous multi-stage aggregation process eliminates idle waiting time between communication phases, maintaining adaptability for distributed deep learning while minimizing processing time through uninterrupted useful action across the hierarchy
Data Source
AI summary
A distributed processing node includes a computing device that calculates gradient data of a loss function from an output result obtained by inputting learning data to a learning target model, an interconnect device that aggregates gradient data between the distributed processing node and other distributed processing nodes, a computing function unit that is provided in a bus device and performs processing of gradient data from the computing device, and a DMA controller that controls DMA transfer of gradient data between the computing device and the bus device and DMA transfer of gradient data between the bus device and the interconnect device.


