Bus Device Computing Function Unit for Gradient Aggregation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed processing systems face communication congestion and bottlenecks during group communication in distributed deep learning, leading to reduced processing efficiency and increased latency.

Innovation Solution

The system incorporates a computing function unit in a bus device connected to computing devices and an interconnect device, with a DMA controller for managing data transfer, and a control unit for allocating learning jobs, enabling high-speed processing of gradient data without congestion by integrating computing functions within the bus device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the number of weights and sample data are increased to improve inference accuracy, then inference accuracy is improved, but communication time and deep learning time increase

Engineering Contradiction:
Improveinference accuracyVSAvoidcommunication time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the aggregation process by introducing intermediate aggregation nodes that collect gradients from multiple computing devices before forwarding to the central aggregation processing node. This segmentation reduces the communication burden on the central node and enables parallel processing of gradient aggregation across multiple intermediate nodes, thereby reducing overall communication time while maintaining the ability to process large numbers of weights and sample data

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical aggregation structure that adds a spatial dimension to the gradient aggregation process. By organizing intermediate aggregation nodes in a hierarchical manner between computing devices and the central aggregation node, the system creates multiple communication pathways and reduces the dimensionality of communication bottlenecks, enabling efficient handling of increased data volumes

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If multiple computing devices are mounted in distributed processing nodes at high density to improve processing performance, then processing performance is improved, but communication congestion occurs at the interconnect device

Engineering Contradiction:
Improveprocessing performanceVSAvoidcommunication congestion
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The patent introduces intermediate aggregation nodes as mediators between multiple computing devices and the central aggregation processing node. These intermediate nodes collect gradients from multiple computing devices locally before forwarding aggregated results to the central node, thereby reducing the communication load on the interconnect device and eliminating congestion while maintaining high processing performance across multiple computing devices

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the aggregation function across multiple intermediate aggregation nodes distributed among the processing nodes. Each intermediate node handles a subset of computing devices, dividing the total communication load into manageable segments that can be processed in parallel, thereby preventing congestion at any single interconnect device while maintaining high overall processing performance

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If integrated communication and distributed communication are performed frequently for aggregation processing, then deep learning processing is enabled, but processing time increases due to communication overhead

Engineering Contradiction:
Improvedistributed deep learning capabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary gradient aggregation at intermediate aggregation nodes before the final aggregation at the central node. By pre-aggregating gradients from multiple computing devices at intermediate nodes, the system reduces the volume of data requiring communication in subsequent steps, thereby enabling distributed deep learning capability while significantly reducing the time consumed by communication overhead

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes a continuous hierarchical aggregation pipeline where intermediate nodes continuously collect and aggregate gradients from computing devices, and the central node continuously receives aggregated results. This continuous multi-stage aggregation process eliminates idle waiting time between communication phases, maintaining adaptability for distributed deep learning while minimizing processing time through uninterrupted useful action across the hierarchy

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12045183B2Distributed processing node and distributed processing system
Publication Date: 2024.07.23 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12045183B2 patent drawing
  • US12045183B2 patent drawing
  • US12045183B2 patent drawing

AI summary

A distributed processing node includes a computing device that calculates gradient data of a loss function from an output result obtained by inputting learning data to a learning target model, an interconnect device that aggregates gradient data between the distributed processing node and other distributed processing nodes, a computing function unit that is provided in a bus device and performs processing of gradient data from the computing device, and a DMA controller that controls DMA transfer of gradient data between the computing device and the bus device and DMA transfer of gradient data between the bus device and the interconnect device.