Multilayered Collective Operations in Network Switches
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As distributed computing systems grow in size and scale, they face bandwidth and hardware strain when collecting or gathering distributed information due to high network traffic and resource demands, even with existing collective operations at network switches.
Innovation Solution
Implementing multilayered collective operations at both network switches and target/response computing nodes, where the network switch receives collective operation requests, identifies target nodes, and sends unicast messages to perform calculations, reducing the number of messages transmitted and the load on cache, memory, and processors by leveraging field programmable gate arrays (FPGAs) for collective logic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If collective operations are implemented at network switch only, then network traffic is reduced, but bandwidth and hardware resources are still strained when systems grow in size
Solution Approach 1:
The patent divides the collective operation functionality into two segments: network switch-based collective operations and compute node-based collective operations. This segmentation allows the system to handle different portions of data gathering at different locations, reducing overall network traffic while distributing the processing load and avoiding concentration of complexity in a single device.
Solution Approach 2:
The patent adds a new dimension of operation by implementing collective operations not only at the network switch but also at the compute nodes. This multi-layered approach creates a hierarchical structure where operations can be performed in parallel at different levels, reducing the burden on any single component and scaling more efficiently as system size increases.
2Productivity
If more target nodes are included in distributed systems, then compute capacity increases, but network bandwidth and hardware availability become strained
Solution Approach 1:
By segmenting the data gathering process into switch-based and node-based operations, the patent enables compute nodes to perform local processing and reduce the volume of data that must be transmitted across the network. This allows the system to scale to more nodes without proportionally increasing network bandwidth requirements.
Solution Approach 2:
Compute nodes are empowered to perform collective operations locally on their own data without requiring all data to be centralized at the network switch. This self-service capability at the node level reduces the demand on shared network resources and allows the system to accommodate more nodes with existing hardware resources.
3Ease of operation
If traditional multicast techniques are used to broadcast requests to multiple target nodes, then all nodes receive the request, but a large amount of network bandwidth is consumed
Solution Approach 1:
The patent extracts the data gathering function from the network switch alone and distributes it to compute nodes as well. This extraction allows the system to perform operations locally at nodes rather than centralizing all data transmission through the network, significantly reducing the bandwidth required for request distribution and data collection.
Solution Approach 2:
The patent transitions from a single-dimensional multicast approach (switch to all nodes) to a multi-dimensional approach where operations occur at both the network switch level and the compute node level. This layered structure enables selective data gathering that reduces overall network traffic while maintaining comprehensive data collection.
Data Source
AI summary
Examples may include techniques for collective operations in a distributed architecture. A collective operation request message from a computing node causes collective operations at one or more target computing nodes communicatively coupled with the computing node through a network switch. The collective operation request message also causes the network switch to perform collective operations on collective operation results received from the one or more target computing nodes.


