Random Next Iteration for Data Center Network Congestion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches to managing network traffic in data centers and cloud computing environments often lead to network congestion due to synchronized communication patterns, resulting in increased latency and packet loss, especially in oversubscribed networks where large, expensive routers are underutilized most of the time.

Innovation Solution

The proposed solution involves dispersing workloads across multiple network switches to maximize buffering capacity and reduce congestion by using Random Next Iteration or Ordered Next Iteration techniques, which randomize or structure the communication order among hosts to minimize convoying behavior and incast events, thereby optimizing network performance with commodity switches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large routers with significant buffer capacity are used to mitigate network congestion, then network reliability is improved, but device cost increases

Engineering Contradiction:
Improvenetwork reliabilityVSAvoiddevice cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent segments the buffering function from centralized large routers and distributes it across multiple commodity switches. Each switch maintains its own buffer, and the system leverages the aggregate buffering capacity of all switches together, eliminating the need for expensive centralized buffer capacity while maintaining network reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces expensive, high-capacity router buffers with multiple inexpensive commodity switch buffers. While individual switch buffers are smaller and less durable, their collective capacity matches or exceeds that of a single large router buffer, achieving the same reliability at lower cost.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Productivity

If synchronized communication patterns are used among hosts, then coordination efficiency is improved, but network congestion increases

Engineering Contradiction:
Improvecoordination efficiencyVSAvoidnetwork congestion
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent implements periodic communication rounds where hosts take turns communicating with each other. In each round, a subset of hosts communicates while others remain quiet, creating periodic rather than continuous synchronized traffic patterns. This distributes network load over time, reducing peak congestion while maintaining coordination through the structured rotation of communication turns.

Inventive Principle:
Principle #19Periodic action

3Ease of manufacture

If commodity switches with small buffers are used instead of large routers, then device cost is reduced, but network reliability deteriorates

Engineering Contradiction:
Improvedevice costVSAvoidnetwork reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent merges the buffering resources of multiple commodity switches into a unified aggregate buffer system. By coordinating communication patterns across the network, the system effectively pools the small buffers of individual switches into a large virtual buffer capacity that matches or exceeds that of expensive dedicated router buffers, achieving high reliability with low-cost hardware.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10148744B2Random next iteration for data update management
Publication Date: 2018.12.04 AMAZON TECH INC
  • US10148744B2 patent drawing
  • US10148744B2 patent drawing
  • US10148744B2 patent drawing

AI summary

Host machines and other devices performing synchronized operations can be dispersed across multiple racks in a data center to provide additional buffer capacity and to reduce the likelihood of congestion. The level of dispersion can depend on factors such as the level of oversubscription, as it can be undesirable in a highly connected network to push excessive host traffic into the aggregation fabric. As oversubscription levels increase, the amount of dispersion can be reduced and two or more host machines can be clustered on a given rack, or otherwise connected through the same edge switch. By clustering a portion of the machines, some of the host traffic can be redirected by the respective edge switch without entering the aggregation fabric. When provisioning hosts for a customer, application, or synchronized operation, for example, the levels of clustering and dispersion can be balanced to minimize the likelihood for congestion throughout the network.