Reinforcement Learning Routing Agents for Network Convergence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional computer networks face challenges in dynamic routing protocols, including increased convergence time and complexity with larger routing domains, leading to delayed failure recovery and substantial computational overhead.

Innovation Solution

The implementation of AI-defined networking using reinforcement learning (RL) trained neural network agents that make routing decisions based on network policies, trained through a digital twin simulation of the network environment with representative network traffic.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional dynamic routing protocols (BGP, EIGRP, OSPF) are used for route selection, then routing decisions can be made based on multiple criteria, but convergence time increases and computational overhead becomes substantial as routing domain size increases

Engineering Contradiction:
Improverouting decision capabilityVSAvoidconvergence time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent pre-trains neural network agents in a simulated environment before deploying them to actual network routers. This preliminary training allows the agents to learn optimal routing decisions through reinforcement learning without requiring real-time complex computations during network convergence events, thus reducing convergence time while maintaining adaptive routing capabilities

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a digital twin - a virtual copy of the network environment - to train routing agents. This simulation environment replicates network topologies, traffic patterns, and failure scenarios, allowing agents to practice routing decisions repeatedly before deployment, thereby reducing the time needed to adapt to real network changes

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If traditional routing protocols perform route selection after complicated router peering and table exchange processes, then routing decisions can be made with multiple criteria, but computational overhead becomes substantial

Engineering Contradiction:
Improverouting decision capabilityVSAvoidcomputational overhead
Core Design Contradiction:
Adaptability or versatilityVSPower

Solution Approach 1:

The patent replaces traditional mechanical routing processes (router peering, table exchange, convergence algorithms) with neural network-based decision making. The trained neural network agents directly determine optimal routes based on learned patterns, eliminating the need for complex iterative computations and reducing computational overhead significantly

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of routing decision-making from traditional protocol-based metrics (hop count, administrative distance) to neural network output based on learned representations of network state and traffic patterns, enabling more efficient computation while maintaining versatile routing capabilities

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If routing protocols are configured by human administrators with manual tuning parameters, then routing behavior can be optimized for business specifications, but administration becomes error-prone and requires substantial effort

Engineering Contradiction:
Improverouting configuration capabilityVSAvoidprotocol administration complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements self-service routing where neural network agents autonomously learn and execute routing decisions based on network conditions and business policies. The system automatically adapts routing behavior without requiring manual intervention for route tuning or traffic engineering, eliminating configuration errors while maintaining optimization capabilities

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where the neural network agents continuously learn from network observations and performance data. This feedback loop allows the system to automatically optimize routing decisions based on actual network behavior and business requirements, replacing manual tuning with automated adaptive learning

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If nodes allocate hardware resources and table space to hold multiple candidate routes, then more routing options become available, but hardware resources become precious and costly

Engineering Contradiction:
Improveroute options availabilityVSAvoidhardware resources
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent changes how routing information is stored and processed by using neural network representations instead of traditional tabular route entries. This allows the system to maintain multiple routing options in a compressed, efficient format that leverages the neural network's ability to handle complex patterns, reducing the hardware resources needed while preserving route options availability

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250124287A1Reinforcemtn-learning modeling interfaces
Publication Date: 2025.04.17 WORLD WIDE TECHNOLOGY HOLDING CO LLC
  • US20250124287A1 patent drawing
  • US20250124287A1 patent drawing
  • US20250124287A1 patent drawing

AI summary

A computer-implemented method comprising transmitting a user interface to be displayed to a user. The user interface can include one or more first interactive elements. The one or more first interactive elements display policy settings of a reinforcement learning model. The one or more first interactive elements are configured to allow the user to update the policy settings of the reinforcement learning model. The method also can include receiving one or more inputs from the user. The inputs include one or more modifications of at least a portion of the one or more first interactive elements of the user interface to update the policy settings of the reinforcement learning model. The method additionally can include training a neural network model using a reinforcement learning model with the policy settings as updated by the user to adjust rewards assigned in the reinforcement learning model. Other embodiments are described.