Global RL Model from Subnetwork Agents for Network Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current software products fail to provide effective guidance on network actions due to the difficulty in codifying domain expertise into explicit rules, especially in complex and multi-vendor scenarios, and expert services are resource-intensive.

Innovation Solution

A Reinforcement Learning (RL) model is created from subnetwork agents, trained on end-to-end metrics independent of specific topology, and applied through an Action Recommendation Engine (ARE) to recommend network actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If expert rules are used to provide network action guidance, then domain expertise can be applied, but codifying collective domain expertise into explicit rules becomes incrementally difficult and expensive in complex scenarios

Engineering Contradiction:
Improveeffectiveness of network action guidanceVSAvoiddifficulty of codifying domain expertise
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical system of explicit rule-based expert systems with a neural network-based machine learning system. The neural network learns network action policies from training data containing network states and corresponding expert actions, eliminating the need to manually codify complex domain expertise into explicit rules while maintaining reliable guidance capability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the representation of domain expertise from discrete symbolic rules to continuous numerical parameters within the neural network. The network learns optimal actions by adjusting internal parameters (weights and biases) based on training data, allowing flexible adaptation to complex scenarios without increasing rule complexity

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If a global RL model is trained on multiple subnetworks, then the model can generalize across different network configurations, but training data requirements and computational resources increase

Engineering Contradiction:
Improvegeneralization capability across network configurationsVSAvoidtraining data requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent segments the global network training problem into multiple independent subnetwork training tasks. Each subnetwork is trained separately on its own training data, and the learned policies are then aggregated or transferred to form the global model. This segmentation reduces the total training data requirement compared to training a single global model on all networks simultaneously

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses policy transfer or fine-tuning approaches where a global model is initialized by copying weights from multiple pre-trained subnetwork models. This allows the global model to inherit learned patterns from various network configurations without requiring extensive new training data, reducing overall data requirements while maintaining generalization capability

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12388719B2Creating a global reinforcement learning (RL) model from subnetwork RL agents
Publication Date: 2025.08.12 CIENA CORP
  • US12388719B2 patent drawing
  • US12388719B2 patent drawing
  • US12388719B2 patent drawing

AI summary

Methods are provided for recommending actions to improve operability of a network. In one implementation, a method includes acknowledging a plurality of subnetworks in a whole network, each subnetwork including multiple nodes and being represented by a tunnel group having multiple end-to-end tunnels through the subnetwork. The method also includes selecting a first group of subnetworks from the plurality of subnetworks and generating a Reinforcement Learning (RL) agent for each subnetwork of the first group. Each RL agent is based on observations of end-to-end metrics of the end-to-end tunnels of the respective subnetwork. The observations are independent of specific topology information of the subnetwork. Also, the method includes training a global model based on the RL agents of the first group of subnetworks and applying the global model to an Action Recommendation Engine (ARE) configured for recommending actions that can be taken to improve a state of the whole network.