RL Training Mediation for Sensitive Mobile Network Data Sharing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 3GPP specifications do not adequately address the training of reinforcement learning (RL) models in mobile networks, particularly in scenarios where consumer and producer entities may belong to different parties, and do not provide sufficient information exchange mechanisms to enable effective RL training, including sensitive information protection.

Innovation Solution

Implement collaborative mechanisms for information exchange between service producers and consumers, including action recommendations, environment states, and reward feedbacks, to facilitate RL model training, with exploration vs. exploitation strategies applied at either the consumer or producer side, depending on the scenario, ensuring sensitive information is protected.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If collaborative mechanisms for information exchange are implemented between service producers and consumers, then effective RL model training is enabled, but system complexity increases

Engineering Contradiction:
ImproveRL model training effectivenessVSAvoidinformation exchange mechanism complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an analytics logical function as an intermediary component that mediates between the model training logical function and the service consumer. This intermediary receives environment states, forwards action indications, and relays feedback information, thereby enabling RL training without requiring direct complex interactions between all parties. The intermediary simplifies the overall system architecture while maintaining training effectiveness.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If sensitive information is protected during information exchange, then security is improved, but information exchange completeness may be reduced

Engineering Contradiction:
Improvesensitive information protectionVSAvoidinformation exchange completeness
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent extracts and separates sensitive information from the information exchange process. The analytics logical function handles only necessary non-sensitive data (environment states, action indications, feedback) while keeping sensitive model parameters and proprietary information localized. This extraction approach maintains security by removing sensitive elements from the exchange while preserving completeness of necessary training information.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If exploration vs. exploitation strategies are applied in RL training, then model learning capability is improved, but training stability may deteriorate

Engineering Contradiction:
Improvemodel learning capabilityVSAvoidtraining stability
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent implements a feedback mechanism where the service consumer provides reward feedback to the analytics logical function, which then relays it to the model training logical function. This feedback loop enables the RL agent to learn from the consequences of its actions through the exploration vs. exploitation strategy. The structured feedback path maintains training stability by ensuring consistent information flow while enabling adaptive learning through reward signals.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20260030546A1Reinforcement learning
Publication Date: 2026.01.29 NOKIA SOLUTIONS & NETWORKS OY
  • US20260030546A1 patent drawing
  • US20260030546A1 patent drawing
  • US20260030546A1 patent drawing

AI summary

Method comprising: monitoring whether a MTLF receives a first state of an environment on which a RL training is to be performed; performing a ML model forward propagation on a first model of the environment having the first state for each of plural actions to obtain a respective expected reward for each of the plural actions: informing a service consumer on the plural actions and their respective expected reward; supervising whether the MTLF receives a RL training result information after the informing the service consumer on the plural actions, wherein the RL training result information comprises an indication of one of the plural actions, a second state of the environment, and a reward feedback; conducting a ML model backward propagation on the first model of the environment having the second state for the one of the plural actions using the reward feedback to obtain a second model of the environment.