RL Training Mediation for Sensitive Mobile Network Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current 3GPP specifications do not adequately address the training of reinforcement learning (RL) models in mobile networks, particularly in scenarios where consumer and producer entities may belong to different parties, and do not provide sufficient information exchange mechanisms to enable effective RL training, including sensitive information protection.
Innovation Solution
Implement collaborative mechanisms for information exchange between service producers and consumers, including action recommendations, environment states, and reward feedbacks, to facilitate RL model training, with exploration vs. exploitation strategies applied at either the consumer or producer side, depending on the scenario, ensuring sensitive information is protected.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If collaborative mechanisms for information exchange are implemented between service producers and consumers, then effective RL model training is enabled, but system complexity increases
Solution Approach 1:
The patent introduces an analytics logical function as an intermediary component that mediates between the model training logical function and the service consumer. This intermediary receives environment states, forwards action indications, and relays feedback information, thereby enabling RL training without requiring direct complex interactions between all parties. The intermediary simplifies the overall system architecture while maintaining training effectiveness.
2Object-affected harmful factors
If sensitive information is protected during information exchange, then security is improved, but information exchange completeness may be reduced
Solution Approach 1:
The patent extracts and separates sensitive information from the information exchange process. The analytics logical function handles only necessary non-sensitive data (environment states, action indications, feedback) while keeping sensitive model parameters and proprietary information localized. This extraction approach maintains security by removing sensitive elements from the exchange while preserving completeness of necessary training information.
3Adaptability or versatility
If exploration vs. exploitation strategies are applied in RL training, then model learning capability is improved, but training stability may deteriorate
Solution Approach 1:
The patent implements a feedback mechanism where the service consumer provides reward feedback to the analytics logical function, which then relays it to the model training logical function. This feedback loop enables the RL agent to learn from the consequences of its actions through the exploration vs. exploitation strategy. The structured feedback path maintains training stability by ensuring consistent information flow while enabling adaptive learning through reward signals.
Data Source
AI summary
Method comprising: monitoring whether a MTLF receives a first state of an environment on which a RL training is to be performed; performing a ML model forward propagation on a first model of the environment having the first state for each of plural actions to obtain a respective expected reward for each of the plural actions: informing a service consumer on the plural actions and their respective expected reward; supervising whether the MTLF receives a RL training result information after the informing the service consumer on the plural actions, wherein the RL training result information comprises an indication of one of the plural actions, a second state of the environment, and a reward feedback; conducting a ML model backward propagation on the first model of the environment having the second state for the one of the plural actions using the reward feedback to obtain a second model of the environment.


