Method performed by a network node for training a reinforcement learning algorithm with multiple reward functions to perform a task in a communication network
The RL algorithm with multiple reward functions efficiently adapts to dynamic network criteria by iteratively updating parameters, addressing the inefficiencies of existing RL algorithms in handling multiple reward functions and reducing computational costs.
Patent Information
- Application Number
- PCT/SE2024/050755
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2026-03-05
AI Technical Summary
Existing reinforcement learning (RL) algorithms are sensitive to changes in reward functions, requiring re-training from scratch, which is computationally costly and inefficient, especially in complex environments, and are not suited for optimizing multiple reward functions or dynamic network criteria.
A method for training and executing an RL algorithm with multiple reward functions using an actor-learner split architecture, where a first network node and a second network node coordinate to iteratively update the RL algorithm parameters based on observations from multiple reward functions, ensuring efficient adaptation to different reward functions.
Enables the RL algorithm to quickly adapt to dynamic network criteria and optimize policies for multiple reward functions, reducing computational overhead and improving performance in communication networks.