Method performed by a network node for training a reinforcement learning algorithm with multiple reward functions to perform a task in a communication network

The RL algorithm with multiple reward functions efficiently adapts to dynamic network criteria by iteratively updating parameters, addressing the inefficiencies of existing RL algorithms in handling multiple reward functions and reducing computational costs.

WO2026049659A1 Publication Date: 2026-03-05TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
0 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2024/050755
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing reinforcement learning (RL) algorithms are sensitive to changes in reward functions, requiring re-training from scratch, which is computationally costly and inefficient, especially in complex environments, and are not suited for optimizing multiple reward functions or dynamic network criteria.

Method used

A method for training and executing an RL algorithm with multiple reward functions using an actor-learner split architecture, where a first network node and a second network node coordinate to iteratively update the RL algorithm parameters based on observations from multiple reward functions, ensuring efficient adaptation to different reward functions.

Benefits of technology

Enables the RL algorithm to quickly adapt to dynamic network criteria and optimize policies for multiple reward functions, reducing computational overhead and improving performance in communication networks.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The disclosure provides methods and apparatus for training and executing a reinforcement learning, RL, algorithm configured with multiple reward functions to perform a task in a communication network (1000) The method (100, 500) comprising, the first network node (1100) transmitting (103, 503) a first message, to the second network node (1200), comprising parameters of the RL algorithm. The first network node (1100) receiving (104, 507) a second message, from the second network node (1200), comprising one or more observations. The first network node (1100) updating (105, 508) the parameters of the RL algorithm based on the received one or more observation. The first network node (1100) performing (106, 509) the transmitting, the receiving and the updating iteratively until a stopping criterion exceeds a threshold and storing (107, 510) the updated parameters of the RL algorithm.
Need to check novelty before this filing date? Find Prior Art