UPF Traffic Management via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The current 5G network architecture lacks flexibility in traffic management, with the User Plane Function (UPF) being limited in its ability to decide on actions and react to changing network conditions, leading to slow optimization and increased CPU and memory load on the Session Management Function (SMF) when dealing with multiple UPFs.

Innovation Solution

Implementing a method where the UPF has access to an observation space of possible network states and an action space of allowed actions, using reinforcement learning to determine optimal actions based on received network states and rewards from the Network Data Analytics Function (NWDAF), allowing the UPF to autonomously manage traffic and improve decision-making without relying solely on the SMF.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the UPF is limited in its ability to decide on actions and relies on the SMF for traffic management decisions, then the control plane has centralized decision-making capability, but the network optimization delay increases and the CPU and memory load on the SMF increases when dealing with multiple UPFs

Engineering Contradiction:
ImproveUPF flexibilityVSAvoidnetwork optimization delay
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent segments the decision-making authority by introducing autonomous agents in the user plane that can independently make traffic management decisions. This divides the centralized control function into distributed intelligent agents, allowing local real-time decisions without constant SMF intervention, thus reducing optimization delay while maintaining adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic decision-making capabilities in the UPF through reinforcement learning agents that can adapt to changing network conditions in real-time. This dynamic approach allows the UPF to flexibly adjust traffic management actions based on current state observations, improving adaptability and reducing the time loss associated with centralized decision-making.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If the UPF is limited in its ability to decide on actions and relies on the SMF for traffic management decisions, then the control plane has centralized decision-making capability, but the CPU and memory load on the SMF increases when dealing with multiple UPFs

Engineering Contradiction:
ImproveUPF flexibilityVSAvoidSMF load
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the decision-making workload by distributing intelligent agents to individual UPFs. Each agent handles local traffic management decisions independently, segmenting the SMF's computational burden and reducing its CPU and memory load while maintaining centralized coordination capabilities when needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables the UPF to serve itself through autonomous reinforcement learning agents that make traffic management decisions without requiring SMF intervention for each action. This self-service capability reduces the SMF's operational complexity and resource consumption while maintaining network-wide coordination through periodic reporting and policy updates.

Inventive Principle:
Principle #25Self-service

3Ease of manufacture

If the current PFCP reporting solution is used for charging with traffic volume metrics, then the reporting mechanism is simple and established, but it cannot provide the load metrics and user plane traffic metrics needed by the NWDAF for analytics

Engineering Contradiction:
Improvereporting mechanism simplicityVSAvoidmetrics coverage
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent makes the PFCP reporting mechanism universal by enhancing it to serve multiple functions: traditional charging based on traffic volume and new analytics requirements for load metrics and user plane traffic metrics. The extended reporting framework can adapt to different NWDAF analytics needs while maintaining backward compatibility with existing charging functions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a dynamic reporting mechanism that can adapt its metrics and granularity based on NWDAF's analytics requirements. The system dynamically adjusts what metrics are collected and reported, transitioning from static charging-oriented reporting to flexible analytics-oriented reporting, thereby increasing metrics coverage while maintaining implementation simplicity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12101662B2Method of managing traffic by a user plane function, UPF, corresponding UPF, session management function and network data analytics function
Publication Date: 2024.09.24 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US12101662B2 patent drawing
  • US12101662B2 patent drawing
  • US12101662B2 patent drawing

AI summary

A method of managing traffic associated with a User Equipment, UE, by a User Plane Function, UPF, in a telecommunication network, said UPF being associated with a Session Management Function, SMF, and a Network Data Analytics Function, NWDAF, wherein said UPF has access to an observation space comprising a list of possible states said network may take and wherein said UPF has access to an action space comprising a list of possible actions that said UPF is allowed to perform, said method comprising the steps of receiving a state of said network, wherein said state is comprised by said list of possible states, receiving a reward, wherein said reward indicates a degree of satisfaction of said network to be in said state, receiving network traffic from said UE and performing, triggered by said received traffic, an action comprised by said list of possible actions based on said received state of said network and based on said received reward.