Ai / ML assisted CU-du flow control optimizations in o-ran networks
AI/ML-based reinforcement learning optimizes CU-DU flow control in O-RAN networks by dynamically adjusting buffer allocation and feedback, addressing throughput challenges in dynamic network conditions.
Patent Information
- Application Number
- PCT/US2025/036128
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Existing O-RAN networks face challenges in optimizing CU-DU flow control parameters to maximize end-to-end throughput, particularly for TCP applications, due to dynamic changes in cell configurations and resource allocation needs, leading to potential packet drops and throughput degradation.
Implementing AI/ML-based reinforcement learning and Markov Decision Process (MDP) modules for CU-DU Flow Control Optimization, utilizing state, action, and cost/reward functions to dynamically adjust buffer allocation and flow control feedback, optimizing parameters such as buffer space and DDDS message timing.
Enhances end-to-end throughput by dynamically optimizing CU-DU flow control parameters, ensuring efficient resource allocation and minimizing packet drops, even in dynamic network conditions.
Smart Images

Figure US2025036128_08012026_PF_FP_ABST
Abstract
Description
AI / ML ASSISTED CU-DU FLOW CONTROL OPTIMIZATIONS IN O-RAN NETWORKSCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application is an International Patent Application claiming foreign priority to Indian Provisional Patent Application No. 202441050136, filed on July 1, 2024, the entirety of each of which is incorporated herein by reference.BACKGROUND OF THE DISCLOSURE1. Field
[0002] The present disclosure is related to Open Radio Access Network (O-RAN) wireless networks and relates more particularly to machine-learning-assisted low energy radio resource management (RRM) policies in O-RAN Networks.2. Description of Related Art
[0003] Next Generation Radio Access Network (NG-RAN) architecture and 5G New Radio (NR) stacks include user and control plane functions with monolithic gNB (gNodeB). For the user plane, PHY (physical), MAC (Medium Access Control), RLC (Radio Link Control), PDCP (Packet Data Convergence Protocol) and SDAP (Service Data Adaptation Protocol) sublayers originate in the UE and are terminated in the gNB 102 on the network side.SUMMARY
[0004] Disclosed are system, methods, and computer program products for Artificial Intelligence and Machine Learning (AI / ML) CU-DU Flow Control Optimization for implementing an AI / ML Flow Control Optimization method.
[0005] In an implementation, the AI / ML Flow Control Optimization method comprises: mapping a CU-DU Flow Control optimization related decision making process (CUDU-FCtrl-Opt) for a machine learning (ML) module, the ML module being areinforcement learning (RL) module or a Markov Decision Process (MDP) module, wherein the machine learning module comprises State function; an Action function; a Cost / Reward function; and Transition Probabilities function; and computing parameters including performance measures from a DU or CU to a CU-DU Flow Control Optimization module (CUDU-FCtrl-Optimization module), the parameters being a set of state variables for the RL module, wherein a range of values taken by each state variable are quantized to n levels, where n is a finite value.
[0006] In another implementation, the AI / ML Flow Control Optimization method comprises: mapping a CU-DU Flow Control optimization related decision making process (CUDU-FCtrl-Opt) for an AI / ML model; and computing parameters including performance measures from a DU or CU to a CU-DU Flow Control Optimization module (CUDU-FCtrl- Optimization module), the parameters being state variables for the AI / ML module.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. la shows an example of a User Plane Stack.
[0008] FIG. lb shows an example of a User Plane block diagram illustrating the user plane protocols stacks for a PDU session.
[0009] FIG. 2 shows an example of a Control Plane Stack.
[0010] FIG. 3 shows an example of high-level NG-RAN including a gNB CU and DU.
[0011] FIG. 4 shows an example of a separation of CU-CP (CU-Control Plane) and CU- UP (CU-User Plane) in a 5G gNB.
[0012] FIG. 5 shows a DL (Downlink) Layer 2 Structure.
[0013] FIG. 6 shows a UL (Uplink) Layer 2 Structure.
[0014] FIG. 7 shows an L2 Data Flow example.
[0015] FIG. 8 shows an example of an O-RAN architecture.2
[0016] FIG. 9 illustrates a PDU Session architecture comprising multiple DRBs and multiple QoS Flows.
[0017] FIG. 10 shows a PDU Session architecture comprising multiple DRBs and multiple QoS Flows.
[0018] FIG. 11 shows a Resource Allocation MAC Scheduler, DL Data, and FlowControl Feedback for 5G Network.
[0019] FIG. 12 illustrates a high-level view of a UE establishing a PDU session with a specific DNN.
[0020] FIG. 13 illustrates a flow for a Network Slice Instance and a Network SliceSubnet Instance.
[0021] FIG. 14 illustrates three resource categories in connection with a Radio Resource Management (RRM) Policy Ratio.
[0022] FIG. 15 illustrates a flow for DL Data, and Flow Control Feedback in 5GNetworks.
[0023] FIG. 16 illustrates a flow for DL Data, and Flow Control Feedback in 5GNetworks
[0024] FIG. 17 illustrates a Markov Decision Process used to formalizeReinforcement Learning.
[0025] FIG. 18 illustrates a flow for the workings of Q-learning
[0026] FIG. 19 illustrates flow control related parameters that are captured or computed to analyze throughput of a DRB.
[0027] FIG. 20 illustrates a flow for flow controlled feedback.
[0028] FIG. 21 shows various parameters which influence end-to-end TCP throughput in 0-RAN networks.3
[0029] FIG. 22 illustrates a flow for communicating performance measures and counters.
[0030] FIG. 23 illustrates a flow for communicating performance measures and counters.
[0031] FIG. 24 illustrates a flow for communicating performance measures and counters.
[0032] FIG. 25 shows an implementation of AI / ML training for a CU-CP.
[0033] FIG. 26 shows an implementation of AI / ML training for a CU-UP.DETAILED DESCRIPTION OF THE DISCLOSURE
[0034] Described are implementations of technology for a cloud-based Radio AccessNetworks (RAN), where a significant portion of the RAN layer processing is performed at a central unit (CU) and a distributed unit (DU). Both CUs and DUs are also known as the baseband units (BBUs). CUs are usually located in the cloud on commercial off the shelf servers, while DUs can be distributed. The RE and real-time critical functions can be processed in the remote radio unit (RU).
[0035] RAN Architectures
[0036] In the following section is an overview of Next Generation Radio Access Network (NG-RAN) architecture and 5G New Radio (NR) stacks. 5G NR (New Radio) user and control plane functions with monolithic gNB (gNodeB) are shown in FIGS, la, lb and 2. For the user plane (shown in FIG. la, which is in accordance with 3GPP TS 38.300), PHY (physical), MAC (Medium Access Control), RLC (Radio Link Control), PDCP (Packet Data Convergence Protocol) and SDAP (Service Data Adaptation Protocol) sublayers originate in the UE 101 and are terminated in the gNB 102 on the network side.
[0037] As shown in FIG. lb, which is a block diagram illustrating the user plane protocols stacks for a PDU session, in accordance with 3GPP TS 23.501, PDU layer 90104corresponds to the PDU carried between the UE 101 and the data network (DM) 9011 over the PDU session. As shown in FIG. lb, UE 101 is connected to the 5G access network (AN) 902, which AN 902 is in turn connected via an N3 interface to the Intermediate UPF (1-UPF) 903a portion of the UPF 903, which I-UPF 903a is in turn connected via anN9 interface to the PDU session anchor 903b portion of the UPF 903, and which PDU session anchor 903b is connected to the DN 9011. CU-UP of AN 902 is connected to UPF 903b via a Backhaul(BH) path. The PDU session can correspond to IPv4, IPv6, or both types of IP packets, when the PDU session is of type IPv4, IPv6 or IPv4v6, respectively. GTP-U shown in FIG. lb supports tunnelling user plane data over N3 and N9 interfaces and provides encapsulation of end user PDUs for N3 and N9 interfaces.
[0038] For the control plane, shown in FIG. 2, which is in accordance with 3GPP TS 38.300, RRC (Radio Resource Control), PDCP, RLC, MAC and PHY sublayers originate in the UE 101 and are terminated in the gNB 102 on the network side, and NAS (Non-Access Stratum) originate in the UE 101 and is terminated in the AMF (Access Mobility Function) 103 on the network side.
[0039] NG-Radio Access Network (NG-RAN) architecture from 3GPP TS 38.401 is shown in FIGS. 3-4. As shown in FIG. 3, the NG-RAN 301 consists of a set of gNBs 302 connected to the 5GC 303 through the NG interface. Each gNB comprises gNB-CU 304 and one or more gNB-DU 305 (see FIG. 3). As shown in FIG. 4 (which illustrates separation of CU-CP (CU-Control Plane) and CU-UP (CU-User Plane)), El is the interface between gNB- CU-CP (CU-Control Plane) 304a and gNB-CU-UP (CU-User Plane) 304b, Fl-C is the interface between gNB-CU-CP 304a and gNB-DU 305, and Fl-U is the interface between gNB-CU-UP 304b and gNB-DU 305. As shown in FIG. 4, gNB 302 can consist of a gNB-CU-CP 304a, multiple gNB-CU-UPs (or gNB-CU-UP instances) 304b and multiple gNB-DUs (or gNB-DU instances) 305. One gNB-DU 305 is connected to one gNB-CU-CP 304a, and gNB-CU-UP 304b is connected to one gNB-CU-CP 304a.
[0040] In this section, an overview of Layer 2 (L2) of 5G NR is disclosed in connection with FIGS. 5-7. L2 of 5G NR is split into the following sublayers (in accordance with 3GPP TS 38.300):51) Medium Access Control (MAC) 501 in FIGS. 5-7: Logical Channels (LCs) are SAPs (Service Access Points) between the MAC and RLC layers. This layer runs a MAC scheduler to schedule radio resources across different LCs (and their associated radio bearers). For the downlink direction, the MAC layer processes and sends RLC PDUs received on LCs to the Physical layer as Transport Blocks (TBs). For the uplink direction, it receives transport blocks (TBs) from the physical layer, processes these and sends to the RLC layer using the LCs.2) Radio Link Control (RLC) 502 in FIGS. 5-7: The RLC sublayer presents RLC channels to the Packet Data Convergence Protocol (PDCP) sublayer. The RLC sublayer supports three transmission modes: RLC-Transparent Mode (RLC-TM), RLC- Unacknowledged Mode (RLC-UM) and RLC-Acknowledgement Mode (RLC-AM). RLC configuration is per logical channel. It hosts ARQ (Automatic Repeat Request) protocol for RLC-AM mode.3) Packet Data Convergence Protocol (PDCP) 503 in FIGS. 5-7: The PDCP sublayer presents Radio Bearers (RBs) to the SDAP sublayer. There are two types of Radio Bearers: Data Radio Bearers (DRBs) for data and Signaling Radio Bearers (SRBs) for control plane.4) Service Data Adaptation Protocol (SDAP) 504 in FIGS. 5-7: The SDAP maps QoS flows within a PDU session to a specific Data Radio Bearer.FIG. 5 is a block diagram illustrating DL L2 structure, in accordance with 3GPP TS 38.300. FIG. 6 is a block diagram illustrating UL L2 structure, in accordance with 3GPP TS 38.300. FIG. 7 is a block diagram illustrating L2 data flow example, in accordance with 3GPP TS 38.300 (in FIG. 7, H denotes headers or sub-headers).
[0041] Open Radio Access Network (0-RAN) is based on disaggregated components which are connected through open and standardized interfaces based on 3GPP NG-RAN. An overview of O-RAN with disaggregated RAN CU (Centralized Unit), DU (Distributed Unit), and RU (Radio Unit), near-real-time Radio Intelligent Controller (R1C) and non-real-time RIC is illustrated in FIG. 8.6
[0042] As shown in FIG. 8, the CU (shown split as O-CU-CP 801a and O-CU-UP 801b) and the DU (shown as 0-DU 802) are connected using the Fl interface (with Fl-C for control plane and Fl-U for user plane traffic) over a mid-haul (MH) path. One DU can host multiple cells (e.g., one DU could host 24 cells) and each cell can support many users. For example, one cell can support 800 Radio Resource Control (RRC) -connected users and out of these 800, there can be 250 Active users (i.e., users that have data to send at a given point of time).
[0043] A cell site can comprise multiple sectors, and each sector can support multiple cells. For example, one site could comprise three sectors and each sector could support eight cells (with each cell being on a different frequency band in a given sector). One CU-CP (CU-Control Plane) could support multiple DUs and thus multiple cells. For example, a CU-CP could support 500 cells and around 100,000 User Equipments (UEs). Each UE could support multiple Data Radio Bearers (DRBs) and there could be multiple instances of CU-UP (CU-User Plane) to serve these DRBs. For example, each UE could support 4 DRBs, and 400,000 DRBs (corresponding to 100,000 UEs) can be served by five CU-UP instances (and one CU-CP instance).
[0044] The DU can be located in a private data center, or it could be located at a cellsite. The CU could also be in a private data center or even hosted on a public cloud system. The DU and CU, which are typically located at different physical locations, could be tens of kilometers apart. The CU communicates with a 5G core system, which could also be hosted in the same public cloud system (or could be hosted by a different cloud provider). A RU (Radio Unit) (shown as 0-RU 803 in FIG. 8) is located at a cell-site and communicates with the DU via a front-haul (FH) interface.
[0045] The E2 nodes (CU and DU) are connected to the near-real-time R1C 132 using the E2 interface. The E2 interface is used to send data (e.g., user and / or cell KPMs) from the RAN, and deploy control actions and policies to the RAN at near-real-time RIC 132. The applications or services at the near-real-time RIC 132 that deploys the control actions and policies to the RAN are called xApps. During the E2 setup procedures, the E2 node advertises the metrics it can expose, and an xApp in the near-RT RIC can send a7subscription message specifying key performance metrics which are of interest. The near- real-time RIC 132 is connected to the non-real-time RIC 133 (which is shown as part of Service Management and Orchestration (SMO) Framework 805 in FIG. 8) using the Al interface. The applications that are hosted at non-RT-RIC are called rApps. Also shown in FIG. 8 are O-eNB 806 (which is shown as being connected to the near-real-time RIC 132 and the SMO Framework 805) and O-Cloud 804 (which is shown as being connected to the SMO Framework 805).
[0046] In this section, PDU sessions, DRBs, and Quality of Service (QoS) flows are discussed. In 5G networks, PDU connectivity service is a service that provides exchange of PDUs between a UE and a Data Network (DN) identified by a Data Network Name (DNN). The PDU Connectivity service is supported via PDU sessions that are established upon request from the UE. The DNN defines the interface to a specific external data network. One or more QoS flows can be supported in a PDU session. All the packets belonging to a specific QoS flow have the same 5QI (5G QoS Identifier). A PDU session consists of the following: Data Radio Bearers which are between UE and CU in RAN; and an NG-U GTP tunnel which is between CU and UPF (User Plane Function) in the core network. FIG. 9 illustrates an example PDU session (in accordance with 3GPP TS 23.501) consisting of multiple DRBs, where each DRB can consist of multiple QoS flows. In FIG. 9, three components are shown for the PDU session 901: UE 101; access network (AN) 902; and UPF 903, which includes Packet Detection Rules (PDRs) 9031.
[0047] The following should be noted for 3GPP 5G network architecture, which is illustrated in FIG. 10 (in the context of multiple PDU sessions involving multiple DRBs and QoS Flow Identifiers (QFIs), which PDU sessions are implemented involving UE 101, gNodeB 102, UPF 903, and DNNs 9011a and 9011b) and FIG. 11 (in the context of Radio Resource Management (RRM) for connecting UE 101 to the network via RU 306 with a MAC Scheduler 1001):1) The transport connection between the base station (i.e., CU-UP 304b of FIG. 11) and the UPF 903 uses a single GTP-U tunnel per PDU session, as shown in FIGS. 10 and 11. The PDU session is identified using GTP-U TEID (Tunnel Endpoint Identifier).82) The transport connection between the DU 305 and the CU-UP 304b of FIG. 11 uses a single GTP-U tunnel per DRB (see also FIG. 10 and FIG. 11). The DU is provided with an UL GTP-U TEID and the CU is provided with the corresponding DL GTP-U TEID to allow for data communication for that DRB between DU and CU-UP.3) SDAP: a) The SDAP (Service Adaptation Protocol) 504 Layer receives downlink data from the UPF 903 across the NG-U interface (see FIG. 11). b) The SDAP 504 maps one or more QoS Flow(s) onto a specific DRB. c) The SDAP header is present between the UE 101 and the CU (when reflective QoS is enabled), and includes a field to identify the QoS flow within a specific PDU session.4) GTP-U protocol includes a field to identify the QoS flow and is present between CU and UPF 903 (in the core network).5) One (logical) DU (or REC) queue exists per DRB (or per logical channel) for REC PDUs that are to be transmitted for the first time, as shown in FIG. 11. Separate logical queues can exist in DU for packets that are to be retransmitted to UE.
[0048] In this section, standardized 5QI to QoS characteristics mapping will be discussed. As per 3GPP TS 23.501, the one-to-one mapping of standardized 5QI values to 5G QoS characteristics is specified in Table 1 shown below. The first column represents the 5Q1 value. The second column lists the different resource types, i.e., as one of Non-GBR, GBR, Delay-critical GBR. The third column (“Default Priority Level") represents the priority level PrioritySQI, for which lower the value the higher the priority of the corresponding QoS flow. The fourth column represents the Packet Delay Budget (PDB), which defines an upper bound for the time that a packet can be delayed between the UE and the N6 termination point at the UPF. The fifth column represents the Packet Error Rate (PER). The sixth column represents the maximimum data burst volume for delay-critical GBR types.9The seventh column represents averaging window for GBR, delay critical GBR types. Note that only a subset of 5QI values defined in 3GPP TS 23.501 are shown in Table 1 below.
[0049] For example, as shown in Table 1, 5QI value 1 is of resource type GBR with the default priority value of 20, PDB of 100ms, PER of 0.01, and averaging widnow of 2000 ms. Conversational voice falls under this catogery. Similarly, as shown in Table 1, 5QI value 7 is of resource type Non-GBR with the default priority value of 70, PDB of 100ms and PER of 0.001. Voice, video (live streaming), and interactive gaming fall under this catogery.10Table 1
[0050] In this section, Radio Resource Management (RRM) is disclosed (a block diagram for an example RRM with a MAC Scheduler is shown in FIG. 11). L211methods (such as MAC scheduler) play a critical role in allocating radio resources to different UEs in a cellular network. For example, the scheduling priority of a logical channel (PLC) could be determined as part of MAC scheduler using one of the following:PLC = WSQI*PSQI+ WGBR*PGBR+WPDB* PPDB +WPF*PPF + WBO*PBO, orPLC = (WSQI*PSQI+ WPF*PPF) * maximum (WGBR*PGBR , WPDB* PPDB) + WBO*PBO , orPLC = (WSQI*PSQI+ WPF*PPF) + maximum (WGBR*PGBR , WPDB* PPDB) + WBO*PBOOnce one of the above methods is used to compute scheduling priority of a logical channel corresponding to a UE in a cell, the same method is used for all other UEs and these scheduling priorities are used to determine the resources to be allocated to each LC in each cell.
[0051] In the above expressions, the parameters are defined as follows: a) PSQI is the priority metric corresponding to the QoS class ( 5 QI) of the logical channel. Incoming traffic from a DRB is mapped to Logical Channel (LC) at RLC level. PSQI is a function of the default 5QI priority value, PrioritysQi, of a QoS flow that is mapped to the current LC. The lower the value of PrioritysQi the higher the priority of the corresponding QoS flow. For example, Voice over New Radio (VoNR) (with 5Q1 of 1) has a higher PSQI compared to web browsing (with 5QI of 9). b) PGBR is the priority metric corresponding to the target bit rate of the corresponding logical channel. The GBR metric PGBR represents the fraction of data that must be delivered to the UE within the time left in the current averaging window Tavg_win (as per 5QI table, default is 2000 msec.) to meet the UE’s GBR requirement. PGBR is calculated as follows:PGBR = remData / targetData where12targetData is the total data bits to be served in each averaging window Tavg_win in order to meet the GFBR (Guaranteed Flow Bit Rate) of the given QoS flow; remData is the amount of data bits remaining to be served within the time left in the current averaging window;PGBR is reset to 1 (or some other suitable value) at the start of each averaging window Tavg_win, and should go down to 0 towards the end of this window if the GBR criterion is met; andPGBR = 0 for non-GBR flows. c) PPDB is the priority metric corresponding to the packet delay budget at DU for the corresponding logical channel. PPDB = 1 if PDBDU<=QDelayRtc and PPDB - 1 / (PDBDU- QDelayRtc) if PDBDU> QDelayRtc where both PDBDU(Packet Delay Budget at DU) and RLC Queuing delay, QDelayRtc, are measured in terms of slots. QDelayRtc = (t -TRLC) is the delay of the oldest RLC packet in the QoS flow that has not been scheduled yet, and it is calculated as the difference in time between the SDU insertion in RLC queue to current time where t := current time instant, TRLC := time instant when oldest SDU was inserted in RLC. d) PPF is the priority metric corresponding to proportional fair metric of the UE. PPFTa is the PF Metric, calculated on a per UE basis as PPFwhere r: It is the UE’s achievable data rate. DU considers CSI (Channel Status Information) which also includes CQI (Channel Quality Indication), reported by UE to compute this;Ravg= a.Ravg + (l-a).b , UE’s average throughput, where b>=0 is the number of bits scheduled in current TTI (Transmission Time Interval) and 0 < a <= 1 is the HR filter coefficient; a and p are configurable parameters. For example, if one sets a=l and p = 0,13the priority metric, PPF, works in greedy way and favors UEs in good channel conditions. This helps to improve cell throughput but need not be fair to individual logical channels and some of these LCs may not meet their QoS requirements. For some existing systems, a. and p are in the range of 0 to 1. Paramater a is allowed to be upper bounded by a_max (and lower bounded by zero). As a LC is eventually selected by the overall scheduling priority of a logical channel (PLC) which has multiple other factors (and not only the PPF metric), a_max is allowed to be even higher than one (for example, a_max = 1.2) to help design and enforce various type of policies (and associated service level agreements at per-cell, per-DU and per- logical channel level). Similarly, p, is upper bounded by p_max, and lower bounded by zero. e) BO is the buffer occupancy in the RLC queue (e.g. at DU for downlink traffic). PBO is the normalized value of buffer occupancy across all DRBs f) In addition, the following weights are defined: WSQI is the weight of PSQI; WGBR is the weight of PGBR; g) WPDB is the weight of PPDB; h) WPF is the weight of PPF and i) WBO is the weight of PBO. For example, each of the above weights could be set to a value between 0 and 1 though other suitable set of values can be chosen too.
[0052] In this section network slicing is discussed. A network slice is a logical network that provides specific network capabilities and network characteristics, supporting various service properties for network slice customers (e.g., as specified in 3GPP TS 28.500). A network slice divides a physical network infrastructure into multiple virtual networks, each with its own (dedicated or shared) resources and service level agreements. An S-NSSAI (Single Network Slice Selection Assistance Information) identifies a network slice in 5G systems. As per 3GPP TS 23.501, S-NSSAI is comprised of: i) a Slice / Service type (SST), which refers to the expected Network Slice behavior in terms of features and services; and ii) a Slice Differentiator (SD), which is optional information that complements the Slice / Service type(s) to differentiate amongst multiple Network Slices of the same Slice / Service type.14
[0053] UE first registers with a 5G cellular network identified by its PLMN ID (Public Land Mobile Network Identifier). UE knows which S-NSSAIs are allowed in a given registration area. It then establishes a PDU session associated with a given S-NSSA1 in that network towards a target Data Network (DN), such as the internet. As in FIG. 10, one or more QoS flows could be activated within this PDU session. UE can perform data transfer using a network slice for a given data network using that PDU session. A high-level view of UE establishing a PDU session with a specific DNN is shown in FIG. 12. As described in 3GPP TS 23.501, an NSSAI is a collection of S-NSSAls. A Network Slice Instance (NSI) consists of set of network function instances and the required resources that are deployed to serve the traffic associated with one or more S-NSSAIs.
[0054] 3GPP TS 28.541 includes information model definitions, referred to asNetwork Resource Model (NRM), for the characterization of network slices. Management representation of a network slice is realized with Information Object Classes (IOCS), named Networkslice and NetworkSliceSubnet, as specified in 5G Network Resource Model (NRM), 3GPP TS 28.541. The Networkslice IOC and the NetworkSliceSubnet IOC represent the properties of a Network Slice Instance (NSI) and a Network Slice Subnet Instance (NSSI), respectively. As shown in FIG. 13, NSI could be composed of a single NSSI (such as RAN NSSI) or multiple NSSIs (such as RAN NSSI, 5G Core NSSI and Transport Network NSSI).
[0055] A resource model for distribution of resources among slices is shown in FIG.15. which shows the structure of RRMPolicyRatio. Three resource categories have been defined in 3GPP TS 28.541 in connection with RRMPolicyRatio: Category I; Category II; and Category III (as shown in FIG. 14).
[0056] Procedures and functionality of the Fl-U interface are defined in 3GPP TS 38.425. This Fl-U interface supports NR User Plane (NR-U) protocol which provides support for flow control and reliability between CU-UP and DU for each DRB. Figures 15 and 16 show DL Data, and Flow Control Feedback (DDDS) in 5G Networks. As in FIG. 15, Downlink User Data (DUD) PDUs are used to carry PDCP PDUs from CU-UP to DU for each DRB. As in FIG. 16, the Downlink Data Delivery Status (DDDS) message conveys Desired15Buffer Size (DBS), Desired Data Rate (DDR) and some other parameters from DU to CU-UP for each DRB as part of flow control feedback.
[0057] In this section, a general overview of reinforcement learning (RL) is provided. Reinforcement Learning is a feedback-based machine learning technique in which an agent learns to behave in an environment by performing actions and seeing the results of the actions. For each good action, the agent gets a positive feedback or reward and for each bad action, the agent gets a negative feedback or penalty. The goal of the agent is to use RL algorithms to learn the best policy as it interacts with the environment so that, given any state, it will always take the most optimal action to produce the least cost (or the maximum reward) in the long run.
[0058] Some of the terms used in connection with reinforcement learning technique are listed below: o AgentQ: An entity that interacts with the environment and acts upon it. o Environment): A situation in which an agent is present or surrounded by. In RL, a stochastic environment is assumed, which means it is random in nature. o ActionQ: Actions are the moves taken by an agent within the environment. o StateQ: State is a situation returned by the environment after each action taken by the agent. o CostQ: A feedback returned to the agent from the environment to evaluate the action of the agent. o PolicyQ: Policy is a strategy applied by the agent for the next action based on the current state. o ValueQ: It is an expected long-term cost with a discount factor. o Q-value() : It is similar to ValueQ but it takes one additional parameter as the current action ‘a’.16
[0059] In this section, Markov Decision Process (MDP) is discussed. MDP is used to formalize the reinforcement learning (RL) problems. If the environment is completely observable, then its dynamics can be modelled as a Markov Process. In MDP, which is illustrated in FIG. 17, the agent constantly interacts with the environment and performs actions. At each action, the environment generates a new state and responds with a reward (or penalty). MDP includes a tuple of four elements, as follows: States, Actions, Costs (or Rewards), and Transition Probabilities. A MDP is a finite MDP is when there are finite number of states, finite costs, and finite number of actions. Provided below is a simple set notation to represent each of the elements in an MDP:0 S: A set of finite States S0 A: A set of finite Actions A o Cfs, a): Immediate cost (or expected immediate cost) incurred after transitioning from state s to state s', due to action ‘a’.0 P: represents the Transition probability matrix corresponds to state space S and action space A.0 P(s |s, a): Transition Probability of landing in state s' when action ‘a’ is taken at state s.
[0060] A finite MDP is when there are finite states, finite costs, and finite actions.Finite MDP is considered in RL.
[0061] In this section, some of the approaches used in RL such as Value-based approach (Value iteration methods), Q-learning, Deep Q Neural Network (DQN), and Policy - based approach (Policy iteration methods) are discussed.
[0062] The value-based approach is about finding the optimal value function which is the optimal value at a state under any policy TL The below system of equations for the state space are called Bellman equations or optimality equations and these characterize the values and the optimal policies in infinite-horizon models:17Here,Vfs): Value at state s,Cfs, a): Immediate cost at state s for action a, y: Discount factor,P(s’|s,a): transition probability of landing in state s’ when action a is taken at state s, andV(s'): Value at state s’.
[0063] Q-learning involves learning the value function Q (s, a), which characterizes how good it is to take an action "a" at a particular state "s”. The main objective of Q- learning is to learn the policy which can inform the agent what actions should be taken for minimizing the overall cost. The goal of the agent in Q-learning is to optimize the value of Q where the value of Q-learning can be derived from the Bellman equation. Instead of using a value at each state, a Q-value, Q(s,a), is used for a pair of state and action. Q-value specifies which action is more beneficial than the other actions and according to the best Q-value, the agent takes its next move.
[0064] After performing an action "a", the agent will incur a cost Cfs, a), and the agent ends up at a certain state. Q-value equation is;
[0065] The flowchart shown in FIG. 18 illustrates the workings of Q-learning. In block 1801, the Q-table is initialized. In block 1802, an action to perform is selected. In block 1803, the selected action is performed. In block 1804, the associated cost for the action is found. In block 1805, the Q-table is updated.18
[0066] Deep Q Neural Network (DQN) is a Q-learning using Neural networks. For a big state space environment, it is a challenging and complex task to define and update a Q- table. To solve such an issue, a DQN algorithm can be used. In this approach, instead of defining a Q-table, neural network approximates the Q-values for each action and state.
[0067] A policy-based approach is used to find the optimal policy for the minimum future cost without using the value function. This approach involves two types of policies: 1) Deterministic policy where the same action is produced by the policy for any given state, and 2) Stochastic policy where for each state there is an inverse (probability) distribution over set of actions possible at that state.
[0068] Peak Cell Throughput Validation
[0069] Peak Cell Throughput Validation: this is one of the tests that operators carry out to evaluate performance of a base station. With this, there is only one UE in the cell and the operator wants to see that this UE can achieve the throughput which is equal to the (theoretical) peak throughput which can be achieved in that cell and there is no dip in this throughput for the duration of the test. Not every UE may have the capability (e.g. in terms of its hardware, software, radio frequency related circuits etc.) to support the peak cell throughput possible in that cell but usually test (and some commercial) UEs are available which can support the peak cell throughput. Also, UEs join a cell and can get handed over to neighboring cells, and this peak cell throughput test can be done on any chosen UE (which supports this capability for peak cell throughput) by temporary reducing traffic (to zero) for other UEs in that cell.
[0070] Maximum Cell Throughput Validation
[0071] Maximum Cell Throughput Validation: There are factors such as Downlink (DL) Mid-Haul (MH) Latency, Uplink MH Latency, DL Backhaul (BL) Latency and UL Latency which also impact end-to-end throughput for TCP applications. As the MH or BH latency increases (and goes above a threshold), it may not always be possible for a UE to get the peak cell throughput especially when it is using TCP type of transport protocols. Also, it may not always be possible to create a scenario where UE always keeps reporting19maximum possible CQI (Channel Quality Information) index and where DL and UL Block Error Rate (BEER) being experienced by the UE is zero (or almost zero). In such cases, the operator wants to verify the maximum possible throughput for that UE. It can be less than the theoretical peak cell throughput possible in that cell, but the intent is to verify and achieve the maximum possible throughput in such scenarios too. The operator also wants to maximize throughput for each DRB in real deployment scenarios.
[0072] CU-DU flow control related parameters
[0073] CU-DU flow control related parameters such as the amount of buffer space allocated to each DRB in the DU and the conditions using which the flow control feedback (i.e. DDDS message) is sent from DU to CU-UP for each DRB, play a critical role in determining end-to-end throughput for a DRB in 0-RAN Networks. If some RLC SDUs are dropped due to lack of buffer space for a DRB in the DU, its end-to-end throughput can degrade (e.g. for TCP type of applications which that DRB may be supporting). Also, if DDDS messages are delayed from DU to CU-UP for a DRB, PDCP PDUs can get delayed in reaching to DU for that DRB and this can also degrade end-to-end throughput. It becomes important to determine the right value of maximum possible buffer space for that DRB and value of the parameters which determine how the DDDS messages should be sent from DU to CU-UP for each such DRB.
[0074] Take a cell g that can support maximum of 'MaxNconn(g)’ RRC-Connected users (or UEs), maximum of 'MaxNactive(g)’ Active users (or UEs) and maximum of 'MaxNdrb(g)’ DRBs. Active users are those users who have data to send in a given time slot. For example, a cell can support activity factor of 30% and thus number of active users can be 30% of the number of RRC-Connected users. Maximum number of allowed DRBs in cell g, MaxNdrb(g), could be equal to 'b * MaxNactive(g)’, for example, with b - 5 for cell g For example, MaxNconn could be 1000 UEs for a cell, MaxNactive could be 300 UEs for that cell and MaxNdrb could be 1500 for this cell.
[0075] The number of RRC-Connected users in cell g at time t is denoted asNconn(g;t), number of active users at time t is denoted as Nactive(g;t) and number of DRBs20at time t is denoted as Ndrb(g;t). Note that Nconn(g;t) < MaxNconn(g), Nactive(g;t) < MaxNactive(g;t) and Ndrb(g;t)<(b * MaxNative(g)), for each cell g for all time slots t.
[0076] The maximum amount of buffer space which can be allocated to this cell g in the DU for DL RLC SDUs at time t is indicated as maxRLCBufferCell(g;t). Note that this can be the same for all time slots t or can change dynamically depending on the policies used for resource management. For example, a DU can support maximum h cells. If some of these cells are powered down (e.g. to save energy), some of the active cells can be given higher amounts of memory.
[0077] The number of RRC-Connected, number of Active users, and number of DRBs can dynamically keep changing in a cell (subject to maximum limits allowed in that cell). DU needs to allocate buffer space for Ndrb(g;t) DRBs for cell g at time t and also needs to keep buffer space reserved for 'MaxNdrb(g) - Ndrb(g;t)’ DRBs at time t. Note that if DU does not keep adequate buffer space reserved for 'MaxNdrb(g) - Ndrb(g;t)’ DRBs at time t, new DRBs (from existing or new UEs) that get admitted in that cell can experience high packet drop in DU due to lack of buffer space, and this can badly degrade throughput for the newly admitted DRBs. Also, some of the existing DRBs can use very high buffer space in the DU and if the buffer management scheme in DU tries to take away some of this buffer space and allocate it to newly admitted DRBs, this can result in throughput degradation for existing DRBs too. Thus, DU needs to allocate buffer space for existing DRBs, but also needs to keep adequate buffer space reserved for newly admitted DRBs. A DU also supports multiple cells and it needs to find the right amount of buffer space which can be allocated to DRBs from different cells in that DU. This imposes a limitation on the maximum possible buffer space which can be allocates to a given DRB.
[0078] A large-scale network uses various types of cell configurations and it becomes difficult to find the right value of the CU-DU flow control parameters which would help maximize end-to-end throughput of a DRB (e.g. for carrying traffic for applications which are running on TCP). This disclosure provides Al / ML assisted methods to find optimal values of CU-DU flow control related parameters for each DRB.21
[0079] This disclosure also extends methods above for the scenarios where optimal values of CU-DU flow control parameters need to be chosen to maximize throughput for several DRBs in a network.
[0080] METHOD IA
[0081] Described are parameters that influence throughput for single or multi-UE scenarios. Parameters related to MH (mid-haul) which impact throughput (of a DRB) specified as follows include: o Average MH DL latency observed over MH for DL traffic (from CU-UP to DU), denoted as avgMhDILatency (t) at time t. o Average MH UL latency observed over MH for UL traffic (from DU to CU- UP), denoted as avgMhUlLatency(t) at time t. o Average MH DL Packet-Error-Rate (PER) observed over MH for DL traffic, denoted as avgMhDlPer(t) at time t. o Average MH UL PER observed over MH for UL traffic, denoted as avgMhUlPer(t) at time t. o Fraction of mid-haul DL capacity used by DRBs corresponding to UEs in a cell, denoted as fracMhDlCapacity(m, t) for cell m at time t. For the case where there is one link (or hop) between CU-UP and DU, data rate of DRBs corresponding to a specific cell can be measured at the exit point of CU-UP for DL traffic and this can be used to compute fraction of MH DL capacity used by DRBs corresponding to UEs in that cell m. Note that there can be multiple links from CU-UP to DU. In such cases, fraction of the MH DL capacity used by DRBs corresponding to cell m can be measured for each link separately and fracMhDlCapacity(m,t) for MH is set to the maximum of these fraction values across all these links for MH.For example, if there are two links between CU-UP and DU, and if DRBs22corresponding to cell m use 30% of the capacity corresponding to the first link and 20% of the capacity corresponding to the second link at time t, fracMhDlCapacity(m.t) is set to 30% (or 0.3 using the fraction notation). o Fraction of mid-haul UL capacity used by DRBs corresponding to UEs in a cell, denoted as fracMhUlCapacityfm, t) for cell m at time t. For the case where there is one link (or hop) between DU and CU-UP, data rate of DRBs corresponding to a specific cell can be measured at the exit point of the DU for the UL traffic and this can be used to compute fraction of MH UL capacity used by DRBs corresponding to UEs in that cell. Note that there can be multiple links (or hops) from DU to CU-UP for UL traffic. In such cases, fraction of the MH UL capacity used by DRBs corresponding to cell m can be measured for each link separately and fracMhUlCapacity(m,t) for MH is set to the maximum of these fraction values across all these links for MH. For example, if there are two links between DU and CU-UP for UL traffic, and if DRBs corresponding to cell m use 20% of the capacity corresponding to the first link and 15 of the capacity corresponding to the second link at time t, fracMhUlCapacity(m,t) is set to 20% (or 0.2 using the fraction notation).
[0082] Parameters related to BH (Backhaul)that impact throughput (of a DRB) are specified as follows. o Average BH DL latency observed over BH for DL traffic (e.g. from UPF to CU-UP in 5G networks or from Packet Gateway to CU-UP in 4G networks). This is denoted as avgBhDlLatency(t) at time t. o Average BH UL latency observed over BH for UL traffic (e.g. from CU-UP to UPF in 5G networks). This is denoted as avgBhUlLatency(t) at time t. o Average BH DL PER observed over BH for DL traffic (e.g. from UPF to CU- UP in 5G networks), denoted as avgBhDlPer(t) at time t.23o Average BH UL PER observed over BH for UL traffic (i.e. from CU-UP to UPF in 5G networks), denoted as avgBhUlPer(t) at time t.
[0083] Parameters related to CU-DU flow control that are either configured with static values or are assigned values where these values can vary dynamically, and that can impact throughput (of a DRB) are specified as follows. As described earlier, DDDS (DL Data Delivery Status) is the flow control feedback message from DU to CU-UP for each DRB which is sending data in DL direction from CU-UP to DU. Various flow control related parameters which are captured or computed to analyze throughput of a DRB that are shown in FIG. 19 and specified as follows include. o Maximum buffer space allowed for a given DRB to store DL RLC SDUs in the DU over a given time interval is governed by the buffer management policies used at the DU. It could be static or can change based on the buffer management schemes implemented in the DU. Maximum allowed buffer space for DL RLC SDUs of DRB d at time t is denoted as maxAllowedRLCBuffer(d;t). o Minimum rate of DDDS (i.e. flow control feedback) messages (from DU to CU-UP) for each DRB: a rate of DDDS messages (for DRB d) is defined as the minimum rate at which DDDS messages are to be sent from DU to CU- UP for a DRB, i.e. in the uplink direction over the mid-haul (irrespective of whether this DRB is using RLC AM or RLC UM). This is denoted as minRateMhUlDDDS(d; t) for DRB d at time t. For example, the flow control module in DU can be designed to send at least 40 DDDS messages for DRB d every second and in this case, minRateMhU!DDDS(d;t)-40. Note that the minimum required rate (of sending DDDS) can vary with time in the method here. o Maximum rate at which DU can send DDDS messages to CU-UP for a DRB can also be bounded by an upper threshold (e.g. due to software and hardware overhead associated with this). This maximum rate for DRB d24at time is denoted as maxRateMhUlDDDS(d;t), where, for example, minRateMhUlDDDS(d;t) < maxRateMhUlDDDS(d;t). o For DRB d which is communicating data with RLC Configured in Acknowledged Mode (i.e. RLC AM): DDDS can be sent (from DU to CU- UP) when DU receives RLC Status PDU for the corresponding DRB from the UE or periodically (i.e. on expiry of a timer). If RLC Status PDU sent from UE to DU is lost (or delayed due to any reason), DU can send DDDS to CU-UP for that DRB upon expiry of a timer which is denoted as DDDSTAM(d) for DRB d. For example, if RLC Status PDU is expected approximately every 40 ms from a UE, value of DDDSTAM(d) could be set to 50 ms. Thus, DU is sending DDDS every time interval DDDSTAM(d) or earlier for DRB d using RLC AM. Value of DDDSTAM(d) is chosen such that the minimum rate at which DDDS messages are sent from DU to CU-UP for DRB d at any given time t, minRateMhUlDDDS(d;t), is satisfied. o For DRB d which is communicating data with RLC configured in Unacknowledged Mode (i.e. RLC UM): DDDS could be sent periodically from DU to CU-UP and this periodic time interval is denoted as DDDSTuM(d), for DRB d. For example, DDDS could be sent every 20 ms for DRB d which is using RLC UM and in this case, DDDSTuM(d) - 20 ms. Value of DDDSTuM(d) is chosen such that the minimum rate at which DDDS messages are sent from DU to CU-UP for DRB d at any given time t, minRateMhUlDDDS(d;t), is satisfied.
[0084] Parameters related to buffer management at the DU which impact throughput for a DRB are specified as follows. Note that these are observed parameters in the DU for each DRB d. o Number of DL RLC SDUs dropped in the DU for DRB d due to lack of buffer space and this is denoted as numRLCDropped(d; (t-0,t)) for the time interval (t-0,t). Here, 0 is less than t. If DRB d started carrying traffic25at time zero (i.e. 6 is zero), numRLCDropped(d; t) denotes the number of dropped DL RLC SDUs due to lack of buffer space in the DU for DRB d by time t. o Average buffer space used by DRB d in the DL RLC SDU queue and this is denoted by avgOccupiedRLCBuffer(d;(t-0,t)) for the time interval (t-0,t). Here, 6 is less than t. If DRB d started carrying traffic at time zero (i.e. 6 is zero), avgOccupiedRLCBuffer (d; t) denotes the average buffer space used by DRB d in the DL RLC SDU queue by time t. o Maximum buffer space used by DRB d in the DL RLC SDU queue in the DU and this is denoted as maxOccupiedRLCBuffer(d; (t-0,t)) for the time interval (t-0,t). Here, 0 is less than t. If DRB d started carrying traffic at time zero (i.e. 0 is zero), maxOccupiedRLCBuffer (d; t) denotes the maximum buffer space used by DRB d in the DL RLC SDU queue by time t. Note that maxOccupiedRLCBuffer(d; t) < maxAllowedRLCBuffer(d; t)
[0085] Parameters related to UE distribution (scenario p) for cell m include: 1) RSRP reported for these UEs, 2) SINR reported from these UEs if available, 3) Location of UEs if available.
[0086] Channel State Information (CSI) reported by each UE, which includes Channel Quality Indicator (CQI), Rank Indicator (RI) and Precoding Matrix Indicator (PMI). Some other air-interface related parameters include the following:BLER (Block Error Rate) over the air-interface for UE u in cell m at time t:DL BLER. This is denoted as dLBler(u,t) for UE u in cell m at time t.UL BLER: This is denoted as uLBler(u,t) for UE u at time tResidual DL BLER: This is the residual DL BLER for UE u after configured number of HARQ retransmissions between DU and UE whenever needed. This is26denoted as dLResidualBler(u;t) for UE u at time t.Residual UL BLER: This is the residual UL BLER for UE u after the configured number of HARQ retransmissions between UE and DU whenever needed. This is denoted as uLResidualBler(u;t) for UE u at time t.
[0087] Average number of RLC retransmissions for UE u by time t:Average number of RLC DL retransmissions for UE u by time t. This is denoted as avgNumRLCDlReTx(u;t)Average number of RLC UL retransmissions for UE u by time t. This is denoted as avgNumRLCUlReTx(u;t)
[0088] Input traffic can be using TCP (or UDP) as transport protocol. Key parameters for TCP which impact end-to-end throughput include window-size of the sending TCP entity, receiver window size (as advertised by the receiver), Congestion window size (which depends on the type of TCP used too, e.g. TCP Reno, TCP Cubic etc.), round-trip-time (which includes BH DL / UL latency, MH DL / UL latency and DL / UL scheduling delay over the air-interface in this 5G scenario where input traffic source is hosted on the same server as UPF or in close proximity to UPF), DL / UL PER over BH and MH, DL / UL BLER, Maximum Segment Size (MSS) indicating maximum data that can be sent in a one TCP segment, CU-DU flow control parameters and other parameters described as follows.
[0089] Parameters related to cell configuration which impact throughput (of a DRB) are specified here and include: o Multiplexing mode (e.g. FDD or TDD) o Configured channel bandwidth of this carrier (for DL / UL), o Number of spatial streams used with this carrier (e.g. 4 in DL and 2 in UL)27
[0090] Peak Cell Throughput Validation: As discussed earlier, this is a test that operators carry out to evaluate performance of a base station. With this, there is only one UE in the cell and the operator wants to see that this UE gets the throughput which is equal to the peak throughput which can be achieved in that cell and there is no dip in this throughput for the duration of the test. Some UEs may join a cell and may get handed over to neighboring cells, and this peak cell throughput test can be done on any chosen UE (which supports this capability for peak cell throughput) by temporary reducing traffic (to zero) for other UEs in that cell.
[0091] Maximum Cell Throughput Validation: As discussed earlier, there are factors such as Downlink (DL) Mid-Haul (MH) Latency, Uplink MH Latency, DL Backhaul (BL) Latency and UL Latency that also impact end-to-end throughput for TCP applications. As the MH or BH latency increases (and goes above a threshold), it is not always possible for a UE to get the peak cell throughput, for example, when it is using TCP type of transport protocols. Also, it is not always possible to create a scenario where UE always keeps reporting maximum possible CQ1 and where DL and UL Block Error Rate (BLER) being experienced by the UE is zero (or almost zero). In such cases, the operator wants to verify the maximum possible throughput for that UE. Note that it can be less than the theoretical peak cell throughput possible in that cell, but the operator wants to maximize throughput for each DRB
[0092] One DRB with 5QI 9 for TCP traffic can be created for this peak or maximum cell throughput test for the chosen UE. Full-buffer traffic model can be used as the input traffic model for this test (i.e. there is always data to send for this DRB from the input source) and this input traffic generator is hosted very close to the UPF (e.g. on the same server where UPF is located) with negligible delay between UPF and this input traffic generator.
[0093] BLER and PER related parameters (i.e. avgDlBler, avgUlBler, avgMhDIPer, avgMhUlPer, avgBhDIPer and avgBhUlPer) are kept at zero (or very close to zero) for the duration of this test. Note that such a test could either be carried out in a laboratory (using a wired or over-the-air test setup) or in the field. If test is being carried out in the field (or28over-the-air in a laboratory), location of UE is carefully chosen such as that it can support the maximum possible MCS (Modulation and Coding Scheme) and the rank (for MIMO system) for the duration of this test.
[0094] Some parameters which impact end-to-end throughput for such scenarios (i.e. peak cell throughput test with zero BLER and zero PER over MH and BH) include 1) CU-DU flow control related parameters such as maximum allowed buffer for RLC SDUs for that DRB in DU (i.e. maxAllowedRLCBuffer) and the minimum rate at which DDDS is sent from DU to CU-UP as part of flow control feedback for that DRB (i.e. minRateMhUlDDDS), 2) average MH and BH latencies for DL and UL (i.e. avgMhDILatency, avgMhUlLatency, avgBhDILatency, avgBhUlLatency) and 3) TCP related parameters (such as TCP window sizes and MSS). TCP sender (e.g. at TCP traffic source with full-buffer traffic model, hosted close to UPF) and receiver (e.g. at UE) window sizes can be set to large enough value for such peak throughput tests. Average MH and BH latencies can be configured to low values (e.g. MH DL / UL latency less than a pre-configured threshold and same, or different preconfigured threshold for BH DL / UL latency) for peak throughput tests. An operator would typically also want to increase these MH / BH latencies to evaluate impact of doing this on the end-to-end throughput in 0-RAN networks.
[0095] FIG. 20 shows the following: maximum buffer space allowed for DRB d at time t, maxAllowedRLCBuffer(d;t), minimum rate at which DDDS messages are sent from DU to CU-UP for this DRB at time t, minRateMhUlDDDS(d;t), observed (average) downlink data rate for DRB d over mid-haul, denoted as observedAvgMhDlDataRate(d;t) at time t (i.e. average data rate over the time interval (0,t) if the data communication for this DRB started at time 0), and numRLCDropped(d;t) which is the number of dropped RLC SDUs (by time t) in the DU due to lack of buffer space for DRB d.29
[0096] In FIG. 21, at block 211, shows various parameters which influence end-to- end TCP throughput in 0-RAN networks including DRB d belonging to UE u in cell m at time t. These parameters include the following:Cell ConfigurationUE DistributionAverage MH DL Latency: avgMhDILatency (t) at time t.Average MH UL Latency: avgMhUlLatency (t) at time t.Average BH DL Latency: avgBhDILatency (t) at time t.Average BH UL Latency: avgBHUlLatency (t) at time t.Average MH DL Packet Error Rate (PER): avgMhDlPer(t) at time tAverage MH UL PER: avgMhUlPer(t) at time tAverage BH DL PER: avgBhDlPer(t) at time tAverage BH UL PER: avgBhUlPer(t) at time tFraction of MH DL capacity used by DRBs corresponding to UEs in the cell m: fracMhDlCapacity(m,t)Fraction of MH UL capacity used by DRBs corresponding to UEs in the cell m: fracMhUlCapacity(m,t)DL BLER (for UE u in cell m): dLBler(u,t) for UE u in cell m at time tUL BLER (for UE u in cell m): uLBler(u,t) for UE u in cell m at time tMaximum buffer space allowed for (DL) RLC SDUs for DRB d in DU: maxAllowedRLCBuffer(d;t) .30Minimum rate of sending flow control (DDDS) messages form DU to CU-UP for DRB d: minRateMhUlDDDS(d; t) for DRB d at time tTCP related parameters (sender window size, MSS, type of congestion control method, receiver window size etc.)Input traffic model for TCP source (e.g. full-buffer traffic model)CSI reported by UEs,Interference from neighboring cells and other relevant parameters
[0097] In FIG. 21, block 212 shows the next level of observed parameters in the DU, including DRB d belonging to UE u in cell m at time t:Number of dropped RLC SDUs for DRB d (by time t) due to lack of buffer space in the DU: numRLCDropped(d; t)Total number of time slots where the buffer space for RLC SDUs for DRB d in the DU was empty (i.e. no RLC SDU to transmit) until time t, numSlotsEmptyRLCBuffer(d;t). For example, t could be equal to 1000 slots (with 1 slot being equal to 1 ms), TCP throughput test starts at time zero and there are 50 time slots when the RLC queue in the DU is empty for DRB d. In this case, numSlotsEmptyRLCBuffer(d;t) = 50. Note that value of numSlotsEmptyRLC Buffer (d;t) can also be impacted due to the MH PER, MH latency and the capacity available for traffic for this DRB over the MH.Observed data rate for DL RLC SDUs over mid-haul: observedAvgMhDlDataRate(d;t) for DRB d at time t. Note that its value could also be impacted due to the capacity available for traffic for this DRB over the MH, the MH PER and MH latency.Average number of RLC DL retransmissions for UE u by time t: avgNumRLCDlReTx(u;t) where DRB d corresponds to UE u in cell m,31Number of RLC UL retransmissions for UE u by time t: avgNumRLCUlReTx(u;t) where DRB d corresponds to UE u in cell m,Residual DL BLER: dLResidualBler(u;t) for UE u at time tResidual UL BLER: uLResidualBler(u;t) for UE u at time tMGS, rank used for the UE u in DL and UL
[0098] As in FIG. 21, parameters shown in blocks 211 and 212 impact end-to-end throughput for a DRB.
[0099] As discussed earlier, BLER, MH PER and BH PER related parameters (i.e. avgDlBler, avgUlBler, avgMhDIPer, avgMhUlPer, avgBhDIPer and avgBhUlPer) are kept equal to zero for some such tests (which are conducted to verify peak cell throughput). It also results in number of RLC retransmissions for UE u by time t (i.e. avgNumRLCDlReTx(u;t), avgNumRLCUlReTx(u;t)), and residual BLER for UE u (i.e. dLResidualBler(u;t) and uLResidualBler(u;t)) to have the value equal to zero. Alternatively, values of BLER, MH PER and BH PER related parameters can be very low and this can result in very low values of avgNumRLCDlReTx(u;t), avgNumRLCUlReTx(u;t), dLResidualBler(u;t) and uLResidualBler(u;t).
[0100] MH and BH latencies are also kept very low for peak throughput test (if possible). For example, each of the following can be kept less than a pre-specified threshold.Average MH DL Latency: avgMhDILatency (t) at time t.Average MH UL Latency: avgMHUlLatency (t) at time t.Average BH DL Latency: avgBhDILatency (t) at time t.Average BH UL Latency: avgBhUlLatency (t) at time t.
[0101] If the right amount of buffer space, maxAllowedRLC Buffer (d;t), is not provided in the DU for this DRB d (at time t), this can result in DU dropping RLC SDUs in32the DU. This can also result in the flow control module at DU giving less opportunity to CU- UP to send PDCP PDUs (i.e. RLC SDUs) towards DU, as DU can indicate low or zero Desired Buffer Size and optionally low or zero value of Desired Data Rate to CU-UP for this DRB d over time intervals when buffer space allocated to DRB d in the DU is fully or almost occupied. Such events can degrade end-to-end TCP throughput for this DRB.
[0102] Note that it is also not possible to allocate unlimited buffer space to a DRB in the DU as a DU needs to support large number of DRBs, and the system also needs to allocate right amount of memory for those DRBs. For example, it may not be feasible to allocate same amount of maximum memory (i.e. maxAllowedRLCBuffer) to all the DRBs (for same QoS class, e.g. 5Q1 9) in the DU. Also, as discussed earlier, UEs join a cell and get handed over to a neighboring cell, and this peak cell throughput test may need to be done on any chosen UE (which supports this capability for peak cell throughput).
[0103] Also, if a good value of minRateMhUlDDDSfd; t) is not chosen for this DRB d (at time t), it can slow down the rate at which CU-UP is sending DL PDCP PDUs to DU. This can increase latency for some packets in the 0-RAN Network and degrade end-to-end TCP throughput.
[0104] Thus, it becomes useful to optimize CU-DU flow control related parameters. For example, provide right amount of buffer space for RLC SDUs for DRB d in the DU and choose an optimal value of the minimum rate at which DU should be sending flow control feedback (i.e. DDDS) to CU-UP for this DRB d.
[0105] CU-DU Flow Control optimization related decisions using reinforcement learning: In this section, the mapping of the CU-DU Flow Control optimization related decision making process (denoted as CUDU-FCtrl-Opt) to a reinforcement learning problem by formalizing it using an MDP is disclosed. As mentioned earlier, MDP involves four elements: States; Actions; Costs / Rewards; and Transition Probabilities. These elements are represented as follows: o S: A set of finite States S330 A: A set of finite Actions A o C(s,a): Immediate cost (or expected immediate cost) incurred after transitioning from state s to state s' , due to action a.0 P: represents the Transition probability matrix corresponding to state space S and action space A. o P(s’|s,a): Transition Probability of landing in state s’ when action a is taken at state s.
[0106] From the various parameters (including performance measures] that are communicated from the DU, CU and other entities to the CU-DU Flow Control Optimization module, denoted as CUDU-FCtrl-Optimization module, several of these are provided as state variables for the reinforcement learning (RL) module. The range of values taken by each state variable is quantized to n levels, where n is a finite value (e.g., n=2 or 4 or 6 or a higher number, depending on the parameter being quantized).
[0107] The set of state variables for DRB d (in cell m) is denoted as 5™ and it includes the following:Cell Configuration. This is kept static for peak or maximum cell throughput test. In general, some cell configuration parameters can change dynamically. For example, the number of allowed spatial streams can be dynamically reduced for energy saving.UE Distribution. As discussed earlier, one UE (in the cell) is used for peak cell throughput test and this test is done in a way where this UE can receive highest possible MCS and the rank (which is needed to achieve the peak cell throughput with that UE).Average MH DL Latency: avgMhDILatency (t) at time t. As discussed earlier, this is kept below a pre-specified threshold (such as 4 ms) for the peak (or the maximum) cell throughout test.34Average MH UL Latency: avgMHUlLatency (t) at time t. This is also kept below a prespecified threshold for the peak (or the maximum) cell throughout test.Average BH DL Latency: avgBhUlLatency (t) at time t. This is also kept below a prespecified threshold for the peak (or the maximum) cell throughout test.Average BH UL Latency: avgBhUlLatency (t) at time t. This is also kept below a prespecified threshold for the peak (or the maximum) cell throughout test.Average MH DL Packet Error Rate (PER): avgMhDlPer(t) at time t. As discussed earlier, this is kept equal to zero (or below a very low pre-configured threshold).Average MH UL PER: avgMhUlPer(t) at time t. This is also kept equal to zero (or below a very low pre-configured threshold) for peak (or the maximum) cell throughput test.Average BH DL PER: avgBhDlPer(t) at time t. This is also kept equal to zero (or below a very low pre-configured threshold) for peak (or the maximum) cell throughput test.Average BH UL PER: avgBhUlPer(t) at time t. This is also kept equal to zero (or below a very low pre-configured threshold) for peak (or the maximum) cell throughput test.Fraction of MH DL capacity used by DRBs corresponding to UEs in the cell m: fracMhDlCapacity(m,t). This can keep changing dynamically. It should be on the lower side (e.g. below a pre-configured threshold) for it to have no adverse impact on end-to-end TCP throughput of DRB d.Fraction of MH UL capacity used by DRBs corresponding to UEs in the cell m: fracMhUlCapacity(m,t). This can keep changing dynamically. It should be on the lower side (e.g. below a pre-configured threshold) for it to have no adverse impact on end-to-end TCP throughput of DRB d.35DL BLER (for UE u in cell m): dLBler(u,t) for UE u in cell m at time t. As discussed earlier, the peak (or the maximum) cell throughput test is conducted in a way such that this stays very low (e.g. almost zero) and below a pre-configured threshold.UL BLER (for UE u in cell m): uLBler(u,t) for UE u in cell m at time t. As discussed earlier, the peak (or the maximum) cell throughput test is conducted in a way such that this stays very low (e.g. almost zero) and below a pre-configured threshold.Maximum buffer space allowed for (DL) RLC SDUs for DRB d in DU: maxAllowedRLCBuffer(d;t). This could be static or can change dynamically based on some policies.Minimum rate of sending flow control (DDDS) messages form DU to CU-UP for DRB d: minRateMhUlDDDS(d; t) for DRB d at time t.TCP related parameters such as sender window size, MSS, type of congestion control method, receiver window size etc. These are set to values such as that TCP source and destination can support the peak cell throughput for DRB d.Input traffic model for TCP source (e.g. full-buffer traffic model)CSI reported by UEs. Peak (or the maximum) cell throughput test is conducted such that this stays equal to the highest (or the maximum) possible CSI.Interference from neighboring cells. Peak (or the maximum) cell throughput test is conducted such that there is no or very low (e. SDUs g. below a pre-configured threshold) interference from neighboring cells.Number of dropped RLC SDUs for DRB d (by time t) due to lack of buffer space in the DU: numRLCDropped(d; t). This can change dynamically and is monitored at the DU.Total number of time slots where the buffer space for RLC SDUs for DRB d in the DU was empty (i.e. no RLC SDU to transmit) until time t, numSlotsEmptyRLCBuffer(d;t).Observed data rate for DL RLC SDUs over mid-haul: observedAvgMhDlDataRate(d;t)36for DRB d at time t. This can change dynamically and is monitored at the DU for DE RLC SDUs.Average number of RLC DL retransmissions for UE u by time t: avgNumRLCDlReTx(u;t) where DRB d corresponds to UE u in cell m. This can change dynamically and is monitored at the DU. It is expected to have a very low value (e.g. below a pre-configured threshold) for the peak (or the maximum) cell throughout test.Number of RLC UL retransmissions for UE u by time t: avgNumRLCUlReTx(u;t) where DRB d corresponds to UE u in cell m. This can change dynamically and is expected to have a very low value (e.g. below a pre-configured threshold) for the peak (or the maximum) cell throughout test.Residual DE BEER: dLResidualBler(u;t) for UE u at time t. This can change dynamically and is monitored at the DU. It is expected to have value (almost) equal to zero for the peak (or the maximum) cell throughout test.Residual UL BEER: uLResidualBler(u;t) for UE u at time t. This can also change dynamically and is expected to have value (almost) equal to zero for the peak (or the maximum) cell throughout test.MCS, rank used for the UE u in DE and UL. These are expected to be the highest possible values for the peak (or the maximum) cell throughput test.Observed (over-the-air) throughput for DRB d, observedThroughput(d;t),
[0108] The state space is finite as there are finite state variables and each state variable takes finite values. Also, note that some of the state variables, specifically those related to TCP window sizes at the sender and receiver, fraction of capacity given to this DRB over MH, and latencies over MH / BH and MH / BH PER, influence other variables such as numSlotsEmptyRLCBuffer and observedAvgMhDIDataRate, and need not directly be used as input to the RE model.37
[0109] Cost (or reward): For DRB d, cost (at time t) is computed as weighted sum of the following components:Cost associated with number of dropped RLC SDUs due to lack of buffer space in DU (for that DRB d), numRLCDropped(d;t), is given as C[(d; t) = (numRLCDropped(d; t))^5. Here p is equal to or greater than 0 (and can be chosen to have value greater than 1 too). Choosing a value of p greater than 1 increases this cost factor more aggressively as the number of dropped RLC SDUs increases and can be used for DRBs carrying data over TCP type of transport protocols. Value of p can be chosen to be less than 1 for DRBs which are using UDP type of transport protocols.Cost associated with Observed data rate for DL RLC SDUs over mid-haul, observedAvgMhDlDataRate(d;t), is given as follows: f -1- - ) when observedAvgMhDlDataRate(d;t) is\observedAvgMHDlDataRate(d;t) not equal to zero. Here e is equal to or greater than 0. If observedAvgMhDlDataRate(d;t) is equal to zero, this cost factor is not used (and is defined to be zero).Cost associated with number of slots where the buffer space for RLC SDUs for DRB d in the DU was not occupied until time t, numSlotsEmptyRLCBuffer(d;t), is given asCm(d; t) = (numSlotsEmptyRLCBuffer(d; t))R. Here p is equal to or greater than 0 (and can be chosen to have value greater than 1 too). Choosing value of p greater than 1 increases this cost factor more aggressively as the number of slots where DU has nothing to send to that UE increases and can be used for DRBs carrying data over TCP type of transport protocols or under light load conditions in the network (e.g. when number of DRBs is below a pre-defined threshold in a cell). Value of p can be chosen to be less than 1 for DRBs which are using UDP type of transport protocols or for the scenarios where number of DRBs in a cell is above a pre-defined threshold.38Cost associated with maximum occupied buffer space in DU for DRB d: CIV(d; t) = (maxAllowedRLCBuffer(d; t) — maxOccupiedRLCBuffer(d; t))v. Here v is equal to or greater than 0. This factor is used as DU buffer space (especially fast memory) is also a critical resource and this method works to increase utilization of this buffer space and reduce wastage of buffer space at the same time while meeting cell-level and per-DRB performance targets. Value of v can be set to less than 1 when difference between maxAllowedRLCBuffer(d;t) and maxOccupiedRLCBuffer(d;t) is before a pre-defined threshold. Otherwise, v can be set to a values higher than 1.Cost (e.g. due to software and hardware overhead of sending DDDS over MH) associated with the rate at which DDDS is sent from DU to CU-UP for DRB d:Cv(d; t) = 0, if minRateMhUlDDDS(d;t) is less than (or equal to) thresholdRateMhUlDDDS(d); Otherwise, Cv(d; t) =( - — - ) . Here, thresholdRateMhUlDDDS(d) is a\maxflateM / iUtoDDS(d)-mi7iRateM / il / ZDDDS(d;t)zv Jpre-specified threshold for DRB d and is chosen such that thresholdRateMhUlDDDS(d) < maxRateMhUlDDDS(d)). Also, Y is in the range [0, 1],Cost associated with the observed (over-the-air) throughput for DRB d, observedThroughput(d;t), is given as: Cvl(d; t) = (targetPeakThroughput(d; t) — observedThroughput(d; t))\ Here A is equal to or greater than 0. For peak throughput validation with SU-MIMO (Single User - M1M0), value of A can be chosen to be equal to 1. For MU-M1M0 (Multi-User M1M0) case (e.g. when peak throughput test is done with MU-MIMO for small number of UEs in a cell), value of A can be chosen to be less than 1.
[0110] Cost for this DRB d (at time t), C(d;t), is given asC(d;t) = w, * C^d; t) + wn* CH(d; t) + wm* Cm(t) + wlv* CiV(d; t) + wv*Cv(d;t) + wVI* CVI(d; t)Here, w7, w7 / , wm, wlv, wvand wvlare weights associated with cost components C;, CH, Cm, CIVCvand CV1respectively. These weights could be in the range of [0,1] though39higher values are also possible.
[0111] This method uses the above cost function to help find optimal values of CU- DU flow control parameters for the peak (or the maximum) cell throughput verification case or to maximize throughput for each DRB for any given scenario.
[0112] For a non-ideal test scenario, (e.g. where MH or BH latency is above a prespecified threshold or where BLER is above a pre-specified threshold), above cost function is used to find optimal values of CU-DU flow control parameters to maximize throughput (for a given DRB). In this case, it is not possible to achieve theoretical value of peak cell throughput and wvlis set equal to zero. Other parts of the cost function are used to find CU-DU flow control parameters which help achieve the maximum possible throughput in such a scenario.
[0113] Action (A): Following parameters are updated as part of the Action taken by the RL module to optimize CU-DU flow control to achieve maximum (or peak) throughput for DRB d corresponding to UE u in cell m:Minimum rate of sending flow control (DDDS) messages from DU to CU-UP for DRB d, minRateMhUlDDDS(d; t) for DRB d at time t : Here, the action could be to keep its value same or increase by a(d) or decrease by a(d)). Value of a(d) is chosen to be greater than zero.Maximum buffer space allowed for (DL) RLC SDUs for DRB d in DU, maxAllowedRLCBuffer(d;t) : Here the action could be to keep its value same or increase by 8(d) or decrease by 8(d). Value of 8(d) is chosen to be greater than zero.
[0114] It is noted that increments or decrements by a(d) and 8(d) are governed by the boundaries of respective state variables (i.e. by boundaries of minRateMhUlDDDS and maxAllowedRLCBuffer respectively). The RL model works on finite state space (with finite values) and finite action space.
[0115] Transition probability matrix is of finite size as the state space is finite and the action space is finite. For the unknown transition probability matrix, the matrix is40initialized with zero and update it in the following manner. From state s, after taking action “a”, if the system moves to state s', an update is implemented as P. Later, at the same state s and taking the same action “a”, if the system moves to state s", an update is implemented as PIf the system moves to the state s' at a later point of time from same state s and taking the same action "a”, an update is implemented as P(s' |s, a) = 2 / 3, and P(s" |s, a) = 1 / 3. The transition probabilities are updated based on i) the different states to which the system moves (from a given state for the same action) and ii) the number of times the system moved to each such state.
[0116] After initial learning of transition probabilities as above, the RL can be run with an exploration and exploitation strategy. In the exploration stage, the action is chosen to not change or to change the values of parameters (i.e. increase or decrease) at random. In the exploitation stage, the action is chosen that incurred the minimum cost until now (i.e., during the training phase). For example, exploration can be chosen with probability € and exploitation can be chosen with probability (l-€). The parameter, € , is used to control the amount of exploration vs. exploitation in the RL method. The value function is computed as explained earlier for the RL method. For a given state, an optimal action is chosen as per the policy learnt using the RL method.
[0117] RL method was used to find the optimal or near-optimal values of flow control parameters above to maximize throughput for any given DRB. It is also possible to use other Al / ML models such as deep neural network (DNN) where inputs include some of the state variables as specified earlier and the output is specified by minRateMhUlDDDS(d; t) and maxAllowedRLCBuffer(d;t).
[0118] The CUDU-FCtrl-Optimization module can be located at CU, DU, Near-RT-RIC, non-RT-RIC or another analytics server.
[0119] METHOD IB
[0120] For the case where the CUDU-FCtrl-Optimization module is located in the non-RT-RIC, various performance measures and counters (as given in FIG. 21 and discussed earlier) are communicated from gNB-DU and gNB-CU to the non-RT-RIC via 0141interface. As shown in FIG. 22, gNB-DU communicates DU related performance measures and counters, and gNB-CU communicates CU related performance measures and counters to the non-RT-RIC. Some parameters that are not directly monitored by CU and DU (such as TCP window sizes etc.) are communicated to the non-RT-RIC as part of contextual information. CU-DU Flow Control Optimization module, CUDU-FCtrl-Optimization module (with Reinforcement learning as specified earlier), runs as part of non-RT-RIC. It analyzes various parameters and decides the values of CU-DU flow control parameters (such as minRateMhUlDDDS(d; t), maxAllowedRLCBuffer(d;t)) to help optimize end-to-end performance for each DRB d. These selected values of CU-DU flow control parameters are communicated from non-RT-RIC to CMS (Centralized Management System), which further communicates these to the gNB-DU.
[0121] METHOD IC
[0122] FIG. 23 shows an implementation where the CUDU-FCtrl-Optimization module is hosted as part of the Near-RT-RIC. The Near-RT-RIC subscribes for various performance measures and counters for CUDU-FCtrl-Optimization xApp from DU and CU. As shown in FIG. 23, the E2 interface (and associated protocol) between the gNB-DU and the Near-RT-RIC is enhanced to communicate performance measures and counters which are shown in FIG. 21 (and were specified earlier). Similarly, as shown in FIG. 23, the E2 interface between the gNB-CU and the Near-RT-RIC is enhanced to communicate performance measures and counters which are shown in FIG. 21 (and were discussed earlier). TCP related parameters (such as TCP window size, and so on) are communicated as part of contextual information to the Near-RT-RIC. The CUDU-FCtrl-Optimization xApp at the Near-RT-RIC analyzes these parameters and decides the optimal values of CU-DU flow control parameters (such as minRateMhUlDDDS(d; t), maxAllowedRLCBuffer(d; t)) to help optimize end-to-end performance of DRB k. These are communicated from the Near- RT-RIC to the gNB-DU as in FIG. 23, and the E2 interface (and the associated protocol), between gNB-DU and the Near-RT-RIC, is enhanced to communicate these chosen flow control parameters. The gNB-DU can apply these parameters directly and inform about these to CMS (centralized management system). Alternatively, gNB-DU can inform about42these parameters to CMS and wait for confirmation from CMS before applying these to itself (i.e. to the gNB-DU).
[0123] METHOD ID
[0124] FIG. 24 shows an implementation where one part of the CUDU-FCtrl- Optimization module, denoted as CUDU-FCtrl-Optimization (II) module, runs on the Near- RT-R1C and subscribes to various parameters from gNB-CU and gNB-DU (as was described above for FIG. 23 also). The CUDU-FCtrl-Optimization (II) module communicates these parameters to another part of CU-DU Flow Control Optimization module, denoted as CUDU- FCtrl-Optimization (I), which is running on the non-RT-RIC. This CUDU-FCtrl-Optimization (I) module and the corresponding rApp, analyze various parameters and decide the optimal values of the CU-DU flow control parameters (such as minRateMhUlDDDS(d; t), maxAllowedRLCBuffer(d;t)) as described earlier. The non-RT- RIC communicates these to CMS (centralized management system) which communicates these to gNB-DU and asks gNB-DU to start using updated CU-DU flow control parameters for DRB k.
[0125] METHOD IE
[0126] FIG. 25 shows an implementation where AI / ML training is done at CU-CP (or at another server which CU-CP can use for AI / ML training) and the CUDU-FCtrl- Optimization module at the CU-CP decides optimal values of the CU-DU flow control parameters. Steps show in FIG. 25 are given below:I) gNB-DU communicates performance measures and counters collected at gNB-DU (as in FIG. 21) to gNB-CU-CP. The F1AP (Fl Application Protocol, 3GPP TS 38.473) running over Fl-C interface is enhanced to carry these parameters by adding new messages to carry these or by adding new fields or using some reserved fields (i.e. which are not used at present) in the existing F1AP messages.II) gNB-CU-UP communicates performance measures and counters collected at gNB- CU-UP (as in FIG. 21) to gNB-CU-CP. The E1AP protocol (El Application Protocol,433GPP TS38.463) running over El interface is enhanced to carry these parameters by adding new messages to carry these or by adding new fields or using some reserved fields (i.e. which are not used at present) in the existing El messages.III) gNB-CU-CP collects performance measures and counters generated locally (i.e. at gNB-CU-CP).IV) gNB-CU-CP continues to train the Al / ML (such as the Reinforcement Learning) model.V) Al / ML model is downloaded to gNB-CU-CP (if training was happening at another entity).VI) gNB-CU-CP uses this AI / ML model to choose the optimal values of CU-DU flow control parameters (such as such as minRateMhUlDDDS(d; t), maxAllowedRLCBuffer(d; t)).VII) gNB-CU-CP communicates these updated CU-DU flow control parameters to gNB-DU. The F1AP (Fl Application Protocol, 3GPP TS 38.473) running over Fl-C interface is enhanced to carry these parameters by adding new messages to carry these or by adding new fields or using some reserved fields (i.e. which are not used at present) in the existing F1AP messages.VIII) gNB-DU can apply these parameters directly and inform about these to CMS. Alternatively, gNB-DU informs about these parameters to CMS and waits for confirmation from CMS and then applies these to the gNB-DU.
[0127] METHOD IF
[0128] FIG. 26 shows an implementation where AI / ML training (for Reinforcement Learning) is done at CU-UP and the CUDU-FCtrl-Optimization module at the CU-CP decides optimal values of the CU-DU flow control parameters. Steps shown in FIG. 25 are given below:44I) gNB-DU communicates performance measures and counters collected at gNB-DU (as in FIG. 21) to gNB-CU-CP as described earlier.II) gNB-CU-CP communicates performance measures and counters received from gNB-DU, and performance measures and counters generated locally at gNB-CU-CP to gNB-CU-UP. The E1AP protocol (El Application Protocol, 3GPP TS38.463) running over El interface is enhanced to carry these parameters by adding new messages to carry these or by adding new fields or using some reserved fields (i.e. which are not used at present) in the existing El messages.Ill) gNB-CU-UP collects performance measures and counters generated locally (i.e. at gNB-CU-UP).IV) gNB-CU-UP continues to train the AI / ML (such as the Reinforcement Learning) model as described earlier.V) AI / ML model is downloaded to gNB-CU-CP.VI) gNB-CU-CP uses this AI / ML model to fond the optimal values of CU-DU flow control parameters (such as such as minRateMhUlDDDS(d; t), maxAllowedRLCBuffer(d; t)).VII) gNB-CU-CP communicates these updated CU-DU flow control parameters to gNB-DU. The F1AP (Fl Application Protocol, 3GPP TS 38.473) running over Fl-C interface is enhanced to carry these parameters by adding new messages to carry these or by adding new fields or using some reserved fields (i.e. which are not used at present) in the existing F1AP messages.VIII) gNB-DU can apply these parameters directly and inform about these to CMS. Alternatively, gNB-DU informs about these parameters to CMS and waits for confirmation from CMS and then applies these to the gNB-DU.
[0129] METHOD II45
[0130] Described is an implementation of a method that enhances above methods for the scenario where the operator needs to find optimal values of CU-DU flow control parameters to optimize performance for large number of DRBs in an 0-RAN network.
[0131] State Variables for the reinforcement learning module: State variable for a given DRB were specified in Method IA above. These are used for each DRB in the DU now.
[0132] Cost (or reward): For cell g, cost at time t, denoted as Ccell(g; t) , is computed as weighted sum of cost of all the DRBs in that cell:Here, ydis the weight assigned to C(d;t) and is in the range [0, 1]. Value of ydcan depend on 5QI (or other QoS indicators) too. For example, value of ydcan be higher for delay sensitive applications compared to non-delay sensitive applications.As before, cost of DRB d at time t, C(d;t), is given asC(d;t) = W] * Ci(d; t) + wH* CH(d; t) + wm* + wlv* CiV(d; t) + wv* Cv(d;t) + wvl* CVI(d; t)In this case, wVIis set equal to zero and other weights (i.e. w;, wH, win, wIV, wv) are in the range [0, 1] .
[0133] For the corresponding DU, cost at time t, denoted a) , is computed as weighted sum of cost of all the DRBs across all the H cells in that DU:Here, zgis the weight associated with Ccel1(g; t) and is in the range [0, 1].
[0134] Action (A): Following parameters are updated as part of the Action taken by the RL module to optimize CU-DU flow control for DRBs across several cells in a DU to optimize performance of each DRB d:46Minimum rate of sending flow control (DDDS) messages form DU to CU-UP for DRB d, minRateMhUlDDDSfd; t) for DRB d at time t : Here, the action could be to keep its value same or increase by a(d) or decrease by a(d)). Value of a(d) is chosen to be greater than zero.Maximum buffer space allowed for (DL) RLC SDUs for DRB d in DU, maxAllowedRLCBuffer(d;t) : Here the action could be to keep its value same or increase by 8(d) or decrease by 6(d). Value of 6(d) is chosen to be greater than zero.
[0135] Note that the total amount of buffer space which can be allocated to all the DRBs in cell g at time t is upper bounded by maxRLCBufferCell(g;t).Ndrb(g;t) maxAllowedRLCBuffer(d; t) < maxRLCBufferCell(g; t) d=l
[0136] Let a DU support maximum of H cells. Total amount of memory for RLC SDUs across all cells of this DU is denoted as maxRLCBufferDU(t) at time t. Following is maintained as different types of actions are taken by the Al / ML module.H maxRLCBufferCell(g; t) < maxRLCBufferDU(t)5=1
[0137] This buffer constraint is also used for search space pruning. If the RL module picks up a point (i.e. action) in the search space which violates total buffer constraint, that action will not be applied (and will not be communicated to the DU). The RL module will be asked to pick up another point in the search space based on its exploration and exploitation strategy. In the end, it gives a solution which does not violate total buffer constraints and is near-optimal for flow control.
[0138] As shown above, the maximum rate at which DU can send DDDS messages to CU-UP for a given DRB is kept bounded by an upper threshold (e.g. due to software and hardware overhead associated with this).47
[0139] Methods shown in FIG. 22 to FIG. 26 are valid for this implementation too, though these are applicable for all the DRBs in the DU (and CU) with this method.
[0140] This method helps find the optimal set of CU-DU flow control parameters for each DRB d in the DU to maximize throughput for all the DRBs in a given system. The previous method (i.e. Method IA) finds the optimal set of CU-DU flow control parameters for a given DRB to maximize throughput of this specific DRB. Both of these methods can be used together for the deployment scenarios where there are large number of UEs in each cell..
[0141] TCP is used as the transport protocol to describe the methods above. In general, these methods are also applicable when other transport protocols (such as QUIC: Quick UDP Internet Connection, UDP, TCP Prague with low-latency, low loss and scalable throughput, etc.) are used in the system.
[0142] Method IA, IB, IC, ID, IE and IF are collectively referred to as Method I in this disclosure.
[0143] Reference is made to Third Generation Partnership Project (3GPP) and the Internet Engineering Task Force (IETF) and related standards bodies in accordance with embodiments of the present disclosure. The present disclosure employs abbreviations, terms and technology defined in accord with Third Generation Partnership Project (3GPP) and / or Internet Engineering Task Force (IETF) technology standards and papers, including the following standards and definitions. 3GPP and IETF technical specifications (TS), standards (including proposed standards), technical reports (TR) and other papers are incorporated by reference in their entirety hereby, define the related terms and architecture reference models that follow.3GPP TS 23.501 V 18.1.0 2024-06-263GPP TS 28.500 V 18.0.0 2024-04-203GPP TS 28.541 V 18.7.0 2024-04-05483GPP TS 38.300 V 18.1.0 4-03-20243GPP TS 38.401 V 18.1.0 2024-03-29
[0144] Abbreviations5GC: 5G Core Network5G NR: 5G New Radio5QI: 5G QoS IdentifierACK: AcknowledgementAl: Artificial IntelligenceAl / ML (or AIML): Artificial Intelligence and Machine LearningAM: Acknowledged ModeAPN: Access Point NameARP: Allocation and Retention PriorityBLER: Block Error RateBO: Buffer OccupancyBS: Base StationBSR: Buffer Status ReportCMS: Centralized (or Configuration) Management SystemCNN: Convolution Neural NetworkCP: Control PlaneCSI: Channel State InformationCU: Centralized UnitCU-CP: Centralized Unit - Control PlaneCU-UP: Centralized Unit - User Plane49DL: DownlinkDDDS: DL Data Delivery StatusDNN: Data Network NameDNN: Deep Neural NetworkDQN: Deep Q NetworkDRB: Data Radio BearerDU: Distributed Unit eNB: evolved NodeBEPC: Evolved Packet CoreEN-DC: E-UTRAN New Radio Dual ConnectivityGBR: Guaranteed Bit Rate gNB: gNodeBGTP-U: GPRS Tunnelling Protocol - User PlaneIP: Internet ProtocolLI: Layer 1L2: Layer 2L3: Layer 3L4S: Low Latency, Low Loss and Scalable ThroughputLC: Logical ChannelLESS: Low Energy Scheduler SolutionLSTM: Long Short-Term MemoryMAC: Medium Access ControlMDP: Markov Decision ProcessMIB: Master Information Block50ML: Machine LearningMR-DC: Multi-RAT Dual ConnectivityMSS: Maximum Segment SizeMU-MIMO: Multi-User Multiple-Input Multiple-OutputNACK: Negative AcknowledgementNAS: Non-Access StratumNG-RAN: Next Generation Radio Access NetworkNR-U: New Radio - User PlaneNS1: Network Slice InstanceNSSI: Network Slice Subnet InstanceNWDAF: Network Data Analytics FunctionO-RAN: Open Radio Access NetworkOAM: Operations, Administration MaintenancePDB: Packet Delay BudgetPDCP: Packet Data Convergence ProtocolPDU: Protocol Data UnitPER: Packet Error RatePF: Proportional FairPHY: Physical LayerPRB: Physical Resource BlockQCI: QoS Class IdentifierQFI: QoS Flow IdentifierQoS : Quality of ServiceRAN: Radio Access Network51RAT: Radio Access TechnologyRB: Resource BlockRD1: Reflective QoS Flow to DRB IndicationRL: Reinforcement LearningRLC: Radio Link ControlRLC-AM: RLC Acknowledged ModeRLC-UM: RLC Unacknowledged ModeRNN: Recurrent Neural NetworksRQ1: Reflective QoS IndicationRRC: Radio Resource ControlRRM: Radio Resource ManagementRTF: Real-Time Transport ProtocolRTCP: Real-Time Transport Control ProtocolRU: Radio UnitSCTP: Stream Control Transmission ProtocolSD: Slice DifferentiatorSDAP: Service Data Adaptation ProtocolSIB: System Information BlockSLA: Service Level AgreementS-NSSAI: Single Network Slice Selection AssistanceSST: Slice / Service TypeSU-MIMO: Single User Multiple-Input Multiple-OutputTB: Transport BlockTCP: Transmission Control Protocol52TE1D: Tunnel Endpoint IdentifierUE: User EquipmentUP: User PlaneUL: UplinkUM: Unacknowledged ModeUPF: User Plane Function53
Claims
CLAIMS1. A method for Artificial Intelligence and Machine Learning (Al / ML) Centralized Unit- Distributed Unit (CU-DU) Flow Control Optimization comprising: mapping a CU-DU Flow Control optimization related decision making process (CUDU-FCtrl-Opt) for a machine learning (ML) module, the ML module being a reinforcement learning (RL) module or a Markov Decision Process (MDP) module, wherein the machine learning module comprises State function; an Action function, a Cost / Reward function, and Transition Probabilities function; and computing parameters including performance measures from a DU or CU to a CU- DU Flow Control Optimization module (CUDU-FCtrl-Optimization module), the parameters being a set of state variables for' the RL module, wherein a range of values taken by each state variable are quantized to n levels, where n is a finite value.
2. The method of claim 1, wherein n~2 or more.3 The method of claim 1, wherein the set of state variables for a Data Radio Bearer (DRB) d in a cell m is S™.
4. The method of claim 1, wherein the Cost / Reward function is a Cost / Reward function for Data Radio Bearer (DRB) d cos t at time t C(d;t) is computed as a weighted sum comprising a plurality of cost components5. The method of claim 4, wherein C(d;t) = wj* C_1 (d;t)+ wjl* CJI (d;t)+ w_III* CJII (t) + wJV* CJV (d;t)+ w_V*C_V (d;t)+w_VI* C_V1 (d;t), wherein C(d;t) is computed to identify optimal values of CU-DU flow control parameters for a peak cell throughput verification case.
6. The method of claim 1, wherein an Action function comprises: a Minimum rate of sending flow control (DDDS) messages from DU to CU-UP for Data Radio Bearer (DRB) d, minRateMhUlDDDS(d; t) for DRB d at time t; and54a Maximum buffer space allowed for (DL) RLC SDUs for DRB d in DU, maxAllowedRLCBuffer(d;t) .
7. The method of claim 1, wherein the Transition Probability function comprises a transition probability matrix, and the method comprises: updating the transition probabilities based on i) different states to which the system moves from a given state for a same action and ii) a number of times the system moved to each such state.
8. The method of claim 7, wherein the method further comprises: initializing the transition probability matrix with a zero value; updating the transition probability matrix by, from state s, after taking action “a", if the system moves to state s', an update is implemented as P(s^sfa) = 1; later, at the same state s and taking the same action "a”, if the system moves to state s"» an update is implemented as P(sf|sfa) = P(s#<|3* a) = 0.5.; and if the system moves to the state s’ at a later point of time from same state s and taking the same action “a”, an update is implemented as P(s' |s, a) = 2 / 3, and P(s" |s, a) = 1 / 3.
9. The method of claim 7, further comprising: after an initial learning of transition probabilities, running an exploration stage of a Reinforcement Learning (RL) module, where an action is chosen to not change or to change the values of parameters at random; running an exploitation stage of the RL module where an action that incurred a minimum cost until during the training phase is chosen; and at a given state, choosing an optimal action accord with the policy learned by RL55module.
10. A method for Artificial Intelligence and Machine Learning (Al / ML) Centralized Unit- Distributed Unit (CU-DU) Flow Control Optimization comprising: mapping a CU-DU Flow Control optimization related decision making process (CUDU-FCtrl-Opt) for an AI / ML model; and computing parameters including performance measures from a DU or CU to a CU- DU Flow Control Optimization module (CUDU-FCtrl-Optimization module), the parameters being state variables for the Al / ML module.
11. The method of claim 10, wherein the AI / ML model comprises a Deep Neural Network (DNN), wherein the output is by minRateMhUlDDDS(d;t) and maxAllowedRLCBuffer(d;t) .
12. The method of claim 10, wherein the CU-DU Flow Control Optimization module (CUDU- FCtrl-Optimization module), is hosted at the non-RT-RIC.
13. The method of claim 12, further comprising: communicating, via an 02 interface, the performance measures and counters from the DU to the CUDU-FCtrl-Optimization module at the non-RT-RIC and from CU to CUDU-FCtrl-Optimization module at the non-RT-RIC.
14. The method of claim 12, further comprising: running AI / ML at the Ctrl-Optimization module to analyze the performance measures and counters and decide values of CU-DU flow control parameters to optimize end-to-end performance for each Data Radio Bearer (DRB) d, wherein the CU-DU flow control parameters have selected values that are communicated from non-RT-RIC to a CMS (Centralized Management System) to further communicates the flow control parameters to the gNB-DU.
15. The method of claim 10, wherein the CU-DU Flow Control Optimization module (CUDU- FCtrl-Optimization module) is hosted at the Near-RT-RIC, the method further comprising: subscribing at Near-RT-RIC for performance measures and counters for a CUDU-FCtrl-Optimization xApp from the DU and the CU;56communicating, via an E2 interface between the gNB-DU and the Near-RT-RIC, the performance measures and counters; and communicating, via the E2 interface between the gNB-CU and the Near-RT- RIC, the performance measures and counters, wherein TCP related parameters are communicated as part of contextual information to the Near-RT-RIC.
16. The method of claim 15, wherein the CUDU-FCtrl-Optimization xApp at the Near-RT- RIC analyzes the TCP parameters and decides optimal values of the CU-DU flow control parameters to optimize end-to-end performance of DRB k, the method further comprising: communicating the optimal values of CU-DU flow control parameters from the Near-RT-RIC to the gNB-DU via the E2 interface; and applying, by the gNB-DU, the optimal values of CU-DU flow control parameters and and informing the CMS.
17. The method of claim 15, wherein the gNB-DU first informs the CMS and waits for confirmation from CMS before applying the optimal values of CU-DU flow control parameters.
18. The method of claim 1, wherein one part of the CUDU-FCtrl-Optimization module (CUDU-FCtrl-Optimization (II)) module runs on the Near-RT-RIC and subscribes to parameters from the gNB-CU and the gNB-DU, and the CUDU-FCtrl-Optimization (II) module communicates these parameters to another part of the CU-DU Flow Control Optimization module (CUDU-FCtrl-Optimization (1)) running on the non-RT-RIC.
19. The method as in claim 10, further comprising: training the Al / ML for the CU-CP or training Al / ML at the CU-UP so that the CDU- FCtrl-Optimization module at the CU-CP decides the optimal values of the CU-DU flow control parameters.
20. The method of claim 19, wherein the training the AI / ML for the CU-CP so that the CUDU-57FCtrl-Optimization module at the CU-CP decides the optimal values of the CU-DU flow control parameters further comprises: communicating, by the gNB-DU, the performance measures and counters collected at the gNB-DU to the gNB-CU-CP via an Fl Application Protocol running over an Fl-C interface. communicating, by the gNB-CU-UP, the performance measures and counters collected at gNB-CU-UP to the gNB-CU-CP via an El Application Protocol running over an El interface; collecting, by the gNB-CU-CP, locally generated performance measures and counters; training the Al / ML model for the gNB-CU-CP; using, by the gNB-CU-CP, the Al / ML model to choose the optimal values of CU-DU flow control parameters; communicating, by the gNB-CU-CP, the updated CU-DU flow control parameters to gNB-DU via the Fl Application Protocol running over the Fl-C interface; and applying, by the gNB-DU, the updated CU-DU flow control parameters and informing the CMS (centralized management system).
21. The method as in claim 19, wherein the training AI / ML at the CU-UP so that the CUDU- FCtrl-Optimization module at the CU-CP decides optimal values of the CU-DU flow control parameters further comprises: communicating, by the gNB-DU, the performance measures and counters collected at the gNB-DU to the gNB-CU-CP via an Fl Application Protocol running over an Fl-C interface; communicating, by the gNB-CU-CP the performance measures and counters received from gNB-DU to the CU-UP;58communicating performance measures and counters generated locally at gNB-CU- CP to the gNB-CU-UP via an El Application Protocol; collecting at the gNB-CU-UP locally generated performance measures and counters generated locally; training the Al / ML model at the gNB-CU-UP; downloading the AI / ML model to the gNB-CU-CP; using, by the gNB-CU-CP, the AI / ML model to choose the optimal values of CU-DU flow control parameters; communicating, by the gNB-CU-CP, the updated CU-DU flow control parameters to gNB-DU via the Fl Application Protocol running over the Fl-C interface; and applying, by the gNB-DU, the updated CU-DU flow control parameters and informing the CMS.59
Citation Information
Patent Citations
Reinforcement learning for multi-access traffic management
US20220014963A1
Method and apparatus for programmable and customized intelligence for traffic steering in 5g networks using open ran architectures
US20230319662A1
A method for network configuration in dense networks
WO2023067610A1