Configuration of reinforcement learning environment in a telecommunication network system
The method allows consumers to specify reinforcement learning environments in telecommunication networks, addressing flexibility and customization issues, resulting in tailored and effective model training aligned with consumer needs.
Patent Information
- Application Number
- PCT/KR2025/010318
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Current mechanisms for AI/ML management in telecommunication networks lack flexibility and customization, failing to allow consumers to specify requirements such as the desired network environment and whether training should be conducted online or offline, which hampers the effectiveness of reinforcement learning-based training.
A method and system enabling consumers to provide information for configuring a reinforcement learning environment, allowing them to specify the learning type, location, time, and nodes, with the producer device creating or identifying the appropriate environment for training, either online or offline, and updating the model based on action effectiveness and rewards.
Enables tailored reinforcement learning environments that align with consumer-specific needs, enhancing the relevance and effectiveness of model training, and facilitating seamless integration with existing network infrastructure.
Smart Images

Figure KR2025010318_22012026_PF_FP_ABST
Abstract
Description
CONFIGURATION OF REINFORCEMENT LEARNING ENVIRONMENT IN A TELECOMMUNICATION NETWORK SYSTEM
[0001] The present application relates to wireless communication more particularly, to method and MnS producer device for configuring a reinforcement learning (RL) environment for model training in a telecommunication network system.
[0002] 5G mobile communication technologies define broad frequency bands such that high transmission rates and new services are possible, and can be implemented not only in "Sub 6GHz" bands such as 3.5GHz, but also in "Above 6GHz" bands referred to as mmWave including 28GHz and 39GHz. In addition, it has been considered to implement 6G mobile communication technologies (referred to as Beyond 5G systems) in terahertz bands (for example, 95GHz to 3THz bands) in order to accomplish transmission rates fifty times faster than 5G mobile communication technologies and ultra-low latencies one-tenth of 5G mobile communication technologies.
[0003] At the beginning of the development of 5G mobile communication technologies, in order to support services and to satisfy performance requirements in connection with enhanced Mobile BroadBand (eMBB), Ultra Reliable Low Latency Communications (URLLC), and massive Machine-Type Communications (mMTC), there has been ongoing standardization regarding beamforming and massive MIMO for mitigating radio-wave path loss and increasing radio-wave transmission distances in mmWave, supporting numerologies (for example, operating multiple subcarrier spacings) for efficiently utilizing mmWave resources and dynamic operation of slot formats, initial access technologies for supporting multi-beam transmission and broadbands, definition and operation of BWP (BandWidth Part), new channel coding methods such as a LDPC (Low Density Parity Check) code for large amount of data transmission and a polar code for highly reliable transmission of control information, L2 pre-processing, and network slicing for providing a dedicated network specialized to a specific service.
[0004] Currently, there are ongoing discussions regarding improvement and performance enhancement of initial 5G mobile communication technologies in view of services to be supported by 5G mobile communication technologies, and there has been physical layer standardization regarding technologies such as V2X (Vehicle-to-everything) for aiding driving determination by autonomous vehicles based on information regarding positions and states of vehicles transmitted by the vehicles and for enhancing user convenience, NR-U (New Radio Unlicensed) aimed at system operations conforming to various regulation-related requirements in unlicensed bands, NR UE Power Saving, Non-Terrestrial Network (NTN) which is UE-satellite direct communication for providing coverage in an area in which communication with terrestrial networks is unavailable, and positioning.
[0005] Moreover, there has been ongoing standardization in air interface architecture / protocol regarding technologies such as Industrial Internet of Things (IIoT) for supporting new services through interworking and convergence with other industries, IAB (Integrated Access and Backhaul) for providing a node for network service area expansion by supporting a wireless backhaul link and an access link in an integrated manner, mobility enhancement including conditional handover and DAPS (Dual Active Protocol Stack) handover, and two-step random access for simplifying random access procedures (2-step RACH for NR). There also has been ongoing standardization in system architecture / service regarding a 5G baseline architecture (for example, service based architecture or service based interface) for combining Network Functions Virtualization (NFV) and Software-Defined Networking (SDN) technologies, and Mobile Edge Computing (MEC) for receiving services based on UE positions.
[0006] As 5G mobile communication systems are commercialized, connected devices that have been exponentially increasing will be connected to communication networks, and it is accordingly expected that enhanced functions and performances of 5G mobile communication systems and integrated operations of connected devices will be necessary. To this end, new research is scheduled in connection with eXtended Reality (XR) for efficiently supporting AR (Augmented Reality), VR (Virtual Reality), MR (Mixed Reality) and the like, 5G performance improvement and complexity reduction by utilizing Artificial Intelligence (AI) and Machine Learning (ML), AI service support, metaverse service support, and drone communication.
[0007] Furthermore, such development of 5G mobile communication systems will serve as a basis for developing not only new waveforms for providing coverage in terahertz bands of 6G mobile communication technologies, multi-antenna transmission technologies such as Full Dimensional MIMO (FD-MIMO), array antennas and large-scale antennas, metamaterial-based lenses and antennas for improving coverage of terahertz band signals, high-dimensional space multiplexing technology using OAM (Orbital Angular Momentum), and RIS (Reconfigurable Intelligent Surface), but also full-duplex technology for increasing frequency efficiency of 6G mobile communication technologies and improving system networks, AI-based communication technology for implementing system optimization by utilizing satellites and AI (Artificial Intelligence) from the design stage and internalizing end-to-end AI support functions, and next-generation distributed computing technology for implementing services at levels of complexity exceeding the limit of UE operation capability by utilizing ultra-high-performance communication and computing resources.
[0008] The principal object of the disclosure herein is to provide a method and Management Services (MnS) producer device for configuring an RL environment for model training in a telecommunication network system.
[0009] Another object of the disclosure herein involves the consumer providing a set of information related to RL to the producer, enabling the producer to identify or create the environment for RL.
[0010] Yet another object of the disclosure herein is to provide an MnS producer device for configuring an RL environment for model training in a telecommunication network system.
[0011] Yet another object of the disclosure herein is to provide the ability for the consumer to provide information that can be used to select / create the environment for reinforcement learning.
[0012] Yet another object of the disclosure herein is to allow the consumer to dictate the particular environment for the RL, making the entire RL process more relevant to the target environment where the RL-based trained model is used.
[0013] In an aspect, the objects are achieved by a MnS Consumer device sending a request to initiate ML training. The consumer will provide information enabling the provider to use RL for the training. The information will include Learning type and Environment Selection Information. A producer will start the training process. The producer will send an acknowledgment to the MnS consumer device. When the Learning type states offline, the producer will create an NDT based on the information provided as part of environment selection information. The NDT MnS producer will send an acknowledgment indicating the successful creation of the NDT. When the Learning type states online, the producer will identify the network environment for RL based on the environment selection information provided. After the environment is established, the producer will run the RL process. The ML training process created will be updated. During this process, the producer will take action(s) on the environment. These actions will be in terms of provisioning modifications. Once the actions are executed, an entity will assess the effectiveness of the actions and provide rewards. Based on the actions and rewards, an optimal policy / model is learned to take subsequent suitable actions serving the training expectations. The producer will provide the information related to the trained ML model to the consumer. The information includes a link for the MnS consumer device to fetch the model, model file, etc.
[0014] In an aspect, the objects are achieved by providing a method and a system for configuring a reinforcement learning (RL) environment for model training in a telecommunication network system. The method includes receiving by an MnS producer device RL information from an RL MnS consumer device. The RL information comprises at least one of (a) a learning type indicating whether the RL is to be performed online or offline, (b) a location defining a geographical area of the RL environment, (c) a time during which the model is to be trained or perform inference, (d) a node identifying network nodes that should be part of the RL environment, and (e) a target network load for the network nodes in the RL environment.
[0015] The method comprises determining by the MnS producer device the RL environment based on the RL information, wherein the RL environment comprises at least one of (a) for online training, nodes of a live network on which actions are performed and the model is learned, (b) for offline training, a simulated network on which actions are performed, (c) the location defined in terms of geo-coordinate tracking areas, (d) a time that defines a specific duration or schedule for training or inference, (e) a node specification indicating which network nodes are included in the RL environment, and (f) the target network load for those nodes. Further, the method comprises configuring by the MnS producer device the RL environment for the model training based on the determined parameters.
[0016] In yet another aspect, the objects are achieved by providing an MnS producer device for configuring an RL environment for model training in a telecommunication network system. The MnS producer device comprises a memory which includes information about an MnS consumer device, a processor, and a Machine Learning Reinforcement Learning (ML-RL) controller. The ML-RL controller is coupled to both the memory and the processor. Upon receiving RL information from the MnS consumer device, the ML-RL controller processes the RL information, wherein the RL information comprises at least one of (a) a learning type indicating whether the RL is to be performed online or offline, (b) a location defining a geographical area of the RL environment, (c) a time during which the model is to be trained or perform inference, (d) a node identifying network nodes that should be part of the RL environment, and (e) a target network load for the network nodes in the RL environment.
[0017] Based on the received RL information, the ML-RL controller determines the RL environment, wherein the RL environment comprises at least one of (a) for online training, nodes of a live network on which actions are performed and the model is learned, (b) for offline training, a simulated network on which actions are performed, (c) the location defined in terms of geo-coordinate tracking areas, (d) a time that defines a specific duration or schedule for training or inference, (e) a node specification indicating which network nodes are included in the RL environment, and (f) the target network load for the specified nodes. Furthermore, the ML-RL controller configures the RL environment based on the determined parameters to facilitate the model training within the telecommunication network system.
[0018] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating preferred embodiments and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein, and the embodiments herein include all such modifications.
[0019] The present disclosure enables the dictation of the particular environment for Reinforcement Learning (RL) based on the RL information from the consumer, wherein the RL information includes requirements such as a learning type, location, time, and nodes.
[0020] The present disclosure is illustrated in the accompanying drawings, where like reference letters indicate corresponding parts. The embodiments will be better understood from the following description with reference to the drawings.
[0021] FIG. 1 is a sequence diagram that illustrates a method for configuring a reinforcement learning (RL) environment for model training in a telecommunication network system according to embodiments as disclosed herein.
[0022] FIG. 2 is a block diagram of an MnS producer device for configuring an RL environment for model training in a telecommunication network system according to embodiments as disclosed herein.
[0023] FIG. 3 is a flow chart that illustrates a method for an MnS producer device for configuring an RL environment for model training in a telecommunication network system according to embodiments as disclosed herein.
[0024] The embodiments and their features are detailed with reference to the non-limiting examples shown in the drawings and described below. Well-known components and techniques are omitted to avoid unnecessary detail. The described embodiments are not mutually exclusive and can be combined to form new embodiments. The term "or" is used in a non-exclusive sense unless stated otherwise. The examples provided are for illustrative purposes to aid understanding and should not be seen as limiting the scope of the embodiments.
[0025] As is existing in the field, embodiments can be described and illustrated in terms of blocks which carry out a described function or functions. These blocks, which can be referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and can optionally be driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block can be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments can be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments can be physically combined into more complex blocks without departing from the scope of the disclosure.
[0026] The accompanying drawings aid in understanding the technical features, but the embodiments are not limited to these drawings. The disclosure extends to any alterations, equivalents, and substitutes beyond those shown.
[0027] Artificial Intelligence (AI) and Machine Learning (ML) techniques have seen widespread adoption across various industries, demonstrating significant success in enhancing operational efficiencies and decision-making processes. These advanced technologies are now being increasingly applied to the telecommunication industry, including mobile networks, to optimize various functions such as network management, resource allocation, and service delivery.
[0028] Despite the maturity of AI / ML techniques, several aspects of these technologies are still evolving, and new complementary techniques continue to emerge. The learning methods employed in AI / ML include supervised learning, semi-supervised learning, unsupervised learning, and reinforcement learning (RL). Each of these methods is tailored to specific inference categories, such as prediction, and necessitates particular types of training data.
[0029] The lifecycle management of AI / ML models within the telecommunication sector is being standardized by the 3GPP SA5 working group. The defined lifecycle stages include ML model training, ML testing, ML emulation, ML entity loading, and inference phase. These stages encompass initial training, re-training, validation, testing, emulation, and deployment of ML models, ensuring that the models perform optimally before being applied to the target network or system.
[0030] Reinforcement Learning (RL) is a distinct type of machine learning where an agent learns to make decisions by interacting with the environment and receiving feedback in the form of rewards. This trial-and-error approach allows the agent to adapt to dynamic environments, making RL particularly suitable for handling the complexities of mobile networks. In RL, the agent's objective is to learn a policy that maximizes the cumulative rewards over time.
[0031] In existing systems, when RL is applied to telecommunication networks, the agent corresponds to the entity performing ML training, and the environment represents the network configuration with specific contexts such as location and time. For instance, a handover optimization ML model can be trained using RL for a particular network environment, where actions taken by the model and the resulting rewards (feedback on provisioning actions) help optimize the handover process.
[0032] However, the current mechanisms for AI / ML management in telecommunication networks have notable limitations. Specifically, they do not provide adequate means for consumers to specify requirements related to RL training, such as the desired network environment and whether the training should be conducted online (on the real network) or offline (using a simulated network). This lack of flexibility and customization hampers the effectiveness of RL-based training and the ability to tailor the training environment to specific consumer needs.
[0033] Thus, it is desired to address the above-mentioned disadvantages, issues, or other shortcomings or at least provide a useful alternative.
[0034] Embodiments disclosed herein provide a method and system for machine learning. A MnS consumer device (100) sends a request to initiate ML training. The MnS consumer device (100) provides information enabling the provider to use RL for the training. This information includes Learning type and Environment Selection Information. A MnS producer device (105) starts the training process and sends an acknowledgment to the MnS consumer device (100).
[0035] When the Learning type states offline, the MnS producer device (105) creates an NDT based on the information provided as part of environment selection information. The MnS NDT producer (110) sends an acknowledgment indicating the successful creation of the NDT. Conversely, when the Learning type states online, the MnS producer device (105) identifies the network environment for RL based on the environment selection information provided. After establishing the environment, the MnS producer device (105) runs the RL process. The ML training process created is updated, and during this process, the MnS producer device (105) takes action(s) on the environment. These actions involve provisioning modifications. Once the actions are executed, an entity assesses the effectiveness of the actions and provides rewards. Based on the actions and rewards, an optimal policy / model is learned to take subsequent suitable actions serving the training expectations. The MnS producer device (105) then provides the information related to the trained ML model to the consumer, which includes a link for the consumer to fetch the model, model file, etc.
[0036] Applying the concept of RL in telecom networks, the agent can be mapped to the entity performing the ML training, and the environment can be represented by a particular configuration of the network with specific context (e.g., location, time). A model can be trained using RL for a particular environment. A MnS consumer device (100) may request a MnS producer device (105) to train a handover optimization ML model for a particular location with a specific amount of load in the network at a particular point in time. In such RL scenarios, the actions will be the provisioning decision(s) that may be taken by the ML model for handover optimization, and the rewards may be the provided feedback of the provisioning actions performed in terms of how efficient they were in order to fulfill the required training expectations like handover optimization.
[0037] When using the RL method in the context of telecom networks, it is crucial to select or create the appropriate environment to enable an agent to execute the actions and get the rewards in order to learn an optimal policy. A MnS consumer device (100) may have specific requirements (e.g., location, time) that should be supported by the environment used for RL. It is desirable to provide mechanisms enabling the consumer to provide information that can be used to create / identify an appropriate network environment for RL during the training phase. The consumer may also wish to declare if the RL-based training should be done online or offline. Online implies that the RL actions are taken on the real network, whereas offline would imply that the MnS producer device (105) will have to create an appropriate simulation of the network on which the RL actions would be taken.
[0038] The current mechanism defined for AI / ML management does not enable the MnS consumer device (100) to provide information related to RL to the MnS producer device (105).
[0039] Referring now to the drawings, and more particularly to FIGS. 1 through 3, there are shown preferred embodiments.
[0040] FIG. 1 is a sequence diagram illustrating a method for configuring a reinforcement learning (RL) environment for model training in a telecommunication network system. The solution involves the MnS consumer device (100) providing a set of information related to RL to the MnS producer device (105), enabling the MnS producer device (105) to identify or create the environment for RL.
[0041] Reinforcement Learning Information includes:
[0042] Learning Type: This defines the type of reinforcement learning to be done, indicating either Online or Offline. In the case of Online RL, the environment contains the nodes of a live network on which actions are performed and the required model is learned. For Offline RL, the environment contains a simulated network on which actions are performed.
[0043] Environment Selection Info: This defines the information that can be used to select, determine, or create the environment for reinforcement learning. The environment may contain nodes in the live network or a network digital twin representing the live network.
[0044] Location: It defines the target geographical location of the environment. When defined, the network node(s) serving the specified location forms the RL environment. This can be defined in terms of geo-coordinate tracking areas, coverage polygon, and similar parameters.
[0045] Time: It defines the timeframe information at which the model is to be trained. This includes a time duration in a day or a time schedule, indicating when the model is expected to be trained or when the trained model is expected to perform its inference.
[0046] Node: It defines the network nodes that should be part of the environment.
[0047] Load: It defines the target load for all the network nodes in the environment. The load is defined from 1 to 10, with 1 being the minimum and 10 being the maximum load. For example, if the target load is defined as 9, the load of the nodes in the environment will be configured to 9.
[0048] At step 1, the MnS consumer device (100) sends a request to initiate ML training. The consumer provides information enabling the provider to use RL for the training, including Learning Type and Environment Selection Information. At step 2, the MnS producer device (105) starts the training process. At step 3, the MnS producer device (105) sends an acknowledgment to the MnS consumer device (100). At step 4, based on the value of Learning Type provided in step 1, if the Learning Type is Offline, the MnS producer device (105) creates a network digital twin (NDT) based on the information provided as part of Environment Selection Info in step 1. The MnS producer device (110) then sends an acknowledgment indicating the successful creation of the NDT. If the Learning Type is Online, the MnS producer device (105) identifies the network environment for RL based on the Environment Selection Information provided in step 1.
[0049] At step 5, after the environment is established, the MnS producer device (105) runs the RL process. The ML training process created in step 2 is updated. During this process, the MnS producer device (105) takes actions on the environment, such as provisioning modifications. Once the actions are executed, an entity assesses the effectiveness of the actions and provides rewards. Based on the actions and rewards, an optimal policy / model is learned to take subsequent suitable actions serving the training expectations. At step 6, the MnS producer device (105) provides information related to the trained ML model to the MnS consumer device (100). This information includes a link for the MnS consumer device (100) to fetch the model, model file, etc.
[0050] In an embodiment, the proposed solution allows the MnS consumer device (100) to dictate the particular environment for RL, making the entire RL process more relevant to the target environment where the RL-based trained model is to be used.
[0051] Examples of the MnS producer device (105) include, but are not limited to, dedicated hardware appliances, virtualized network functions (VNFs), and cloud-based instances. These examples are illustrative and not intended to limit the scope of the invention as the MnS producer device (105) (Agent) can be implemented in various other configurations depending on the requirements of the network architecture. For instance, in a cloud-based implementation, the MnS producer device (105) can leverage distributed computing resources to scale up the model training process dynamically. In a virtualized environment, it can be deployed as a VNF that runs on standard server hardware, providing flexibility and ease of integration with existing network infrastructure.
[0052] Examples of the MnS consumer device (100) include, but are not limited to, user equipment (UE) such as smartphones, tablets, and IoT devices that interact with the telecommunication network. The MnS consumer device (100) can also be network elements like base stations, routers, and switches that utilize the trained RL models to optimize their operations. The MnS consumer device (100) can also be any OSS (Operation Support System) entities that utilize the trained RL models to optimize network operations. These devices receive the model outputs and apply them to real-time decision-making processes, such as dynamic resource allocation, traffic management, and fault detection. The MnS consumer device (100) may also include software components that facilitate communication with the MnS producer device (105) and other network entities.
[0053] Examples of the NDT MnS Producer device (110) include, but are not limited to, specialized network diagnostic tools and platforms that simulate the network and generate network diagnostic and testing (NDT) data. These devices are designed to produce detailed metrics and logs that are essential for training RL models aimed at network troubleshooting and performance optimization. The NDT MnS Producer device (110) can be integrated with network monitoring systems to continuously collect data on network health, latency, throughput, and error rates. This data is then fed into the RL environment to enhance the accuracy and reliability of the trained models. The device may also support advanced features like anomaly detection and predictive maintenance.
[0054] In an embodiment, the MnS producer device (105) includes a processor (201), a memory (202), a communicator (204), and a Machine learning Reinforcement learning (ML-RL) controller (205). The processor (201) is coupled with the memory (202), the communicator (204), and the ML-RL controller (205).
[0055] The processor (201) is responsible for configuring a RL environment for model training in a telecommunication network system. The processor (201) communicates with the memory (202), the communicator (204), and the ML-RL controller (205). The processor (201) is configured to execute instructions stored in the memory (102) and to configuring a RL environment for model training in a telecommunication network system. The processor (201) includes one or a plurality of processors, maybe a general-purpose processor such as a Central Processing Unit (CPU), an Application Processor (AP), or the like, a graphics-only processing unit such as a Graphics Processing Unit (GPU), a Visual Processing Unit (VPU), and / or an Artificial Intelligence (AI) dedicated processor such as a Neural Processing Unit (NPU).
[0056] The memory (202) stores the operating system, application software, and temporary data used by the processor (201). The memory (202) stores instructions to be executed by the processor (201). The memory (202) includes non-volatile storage elements. Examples of such non-volatile storage elements includes magnetic hard disks, optical disks, floppy disks, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the memory (202) may in some examples be considered a non-transitory storage medium. The term non-transitory may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term non-transitory should not be interpreted that the memory (202) is non-movable. The memory (202) includes MnS consumer device (100)information. The MnS consumer device (100) information can be : a. a learning type indicating whether the RL is to be performed online or offline, b. a location defining a geographical area of the RL environment, c. a time during which the model is to be trained or perform inference, d. a node identifying network nodes that should be part of the RL environment, and e. a target network load for the network nodes in the RL environment, ensuring seamless connectivity and robust communication within the network.
[0057] The communicator (204) facilitates communication between the processor (201) and the memory (202), supporting various communication protocols such as Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Transport Layer Security (TLS), Internet Protocol Security (IPSec), Hypertext Transfer Protocol Secure (HTTPS). Further, the communicator (204) is configured for communicating internally between internal hardware components. The communicator (204) includes an electronic circuit specific to a standard that enables wired or wireless communication. The communicator (204) facilitates the transmission of messages.
[0058] The ML-RL controller (205) is specialized hardware designed for configuring a RL environment for model training in a telecommunication network system. In an embodiment, the structure of such an innovative integrated circuit of the ML-RL controller (205) can include a multi-core architecture that enables the configuring a RL environment for model training in a telecommunication network system. Each core is optimized for specific tasks such as RL is to be performed online or offline response messages. Store data location defining a geographical area of the RL environment, data size, and time during which the model is to be trained or perform inference, identifying network nodes that should be part of the RL environment, etc. The integrated circuit of the ML-RL controller (205) is made of a combination of analog and digital components designed to enable leveraging S&F information in a communication system. The analog components include a low-noise amplifier. The digital components include a microcontroller unit (MCU) and a digital signal processor (DSP) that work in tandem to introduce one new attribute that includes information related to the storage format to be used by the MnS producer device (105) and MnS consumer device (100) respectively. In an embodiment, the multi-core architecture may incorporate specialized memory management units (MMUs) to handle large datasets, ensuring rapid access and processing of RL information. The analog components may also include phase-locked loops (PLLs) for precise timing control, which is crucial for synchronizing data transmission and reception in the telecommunication network. The digital components may further include field-programmable gate arrays (FPGAs) to allow for customizable logic operations, enhancing the flexibility and scalability of the RL environment configuration.
[0059] The ML-RL controller (205) receives (301) by an MnS producer device (105) from an RL MnS consumer device (100) RL information. Further, the RL information comprises one or more of: a) a learning type indicating whether the RL is to be performed online or offline, b) a location defining a geographical area of the RL environment, c) a time during which the model is to be trained or perform inference, d) a node identifying network nodes that should be part of the RL environment, and e) a target network load for the network nodes in the RL environment. The RL information may also include metadata such as the historical performance data of the network nodes, which can be used to fine-tune the RL model. The learning type parameter may specify additional details such as the frequency of online updates or the batch size for offline training. In an embodiment, the location parameter includes not only geographical coordinates but also network topology information, which helps in understanding the connectivity and interaction between different nodes. The time parameter may be broken down into training intervals and inference windows, allowing for precise scheduling and resource allocation.
[0060] The ML-RL controller (205) determines (302) by the MnS producer device (105) the RL environment based on the RL information. Further, the RL environment comprises one or more of: a) for online training, the RL environment comprises nodes of a live network on which actions are performed and the model is learned, b) for offline training, the RL environment comprises a simulated network on which actions are performed, c) the location is defined in terms of geo-coordinate tracking areas, d) time that defines a specific duration or schedule for training or inference, e) node specifies which network nodes are included in the RL environment, and f) the target network load. The determination process may involve complex techniques to analyze the RL information and predict the optimal configuration for the RL environment. In an embodiment, for online training, the RL environment includes mechanisms for real-time feedback and adjustment based on network performance metrics. For offline training, the simulated network may be designed to mimic real-world conditions as closely as possible, including traffic patterns and node behaviors. The geo-coordinate tracking areas may be dynamically adjusted based on network demand and user distribution, ensuring that the RL model is trained in relevant and high-impact regions.
[0061] The ML-RL controller (205) configures (303) by the MnS producer device (105) the RL environment for the model training. The configuration process may involve setting up communication protocols, data storage formats, and processing pipelines to ensure handling of RL information. The controller may also allocate computational resources such as CPU cycles and memory bandwidth to different cores based on their specific tasks. In an embodiment, the configuration includes security measures to protect sensitive RL information and ensure compliance with data privacy regulations. The controller may use machine learning techniques to continuously monitor and adjust the configuration based on real-time network conditions and performance metrics.
[0062] The comprehensive management capabilities of the ML-RL controller (205) ensure seamless configuring a RL environment for model training in a telecommunication network system. This includes real-time monitoring, data analytics, and adaptive response mechanisms to optimize network performance. This ensures the receiving by an MnS producer device (105) from an RL MnS consumer device (100) of RL information. The ML-RL controller (205) configuring a RL environment for model training in a telecommunication network system for application enablement and optimized service delivery, facilitating data handling, reduced latency, and enhanced user experience across the network. In an embodiment, the real-time monitoring capabilities includes advanced visualization tools to track network performance and RL model progress. Data analytics may involve sophisticated techniques to identify patterns and anomalies in network behavior, providing insights for further optimization. In an embodiment, adaptive response mechanisms includes automated adjustments to network configurations based on predictive analytics, ensuring that the network remains robust and under varying conditions. The overall system may be designed to support scalability, allowing for the integration of additional network nodes and expansion of the RL environment as needed.
[0063] FIG 3 is a flow chart illustrating a method for an MnS producer device (105) to configure a RL environment for model training in a telecommunication network system. At step 301, the method includes receiving RL information by the MnS producer device (105) from an RL MnS consumer device (100). In an embodiment, the RL information can be transmitted using a secure communication protocol to ensure data integrity and confidentiality.
[0064] In an embodiment, the RL information comprises at least one of the following: a learning type indicating whether the RL is to be performed online or offline, a location defining a geographical area of the RL environment, a time during which the model is to be trained or perform inference, a node identifying network nodes that should be part of the RL environment, and a target network load for the network nodes in the RL environment. The learning type may further specify the technique or framework to be used for RL, such as deep reinforcement learning. In an embodiment, the location information can include additional parameters like signal strength and network coverage maps. The node identification may involve unique identifiers or IP addresses of the network nodes, and the target network load may be dynamically adjusted based on real-time network conditions.
[0065] At step 302, the MnS producer device (105) determines the RL environment based on the RL information. In an embodiment, the RL environment includes one or more of the following: for online training, the RL environment includes nodes of a live network on which actions are performed and the model is learned; for offline training, the RL environment includes a simulated network on which actions are performed. The location is defined in terms of geo-coordinate tracking areas; time defines a specific duration or schedule for training or inference; node specifies which network nodes are included in the RL environment; and the target network load. The determination process may involve complex techniques to optimize the RL environment for model training. The MnS producer device (105) can consider historical data and predictive analytics to refine the RL environment settings. Furthermore, the simulated network may incorporate advanced features like network topology emulation and traffic pattern simulation.
[0066] At step 303, the MnS producer device (105) configures the RL environment for the model training. In an embodiment, the target network load is defined on a scale from 1 to 10, with 1 being the minimum load and 10 being the maximum load, and the network nodes are configured to the target network load. The configuration process can include adjusting network parameters such as bandwidth allocation, latency settings, and resource distribution. The MnS producer device (105) may also deploy monitoring tools to track the performance of the RL environment during the training phase. Further, the configuration can include setting up failover mechanisms to ensure network stability in case of unexpected disruptions.
[0067] In an embodiment, the time frame information comprises a time duration within a day or a scheduled time window for training or inference. The RL consumer (100) communicates the RL information through an MLTrainingRequest Information Object Class (IOC). Further, the simulated network comprises a digital replica of the live network used for offline training. The time frame information may be synchronized with network maintenance schedules to minimize impact on live network operations. The MLTrainingRequest IOC can include detailed specifications for the training process, such as data sampling rates and model evaluation criteria. The digital replica of the live network may be continuously updated to reflect changes in the live network, ensuring accurate and effective offline training.
[0068] The description of the specific embodiments provided here will clearly reveal their general nature, allowing others to modify or adapt them for various applications without straying from the core concept. Such adaptations and modifications are intended to be included within the scope of the disclosed embodiments. The terminology used is for descriptive purposes only and not meant to limit the scope. Therefore, while preferred embodiments are described, those skilled in the art will understand that modifications can be made within the scope of the described embodiments.
Claims
1.A method performed by a management service (MnS) producer, the method comprising:receiving, from an MnS consumer, a request for training a machine learning (ML) model, the request comprising at least one requirement for training the ML model for reinforcement learning (RL);determining an environment for performing the RL based on the at least one requirement for training the ML model for the RL; andtraining the ML model based on the environment for performing the RL.2.The method of claim 1, wherein the at least one requirement includes:type information indicating whether the ML model is to be trained in a real-network environment or a simulation network environment.3.The method of claim 2, wherein the determining of the environment comprises:selecting the environment based on the at least one requirement for training the ML model for the RL, in case that the type information indicates that the ML model is to be trained in the real-network environment; andcreating the environment based on the at least one requirement for training the ML model for the RL, in case that the type information indicates that the ML model is to be trained in the simulation network environment.4.The method of claim 1, wherein the at least one requirement includes environment scope information, the environment scope information including at least one of:location information indicating a target geographical location of the environment;time information indicating a timeframe at which the ML model is to be trained; ornode information indicating at least one node to be included in the environment.5.A method performed by a management service (MnS) consumer, the method comprising:transmitting, to an MnS producer, a request for training a machine learning (ML) model, the request comprising at least one requirement for training the ML model for reinforcement learning (RL),wherein an environment for performing the RL is determined based on the at least one requirement for training the ML model for the RL, andwherein the ML model is trained based on the environment for performing the RL.6.The method of claim 5, wherein the at least one requirement includes:type information indicating whether the ML model is to be trained in a real-network environment or a simulation network environment.7.The method of claim 5, wherein the at least one requirement includes environment scope information, the environment scope information including at least one of:location information indicating a target geographical location of the environment;time information indicating a timeframe at which the ML model is to be trained; ornode information indicating at least one node to be included in the environment.8.A management service (MnS) producer, the MnS producer comprising:at least one communicator;at least one processor communicatively coupled to the at least one communicator; andmemory, communicatively coupled to the at least one processor, storing instructions executable by the at least one processor individually or in any combination to cause the MnS producer to:receive, from an MnS consumer, a request for training a machine learning (ML) model, the request comprising at least one requirement for training the ML model for reinforcement learning (RL);determine an environment for performing the RL based on the at least one requirement for training the ML model for the RL; andtrain the ML model based on the environment for performing the RL.9.The MnS producer of claim 8, wherein the at least one requirement includes:type information indicating whether the ML model is to be trained in a real-network environment or a simulation network environment.10.The MnS producer of claim 9, wherein the instructions cause the MnS producer to determine the environment by:selecting the environment based on the at least one requirement for training the ML model for the RL, in case that the type information indicates that the ML model is to be trained in the real-network environment; andcreating the environment based on the at least one requirement for training the ML model for the RL, in case that the type information indicates that the ML model is to be trained in the simulation network environment.11.The MnS producer of claim 8, wherein the at least one requirement includes environment scope information, the environment scope information including at least one of:location information indicating a target geographical location of the environment;time information indicating a timeframe at which the ML model is to be trained; ornode information indicating at least one node to be included in the environment.12.A management service (MnS) consumer, the MnS consumer comprising:at least one communicator;at least one processor communicatively coupled to the at least one communicator; andmemory, communicatively coupled to the at least one processor, storing instructions executable by the at least one processor individually or in any combination to cause the MnS consumer to:transmit, to an MnS producer, a request for training a machine learning (ML) model, the request comprising at least one requirement for training the ML model for reinforcement learning (RL),wherein an environment for performing the RL is determined based on the at least one requirement for training the ML model for the RL, andwherein the ML model is trained based on the environment for performing the RL.13.The MnS consumer of claim 12, wherein the at least one requirement includes:type information indicating whether the ML model is to be trained in a real-network environment or a simulation network environment.14.The MnS consumer of claim 12, wherein the at least one requirement includes environment scope information, the environment scope information including at least one of:location information indicating a target geographical location of the environment;time information indicating a timeframe at which the ML model is to be trained; ornode information indicating at least one node to be included in the environment.
Citation Information
Patent Citations
Communication network
EP4366258A1
Network resource model based solutions for ai-ML model training
WO2023114017A1
Reinforcement learning
WO2024027921A1