Reinforcement learning model training method, storage medium, electronic apparatus, and product
By providing a training method for reinforcement learning models, the problem of separating reinforcement learning model training and inference functions is solved, enabling RL agents to learn autonomously in dynamic environments and optimize network performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2025-12-18
- Publication Date
- 2026-07-30
AI Technical Summary
In existing technologies, the design of separating the training and inference functions of reinforcement learning models cannot support the functional definition of RL Agents and cannot dynamically select entities to participate in reinforcement learning training.
A reinforcement learning model training method is provided, in which a first network element responds to a query request from a second network element, sends a model training method, receives and determines the entities participating in the reinforcement learning model training and starts the training process, including configuration management and performance monitoring, and supports the functions of an RL Agent.
It enables autonomous exploration and learning of optimal decisions in dynamic environments, rapidly optimizes network performance metrics, and supports the functional definition of RL Agents and network resource management.
Smart Images

Figure CN2025143650_30072026_PF_FP_ABST
Abstract
Description
A reinforcement learning model training method, storage medium, electronic device, and product.
[0001] Cross-references to related applications
[0002] This disclosure is based on and claims priority to Chinese patent application CN202510117500.1, filed on January 23, 2025, entitled “A reinforcement learning model training method, storage medium, electronic device and product”, and incorporates the entire contents of that patent application by reference. Technical Field
[0003] This disclosure relates to the field of communications, and more specifically, to a reinforcement learning model training method, storage medium, electronic device, and product. Background Technology
[0004] Reinforcement learning enables intelligent agents (RL agents) to efficiently learn optimal policies through repeated trial and error and environmental feedback. It can autonomously explore and learn optimal decisions in unknown and dynamic environments, and compared to other machine learning methods, it excels at solving complex decision-making problems in existing networks. Given the complex and ever-changing 5G network environment, reinforcement learning can quickly optimize network performance metrics and intelligently adjust network resource allocation and management strategies based on real-time feedback.
[0005] Currently, 3GPP models use separate logical functions for training and inference: model training is handled by the "MLTrainingFunction," while model inference is handled by the "AIMLInferenceFunction." However, RL Agents possess both training and inference capabilities and need to be aware of the environmental impact of their actions. The current separation of training and inference does not support the functional definition of RL Agents. Summary of the Invention
[0006] This disclosure provides a reinforcement learning model training method, storage medium, electronic device, and product to at least solve the problem in related technologies where entities cannot be dynamically selected to participate in reinforcement learning training.
[0007] According to one embodiment of this disclosure, a reinforcement learning model training method is provided, comprising: a first network element responding to a query request sent by a second network element and sending a model training method supported by the first network element to the second network element; the first network element receiving a reinforcement learning model training request sent by the second network element; and the first network element determining the entities participating in the reinforcement learning model training process and initiating reinforcement learning model training based on the reinforcement learning model training request.
[0008] According to yet another embodiment of this disclosure, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in the above method embodiments when executed.
[0009] According to yet another embodiment of this disclosure, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in the above method embodiments.
[0010] According to yet another embodiment of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments. Attached Figure Description
[0011] Figure 1 is a hardware structure block diagram of a computer terminal for a reinforcement learning model training method according to an embodiment of the present disclosure;
[0012] Figure 2 is a flowchart of a reinforcement learning model training method according to an embodiment of the present disclosure;
[0013] Figure 3 is a structural block diagram of a reinforcement learning model training device according to an embodiment of the present disclosure;
[0014] Figure 4 is a schematic diagram of the network management architecture;
[0015] Figure 5 is a schematic diagram of reinforcement learning according to an embodiment of the present disclosure;
[0016] Figure 6 is a schematic diagram of a reinforcement learning scenario according to an embodiment of the present disclosure (I);
[0017] Figure 7 is a schematic diagram of a reinforcement learning scenario (II) according to an embodiment of the present disclosure;
[0018] Figure 8 is a schematic diagram of a reinforcement learning scenario (III) according to an embodiment of the present disclosure;
[0019] Figure 9 is a flowchart of reinforcement learning according to an embodiment of the present disclosure (I);
[0020] Figure 10 is a schematic diagram of a reinforcement learning scenario (IV) according to an embodiment of the present disclosure;
[0021] Figure 11 is a flowchart (II) of reinforcement learning according to an embodiment of the present disclosure;
[0022] Figure 12 is a schematic diagram of a reinforcement learning scenario according to an embodiment of the present disclosure (V);
[0023] Figure 13 is a schematic diagram of a reinforcement learning scenario (six) according to an embodiment of the present disclosure;
[0024] Figure 14 is a flowchart of reinforcement learning according to an embodiment of the present disclosure (III);
[0025] Figure 15 is a schematic diagram of a reinforcement learning scenario according to an embodiment of the present disclosure (VII);
[0026] Figure 16 is a flowchart of reinforcement learning according to an embodiment of the present disclosure (IV);
[0027] Figure 17 is a flowchart (V) of reinforcement learning according to an embodiment of the present disclosure. Detailed Implementation
[0028] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings and examples.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0030] The methods and embodiments provided in this disclosure can be executed in a mobile terminal, a computer terminal, or a similar computing device. Taking a computer terminal as an example, FIG1 is a hardware structure block diagram of a computer terminal in which the methods and embodiments of this disclosure are run. As shown in FIG1, the computer terminal may include one or more (only one is shown in FIG1) processors 102 (processors 102 may include, but are not limited to, microprocessors MCUs or programmable logic devices FPGAs, etc.) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that the structure shown in FIG1 is only illustrative and does not limit the structure of the computer terminal. For example, the computer terminal may also include more or fewer components than shown in FIG1, or have a different configuration than shown in FIG1.
[0031] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the reinforcement learning model training method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0032] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0033] This embodiment provides a reinforcement learning model training method running on the aforementioned computer terminal. Figure 2 is a flowchart of the reinforcement learning model training method according to an embodiment of this disclosure. As shown in Figure 2, the process includes the following steps:
[0034] In step S202, the first network element responds to the query request sent by the second network element by sending the model training method supported by the first network element to the second network element.
[0035] Step S204: The first network element receives the reinforcement learning model training request sent by the second network element.
[0036] Step S206: The first network element determines the entities participating in the reinforcement learning model training process and starts the reinforcement learning model training based on the reinforcement learning model training request.
[0037] In an exemplary embodiment of this disclosure, the entity participating in the reinforcement learning model training process includes at least one of the following: a reinforcement learning model training execution entity, and a reinforcement learning model training affected entity.
[0038] In an exemplary embodiment of this disclosure, after the first network element determines the entities participating in the reinforcement learning model training process based on the reinforcement learning model training request and starts the reinforcement learning model training, it manages the entities participating in the reinforcement learning model training process, including at least one of the following: configuration management and performance monitoring.
[0039] In an exemplary embodiment of this disclosure, the reinforcement learning model training request includes at least one of the following: reinforcement learning model training instruction information; reinforcement learning model training environment selection information; reinforcement learning model training configuration information; and reinforcement learning model training performance requirement information.
[0040] In one embodiment, reinforcement learning model training instruction information is used to instruct reinforcement learning model training to be performed.
[0041] In one embodiment, the reinforcement learning model training configuration information includes the range of configurations allowed to be issued during the reinforcement learning model training process, including configuration objects, configuration attributes, and configuration parameters. The configuration objects can be entity identifiers or conditions that entities need to meet, such as regional ranges or time ranges; specific conditions, such as performance conditions, or PM / KPIs meeting a certain threshold condition.
[0042] In one embodiment, the reinforcement learning model training performance requirements information includes the performance requirements that the reinforcement learning model must meet during training, comprising two parts: network performance requirements and model performance requirements. Network performance requirements refer to the PM / KPIs that need to be monitored during reinforcement learning and the threshold ranges that must be guaranteed; model performance requirements refer to the performance requirements that the model must meet in each round or throughout the entire reinforcement learning process.
[0043] In an exemplary embodiment of this disclosure, the reinforcement learning model training environment selection information includes at least one of the following: environment type, which includes simulation environment, physical environment, and simulation + physical environment; execution entity information, i.e., information of reinforcement learning execution entities, used to select and determine the entities that generate reinforcement learning model training behavior, which may include, but is not limited to, the following forms: execution entity identifier (DN), conditions that the execution entity must meet (e.g., regional range, time range; specific conditions, such as performance conditions, PM / KPI meeting a certain threshold condition); and affected entity information, i.e., information of entities affected by the reinforcement learning model training process, used to select, determine, and restrict entities affected by the reinforcement learning model training behavior, which may include, but is not limited to, the following forms: affected entity identifier (DN), conditions that the affected entity must meet (e.g., regional range, time range; specific conditions, such as performance conditions, PM / KPI meeting a certain threshold condition).
[0044] In an exemplary embodiment of this disclosure, after the first network element initiates reinforcement learning model training, it generates reinforcement learning model training results, including at least one of the following: intermediate results of reinforcement learning model training and final results of reinforcement learning model training; wherein, the intermediate results of reinforcement learning model training include at least one of the following: reinforcement learning model training records and suggested reinforcement learning behaviors.
[0045] In one embodiment, the final result of training a reinforcement learning model includes the model's final address, energy consumption, performance information, etc.
[0046] In one embodiment, the proposed reinforcement learning behavior includes configuration objects, configuration attributes, configuration parameters, etc., such as configuration object: base station, configuration attribute: energy saving state, configuration parameter: enter energy saving state.
[0047] In an exemplary embodiment of this disclosure, the reinforcement learning model training record includes at least one of the following: the number of training rounds of the reinforcement learning model, i.e., the current round number of the reinforcement learning model training process; reinforcement learning model training behavior information, i.e., the configuration behavior generated by the reinforcement learning model in this round of training, with specific parameters similar to the reinforcement learning behavior suggested above; reinforcement learning model training status information, which is the information collected from the environment during the reinforcement learning model training in this round, such as PM and KPI; and reinforcement learning model training reward information, which is the reward value set for the reinforcement learning model training in this round, the environmental information and basis for setting the reward value, etc.
[0048] In an exemplary embodiment of this disclosure, after the first network element determines the entities participating in the reinforcement learning model training process and starts reinforcement learning model training based on the reinforcement learning model training request, the method further includes: the first network element selecting entities affected by the reinforcement learning model training behavior and performing configuration on the entities, wherein the reinforcement learning model training behavior is generated during the reinforcement learning model training process.
[0049] In one embodiment, the first network element selects an entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity based on at least one of the following: reinforcement learning model training behavior information, reinforcement learning model training configuration information, and reinforcement learning model training performance requirement information.
[0050] In an exemplary embodiment of this disclosure, after the first network element receives the reinforcement learning model training request sent by the second network element, the method further includes: the first network element creating a machine learning training request instance, a machine learning training process instance, and a machine learning training report instance based on the reinforcement learning model training request; wherein, the machine learning training request instance includes at least one of the following: reinforcement learning model training instruction information, reinforcement learning model training environment selection information, reinforcement learning model training configuration information, and reinforcement learning model training performance requirement information; the machine learning training report instance and / or the machine learning training process instance includes at least one of the following: reinforcement learning model training records and suggested reinforcement learning behaviors.
[0051] In an exemplary embodiment of this disclosure, after the first network element selects an entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity, the method further includes: the first network element performing environmental state monitoring and updating the reinforcement learning model training record based on the environmental state monitoring results.
[0052] In one embodiment, the scope of environmental state monitoring is determined by the entities affected by reinforcement learning training, and the content is determined based on the reinforcement learning model training performance requirements.
[0053] In an exemplary embodiment of this disclosure, before the first network element selects an entity affected by the reinforcement learning model training behavior and performs configuration on the entity, the method further includes: the first network element receiving a reinforcement learning model training request sent by the second network element; the first network element performing inference and determining the reinforcement learning model training behavior based on the inference result.
[0054] In an exemplary embodiment of this disclosure, before the first network element selects an entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity, the method further includes: the first network element receiving a reinforcement learning model training request sent by the second network element; and the first network element selecting a third network element as the reinforcement learning model training execution entity.
[0055] In an exemplary embodiment of this disclosure, the first network element selecting a third network element as the execution entity for training the reinforcement learning model includes at least one of the following: selecting a third network element as the object entity for loading the execution model; or selecting a third network element as the object entity for performing inference.
[0056] In an exemplary embodiment of this disclosure, after the first network element selects the third network element as the entity to perform reinforcement learning model training, the method further includes at least one of the following: the third network element selects an entity affected by the reinforcement learning model training behavior according to the reinforcement learning model training request, and the first network element selects an entity affected by the reinforcement learning model training behavior according to the reinforcement learning model training request.
[0057] In one embodiment, the third network element is an artificial intelligence machine learning inference producer.
[0058] In an exemplary embodiment of this disclosure, before the first network element selects an entity affected by the reinforcement learning model training behavior and performs configuration on the entity, the method further includes: the first network element receiving a reinforcement learning model training request sent by the second network element; the first network element creating a reinforcement learning model training proxy instance, wherein the reinforcement learning model training proxy instance includes at least one of the following: reinforcement learning model training environment selection information, reinforcement learning model training configuration information, reinforcement learning model training performance requirement information, reinforcement learning model training execution subject index identifier, and reinforcement learning model training record.
[0059] In one embodiment, the first network element treats the reinforcement learning model training agent as a separate managed entity. The reinforcement learning agent is used to independently perform reinforcement learning model training, environmental state monitoring, and execution of reinforcement learning model training behaviors. The reinforcement learning model training agent can be deployed, loaded, and transmitted to other network elements.
[0060] In one embodiment, the reinforcement learning model trains agent instances as RLAgent MOI.
[0061] In an exemplary embodiment of this disclosure, the first network element updates the reinforcement learning model training agent instance during the reinforcement learning model training process.
[0062] In an exemplary embodiment of this disclosure, before the first network element selects an entity affected by the reinforcement learning model training behavior and performs configuration on the entity, the method further includes: the first network element receiving a reinforcement learning model training request sent by a second network element, wherein the reinforcement learning model training request includes execution entity information; the first network element sending a reinforcement learning model training agent creation request to a fourth network element, wherein the fourth network element creates a reinforcement learning model training agent instance based on the reinforcement learning model training agent creation request.
[0063] In one embodiment, the fourth network element is a reinforcement learning producer.
[0064] In an exemplary embodiment of this disclosure, the reinforcement learning model training agent creation request includes at least one of the following: information about the affected entity; reinforcement learning model training configuration information; reinforcement learning model training performance requirement information; and reinforcement learning execution entity index identifier.
[0065] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0066] This embodiment also provides a reinforcement learning model training device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0067] Figure 3 is a structural block diagram of a reinforcement learning model training device according to an embodiment of the present disclosure. As shown in Figure 3, the device includes a sending module 10, a receiving module 20, and a determining module 30.
[0068] The sending module 10 is configured to respond to a query request sent by the second network element and send the model training method supported by the first network element to the second network element;
[0069] The receiving module 20 is configured to receive reinforcement learning model training requests sent by the second network element;
[0070] Module 30 is configured to determine the entities participating in the reinforcement learning model training process and initiate reinforcement learning model training based on the reinforcement learning model training request.
[0071] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0072] To facilitate understanding of the technical solutions provided in the embodiments of this disclosure, the following description is based on specific scenarios.
[0073] Figure 4 illustrates the network management architecture, which includes the Business Support System (BSS), cross-domain management system, domain management system (including access network management and core network management), base stations, and network elements. The cross-domain management system manages the access network management and core network management. The access network management system can manage one or more base stations. The core network management system can manage one or more network elements.
[0074] A business support system (BSS) is oriented towards communication services and provides functions and management services such as billing, settlement, accounting, customer service, sales, network monitoring, communication service lifecycle management, and service intent translation. The BSS can be an operator's operating system or a vertical OT system.
[0075] A cross-domain management system, also known as a Network Management Function (NMF), can be a network management entity such as a Network Management System (NMS), a Network Management Service Producer (MnS Producer), a Network Management Service Consumer (MnS Consumer), and a Network Function Management Service Consumer (NFMS_C). The NMF provides one or more of the following management functions or services: network lifecycle management, network deployment, network fault management, network performance management, network configuration management, network assurance, network optimization, and translation of network intents from communication service providers (Intent-CSPs). The network referred to in the above management functions or services can include one or more network elements or subnetworks, or it can be a network slice. In other words, a network management function unit can be a Network Slice Management Function (NSMF), a Management Data Analytical Function (MDAF), a Self-Organization Network Function (SON), an Intent-Driven Management Service (Intent Driven MnS), or a Close Control Loop Management Service.
[0076] The domain management system, also known as the domain management function unit, network subnet management function unit (NSMF), or network element management function unit, includes access network management and core network management. It can be a wireless automation engine (MBB automation engine, MAE), a network element management system (EMS), a network function management service provider (NFMS_P), a network slice subnet management function (NSSMF), a domain management data analysis function (Domain MDAF), a self-organization network function (SON Function), a domain intent management function unit, or a close control loop management service, MnS Producer, MnS Consumer, and other network element management entities. Domain management function units can be classified in the following ways: By network type, they can be divided into: Radio access network (RAN) domain management function units (RAN domain MnF), Core network domain management function units (CN domain MnF), and Transport network domain management function units (TN domain MnF), etc. It should be noted that a domain management function unit can also be a domain network management system, managing one or more of the access network, core network, or transport network. By administrative region, they can be divided into: domain management function units for a specific region, such as the domain management function unit for city A, the domain management function unit for city B, etc.The domain management functional unit provides one or more of the following functions or management services: subnetwork or network element lifecycle management, subnetwork or network element deployment, subnetwork or network element fault management, subnetwork or network element performance management, subnetwork or network element assurance, subnetwork or network element optimization functions, and translation of subnetwork or network element intents (Intent from Network Operator, Intent-NOP), etc. Here, a subnetwork includes one or more network elements. A subnetwork can also include subnetworks, that is, one or more subnetworks forming a larger subnetwork. A subnetwork can also be a network slice subnetwork.
[0077] A network element is an entity that provides network services. Network elements include core network elements, radio access network elements, and transport network elements. Specifically, core network elements may include, but are not limited to, Access and Mobility Management Function (AMF) entities, Session Management Function (SMF) entities, Policy Control Function (PCF) entities, Network Data Analysis Function (NWDAF) entities, Network Repository Function (NRF) entities, and gateways. Wireless access network elements may include, but are not limited to: various base stations (e.g., Generation Node B (gNB), Evolved Node B (eNB), Central Unit Control Panel (CUCP), Central Unit (CU), Distributed Unit (DU), Central Unit User Panel (CUUP), etc.). In this disclosure, network function (NF) is also referred to as network element (NE). A network element can provide one or more of the following management functions or services: network element lifecycle management, network element deployment, network element fault management, network element performance management, network element assurance, network element optimization functions, and translation of network element intents, etc.
[0078] Management services include network management services: the basic component of Service Based Management Architecture (SBMA) is the management service (MnS). MnS is a set of capabilities for managing and orchestrating networks and services. Entities that provide MnS are called MnS producers, and entities that consume MnS are called MnS consumers. Any entity with proper authorization and authentication can consume MnS provided by MnS producers (as shown in the units listed in Figure 1). MnS producers provide their services through standardized service interfaces composed of individually designated MnS components.
[0079] MnS is specified using different independent components. A specific MnS consists of at least two of these components. Currently, three different component types are defined, referred to as MnS component type A, MnS component type B, and MnS component type C.
[0080] MnS Component Type A: MnS Component Type A is a set of management operations and / or notifications that are independent of the managed entity. These operations and notifications themselves do not involve any information related to the managed network. These operations and notifications are referred to as generic or network-independent. For example, operations that create, read, update, and delete instances of managed objects, where the managed object instance to be operated on is specified only in the operation's signature, are generic.
[0081] MnS Component Type B: MnS Component Type B refers to management information represented by an information model, representing the managed entity. MnS Component Type B is also known as the Network Resource Model (NRM). Examples of MnS Component Type B include the Network Resource Model defined in TS28.622 and the Network Resource Model defined in TS28.541.
[0082] MnS Component Type C: MnS Component Type C contains performance and fault information for the managed entity. Examples of Management Service Component Type C include: alarm information defined in TS28.532 and TS28.545; and performance data defined in TS28.552, TS28.554, and TS32.425.
[0083] 3GPP defines four types of management services (MnSType): ProvMnS, FaultSupervisionMnS, StreamingDataReportingMnS, and FileDataReportingMnS. When MnSType is ProvMnS, MnsCapability includes MANAGEMENT_DATA_CONTROL, FAULT_MANAGEMENT, FILE_MANAGEMENT, NR_PROVISIONING, 5GC_PROVISIONING, NETWORK_SLICING_PROVISIONING, EDGE_COMPUTING_PROVISIONING, AI / ML_MANAGEMENT, MDA, SON_POLICY, RANSC_MANAGEMENT, INTENT_DRIVEN_MANAGEMENT, MNS_REGISTRY_AND_DISCOVERY, COMMUNICATION_SERVICE_ASSURANCE, NSOEU, MSAC_MANAGEMENT, and CCL.
[0084] Figure 5 is a schematic diagram of reinforcement learning according to an embodiment of this disclosure. As shown in Figure 5, the RL Agent (Reinforcement Learning Model Training Agent) is defined in one round of reinforcement learning. During this round, the RL Agent generates Actions and applies them to the environment, then perceives the State of the RL Environment and receives the Reward based on the RL Environment. As can be seen from Figure 5, the RL Agent actually possesses both model training (iteratively optimizing the model based on the Reward) and model inference (generating Actions through inference). Currently, in 3GPP, model training and inference are implemented by different logical functions: model training is implemented by the "MLTrainingFunction," while model inference is implemented by the "AIMLInferenceFunction." Therefore, to support reinforcement learning, it is necessary to extend the existing logical functions to support the functions of the RL Agent.
[0085] Scenario Example 1
[0086] In this scenario, MLTrainingFunction acts as the main entity performing reinforcement learning, i.e., the RL Agent. As shown in Figure 6, MLTrainingFunction is located in RAN-OAM, and in the reinforcement learning environment, it is an affected gNB; as shown in Figure 7, MLTrainingFunction is located in gNB, and in the reinforcement learning environment, it is an affected gNB; as shown in Figure 8, MLTrainingFunction is located in NWDAF, and in the reinforcement learning environment, it is an affected NF.
[0087] Scenario Example 1 is illustrated with MLTrainingFunction located in RAN-OAM in Figure 6. Figures 7 and 8 can be referred to the process shown in Figure 9, the only difference being that ML Training MnS Producer is changed to gNB and NWDAF respectively.
[0088] Figure 9 is a flowchart of reinforcement learning according to an embodiment of the present disclosure. As shown in Figure 9, it includes the following steps:
[0089] Step S901: The ML Training MnS Consumer sends a machine learning training capability query to the ML Training MnS Producer.
[0090] Specifically, query the supported model training methods to see if reinforcement learning is supported. The ML Training MnS Producer responds to the ML Training MnS Consumer's query request, and the response includes the supported model training methods.
[0091] In step S902, the ML Training MnS Consumer sends a reinforcement learning model training request to the ML Training MnS Producer.
[0092] Specifically, the reinforcement learning model training request includes the following parameters:
[0093] (1) Reinforcement learning model training instruction information, used to instruct the ML Training MnS Producer to train the reinforcement learning model;
[0094] (2) Information on the selection of training environment for reinforcement learning models, including at least one of the following:
[0095] a. Environment type, which includes simulation environment, physical environment, and simulation + physical environment;
[0096] b. Execution Entity Information, i.e., information about the reinforcement learning execution entity, is used to select and determine the entity that generates the training behavior of the reinforcement learning model. This information may include, but is not limited to, the following definitions: Execution Entity Identifier (DN), conditions that the execution entity must meet (e.g., regional range, time range; specific conditions, such as performance conditions, PM / KPI meeting a certain threshold condition). This information is used to select and determine the entity that generates Actions in the environment. In this implementation, the execution entity is the MLTrainingFunction. When it is located in the OAM, it is considered as the ML Training MnS Consumer specifying the entity to execute reinforcement learning. When it is not located in the OAM, this information can be used to determine the entity executing reinforcement learning, i.e., to determine the gNB or NWDAF where the MLTrainingFunction executing reinforcement learning resides.
[0097] c. Affected entity information, i.e., information about entities affected by the reinforcement learning model training process, is used to select, determine, and restrict entities affected by the reinforcement learning model training behavior. It may be defined in the following forms, including but not limited to: affected entity identifier (DN), conditions that the affected entity must meet (such as regional range, time range; specific conditions, such as performance conditions, PM / KPI meeting a certain threshold condition), used to select, determine, and restrict entities affected by Action in the environment; in this embodiment, it is used to determine the gNBs affected by reinforcement learning in the environment.
[0098] (3) Reinforcement learning model training configuration information, including the scope of configurations allowed to be issued during the reinforcement learning process, including configuration objects, configuration attributes, and configuration parameters. The configuration objects can be entity identifiers or conditions that entities need to meet, such as regional range or time range; specific conditions, such as performance conditions, or PM / KPI meeting a certain threshold condition.
[0099] (4) Reinforcement learning model training performance requirements information, including the performance requirements that need to be met during the reinforcement learning process, which includes two parts: network performance requirements and model performance requirements. Network performance requirements are the PM / KPIs that need to be monitored during the reinforcement learning process and the threshold ranges that need to be guaranteed; model performance requirements are the performance requirements that the model needs to meet in each round or throughout the entire reinforcement learning process.
[0100] Step S903: The ML Training MnS Producer creates an instance of a Machine Learning Training Request (MLTrainingRequest), an instance of a Machine Learning Training Process (MLtrainingProcess), and an instance of a Machine Learning Training Report (MLTrainingReport).
[0101] Specifically, the machine learning training request instance includes the same reinforcement learning model training instruction information, reinforcement learning model training environment selection information, reinforcement learning model training configuration information, and reinforcement learning model training performance requirement information as in step S902 above.
[0102] Step S904: The ML Training MnS Producer performs reinforcement learning and generates reinforcement learning model training results. The reinforcement learning model training results include one of the following: intermediate results of reinforcement learning model training and final results of reinforcement learning model training.
[0103] Specifically, the intermediate results of reinforcement learning model training include at least one of the following: reinforcement learning model training records and suggested reinforcement learning behaviors, which include configuration objects, configuration attributes, configuration parameters, etc., such as configuration object: base station, configuration attribute: energy-saving state, configuration parameter: enter energy-saving state; reinforcement learning model training records include at least one of the following: reinforcement learning model training round number, i.e., the current round number of the reinforcement learning model training process; reinforcement learning model training behavior information, i.e., the configuration behavior generated in this round of reinforcement learning model training, with specific parameters similar to the suggested reinforcement learning behaviors mentioned above; reinforcement learning model training status information, which is the information collected from the environment during this round of reinforcement learning model training, such as PM and KPI; reinforcement learning model training reward information, which is the reward value set for this round of reinforcement learning model training, and the environmental information and basis for setting the reward value; the final results of reinforcement learning model training include the model's final address, energy consumption, performance information, etc.
[0104] In steps S905-S906, the ML Training MnS Producer selects entities affected by the training behavior of the reinforcement learning model based on the RL environment selection information and performs configuration on the determined entities.
[0105] Specifically, select entities affected by the training behavior of the reinforcement learning model and perform configuration on the entities based on at least one of the following: reinforcement learning model training behavior information, reinforcement learning model training configuration information, and reinforcement learning model training performance requirement information, and perform configuration.
[0106] In steps S907-S909, the ML Training MnS Producer requests the Threshold Monitor Producer to perform environmental state monitoring. Based on the monitoring results in the state monitoring report, the ML Training MnS Producer may dynamically adjust the affected entities and the configurations issued to the affected entities, specifically by modifying the issued configurations and modifying the reinforcement learning model training requests.
[0107] The scope of environmental status monitoring is determined by the entities affected by reinforcement learning training, and the content is determined based on the performance requirements of the reinforcement learning model training.
[0108] In step S910, the ML Training MnS Producer configures a machine learning training process instance and / or a machine learning training report instance (MLTrainingReport MOI) based on the environmental status monitoring results. The machine learning training process instance and the machine learning training report instance include reinforcement learning model training records.
[0109] Step S911: The ML Training MnS Producer evaluates whether to continue reinforcement learning.
[0110] When the Producer decides to continue reinforcement learning, proceed to step S912.
[0111] Step S912, repeat steps S906 to S911.
[0112] Step S913: Based on the training results of this round, update MLtrainingProcess and MLTrainingReport MOI.
[0113] When the Producer decides to terminate reinforcement learning, proceed to step S914.
[0114] In step S914, the ML Training MnS Producer sends a machine learning training report (ML Training Report) to the ML Training MnS Consumer.
[0115] Specifically, this includes the total number of reinforcement learning rounds and the information recorded in each round of reinforcement learning.
[0116] In this embodiment, MLTrainingFunction performs reinforcement learning independently. MLTrainingFunction is located in RAN-OAM, and reinforcement learning is supported by extending MLTrainingFunction.
[0117] Scenario Example 2
[0118] Figure 10 is a schematic diagram of a reinforcement learning scenario according to an embodiment of the present disclosure. As shown in Figure 10, both MLTrainingFunction and AIMLInferenceFunction are located in RAN-OAM.
[0119] Figure 11 is a flowchart (II) of reinforcement learning according to an embodiment of the present disclosure. As shown in Figure 11, the process includes the following steps:
[0120] Step S1101: The ML Training MnS Consumer sends a reinforcement learning model training request to the ML Training MnS Producer.
[0121] Step S1102: ML Training MnS Producer creates machine learning training requests, machine learning training processes, and machine learning training report instances.
[0122] Step S1103, Loading the machine learning model;
[0123] Specifically, the ML Training MnS Producer sends a Model Loading request, containing the learning model for reinforcement learning, to the AIML Inference Producer. The AIML Inference Producer performs the Model Loading and sends a Model Loading Complete response.
[0124] Step S1104, Artificial intelligence machine learning inference;
[0125] Specifically, the ML Training MnS Producer sends an inference request to the AIML Inference Producer, requesting the AIML Inference Producer to perform inference and obtain an Action.
[0126] Steps S1105-S1114 are the same as steps S905-S914 in Scenario Example 1.
[0127] In this embodiment, MLTrainingFunction and AIMLInferenceFunction work together to perform reinforcement learning. This embodiment extends MLTrainingFunction and AIMLInferenceFunction to support reinforcement learning.
[0128] Scenario Example 3
[0129] In this embodiment, MLTrainingFunction and AIMLInferenceFunction work together to perform reinforcement learning, as shown in Figure 12. MLTrainingFunction is located in RAN-OAM, and AIMLInferenceFunction is located in gNB. As shown in Figure 13, MLTrainingFunction is located in Core-OAM, and AIMLInferenceFunction is located in NWDAF.
[0130] Figure 14 is a flowchart (III) of reinforcement learning according to an embodiment of the present disclosure. As shown in Figure 14, the process includes the following steps:
[0131] Step S1401: The ML Training MnS Consumer sends a reinforcement learning model training request to the ML Training MnS Producer.
[0132] Step S1402: ML Training MnS Producer creates machine learning training requests, machine learning training processes, and machine learning training report instances.
[0133] Step S1403: ML Training MnS Producer selects the execution subject;
[0134] Specifically, selecting AIML Inference Producer as the execution entity for training the reinforcement learning model includes at least one of the following: selecting AIML Inference Producer as the object entity for loading the execution model; or selecting AIML Inference Producer as the object entity for performing inference.
[0135] Step S1404, Loading the machine learning model;
[0136] Specifically, the ML Training MnS Producer sends a Model Loading request, containing the model to be learned for reinforcement learning, to the AIML Inference Producer. The AIML Inference Producer performs the Model Loading and sends a Model Loading Complete response.
[0137] Step S1405, Artificial Intelligence Machine Learning Inference;
[0138] Specifically, the ML Training MnS Producer sends an inference request to the AIML Inference Producer, requesting the AIML Inference Producer to perform inference and obtain an Action.
[0139] Step S1406: Select the affected entities.
[0140] Specifically, the AIML Inference Producer selects entities affected by the reinforcement learning model's training behavior based on the reinforcement learning model's training request, or the ML Training MnS Producer selects entities affected by the reinforcement learning model's training behavior based on the reinforcement learning model's training request.
[0141] Step S1407: Configure the managed entity;
[0142] Specifically, the entities affected by an Action include the executing entity and the affected entity; that is, the configuration object includes two parts: the executing entity and the affected entity.
[0143] Steps S1408-S1415 are the same as steps S907-S914 in Scenario Example 1.
[0144] Scenario Example 4
[0145] In this embodiment, as shown in Figure 15, the RL Agent Function is located in RAN-OAM, and the RL Agent performs reinforcement learning independently.
[0146] Figure 16 is a flowchart (four) of reinforcement learning according to an embodiment of the present disclosure. As shown in Figure 16, the process includes the following steps:
[0147] In step S1601, the RL Consumer sends a reinforcement learning model training request to the RL Producer.
[0148] Step S1602: Create a reinforcement learning model training agent instance (RLAgent MOI).
[0149] Specifically, the reinforcement learning model training agent instance includes at least one of the following: reinforcement learning model training environment selection information, reinforcement learning model training configuration information (e.g., RLConfigScope), reinforcement learning model training performance requirement information (e.g., RLPerfReq), reinforcement learning model training execution subject index identifier, and reinforcement learning model training record.
[0150] Specifically, RL Producer treats the reinforcement learning agent (RLAgent) as a separate managed entity. The reinforcement learning agent is used to independently train reinforcement learning models, monitor environmental states, and execute reinforcement learning model training behaviors. The reinforcement learning model training agent can be deployed, loaded, and transmitted to other network elements.
[0151] Step S1603: Perform reinforcement learning model training and generate reinforcement learning model training behavior.
[0152] Step S1604: Select entities affected by the training behavior of the reinforcement learning model.
[0153] S1605, execute reinforcement learning model training;
[0154] Specifically, the reinforcement learning model training agent instance is updated during the reinforcement learning model training process.
[0155] S1606, the RL Producer sends a reinforcement learning model training response to the RL Consumer.
[0156] In this embodiment, RLAgentFunction can be an independent Function or MLTrainingFunction, which is RLAgentFunction.
[0157] Scenario Example 5
[0158] Figure 17 is a flowchart (V) of reinforcement learning according to an embodiment of the present disclosure. As shown in Figure 17, the process includes the following steps:
[0159] Step S1701: The ML Training MnS Consumer sends a reinforcement learning model training request to the ML Training MnS Producer.
[0160] Specifically, reinforcement learning model training requests include execution entity information.
[0161] Step S1702: ML Training MnS Producer creates a reinforcement learning model training agent creation request;
[0162] Specifically, the reinforcement learning model training agent creation request includes at least one of the following: information about the affected entity; reinforcement learning model training configuration information; reinforcement learning model training performance requirements information; and reinforcement learning execution entity index identifier.
[0163] In step S1703, the ML Training MnS Producer sends a reinforcement learning model training agent creation request to the RL Producer.
[0164] Step S1704: The RL Producer creates a reinforcement learning model training agent instance based on the reinforcement learning model training agent creation request.
[0165] Step S1705: Perform reinforcement learning model training and generate reinforcement learning model training behavior.
[0166] Step S1706: Identify the entities affected by the training behavior of the reinforcement learning model.
[0167] Step S1707: Perform reinforcement learning model training.
[0168] Step S1708: Generate the training response for the reinforcement learning model.
[0169] Step S1709: Generate a reinforcement learning model training report.
[0170] Embodiments of this disclosure also provide a computer-readable storage medium storing a computer program configured to perform the steps in any of the above method embodiments when executed.
[0171] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0172] Embodiments of this disclosure also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0173] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0174] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0175] It is obvious to those skilled in the art that the modules or steps of this disclosure described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this disclosure is not limited to any particular combination of hardware and software.
[0176] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A reinforcement learning model training method, comprising: In response to the query request sent by the second network element, the first network element sends the model training method supported by the first network element to the second network element; The first network element receives the reinforcement learning model training request sent by the second network element; The first network element determines the entities participating in the reinforcement learning model training process and initiates reinforcement learning model training based on the reinforcement learning model training request.
2. The method according to claim 1, wherein, The entities participating in the reinforcement learning model training process include at least one of the following: Reinforcement learning models train executive entities and affected entities.
3. The method according to claim 1, characterized in that, After the first network element determines the entities participating in the reinforcement learning model training process based on the reinforcement learning model training request and starts the reinforcement learning model training, it manages the entities participating in the reinforcement learning model training process, including at least one of the following: configuration management and performance monitoring.
4. The method according to claim 1 or 2, wherein, The reinforcement learning model training request includes at least one of the following: Training instructions for reinforcement learning models; Information on the selection of training environment for reinforcement learning models; Reinforcement learning model training configuration information; Information on the performance requirements for training reinforcement learning models.
5. The method according to claim 4, wherein, The reinforcement learning environment selection information includes at least one of the following: Environment type; Execution entity information; Information about the affected entities.
6. The method according to claim 1, wherein, After the first network element starts training the reinforcement learning model, it generates the reinforcement learning model training result, including at least one of the following: intermediate results of reinforcement learning model training and final results of reinforcement learning model training; The intermediate results of the reinforcement learning model training include at least one of the following: reinforcement learning model training records, and suggested reinforcement learning behaviors.
7. The method according to claim 6, wherein, The reinforcement learning model training records include at least one of the following: Number of training rounds for the reinforcement learning model; Reinforcement learning model training behavior information; Reinforcement learning model training state information; Reinforcement learning model training reward information.
8. The method according to claim 1, wherein, After the first network element determines the entities participating in the reinforcement learning model training process and initiates reinforcement learning model training based on the reinforcement learning model training request, the method further includes: The first network element selects an entity affected by the reinforcement learning model training behavior and performs configuration on the entity, wherein the reinforcement learning model training behavior is generated during the reinforcement learning model training process.
9. The method according to claim 1, wherein, After receiving the reinforcement learning model training request sent by the second network element, the first network element further includes: The first network element creates a machine learning training request instance, a machine learning training process instance, and a machine learning training report instance based on the reinforcement learning model training request. The machine learning training request instance includes at least one of the following: reinforcement learning model training instruction information, reinforcement learning model training environment selection information, reinforcement learning model training configuration information, and reinforcement learning model training performance requirement information; The machine learning training report instance and / or machine learning training process instance include at least one of the following: reinforcement learning model training records and suggested reinforcement learning behaviors.
10. The method according to claim 8, wherein, After the first network element selects the entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity, it further includes: The first network element performs environmental status monitoring and updates the reinforcement learning model training records based on the environmental status monitoring results.
11. The method according to claim 8, wherein, Before the first network element selects an entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity, the method further includes: The first network element receives a reinforcement learning model training request sent by the second network element; The first network element performs inference and determines the training behavior of the reinforcement learning model based on the inference result.
12. The method according to claim 8, wherein, Before the first network element selects an entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity, the method further includes: The first network element receives a reinforcement learning model training request sent by the second network element; The first network element selects the third network element as the execution entity for training the reinforcement learning model.
13. The method according to claim 2 or 12, wherein, The first network element selects a third network element as the execution entity for training the reinforcement learning model, including at least one of the following: Select the third network element as the object entity for loading the execution model; The third network element is selected as the object entity for performing inference.
14. The method according to claim 13, wherein, After the first network element selects the third network element as the execution entity for training the reinforcement learning model, the method further includes at least one of the following: the third network element selects an entity affected by the training behavior of the reinforcement learning model, and the first network element selects an entity affected by the training behavior of the reinforcement learning model.
15. The method according to claim 8, wherein, Before the first network element selects an entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity, the method further includes: The first network element receives a reinforcement learning model training request sent by the second network element; The first network element creates a reinforcement learning model training proxy instance, wherein the reinforcement learning model training proxy instance includes at least one of the following: reinforcement learning model training environment selection information, reinforcement learning model training configuration information, reinforcement learning model training performance requirement information, reinforcement learning model training execution subject index identifier, and reinforcement learning model training record.
16. The method according to claim 15, wherein, The first network element will update the reinforcement learning model training agent instance during the reinforcement learning model training process.
17. The method according to claim 8, wherein, Before the first network element selects an entity affected by the training behavior of the reinforcement learning model and performs configuration on the entity, the method further includes: The first network element receives a reinforcement learning model training request sent by the second network element, wherein the reinforcement learning model training request includes execution entity information; The first network element sends a reinforcement learning model training agent creation request to the fourth network element, wherein, The fourth network element creates a reinforcement learning model training agent instance based on the reinforcement learning model training agent creation request.
18. The method according to claim 17, wherein, The reinforcement learning model training agent creation request includes at least one of the following: Information about the affected entities; Reinforcement learning model training configuration information; Information on the performance requirements for training reinforcement learning models; Reinforcement learning execution subject index identifier.
19. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 18.
20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the steps of the method of any one of claims 1 to 18.
21. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1 to 18.