Communication method based on reinforcement learning, and proxy node, management network element, system and medium

By introducing reinforcement learning methods into 5G networks, the collaborative operation between agent nodes and management network elements solves the problem of low management efficiency caused by the complexity of the network environment, and realizes network performance optimization and intelligent resource allocation.

WO2026011821A1PCT designated stage Publication Date: 2026-01-15ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/082525
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-03-14
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

The 5G network environment is complex and ever-changing, and existing technologies cannot effectively utilize reinforcement learning to optimize network performance indicators and resource allocation, resulting in low network management efficiency.

Method used

By implementing reinforcement learning methods between agent nodes and management network elements, including receiving and sending reinforcement learning initiation requests and reports, and combining reward strategies and backoff mechanisms, network performance can be optimized.

Benefits of technology

It enables autonomous exploration of optimal decisions in unknown dynamic environments, improves the efficiency of network resource allocation and network performance optimization capabilities, and enhances the level of intelligence in network management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025082525_15012026_PF_FP_ABST
    Figure CN2025082525_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a communication method based on reinforcement learning, and a proxy node, a management network element, a system and a medium. The method comprises: receiving a reinforcement learning start request sent by a first management network element or a second management network element (110); on the basis of the reinforcement learning start request, executing reinforcement learning (120); and sending a reinforcement learning report to the first management network element or the second management network element (130).
Need to check novelty before this filing date? Find Prior Art

Description

Reinforcement learning-based communication methods, agent nodes, management network elements, systems, and media. Technical Field

[0001] This application relates to the field of wireless communication technology, and for example to a communication method, agent node, management network element, system, and medium based on reinforcement learning. Background Technology

[0002] 3GPP SA5 (Service and System Aspects Working Group 5) is researching management specifications for Artificial Intelligence (AI) and Machine Learning (ML), with a particular focus on the management and orchestration of AI or ML functions in 5G systems, such as Management Data Analytics (MDA), Network Data Analytics Function (NWDAF), and Next Generation Radio Access Network (NG-RAN). The complex and dynamic 5G network environment presents significant challenges in rapidly optimizing network performance metrics based on real-time feedback and intelligently adjusting network resource allocation and management strategies. Reinforcement Learning (RL) allows agents to efficiently learn optimal strategies through repeated trial and error and environmental feedback. RL can autonomously explore and learn optimal decisions in unknown dynamic environments, and compared to other machine learning methods, it excels at solving complex decision-making problems in existing networks. How to integrate reinforcement learning into communication systems is a pressing issue that needs to be addressed. Summary of the Invention

[0003] This application provides a communication method, proxy node, management network element, system, and medium based on reinforcement learning.

[0004] This application provides a reinforcement learning-based communication method applied to a proxy node, including:

[0005] Receive reinforcement learning initiation request sent by the first management network element or the second management network element;

[0006] Execute reinforcement learning according to the reinforcement learning initiation request;

[0007] Send a reinforcement learning report to the first management network element or the second management network element.

[0008] This application also provides a reinforcement learning-based communication method applied to management network elements, including:

[0009] Send a reinforcement learning initiation request to the agent node, the reinforcement learning initiation request being used to instruct the agent node to perform reinforcement learning;

[0010] Receive the reinforcement learning report from the agent node.

[0011] This application embodiment also provides a proxy node, including: a memory, and one or more processors;

[0012] The memory is configured to store one or more programs;

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the reinforcement learning-based communication method applied to the agent node.

[0014] This application embodiment also provides a management network element, including a memory and one or more processors;

[0015] The memory is configured to store one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the reinforcement learning-based communication method applied to management network elements as described above.

[0017] This application also provides a communication system based on reinforcement learning, including: the aforementioned proxy node and the aforementioned management network element.

[0018] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described reinforcement learning-based communication method. Attached Figure Description

[0019] Figure 1 is a schematic diagram of a service-oriented management architecture provided in one embodiment;

[0020] Figure 2 is a schematic diagram of a general operation flow in the life cycle of an ML model;

[0021] Figure 3 is a schematic diagram illustrating the principle of a reinforcement learning-based communication method according to an embodiment;

[0022] Figure 4 is a flowchart of a reinforcement learning-based communication method provided in one embodiment;

[0023] Figure 5 is a flowchart of a reinforcement learning-based communication method provided in one embodiment;

[0024] Figure 6 is a schematic diagram of triggering reinforcement learning according to an embodiment;

[0025] Figure 7 is a schematic diagram of triggering reinforcement learning according to an embodiment;

[0026] Figure 8 is a schematic diagram of triggering reinforcement learning according to an embodiment;

[0027] Figure 9 is a schematic diagram of a communication process based on reinforcement learning provided in one embodiment;

[0028] Figure 10 is a schematic diagram of another reinforcement learning-based communication process provided in one embodiment;

[0029] Figure 11 is a schematic diagram of another reinforcement learning-based communication process provided in one embodiment;

[0030] Figure 12 is a schematic diagram of another communication process based on reinforcement learning provided in one embodiment;

[0031] Figure 13 is a schematic diagram of the structure of a reinforcement learning-based communication device according to an embodiment;

[0032] Figure 14 is a schematic diagram of the structure of a reinforcement learning-based communication device according to an embodiment;

[0033] Figure 15 is a schematic diagram of the hardware structure of a proxy node provided in one embodiment;

[0034] Figure 16 is a schematic diagram of the hardware structure of a management network element according to an embodiment;

[0035] Figure 17 is a schematic diagram of the structure of a reinforcement learning-based communication system provided in one embodiment. Detailed Implementation

[0036] The present application will now be described in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. It should be noted that, unless otherwise specified, the embodiments and features described herein can be arbitrarily combined with each other. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present application, not the entire structure.

[0037] Figure 1 is a schematic diagram of a service-based management architecture provided in one embodiment. As shown in Figure 1, the service-based management architecture includes a Business Support System (BSS), a Cross Domain Management Function (CD-MnF), a Domain Management Function (Domain-MnF), and a Network Element (NE). The CD-MnF manages one or more Domain Management Functions. The Domain Management Functions can manage one or more Network Elements.

[0038] A business support system (BSS) is oriented towards communication services and provides functions and management services such as billing, settlement, accounting, customer service, sales, network monitoring, communication service lifecycle management, and service intent translation. The BSS can be an operator's operating system or a vertical OT system.

[0039] The cross-domain management function unit, also known as the network management function unit (NMF) or management network element, can be a network management entity such as a network management system (NMS), a network management service producer (MnS Producer), a network management service consumer (MnS Consumer), or a network function management service consumer (NFMS_C). The cross-domain management function unit provides one or more of the following management functions or services: network lifecycle management, network deployment, network fault management, network performance management, network configuration management, network assurance, network optimization, and translation of network intents from communication service providers (Intent-CSPs). The network referred to in the above management functions or services can include one or more network elements or subnetworks, or it can be a network slice. In other words, the network management function unit can be a network slice management function (NSMF), a management data analysis function (MDAF), a self-organization network function (SON function), or an intent-driven management service (Intent Driven MnS).

[0040] The domain management function unit, also known as the network subnet management function (NSMF) or network element management function unit, can be a wireless automation engine (MBB automation engine, MAE), an element management system (EMS), a network function management service provider (NFMS_P), a network slice subnet management function (NSSMF), a domain management data analysis function (Domain MDAF), a self-organization network function (SON function), a domain intent management function unit, an MnS producer, an MnS consumer, and other network element management entities. Domain management function units can be classified in the following ways: By network type, they can be divided into: Radio Access Network (RAN) Domain Management Function (RAN domain MnF), Core Network Domain Management Function (CN domain MnF), and Transport Network Domain Management Function (TN domain MnF), etc. It is important to note that a domain management function unit can also be a domain network management system, managing one or more of the access network, core network, or transport network. By administrative region, they can be divided into: domain management function units for a specific region, such as the domain management function unit for city A, the domain management function unit for city B, etc. Domain management function units provide one or more of the following functions or management services: subnetwork or network element lifecycle management, subnetwork or network element deployment, subnetwork or network element fault management, subnetwork or network element performance management, subnetwork or network element assurance, subnetwork or network element optimization functions, and translation of subnetwork or network element intents (Intent from Network Operator, Intent-NOP), etc. The subnetwork here includes one or more network elements.A subnetwork can also contain other subnetworks, meaning one or more subnetworks can form a larger subnetwork. Here, a subnetwork can also be a network slice subnetwork.

[0041] A network element is an entity that provides network services. Network elements may include core network elements, radio access network elements, or transport network elements. Specifically, core network elements may include, but are not limited to, Access and Mobility Management Function (AMF) entities, Session Management Function (SMF) entities, Policy Control Function (PCF) entities, Network Data Analysis Function (NWDAF) entities, Network Repository Function (NRF) entities, and gateways. Wireless access network elements may include, but are not limited to: various base stations (e.g., Generation Node B (gNB), Evolved Node B (eNB), Central Unit Control Panel (CUCP), Central Unit (CU), Distributed Unit (DU), Central Unit User Panel (CUUP), etc.). In this application, network function (NF) is also referred to as network element (NE). A network element can provide one or more of the following management functions or services: network element lifecycle management, network element deployment, network element fault management, network element performance management, network element assurance, network element optimization functions, and translation of network element intent, etc.

[0042] The cross-domain management function unit can be used for model training of AI or ML models, and for model inference of AI or ML models; the domain management function unit can be used for model training of AI or ML models, and for model inference of AI or ML models; the network element is an entity that provides network services, including core network elements and access network elements (base stations), which can provide at least one of model training of AI or ML models and model inference of AI or ML models.

[0043] AI or ML technologies are widely used in fifth-generation systems (5GS), including 5GC, NG-RAN, and Operation Administration and Maintenance (OAM). Figure 2 illustrates a general operational flow in the lifecycle of an ML model. As shown in Figure 2, the ML model lifecycle includes model training, model testing, inference simulation, model deployment, and inference. Each operational step can be supported by one or more AI or ML management functions.

[0044] ML model training includes initial training and retraining of a single ML model or a set of ML models. It also includes validating the ML model to evaluate its performance on both training and validation data. If the validation results are not as expected (e.g., unacceptable variance), the ML model needs to be retrained.

[0045] - ML Model Testing: Test the validated ML model to evaluate its performance when executed on the test data. If the test results meet expectations, the ML model may proceed to the next step. If the test results do not meet expectations, the ML model needs to be retrained.

[0046] - AI or ML Inference Simulation: Run the ML model for inference in a simulation environment. The purpose is to evaluate the inference performance of the ML model in the simulation environment before applying it to the target network or system. Note: AI or ML inference simulation is optional and can be skipped in the AI ​​or ML operation workflow.

[0047] ML Model Deployment: ML model deployment involves the ML model loading process (also known as a series of atomic operations) to make the trained ML model available for the target AI or ML inference function. Note: In some cases, deploying the ML model may not be necessary, for example, when the training and inference functions are located in the same place.

[0048] AI or ML Inference: Use the AI ​​or ML inference function to perform inference using a trained ML model. AI or ML inference can also trigger model retraining or updates based on performance monitoring and evaluation.

[0049] Figure 3 is a schematic diagram illustrating the principle of a reinforcement learning-based communication method according to an embodiment. As shown in Figure 3, the reinforcement learning-based communication method of this embodiment defines OAM, which can collect and store a large amount of monitoring indicator data, supports network environment monitoring, and is applicable to and supports reinforcement learning. 5G network optimization requires balancing multiple performance indicators such as throughput, latency, and reliability. Based on the monitored values ​​of the indicators, the scenario, and user needs, model state and reward strategies can be designed. Simultaneously, the need for fallback can be assessed based on the state of the network environment to avoid impacting network performance.

[0050] Figure 4 is a flowchart of a reinforcement learning-based communication method provided in one embodiment. This method can be applied to proxy nodes. As shown in Figure 4, the method provided in this embodiment includes the following steps:

[0051] Step 110: Receive a reinforcement learning start request sent by the first management network element or the second management network element.

[0052] Step 120: Execute reinforcement learning according to the reinforcement learning initiation request.

[0053] Step 130: Send a reinforcement learning report to the first management network element or the second management network element.

[0054] In this embodiment, the first management network element can be a consumer of machine learning reinforcement learning (Cross-domain OAM), the second management network element can be a producer of machine learning reinforcement learning in RAN OAM (MLT Function of RAN OAM), and the agent node can be an AIML inference function in RAN OAM (AIMLInferenceFunction of RAN OAM), such as MDAF.

[0055] The first or second management network element can send a reinforcement learning initiation request to the agent node. This request can include reinforcement learning-related instructions, such as indicating the model's reinforcement learning status, process, rewards, and rollback mechanisms. This allows the agent node to perform reinforcement learning on the machine learning model and send a reinforcement learning report to the first or second management network element based on the information and results of each round of reinforcement learning. During this process, the first and second management network elements can cooperate to train the machine learning system, with the agent node executing the reinforcement learning, thus integrating reinforcement learning into the communication system.

[0056] In one embodiment, the reinforcement learning initiation request includes at least one of the following: reinforcement learning instruction information, reinforcement learning policy information (RL Policy), reinforcement learning control information (RL Control), reinforcement learning fallback instruction information (RL Fallback), reinforcement learning reward instruction information (RL Reward), reinforcement learning stop instruction information (RL Stop), and the address of a first management network element.

[0057] In one embodiment, reinforcement learning indication information can be used to indicate whether reinforcement learning should be performed, and can also further refine the reinforcement learning request. The reinforcement learning indication information includes at least one of the following:

[0058] To proceed, i.e., to execute RL (explore);

[0059] Fallback means reverting to a previous state.

[0060] Termination, i.e., stopping RL (terminate).

[0061] In one embodiment, the reinforcement learning rollback instruction information includes at least one of the following:

[0062] Revert this instruction to indicate that this round of reinforcement learning needs to be reverted;

[0063] Fallback target round indicator, used to indicate the identifier of the RL round to which to return (FallbackRLID).

[0064] In one embodiment, the reinforcement learning reward indication information includes at least one of the following:

[0065] Current reinforcement learning round;

[0066] The reward value for this round of reinforcement learning, i.e. the reward information for this round of RL, can be 1 / 0 or a specific reward value;

[0067] The reason for this reinforcement learning reward (Context) is the reason for assigning the reward value in this RL.

[0068] In one embodiment, the reinforcement learning report includes at least one of the following:

[0069] Based on the monitoring metrics data reported by the reinforcement learning strategy information, such as the monitored performance measurement (PM) and / or key performance indicators (KPIs);

[0070] Reinforcement learning state (RL State) of the model;

[0071] Reinforcement learning records can be stored in a list to record information for each round of reinforcement learning.

[0072] Fallback Recommendation;

[0073] Reinforcement learning goal fulfillment status (RLGoalFuilfilment), such as indicating whether the goal has been achieved, the percentage of progress achieved, the current value of goal PM, etc.

[0074] In one embodiment, the model reinforcement learning state includes at least one of the following: an executable reinforcement learning state, a reinforcement learning in progress state, a reinforcement learning completed state, and a reinforcement learning rollback state.

[0075] In one embodiment, the rollback recommendation information includes at least one of the following:

[0076] The rollback target round indicator is used to indicate the reinforcement learning rounds that need to be rolled back.

[0077] Fallback Reason indicates the reason for performing a fallback in reinforcement learning, such as potential risks or deterioration of the monitored PM metric.

[0078] In one embodiment, the reinforcement learning record includes at least one of the following:

[0079] Current reinforcement learning round number;

[0080] The reinforcement learning actions in this round are the operations performed in this RL session, such as the configuration parameters of the base station.

[0081] Reward value for this round of reinforcement learning;

[0082] The reason for this reinforcement learning reward.

[0083] In one embodiment, when there is only one proxy node, the first management network element calculates the reward value according to the reward strategy in the reinforcement learning initiation request, and determines whether to continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning. Based on this,:

[0084] While continuing reinforcement learning, the agent node updates the reinforcement learning reward indication information and reinforcement learning records, and sends the corresponding reinforcement learning report.

[0085] When performing reinforcement learning rollback, the agent node updates the reinforcement learning rollback instruction information and the model reinforcement learning status, performs the rollback operation for the corresponding number of rounds, and sends the corresponding reinforcement learning report.

[0086] In the event that reinforcement learning is terminated, the agent node updates the reinforcement learning termination indication information and the model reinforcement learning status, and sends the corresponding reinforcement learning report.

[0087] In one embodiment, the reinforcement learning initiation request further includes at least one of the following information for determining the proxy node:

[0088] Agent selection information is used to indicate the information required to select an agent node;

[0089] Proxy grouping information is used to indicate information on grouping proxy nodes;

[0090] Agent selection information includes at least one of the following:

[0091] Agent ID, used to indicate the agent node performing reinforcement learning, such as the gNB identifier;

[0092] Agent Condition is used to indicate the basis for selecting agent nodes, such as gNB geographical location information, gNB energy consumption threshold information, etc.

[0093] The proxy group information includes at least one of the following:

[0094] Agent Group information is used to indicate each agent group and the agent node corresponding to each agent group, including, for example, the Agent Group ID and the Agent ID under each Agent Group ID;

[0095] Agent Group Condition is used to indicate the basis for grouping agent nodes, such as geographical location information, where gNBs located in the same region are grouped together; or threshold information, where gNBs that meet the same PM or KPI threshold condition are grouped together.

[0096] In one embodiment, when the number of proxy nodes is at least two, the method further includes:

[0097] The master agent node calculates the reward value for each agent group and determines whether each agent group should continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning.

[0098] While continuing reinforcement learning, the master agent node requests to configure reinforcement learning reward indication information, and the agent nodes in each agent group update the corresponding reinforcement learning reward indication information and the corresponding reinforcement learning records, and send the corresponding reinforcement learning report.

[0099] When performing reinforcement learning rollback, the master agent node requests to configure reinforcement learning rollback instruction information or reinforcement learning termination instruction information. Each agent group updates the corresponding reinforcement learning rollback instruction information or the corresponding reinforcement learning termination instruction information according to the rollback strategy. Each agent group updates the corresponding model reinforcement learning status and sends the corresponding reinforcement learning report.

[0100] In the event of termination of reinforcement learning, the master agent node requests the configuration of reinforcement learning termination instruction information, and each agent group updates the corresponding reinforcement learning termination instruction information and model reinforcement learning status, and sends the corresponding reinforcement learning report.

[0101] In one embodiment, the method further includes:

[0102] Determine the master agent node;

[0103] The master agent node is the second management network element, or it is an agent node selected from at least two agent nodes based on agent selection information.

[0104] In one embodiment, when the number of proxy nodes is at least two, the method further includes:

[0105] Each agent group calculates its own reward value and determines whether the agent group should continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning.

[0106] While continuing reinforcement learning, the agent nodes in each agent group update the corresponding reinforcement learning reward indication information and the corresponding reinforcement learning records, and send the corresponding reinforcement learning report.

[0107] When performing reinforcement learning rollback, the agent nodes in each agent group update the corresponding reinforcement learning rollback instruction information or the corresponding reinforcement learning termination instruction information according to the rollback strategy. The agent nodes in each agent group update the corresponding model reinforcement learning status and send the corresponding reinforcement learning report.

[0108] In the event of termination of reinforcement learning, the agent nodes in each agent group update the corresponding reinforcement learning termination indication information and the corresponding model reinforcement learning status, and send the corresponding reinforcement learning report.

[0109] In one embodiment, when there are at least two agent nodes, the second management network element configures corresponding reward strategies for different agent groups and sends reinforcement learning start requests to each agent group respectively.

[0110] In one embodiment, performing reinforcement learning based on a reinforcement learning initiation request includes:

[0111] Each agent node in the agent group executes reinforcement learning according to the corresponding reinforcement learning initiation request.

[0112] In one embodiment, before receiving a reinforcement learning initiation request sent by a first management network element or a second management network element, the method further includes:

[0113] The first management network element sends a Machine Learning Training (MLT) request to the second management network element;

[0114] The second management network element executes machine learning training according to the machine learning training request to obtain a machine learning model (ML Model).

[0115] In one embodiment, the machine learning training request includes at least one of the following information: reinforcement learning instruction information, reinforcement learning policy information, reinforcement learning control information, and the address of a first management network element.

[0116] In one embodiment, the reinforcement learning policy information includes at least one of the following:

[0117] The reward policy is used to indicate the information required to calculate the reward. This includes network performance indicators such as PM and / or KPIs to be monitored, such as gNB energy consumption indicators defined in 3GPP TS28.552. After each round of Action execution, the Agent can report the specific values ​​of the monitored network performance indicators. It can also include thresholds for network performance indicators, such as PM and / or KPI thresholds. After each round of Action execution, the Agent can report PM and / or KPI indicators that exceed the corresponding thresholds and their corresponding specific values.

[0118] A fallback policy is used to indicate the conditions for fallback, including performance-based conditions and reward-based conditions. For example, fallback may occur if a certain threshold is exceeded, or if a reward condition is not met, such as if the positive reward does not reach a certain value or the negative reward exceeds a certain value.

[0119] In one embodiment, the reinforcement learning control information includes at least one of the following: a reinforcement learning goal (RL goal), indicating the final performance target of this RL session. This could be the value achieved or gain obtained by one or more PMs or KPIs after inference using the applied model following RL. RL can be stopped when the goal is achieved; a reinforcement learning interval (RL interval), the time interval between each RL round; and a reinforcement learning round number, the total number of RL rounds executed.

[0120] In one embodiment, it further includes:

[0121] The second management network element creates a machine learning model instance (ML model instance) based on the machine learning training request. The machine learning model instance includes at least one of the following information: model reinforcement learning state, reinforcement learning round number, and reinforcement learning record.

[0122] In one embodiment, the method further includes:

[0123] The second management network element sends a machine learning training report (MLT Report) to the first management network element;

[0124] A machine learning training report includes at least one of the following: the model's reinforcement learning state and the number of reinforcement learning rounds.

[0125] In one embodiment, when the number of proxy nodes is at least two, the machine learning training request further includes at least one of the following information for determining the proxy nodes:

[0126] Agent selection information is used to indicate the information required to select an agent node;

[0127] Proxy grouping information is used to indicate information on grouping proxy nodes;

[0128] Agent selection information includes at least one of the following:

[0129] Agent identifier, used to indicate the agent node performing reinforcement learning;

[0130] The proxy conditions are used to indicate the basis for selecting a proxy node.

[0131] The proxy group information includes at least one of the following:

[0132] Agent group information is used to indicate each agent group and the agent node corresponding to each agent group;

[0133] Agent group conditions are used to indicate the basis for grouping agent nodes.

[0134] Figure 5 is a flowchart of a reinforcement learning-based communication method provided in one embodiment. This method can be applied to management network elements, which can be used for reinforcement learning-based communication. As shown in Figure 5, the method provided in this embodiment includes the following steps:

[0135] Step 210: Send a reinforcement learning initiation request to the agent node, the reinforcement learning initiation request being used to instruct the agent node to perform reinforcement learning.

[0136] Step 220: Receive the reinforcement learning report from the agent node.

[0137] In this embodiment, the management network element can send a reinforcement learning start request to the agent node after completing machine learning training. The reinforcement learning start request can include reinforcement learning-related instruction information, which can be used to indicate the reinforcement learning status of the model, the reinforcement learning process, the reinforcement learning reward, and the reinforcement learning rollback, so that the agent node can perform reinforcement learning on the machine learning model and send a reinforcement learning report to the management network element based on the information of each round of reinforcement learning and the results of the reinforcement information.

[0138] In one embodiment, the management network element is a first management network element or a second management network element; when the number of agent nodes is one, the method further includes:

[0139] The reward value is calculated based on the reward policy in the reinforcement learning initiation request, and it is determined whether to continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning.

[0140] In one embodiment, the management network element is a first management network element or a second management network element; when the number of agent nodes is at least two, a reinforcement learning initiation request is sent to the agent nodes, including:

[0141] Configure corresponding reward strategies for different agent groups, and send reinforcement learning initiation requests to each agent group according to the corresponding reward strategies.

[0142] In one embodiment, the management network element includes a first management network element and a second management network element; before sending a reinforcement learning initiation request to the agent node, the method further includes:

[0143] The first management network element sends a machine learning training request to the second management network element;

[0144] The second management network element executes machine learning training based on the machine learning training request to obtain a machine learning model.

[0145] In one embodiment, the method further includes:

[0146] The second management network element creates machine learning model instances based on machine learning training. The machine learning model instance includes at least one of the following information: model reinforcement learning state, reinforcement learning round number, reinforcement learning record, current reinforcement learning round number, reinforcement learning operation in this round, reinforcement learning reward value in this round, and reason for this reinforcement learning reward.

[0147] In one embodiment, the reinforcement learning report of the agent node is received by a second management network element; the method further includes:

[0148] The second management unit sends a machine learning training report to the first management unit.

[0149] Figure 6 is a schematic diagram of triggering reinforcement learning according to an embodiment. As shown in Figure 6, the first management network element (consumer) and the second management network element (producer) cooperate to complete the machine learning training process. On this basis, the second management network element can send a reinforcement learning start request to the agent node to instruct the agent node to execute reinforcement learning. After receiving the reinforcement learning report from the agent node, the second management network element can send a machine learning training report to the first management network element.

[0150] Figure 7 is a schematic diagram of triggering reinforcement learning according to an embodiment. As shown in Figure 7, the first management network element (consumer) and the second management network element (producer) cooperate to complete the machine learning training process. On this basis, the first management network element can send a reinforcement learning start request to the agent node to instruct the agent node to execute reinforcement learning. The agent node sends a reinforcement learning report to the first management network element.

[0151] Figure 8 is a schematic diagram illustrating a method for triggering reinforcement learning according to an embodiment. When there are multiple agent nodes, one of the agent nodes or the second management network element (producer) can act as the master agent node. Figure 8 uses the second management network element as the master agent node as an example. The first management network element (consumer) and the second management network element cooperate to complete the machine learning training process. Based on this, the first management network element can send a reinforcement learning initiation request to the master agent node. Then, the master agent node can send reinforcement learning initiation requests to the other agent nodes respectively, instructing each agent node to perform reinforcement learning. Each agent node can send a corresponding reinforcement learning report to the master agent node, and then the master agent node sends the reinforcement learning report to the first management network element.

[0152] In some embodiments, multiple agent nodes may be divided into different agent groups. The second management network element may send reinforcement learning initiation requests to each agent group to instruct each agent group to perform reinforcement learning. Each agent node may send a reinforcement learning report to the second management network element. The following examples illustrate the reinforcement learning-based communication method.

[0153] One scenario is that the first management network element (Consumer) calculates the reward value based on the reward policy and assesses whether to request a rollback.

[0154] Example 1

[0155] Figure 9 is a schematic diagram of a reinforcement learning-based communication process according to an embodiment. As shown in Figure 9, in this embodiment, inference is located in RAN-OAM, and execution is located in gNB; the AIML inference function of RAN-OAM acts as an agent node, and the reward is defined by the consumer (Cross-domain OAM); there is one agent node. As shown in Figure 9, the reinforcement learning-based communication process mainly includes:

[0156] Step 1: The ML RL consumer (Cross-domain OAM) sends a machine learning training request, requesting the creation of an MLT Request instance on the ML RL producer (RAN OAM), and requesting the training of an ML model for base station energy saving using reinforcement learning. The machine learning training request may include at least one of the following information: reinforcement learning indication, reinforcement learning policy, and reinforcement learning control. The RL indication may include explore, fallback, and / or terminate. The RL policy may include a reward policy and / or a fallback policy. The RL control may include a reinforcement learning goal, a reinforcement learning interval, and / or a reinforcement learning round number. In practical applications, multiple parameters can be combined; the specific definitions of these parameters are detailed in Step 7.

[0157] Step 2: The ML RL producer creates an MLT Request instance and an MLT Report instance.

[0158] Step 3: The MLRL producer notifies the MLRL consumer that the MLT Request instance and MLT Report instance have been created successfully.

[0159] Step 4: The ML RL producer performs Machine Learning Training (MLT) to train the ML model subsequently used for reinforcement learning. After training, an ML model instance is created, containing at least one of the following parameters:

[0160] 1) The reinforcement learning state (RL State) indicates whether the model is currently in the reinforcement learning available state, the reinforcement learning ongoing state, the reinforcement learning completed state, or the reinforcement learning fallback state.

[0161] 2) Current Reinforcement Learning Rounds (RL ROUND): This indicates the number of reinforcement learning rounds the model has completed. When RL State is RL Available, this value should be 0; when RL State is RL Ongoing or RL Fallback, this value is the current RL round or the RL round to which the model needs to fall back; when RL State is RL Completed, this value is the number of RL rounds the model has completed.

[0162] 3) Reinforcement Learning Record (RL Record): This is a list that records information from each round of reinforcement learning and can include at least one of the following parameters:

[0163] RL ROUND, at this point RL ROUND should be 1;

[0164] This round of reinforcement learning actions refers to the operations performed in this RL session, such as configuration information.

[0165] The reward value for this round of reinforcement learning indicates the reward information for this round of RL, which can be 1 / 0 or a specific reward value;

[0166] The context of this reinforcement learning reward indicates the reason for this RL reward.

[0167] Step 5: The ML RL producer sends an MLT Report to the ML RL consumer, containing at least one of the following information: RL State (RL Available at this time); RL ROUND.

[0168] Step 6 includes the following two cases:

[0169] Step 6a: The ML RL producer (RAN OAM, MLTFunction) sends a reinforcement learning initiation request to the Agent (AIMLInferenceFunction), which includes the RL Policy and RL Control from Step 1, as well as the address of the ML RL consumer.

[0170] Step 6b: The ML RL consumer (Cross-domain OAM) may also directly send a reinforcement learning initiation request to the Agent (AIMLInferenceFunction). In this case, the Agent can be an ML RL producer. The reinforcement learning initiation request includes at least one of the following information: reinforcement learning instruction information, reinforcement learning policy information, and reinforcement learning control information.

[0171] Step 7: The Agent (RAN-OAM) performs reinforcement learning and monitors the network environment. Based on the model deployment scenario, in this embodiment, inference is performed in RAN-OAM, and execution is performed in gNBs. The reinforcement learning process specifically includes the following steps (the order is not limited):

[0172] a. Start the reinforcement learning process, create an RLRequest instance, and create an RL Report instance.

[0173] The RLRequest instance contains at least one of the following information:

[0174] RL indication, RL Policy, RL Control, RL Fallback, RL Reward.

[0175] The RL Policy includes at least one of the following: Reward Policy, which instructs the Consumer on the information used to calculate the reward, including network performance indicators such as PM and KPI that need to be monitored, such as gNB energy consumption related indicators defined in 3GPP TS28.552. After each round of action, the Agent needs to report the specific values ​​of the monitored PM and KPI.

[0176] PM / KPI threshold: After each round of action execution, the Agent reports the PM metrics that exceed the threshold and their specific values.

[0177] The RL Control includes at least one of the following parameters:

[0178] The RL goal indicates the final performance objective of this RL session, such as the value achieved by one or more PMs or KPIs after the application of the model for inference following RL. RL can be stopped when this objective is achieved.

[0179] RL interval, the time interval between each round of RL;

[0180] RL Round number.

[0181] The RL indication is used to indicate the state of RL learning, and the allowed values ​​are explore (start RL), fallback (revert to a previous state), and termination (termination of RL). In this case, the RL State should be Explore.

[0182] The RL Fallback parameter is used to indicate the RL Fallback request process. In addition to the indication information, it is used to configure the RL Fallback request initiated by the Consumer and may include at least one of the following parameters: the current fallback indication; and FallbackRLID, the ID of the RL round to which the user needs to return.

[0183] The RL Reward is used to indicate the RL Reward request process. In addition to the indication information, it can be a list recording the information for each round of RL, which may include at least one of the following parameters:

[0184] RL ROUND indicates the RL round number;

[0185] Reward indicates the reward information for this round of RL, which can be 1 / 0 or a specific reward value;

[0186] Context indicates the reason for assigning the reward value in this RL.

[0187] RL Stop is used to indicate the termination of the RL process, indicating a request to stop the RL process.

[0188] b. The energy-saving ML model is deployed in RAN-OAM, performs model inference, generates energy-saving strategies, and configures the corresponding gNB;

[0189] c. Execution and monitoring of inference results: Execute base station energy-saving inference strategies; create threshold monitoring instances based on the Reward Policy to monitor network performance indicators, assess whether fallback is necessary, and provide fallback recommendations.

[0190] Step 8: gNBs send RL reports to ML RL consumers. RL Report instances must contain information related to each round of RL, including:

[0191] According to the Reward Policy in RL Policy (in this example, the monitored PM and KPI), report the monitored PM and / or KPI data;

[0192] RL Record can be a list that records information for each round of RL. Specifically, it can include at least one of the following parameters: RL ROUND (RL ROUND should be 1 at this time); Actions (the operations performed in this RL, such as configuration information); Reward (indicating the reward information for this round of RL, which can be 1 / 0 or a specific reward value); Context (indicating the reason for the reward in this RL).

[0193] FallBackRecommendation, when the Agent evaluates the FallBackRecommendation report, can specifically include at least one of the following parameters: FallbackRLID (the ID of the RL round to which the fallback needs to be performed); Fallback Reason (indicating the reason for performing the fallback, such as potential risks, deterioration of the PM metrics being monitored, etc.);

[0194] RL state: The current reinforcement learning state of the model, which is currently RL ongoing.

[0195] Step 9: The ML RL consumer calculates the reward based on the reward expression or by itself, according to the monitoring data reported by the ML RL producer, and evaluates whether to continue executing RL, execute fallback, or terminate RL.

[0196] There are three possible scenarios:

[0197] Scenario 1: If the RL process continues, perform the following steps:

[0198] Step 10a: The ML RL consumer sends RL ROUND, Reward value.

[0199] Step 11a: The Agent continues the RL process, configuring and updating the RL Reward in the RLRequest instance, and configuring the RL Record parameters. The RL State in the ML Model instance and RL Report instance is updated to RL ongoing, and the process returns to Step 8.

[0200] Scenario 2: If RL fallback is executed, perform the following steps:

[0201] Step 10b: The ML RL consumer sends a fallback request, which includes the fallback RL ROUND;

[0202] Step 11b: The Agent executes the RL fallback, configures and updates the RL Fallback in the RLRequest instance, and executes the action corresponding to the RL ROUND rollback. The RL State in the ML Model instance and RL Report instance is updated to the RL fallback, and the process returns to step 8.

[0203] Scenario 3: If the RL process is terminated, perform the following steps:

[0204] Step 10c: The ML RL consumer requests to terminate the RL process;

[0205] Step 11c: The Agent executes RL stop, configures and updates RL stop in the RLRequest instance, and updates RL State in the ML Model instance and RL Report instance to RL termination;

[0206] Step 12c: The Agent sends an RL Report to the ML RL consumer, which contains RL process information and may include at least one of the following parameters: RL State; RL Record; RLGoalFuilfilment.

[0207] Example 2

[0208] Figure 10 is a schematic diagram of another reinforcement learning-based communication process provided in one embodiment. As shown in Figure 10, in this embodiment, inference is located in RAN-OAM, and execution is located in the AIML Inference Emulation Function.

[0209] The communication process based on reinforcement learning in this embodiment is basically the same as that in Embodiment 1, but the Agent is AIMLInferenceEmulationFunction. This MnF can be located in RAN-OAM and is used to perform simulation of AIMLInference results, simulating the possible impact of ML Model inference results on the network.

[0210] One scenario is that the second management network element (Producer) calculates the reward value based on the reward policy and automatically rolls it back.

[0211] Example 3

[0212] Figure 11 is a schematic diagram of another reinforcement learning-based communication process provided in one embodiment. As shown in Figure 11, in this embodiment, both inference and execution are located at gNB, that is, gNB acts as Agent, and Producer (RAN-OAM) defines Reward; there are multiple Agents (gNBs). As shown in Figure 11, the reinforcement learning-based communication process mainly includes:

[0213] Step 1: The ML RL consumer (taking Cross-domain OAM as an example here) requests the creation of an MLT Request instance on the ML RL producer (taking RAN OAM as an example here), requesting the training of the ML Model for base station energy saving through reinforcement learning. The reinforcement learning initiation request information includes at least the following: RL indication (including at least one of the following indications: explore; fallback; terminate); RL Policy (including at least one of the following information: Reward Policy; Fallback Policy); RL Control (including at least one of the following parameters: RL goal; RL interval; RL Round number).

[0214] Furthermore, in this embodiment, both inference and execution are located at the gNB, which involves multiple agents. The reinforcement learning initiation request information may additionally contain: agent selection information for the RL process, and / or agent grouping information for agent grouping.

[0215] The agent selection information may include at least one of the following parameters:

[0216] Agent ID, Agent identification information, such as gNB identifier;

[0217] Agent Conditions, such as gNB geographic location information and gNB energy consumption threshold information.

[0218] The proxy group information may include at least one of the following parameters:

[0219] Agent Group, including Agent Group ID and Agent identifier under each Agent Group ID;

[0220] Agent Group Conditions include, for example, geographical location information (gNBs located in the same area are grouped together); PM threshold information (gNBs that meet a certain PM / KPI threshold condition are grouped together); equipment vendor information (gNBs from different equipment vendors are grouped together); or gNBs in different energy-saving states are grouped together.

[0221] Steps 2-5: See Example 1

[0222] Step 6: The ML RL producer (RAN OAM, MLTFunction) confirms the Agent (gNB) based on the Agent selection information, classifies the gNB to determine different Agent Groups, configures different Reward Policies for different Agent Groups, and sends reinforcement learning initiation requests, which may include the RL Policy and RL Control from Step 1.

[0223] Step 7: Multiple agent groups collaborate to perform reinforcement learning and monitor the network environment, specifically including the following steps (in any order):

[0224] a. Each Agent Group initiates the reinforcement learning process, creates an RLRequest instance, and creates an RLReport instance.

[0225] The information contained in the RLRequest instance can be found in Example 1.

[0226] b. Each Agent Group deploys an ML model (different groups may deploy different models), performs model inference, generates energy-saving strategies, and configures the corresponding energy-saving strategy parameters;

[0227] c. Execution and monitoring of inference results: Execute base station energy-saving inference strategies; create Threshold Monitor instances based on Reward Policy and Fallback Policy to monitor network performance metrics.

[0228] Step 8: Calculate the reward based on the reward policy; evaluate whether to continue execution of RL, execute fallback, or terminate RL based on the fallback policy.

[0229] It can be divided into the following three situations:

[0230] Scenario 1: If the RL process continues, perform the following steps:

[0231] Step 9a: Each Agent Group continues the RL process, updates the RLRequest instance, configures the Reward for different Agent Groups, and configures the RL Reward parameters.

[0232] RL Reward is used to indicate the RL Reward request process. In addition to the indication information, it can be a list recording the information for each round of RL, and may include at least one of the following parameters:

[0233] Agent Group ID / Agent ID; RL ROUND; Reward; Context.

[0234] Step 10a: Each Agent Group sends an RL report to the ML RL producer. The RL Report instance must contain RL-related information for each Agent Group in each round, including at least one of the following:

[0235] Based on the Reward Policy / Fallback Policy in the RL Policy (in this example, the monitored PM and KPI), report the monitored PM and KPI data;

[0236] RL Record can be a list that records information for each Agent Group in each round of RL. It can contain at least one of the following parameters: Agent Group ID or Agent ID; RL ROUND (RL ROUND should be 1 in this case); Actions; Reward; Context.

[0237] The RL State in the ML Model instance and RL Report instance of each Agent Group is updated to RL ongoing.

[0238] Go back to step 8.

[0239] Scenario 2: If RL Fallback is executed, perform the following steps:

[0240] Step 9b: Determine the Agent Group or Agent that needs to be fallback / stopped based on the Fallback policy, execute the RL fallback process, update the RLRequest instance for each Agent Group, and configure the RL fallback / RL stop parameters (some Agent Groups fallback, some stop RL).

[0241] Step 10b: Each Agent Group sends an RL report to the ML RL producer.

[0242] The RL State in the ML Model instances and RL Report instances of each Agent Group is updated to RL Fallback.

[0243] Go back to step 8.

[0244] Scenario 2: If the RL process is terminated, perform the following steps:

[0245] Step 9c: All agents execute RL stop, and the RL State is updated to termination;

[0246] Step 10c: The Agent sends an RL Report to the ML RL producer, which contains RL process information and may include at least one of the following parameters: RL State; RL Record; RLGoalFuilfilment.

[0247] Step 11: The ML RL producer sends the MLT Report to the ML RL consumer, containing all the contents of 10c.

[0248] The RL State in the ML Model instance and RL Report instance of each Agent Group is updated to RL Completed.

[0249] Example 4

[0250] Figure 12 is a schematic diagram of another reinforcement learning-based communication process provided in one embodiment. As shown in Figure 12, both inference and execution are located at the gNB, that is, the gNB acts as the Agent, and the Producer (RAN-OAM) defines the Reward as the General Agent. It should be noted that in some scenarios, an agent node can also be selected as the General Agent.

[0251] As shown in Figure 12, the communication process based on reinforcement learning mainly includes:

[0252] Steps 1-7 are described in Example 3.

[0253] Step 8: Each Agent Group sends an RL report to the General Agent, including monitoring data based on the RL Policy;

[0254] Step 9: The General Agent calculates the reward for each Agent Group based on the PMs of multiple Agent Groups, and evaluates whether to continue executing the RL, execute the fallback, or terminate the RL. This can be divided into the following three cases:

[0255] Scenario 1: If you continue with RL, perform the following steps:

[0256] Step 10a: The General Agent requests the configuration of RL Rewards for each Agent Group;

[0257] See 9a-10a of Example 3 for 11a-12a.

[0258] Scenario 2: If RL Fallback is executed, perform the following steps:

[0259] Step 10a: The General Agent requests to configure RL Fallback and / or RL Stop for each Agent Group;

[0260] See 9b-10b of Example 3 for 11b-12b.

[0261] Scenario 3: If the RL process is terminated, perform the following steps:

[0262] Step 10c: The General Agent requests the configuration of RL Stop for each Agent Group;

[0263] See 9c-11c of Example 3 for 11c-13c.

[0264] This application also provides a reinforcement learning-based communication device. Figure 13 is a schematic diagram of the structure of a reinforcement learning-based communication device according to an embodiment. As shown in Figure 13, the reinforcement learning-based communication device includes:

[0265] The startup request sending module 310 is configured to receive reinforcement learning startup requests sent by the first management network element or the second management network element.

[0266] The execution module 320 is configured to perform reinforcement learning in accordance with the reinforcement learning initiation request.

[0267] The reporting module 330 is configured to send reinforcement learning reports to the first management network element or the second management network element.

[0268] In one embodiment, the reinforcement learning initiation request includes at least one of the following: reinforcement learning instruction information, reinforcement learning strategy information, reinforcement learning control information, reinforcement learning rollback instruction information, reinforcement learning reward instruction information, reinforcement learning termination instruction information, and the address of a first management network element.

[0269] In one embodiment, the reinforcement learning instruction information includes at least one of the following: proceed, rollback, or terminate.

[0270] In one embodiment, the reinforcement learning rollback instruction information includes at least one of the following: rollback instruction for the current round, rollback instruction for the target round.

[0271] In one embodiment, the reinforcement learning reward indication information includes at least one of the following: the current reinforcement learning round number, the reinforcement learning reward value for this round, and the reason for this reinforcement learning reward.

[0272] In one embodiment, the reinforcement learning report includes at least one of the following:

[0273] Based on the monitoring metrics reported by the reinforcement learning strategy information; the model's reinforcement learning status; reinforcement learning records; rollback recommendation information; and the achievement of reinforcement learning objectives.

[0274] In one embodiment, the model reinforcement learning state includes at least one of the following: an executable reinforcement learning state, a reinforcement learning in progress state, a reinforcement learning completed state, and a reinforcement learning rollback state.

[0275] In one embodiment, the fallback recommendation information includes at least one of the following:

[0276] The rollback target round indicator is used to indicate the reinforcement learning rounds that need to be rolled back.

[0277] The rollback reason indicates the reason for performing a rollback in reinforcement learning.

[0278] In one embodiment, the reinforcement learning record includes at least one of the following: the current reinforcement learning round number, the reinforcement learning operation in this round, the reinforcement learning reward value in this round, and the reason for the reinforcement learning reward.

[0279] In one embodiment, when the number of agent nodes is one, the first management network element calculates the reward value according to the reward strategy in the reinforcement learning initiation request, and determines whether to continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning; the device further includes an update module configured to:

[0280] While continuing reinforcement learning, update the reinforcement learning reward indication information and reinforcement learning records, and send the corresponding reinforcement learning report;

[0281] In the event of reinforcement learning rollback, update the reinforcement learning rollback instruction information and the model reinforcement learning status, perform the rollback operation for the corresponding number of rounds, and send the corresponding reinforcement learning report;

[0282] If reinforcement learning is terminated, update the reinforcement learning termination instruction information and the model reinforcement learning status, and send the corresponding reinforcement learning report.

[0283] In one embodiment, the reinforcement learning initiation request further includes at least one of the following information for determining the proxy node:

[0284] Agent selection information is used to indicate the information required to select an agent node;

[0285] Proxy grouping information is used to indicate information on grouping proxy nodes;

[0286] The agent selection information includes at least one of the following:

[0287] Agent identifier, used to indicate the agent node performing reinforcement learning;

[0288] The proxy conditions are used to indicate the basis for selecting a proxy node.

[0289] The proxy group information includes at least one of the following:

[0290] Agent group information is used to indicate each agent group and the agent node corresponding to each agent group;

[0291] Agent group conditions are used to indicate the basis for grouping agent nodes.

[0292] In one embodiment, when the number of proxy nodes is at least two, the device further includes:

[0293] The module is set up so that the main agent node calculates the reward value for each agent group and determines whether each agent group should continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning.

[0294] Update the module and set it to:

[0295] While continuing reinforcement learning, the master agent node requests the configuration of reinforcement learning reward indication information, and the agent nodes in each agent group update the corresponding reinforcement learning reward indication information and the corresponding reinforcement learning record, and send the corresponding reinforcement learning report.

[0296] In the case of reinforcement learning rollback, the main agent node requests the configuration of reinforcement learning rollback instruction information or reinforcement learning termination instruction information. Each agent group updates the corresponding reinforcement learning rollback instruction information or the corresponding reinforcement learning termination instruction information according to the rollback strategy. Each agent group updates the corresponding model reinforcement learning status and sends the corresponding reinforcement learning report.

[0297] In the event of termination of reinforcement learning, the master agent node requests the configuration of reinforcement learning termination indication information, and each agent group updates the corresponding reinforcement learning termination indication information and model reinforcement learning status, and sends the corresponding reinforcement learning report.

[0298] In one embodiment, the device further includes:

[0299] The module for determining the general agent is set up to identify the general agent node.

[0300] The general agent node is the second management network element, or it is an agent node selected from at least two agent nodes based on agent selection information.

[0301] In one embodiment, when the number of proxy nodes is at least two, the device further includes:

[0302] The module is set up so that each agent group calculates its own reward value and determines whether to continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning for that agent group.

[0303] Update the module and set it to:

[0304] While continuing reinforcement learning, the agent nodes in each agent group update the corresponding reinforcement learning reward indication information and the corresponding reinforcement learning records, and send the corresponding reinforcement learning report.

[0305] When performing reinforcement learning rollback, the agent nodes in each agent group update the corresponding reinforcement learning rollback instruction information or the corresponding reinforcement learning termination instruction information according to the rollback strategy. The agent nodes in each agent group update the corresponding model reinforcement learning status and send the corresponding reinforcement learning report.

[0306] In the event of termination of reinforcement learning, the agent nodes in each agent group update the corresponding reinforcement learning termination indication information and the corresponding model reinforcement learning status, and send the corresponding reinforcement learning report.

[0307] In one embodiment, when the number of agent nodes is at least two, the second management network element configures corresponding reward strategies for different agent groups and sends reinforcement learning initiation requests to each agent group respectively.

[0308] In one embodiment, performing reinforcement learning according to the reinforcement learning initiation request includes:

[0309] Each agent node in the agent group executes reinforcement learning according to the corresponding reinforcement learning initiation request.

[0310] In one embodiment, before receiving a reinforcement learning initiation request sent by a first management network element or a second management network element, the apparatus further includes:

[0311] The training request sending module is configured to send machine learning training requests from the first management network element to the second management network element;

[0312] The training module is configured to have the second management network element perform machine learning training according to the machine learning training request to obtain a machine learning model.

[0313] In one embodiment, the machine learning training request includes at least one of the following information: reinforcement learning instruction information, reinforcement learning strategy information, reinforcement learning control information, and the address of a first management network element.

[0314] In one embodiment, the reinforcement learning policy information includes at least one of the following:

[0315] Reward strategy, used to indicate the information needed to calculate rewards;

[0316] Rollback strategies are used to indicate the conditions for rollback, including performance-based conditions and reward-based conditions.

[0317] In one embodiment, the reinforcement learning control information includes at least one of the following: reinforcement learning objective; reinforcement learning interval; reinforcement learning rounds.

[0318] In one embodiment, the device further includes:

[0319] The model creation module is configured to allow the second management network element to create a machine learning model instance based on the machine learning training request. The machine learning model instance includes at least one of the following information: model reinforcement learning status, reinforcement learning rounds, and reinforcement learning records.

[0320] In one embodiment, the device further includes:

[0321] The training report module is configured to send machine learning training reports from the second management network element to the first management network element;

[0322] The machine learning training report includes at least one of the following information: the model's reinforcement learning status and the number of reinforcement learning rounds.

[0323] In one embodiment, when the number of proxy nodes is at least two, the machine learning training request further includes at least one of the following information for determining the proxy nodes:

[0324] Agent selection information is used to indicate the information required to select an agent node;

[0325] Proxy grouping information is used to indicate information on grouping proxy nodes;

[0326] The agent selection information includes at least one of the following:

[0327] Agent identifier, used to indicate the agent node performing reinforcement learning;

[0328] The proxy conditions are used to indicate the basis for selecting a proxy node.

[0329] The proxy group information includes at least one of the following:

[0330] Agent group information is used to indicate each agent group and the agent node corresponding to each agent group;

[0331] Agent group conditions are used to indicate the basis for grouping agent nodes.

[0332] The reinforcement learning-based communication device proposed in this embodiment belongs to the same inventive concept as the reinforcement learning-based communication method proposed in the above embodiments. Technical details not described in detail in this embodiment can be found in any of the above embodiments. Furthermore, this embodiment has the same beneficial effects as the reinforcement learning-based communication method.

[0333] This application also provides a reinforcement learning-based communication device. Figure 14 is a schematic diagram of the structure of a reinforcement learning-based communication device according to an embodiment. As shown in Figure 14, the reinforcement learning-based communication device includes:

[0334] The startup request receiving module 410 is configured to send a reinforcement learning startup request to the agent node, the reinforcement learning startup request being used to instruct the agent node to perform reinforcement learning.

[0335] The report receiving module 420 is configured to receive reinforcement learning reports from the agent node.

[0336] In one embodiment, the management network element is a first management network element or a second management network element; when the number of agent nodes is one, the device further includes:

[0337] The determination module is configured to calculate the reward value based on the reward strategy in the reinforcement learning initiation request, and determine whether to continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning.

[0338] In one embodiment, the management network element is a first management network element or a second management network element; when the number of agent nodes is at least two, the device further includes:

[0339] The request sending module is configured to provide corresponding reward strategies for different agent groups, and send reinforcement learning initiation requests to each agent group according to the corresponding reward strategies.

[0340] In one embodiment, the management network element includes a first management network element and a second management network element; before sending a reinforcement learning initiation request to the agent node, the device further includes:

[0341] The training request sending module is configured to send a machine learning training request from the first management network element to the second management network element;

[0342] The training module is configured to allow the second management network element to perform machine learning training based on the machine learning training request, thereby obtaining a machine learning model.

[0343] In one embodiment, the device further includes:

[0344] The model creation module is configured to allow the second management network element to create a machine learning model instance based on the machine learning training. The machine learning model instance includes at least one of the following information: model reinforcement learning status, reinforcement learning round number, reinforcement learning record, current reinforcement learning round number, reinforcement learning operation in this round, reinforcement learning reward value in this round, and reason for this reinforcement learning reward.

[0345] In one embodiment, the reinforcement learning report of the proxy node is received by the second management network element; the method further includes:

[0346] The training report module is configured to send machine learning training reports from the second management element to the first management network element.

[0347] The reinforcement learning-based communication device proposed in this embodiment belongs to the same inventive concept as the reinforcement learning-based communication method proposed in the above embodiments. Technical details not described in detail in this embodiment can be found in any of the above embodiments. Furthermore, this embodiment has the same beneficial effects as the reinforcement learning-based communication method.

[0348] This application also provides a proxy node. Figure 15 is a schematic diagram of the hardware structure of a proxy node provided in an embodiment. As shown in Figure 15, the proxy node provided in this application includes a processor 510 and a memory 520. The processor 510 in the proxy node can be one or more, and Figure 15 shows one processor 510 as an example. The memory 520 is configured to store one or more programs. The one or more programs are executed by the one or more processors 510, so that the one or more processors 510 implement the reinforcement learning-based communication method as described in the embodiment of this application.

[0349] The agent node also includes: communication device 530, input device 540 and output device 550.

[0350] The processor 510, memory 520, communication device 530, input device 540 and output device 550 in the agent node can be connected by a bus or other means. Figure 15 shows an example of connection via a bus.

[0351] Input device 540 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the agent node. Output device 550 may include display devices such as a display screen.

[0352] The communication device 530 may include a receiver and a transmitter. The communication device 530 is configured to perform information transmission and reception communication under the control of the processor 510.

[0353] The memory 520, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the reinforcement learning-based communication method described in the embodiments of this application (e.g., the start request receiving module 310, execution module 320, and report 330 in the reinforcement learning-based communication device). The memory 520 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the agent node, etc. Furthermore, the memory 520 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 520 may further include memory remotely located relative to the processor 510, and these remote memories can be connected to the agent node via a network. Examples of such networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0354] This application also provides a proxy node. Figure 16 is a schematic diagram of the hardware structure of a proxy node provided in one embodiment. As shown in Figure 16, the proxy node provided in this application includes a processor 610 and a memory 620. The processor 610 in the proxy node can be one or more, and Figure 16 shows one processor 610 as an example. The memory 620 is configured to store one or more programs. The one or more programs are executed by the one or more processors 610, so that the one or more processors 610 implement the reinforcement learning-based communication method as described in the embodiment of this application.

[0355] The agent node also includes: a communication device 630, an input device 640, and an output device 650.

[0356] The processor 610, memory 620, communication device 630, input device 640 and output device 650 in the agent node can be connected by a bus or other means. Figure 16 shows an example of connection by bus.

[0357] Input device 640 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the agent node. Output device 650 may include display devices such as a display screen.

[0358] The communication device 630 may include a receiver and a transmitter. The communication device 630 is configured to perform information transmission and reception communication under the control of the processor 610.

[0359] The memory 620, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the reinforcement learning-based communication method described in the embodiments of this application (e.g., the start request sending module 410 and the report receiving module 420 in the reinforcement learning-based communication device). The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the agent node, etc. Furthermore, the memory 620 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 620 may further include memory remotely located relative to the processor 610, and these remote memories can be connected to the agent node via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0360] Figure 17 is a schematic diagram of a reinforcement learning-based communication system provided in an embodiment of this application. As shown in Figure 17, the system includes: a proxy node 710 as described in any of the above embodiments and a management network element 720 as described in any of the above embodiments. The management network element 720 can be divided into a first management network element and a second management network element.

[0361] This application embodiment also provides a storage medium storing a computer program. When executed by a processor, the computer program implements any of the reinforcement learning-based communication methods described in this application embodiment. The method includes: receiving a reinforcement learning initiation request sent by a first management network element or a second management network element; performing reinforcement learning according to the reinforcement learning initiation request; and sending a reinforcement learning report to the first management network element or the second management network element. Alternatively, the method includes: sending a reinforcement learning initiation request to a proxy node, the reinforcement learning initiation request being used to instruct the proxy node to perform reinforcement learning; and receiving a reinforcement learning report from the proxy node.

[0362] This application also provides a stored computer program product, including a computer program / instructions, which, when executed by a processor, implement any of the reinforcement learning-based communication methods described in this application. The method includes: receiving a reinforcement learning initiation request sent by a first management network element or a second management network element; performing reinforcement learning according to the reinforcement learning initiation request; and sending a reinforcement learning report to the first management network element or the second management network element. Alternatively, the method includes: sending a reinforcement learning initiation request to a proxy node, the reinforcement learning initiation request being used to instruct the proxy node to perform reinforcement learning; and receiving the reinforcement learning report from the proxy node.

[0363] The computer storage medium in this application embodiment can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable CD-ROM, optical storage device, magnetic storage device, or any suitable combination thereof. The computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0364] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device.

[0365] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, radio frequency (RF), etc., or any suitable combination thereof.

[0366] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0367] This application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the reinforcement learning-based communication method as described in any of the above embodiments.

[0368] The above description is merely an exemplary embodiment of this application and is not intended to limit the scope of protection of this application.

[0369] Those skilled in the art will understand that the term user terminal encompasses any suitable type of wireless user equipment, such as mobile phones, portable data processing portable web browsers, or vehicle-mounted mobile stations.

[0370] Generally, the various embodiments of this application can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, although this application is not limited thereto.

[0371] Embodiments of this application can be implemented by executing computer program instructions through the data processor of a mobile device, for example, in a processor entity, or through hardware, or through a combination of software and hardware. The computer program instructions can be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages.

[0372] Any block diagram of logical flow in the accompanying drawings of this application may represent program steps, or may represent interconnected logic circuits, modules, and functions, or may represent a combination of program steps and logic circuits, modules, and functions. The computer program may be stored in memory. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as, but not limited to, read-only memory (ROM), random access memory (RAM), optical storage devices and systems (Digital Video Disc (DVD) or Compact Disk (CD), etc.). Computer-readable media may include non-transitory storage media. The data processor may be of any type suitable to the local technical environment, such as, but not limited to, general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and processors based on multi-core processor architectures.

Claims

1. A reinforcement learning-based communication method applied to agent nodes, comprising: Receive reinforcement learning initiation request sent by the first management network element or the second management network element; Execute reinforcement learning according to the reinforcement learning initiation request; Send a reinforcement learning report to the first management network element or the second management network element.

2. The method according to claim 1, wherein, The reinforcement learning initiation request includes at least one of the following: Reinforcement learning instruction information, reinforcement learning strategy information, reinforcement learning control information, reinforcement learning rollback instruction information, reinforcement learning reward instruction information, reinforcement learning termination instruction information, and the address of the first management network element.

3. The method according to claim 2, wherein, In response to determining that the reinforcement learning initiation request includes the reinforcement learning instruction information, the reinforcement learning instruction information includes at least one of the following: proceed, rollback, terminate.

4. The method according to claim 2, wherein, In response to determining that the reinforcement learning start request includes the reinforcement learning rollback instruction information, the reinforcement learning rollback instruction information includes at least one of the following: rollback instruction for this round, rollback instruction for the target round.

5. The method according to claim 2, wherein, In response to determining that the reinforcement learning initiation request includes the reinforcement learning reward indication information, the reinforcement learning reward indication information includes at least one of the following: the current reinforcement learning round number, the reinforcement learning reward value for this round, and the reason for this reinforcement learning reward.

6. The method according to claim 1, wherein, The reinforcement learning report includes at least one of the following: Based on the monitoring metrics reported by the reinforcement learning strategy information; the model's reinforcement learning status; reinforcement learning records; and rollback recommendation information; Strengthen the achievement of learning objectives.

7. The method according to claim 6, wherein, In response to determining that the reinforcement learning report includes the model reinforcement learning state, the model reinforcement learning state includes at least one of the following: The states that can be executed include reinforcement learning state, reinforcement learning in progress state, reinforcement learning completed state, and reinforcement learning rollback state.

8. The method according to claim 6, wherein, In response to determining that the reinforcement learning report includes the fallback recommendation information, the fallback recommendation information includes at least one of the following: The rollback target round indicator is used to indicate the reinforcement learning rounds that need to be rolled back. The rollback reason indicates the reason for performing a rollback in reinforcement learning.

9. The method according to claim 6, wherein, In response to determining that the reinforcement learning report includes the reinforcement learning record, the reinforcement learning record includes at least one of the following: the current reinforcement learning round number, the reinforcement learning operation in this round, the reinforcement learning reward value in this round, and the reason for the reinforcement learning reward.

10. The method according to claim 1, wherein, In response to determining that the number of agent nodes is one, the first management network element calculates the reward value according to the reward strategy in the reinforcement learning initiation request, and determines whether to continue executing reinforcement learning, execute reinforcement learning rollback, or terminate the execution of reinforcement learning. The method further includes: In response to the decision to continue reinforcement learning, update the reinforcement learning reward indication information and reinforcement learning records, and send the corresponding reinforcement learning report; In response to the determination to perform reinforcement learning rollback, update the reinforcement learning rollback instruction information and the model reinforcement learning status, perform the rollback operation for the corresponding number of rounds, and send the corresponding reinforcement learning report; In response to determining to terminate reinforcement learning, update the reinforcement learning termination indication information and the reinforcement learning status of the model, and send the corresponding reinforcement learning report.

11. The method according to claim 1, wherein, The reinforcement learning initiation request also includes at least one of the following pieces of information for determining the proxy node: Agent selection information is used to indicate the information required to select an agent node; Proxy grouping information is used to indicate information on grouping proxy nodes; The agent selection information includes at least one of the following: Agent identifier, used to indicate the agent node performing reinforcement learning; The proxy conditions are used to indicate the basis for selecting a proxy node. The proxy group information includes at least one of the following: Agent group information is used to indicate each agent group and the agent node corresponding to each agent group; Agent group conditions are used to indicate the basis for grouping agent nodes.

12. The method according to claim 1, wherein, In response to determining that the number of proxy nodes is at least two, the method further includes: The master agent node calculates the reward value for each agent group and determines whether each agent group should continue to perform reinforcement learning, perform reinforcement learning rollback, or terminate reinforcement learning. In response to the decision to continue reinforcement learning, the master agent node requests the configuration of reinforcement learning reward indication information, and the agent nodes in each agent group update the corresponding reinforcement learning reward indication information and the corresponding reinforcement learning record, and send the corresponding reinforcement learning report. In response to the determination to perform reinforcement learning rollback, the master agent node requests to configure reinforcement learning rollback instruction information or reinforcement learning termination instruction information. Each agent group updates the corresponding reinforcement learning rollback instruction information or the corresponding reinforcement learning termination instruction information according to the rollback strategy. Each agent group updates the corresponding model reinforcement learning status and sends the corresponding reinforcement learning report. In response to the determination to terminate reinforcement learning, the master agent node requests the configuration of reinforcement learning termination indication information, and each agent group updates the corresponding reinforcement learning termination indication information and the corresponding model reinforcement learning status, and sends the corresponding reinforcement learning report.

13. The method according to claim 1, wherein, In response to determining that the number of proxy nodes is at least two, the method further includes: Each agent group calculates its corresponding reward value and determines whether the agent group should continue to perform reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning. In response to the decision to continue reinforcement learning, the agent nodes in each agent group update the corresponding reinforcement learning reward indication information and the corresponding reinforcement learning record, and send the corresponding reinforcement learning report. In response to determining to perform reinforcement learning rollback, the agent nodes in each agent group update the corresponding reinforcement learning rollback indication information or the corresponding reinforcement learning termination indication information according to the rollback strategy. The agent nodes in each agent group update the corresponding model reinforcement learning status and send the corresponding reinforcement learning report. In response to determining to terminate reinforcement learning, the agent nodes in each agent group update the corresponding reinforcement learning termination indication information and the corresponding model reinforcement learning status, and send the corresponding reinforcement learning report.

14. The method according to claim 1, wherein, In response to determining that the number of agent nodes is at least two, the second management network element configures corresponding reward strategies for different agent groups and sends a reinforcement learning start request to each agent group.

15. The method according to claim 14, wherein, The step of performing reinforcement learning according to the reinforcement learning initiation request includes: The agent nodes in each agent group execute reinforcement learning according to the corresponding reinforcement learning initiation request.

16. The method according to claim 1, further comprising, before receiving the reinforcement learning initiation request sent by the first management network element or the second management network element: The first management network element sends a machine learning training request to the second management network element; The second management network element performs machine learning training according to the machine learning training request to obtain a machine learning model.

17. The method according to claim 16, wherein, The machine learning training request includes at least one of the following information: reinforcement learning instruction information, reinforcement learning strategy information, reinforcement learning control information, and the address of the first management network element.

18. The method according to claim 17, wherein, In response to determining that the machine learning training request includes the reinforcement learning policy information, the reinforcement learning policy information includes at least one of the following: Reward strategy, used to indicate the information needed to calculate rewards; Rollback strategies are used to indicate the conditions for rollback, including performance-based conditions and reward-based conditions.

19. The method according to claim 17, wherein, In response to determining that the machine learning training request includes the reinforcement learning control information, the reinforcement learning control information includes at least one of the following: reinforcement learning objective; reinforcement learning interval; reinforcement learning rounds.

20. The method of claim 16, further comprising: The second management network element creates a machine learning model instance based on the machine learning training request. The machine learning model instance includes at least one of the following information: model reinforcement learning state, reinforcement learning round number, and reinforcement learning record.

21. The method of claim 16, further comprising: The second management network element sends a machine learning training report to the first management network element; The machine learning training report includes at least one of the following information: model reinforcement learning status, number of reinforcement learning rounds, and reinforcement learning records.

22. The method according to claim 16, wherein, In response to determining that the number of proxy nodes is at least two, the machine learning training request further includes at least one of the following information for determining the proxy nodes: Agent selection information is used to indicate the information required to select an agent node; Proxy grouping information is used to indicate information on grouping proxy nodes; The agent selection information includes at least one of the following: Agent identifier, used to indicate the agent node performing reinforcement learning; The proxy conditions are used to indicate the basis for selecting a proxy node. The proxy group information includes at least one of the following: Agent group information is used to indicate each agent group and the agent node corresponding to each agent group; Agent group conditions are used to indicate the basis for grouping agent nodes.

23. The method according to claim 11 or 22, further comprising: Determine the master agent node; The general agent node is either the second management network element or an agent node selected from at least two agent nodes based on agent selection information.

24. A reinforcement learning-based communication method applied to management network elements, comprising: Send a reinforcement learning initiation request to the agent node, the reinforcement learning initiation request being used to instruct the agent node to perform reinforcement learning; Receive the reinforcement learning report from the agent node.

25. The method according to claim 24, wherein, The management network element is either the first management network element or the second management network element; In response to determining that the number of proxy nodes is one, the method further includes: The reward value is calculated based on the reward strategy in the reinforcement learning initiation request, and it is determined whether to continue reinforcement learning, roll back reinforcement learning, or terminate reinforcement learning.

26. The method according to claim 24, wherein, The management network element is either the first management network element or the second management network element; In response to determining that the number of proxy nodes is at least two, sending a reinforcement learning initiation request to the proxy nodes includes: Configure corresponding reward strategies for different agent groups, and send reinforcement learning initiation requests to each agent group according to the corresponding reward strategies.

27. The method according to claim 24, wherein the management network element includes a first management network element and a second management network element; before sending a reinforcement learning initiation request to the agent node, the method further includes: The first management network element sends a machine learning training request to the second management network element; The second management network element performs machine learning training according to the machine learning training request to obtain a machine learning model.

28. The method of claim 27, further comprising: The second management network element creates a machine learning model instance based on the machine learning training. The machine learning model instance includes at least one of the following information: model reinforcement learning state, reinforcement learning round number, reinforcement learning record, current reinforcement learning round number, reinforcement learning operation in this round, reinforcement learning reward value in this round, and reason for this reinforcement learning reward.

29. The method according to claim 27, wherein, The reinforcement learning report of the proxy node is received by the second management network element; the method further includes: The second management sends a machine learning training report to the first management network element.

30. A proxy node, comprising: Memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the reinforcement learning-based communication method as described in any one of claims 1-23.

31. A management network element, comprising: Memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the reinforcement learning-based communication method as described in any one of claims 24-29.

32. A communication system based on reinforcement learning, comprising: The proxy node as described in claim 30 and the management network element as described in claim 31.

33. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the reinforcement learning-based communication method as described in any one of claims 1-29.

Citation Information

Patent Citations

  • Training method for reinforcement learning network, device, training apparatus and storage medium

    CN109242099A

  • Reinforcement learning apparatus and method based on user learning environment

    US20230088699A1