Reinforcement learning processing method, storage medium, electronic device, and computer program product
By introducing reinforcement learning processing methods into the core network and utilizing environment interpreters and action feedback mechanisms, the problem of analysis quality monitoring caused by the core network's lack of support for reinforcement learning was solved. Real-time analysis quality feedback and model enhancement were achieved, improving the accuracy and efficiency of the analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2025-08-21
- Publication Date
- 2026-05-15
AI Technical Summary
The core network does not support reinforcement learning-related technologies, resulting in the lack of a real-time analysis quality feedback mechanism, making it impossible to perform model enhancement and analysis quality monitoring.
A reinforcement learning processing method is provided, which obtains the result feedback of the machine learning model through an environment interpreter, generates rewards and/or states, and obtains action feedback from a first entity to achieve dynamic adjustment and quality monitoring of the model.
It implements a real-time analysis quality feedback mechanism in the core network, supports model enhancement and analysis quality monitoring, improves the accuracy and efficiency of analysis, and is suitable for various data analysis scenarios.
Smart Images

Figure CN2025116145_15052026_PF_FP_ABST
Abstract
Description
Reinforcement learning processing methods, storage media, electronic devices, and computer program products
[0001] Cross-reference to related applications
[0002] This disclosure is based on and claims priority to Chinese patent application CN202411594031.4, filed on November 8, 2024, entitled “Reinforcement Learning Processing Method, Storage Medium, Electronic Device and Computer Program Product”, and incorporates the entire contents of that patent application by reference. Technical Field
[0003] This disclosure relates to the field of communication technology, and more specifically, to a reinforcement learning processing method, storage medium, electronic device, and computer program product. Background Technology
[0004] Reinforcement learning techniques are applicable to various scenarios, such as monitoring the analysis quality provided by network elements of the Analytics Logic Function (AnLF) and Network Data Analysis Function (NWDAF), and monitoring the quality of machine learning models trained by the Model Training Logic Function (MTLF) (NWDAF). However, the core network does not support reinforcement learning-related techniques, resulting in the lack of a real-time analysis quality feedback mechanism, making it impossible to perform model enhancement and analysis quality monitoring.
[0005] No solution has yet been proposed to address the issue that the core network in related technologies does not support reinforcement learning, resulting in the lack of a real-time analysis quality feedback mechanism and the inability to perform model enhancement and analysis quality monitoring. Summary of the Invention
[0006] This disclosure provides a reinforcement learning processing method, storage medium, electronic device, and computer program product to at least solve the problem in related technologies where the core network does not support reinforcement learning-related technologies, resulting in the lack of a real-time analysis quality feedback mechanism and the inability to perform model enhancement and analysis quality monitoring.
[0007] According to one embodiment of this disclosure, a reinforcement learning processing method is provided, applied to an environment interpreter, the method comprising: obtaining result feedback based on model inference of a machine learning model used for output analysis, and generating a reward and / or state required for reinforcement learning based on the result feedback; and obtaining action feedback from MTLF based on the reward and / or the state.
[0008] According to another embodiment of this disclosure, a reinforcement learning processing method is provided, applied to a first entity, the method comprising: receiving a reward and / or state required for reinforcement learning sent by an environment interpreter, wherein the reward and / or the state is generated by the environment interpreter based on the result feedback of model inference of a machine learning model; determining an action feedback based on the reward and / or the state; and sending the action feedback to the environment interpreter.
[0009] According to yet another embodiment of this disclosure, a computer program product is also provided, including computer program instructions, wherein the computer program instructions cause a computer to perform the steps in any of the above method embodiments.
[0010] According to yet another embodiment of this disclosure, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.
[0011] According to yet another embodiment of this disclosure, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments. Attached Figure Description
[0012] Figure 1 is a block diagram of a 5G system architecture including ADRF according to an embodiment of the present disclosure;
[0013] Figure 2 is a schematic diagram of a reinforcement learning algorithm according to an embodiment of the present disclosure;
[0014] Figure 3 is a flowchart of a reinforcement learning processing method according to an embodiment of the present disclosure;
[0015] Figure 4 is a flowchart of a reinforcement learning processing method according to an embodiment of the present disclosure;
[0016] Figure 5 is a flowchart of reinforcement learning initiated by AnLF for per analytics quality monitoring according to an embodiment of the present disclosure;
[0017] Figure 6 is a flowchart of reinforcement learning initiated by a consumer network element for per analytics quality monitoring according to an embodiment of the present disclosure;
[0018] Figure 7 is a flowchart of the reinforcement learning architecture construction according to an embodiment of the present disclosure;
[0019] Figure 8 is a flowchart of reinforcement learning initiated by MTLF for per ML Model quality monitoring according to an embodiment of the present disclosure;
[0020] Figure 9 is a flowchart of reinforcement learning for per ML Model quality monitoring initiated by MTLF according to an embodiment of the present disclosure;
[0021] Figure 10 is a block diagram of a reinforcement learning processing apparatus according to an embodiment of the present disclosure;
[0022] Figure 11 is a second block diagram of a reinforcement learning processing apparatus according to an embodiment of the present disclosure. Detailed Implementation
[0023] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings and examples.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0025] Figure 1 is a block diagram of a 5G system architecture including ADRF according to an embodiment of the present disclosure. As shown in Figure 1, the 5G system architecture supports the Analytics Data Repository Function (ADRF) to store and analyze collected data. ADRF discloses the Nadrf service for information storage and data retrieval for other 5GC network functions (such as NWDAF).
[0026] Based on the configuration of the Data Collection Coordinating Function (DCCF) on the network function, the DCCF can identify the ADRF and interact with it directly or indirectly to request or store data. Interaction methods include:
[0027] Directly: DCCF requests data to be stored in ADRF via the Nadrf service or Ndccf_DataManagement_Notify (e.g., when ADRF requests DCCF to perform a data collection notification). Additionally, DCCF retrieves data from ADRF via the Nadrf service.
[0028] Indirectly: DCCF requires the message frame to store data in ADRF via the Nadrf service or Nmfaf_3daDataManagement_Configure. The message frame may contain one or more adapters for conversion between 3GPP-defined protocols.
[0029] Consumer network functions can be specified in the request sent to DCCF, and data provided by the data source needs to be stored in ADRF.
[0030] ADRF stores data received directly from the Nadrf_DataManagement_StorageRequest sent by the network function, or data received from the Ndccf_DataManagement_Notify / Nmfaf_3caDataManagement_Notify or Nnwdaf_DataManagement_Notify sent from the DCCF, MFAF or NWDAF.
[0031] Reinforcement learning (RL) is a machine learning method that enables AI devices to learn how to make decisions through interaction with their environment. In reinforcement learning, an agent or AI agent influences its environment by performing actions and receives feedback, or rewards, based on the results of these actions. The agent's goal is to maximize its cumulative reward, which typically means learning an optimal policy to guide it in making the best decisions in various situations.
[0032] In general reinforcement learning problems, the state may change each time the agent applies a new action. This can be represented as follows: At a certain moment (s), the agent receives the state of the environment through the interpreter. The agent then selects an action (a) and applies it to the environment. When this action is applied, the environment provides a reward (r) and changes to a new state (s'). The reward and / or state are ultimately provided to the agent by the interpreter. Figure 2 is a schematic diagram of a reinforcement learning algorithm according to an embodiment of this disclosure. As shown in Figure 2, reinforcement learning allows an AI agent to efficiently learn the optimal policy through repeated trial and error and environmental feedback. Reinforcement learning can autonomously explore and learn optimal decisions in unknown dynamic environments, and compared to other machine learning methods, it excels at solving complex decision-making problems in the current network.
[0033] This disclosure provides a reinforcement learning processing method. Figure 3 is a flowchart of a reinforcement learning processing method according to an embodiment of this disclosure. As shown in Figure 3, the method is applied to an environment interpreter, and the process includes the following steps:
[0034] Step S302: Obtain result feedback based on model inference of the machine learning model used for output analysis, and generate the reward and / or state required for reinforcement learning based on the result feedback;
[0035] Step S304: Obtain action feedback from the first entity based on the reward and / or status.
[0036] Through the above steps S302 to S304, the problem that the core network in the related technology does not support reinforcement learning-related technologies, resulting in the lack of a real-time analysis quality feedback mechanism and the inability to perform model enhancement and analysis quality monitoring can be solved. This enables the core network to support reinforcement learning-related technologies, and allows for a real-time analysis quality feedback mechanism, as well as model enhancement and analysis quality monitoring.
[0037] The embodiments disclosed herein can automatically adjust the model to adapt to the ever-changing network environment, improve the accuracy and efficiency of analysis, and are applicable to various data analysis scenarios in 5G core networks, such as network performance optimization and user experience enhancement.
[0038] In one embodiment, the environment interpreter applies action feedback to ensure the quality of the analysis output or the model quality of the machine learning model. By dynamically adjusting the machine learning model, the analysis results can be optimized in real time, improving the utilization efficiency of network resources, especially in high-density user scenarios, which can effectively improve network service quality and user experience.
[0039] For example, applying action feedback to ensure the quality of analytical output or the model quality of a machine learning model can specifically include performing at least one of the following processes on the machine learning model based on the instructions from the action feedback: using a new machine learning model for model inference, updating the machine learning model and performing model inference, continuing to use the machine learning model for inference, or obtaining a new machine learning model from other network elements and performing model inference. This dynamic update mechanism enables the model to quickly adapt to network changes, especially in scenarios such as network fault recovery or new device access, allowing for rapid adjustment of analytical strategies to ensure service continuity and stability.
[0040] In this embodiment, the action feedback instruction information includes at least one of the following: a first instruction indicating the use of a new machine learning model, wherein the first instruction includes at least one of the following: a new model file, a link to the model file of the new machine learning model, or the ID of the new machine learning model; a second instruction indicating the updating of the machine learning model, wherein the second instruction carries the model parameter information to be updated and its corresponding values; a third instruction indicating continued use of the machine learning model; and a fourth instruction indicating the acquisition of a new machine learning model from other network elements, wherein the fourth instruction includes the identification information or address information of the other network elements. By refining the action feedback information, more precise model control can be achieved, which is suitable for scenarios requiring highly accurate analysis, such as network resource scheduling and traffic prediction.
[0041] In this embodiment of the disclosure, step S304 may specifically include: sending a reward and / or status to a first entity; and receiving action feedback from the first entity based on the machine learning model determined by the reward and / or status. This mechanism enables efficient interaction between the environment interpreter and the first entity, and is suitable for scenarios requiring cross-network element collaborative analysis, such as multi-network element joint optimization of network performance.
[0042] In this embodiment of the disclosure, step S302 may specifically include: when the environment interpreter is set in the second entity, performing model inference using a machine learning model and sending the model inference result to the 5GC network element; receiving the result feedback of the model inference result generated by the 5GC network element; and converting the result feedback into a reward and / or a state. This setup can separate the execution of reinforcement learning from the data source, improve the flexibility of the system, and is suitable for large-scale network data analysis scenarios, such as network anomaly detection and user behavior analysis.
[0043] In one embodiment, the method further includes: receiving an analysis quality monitoring request sent by a 5GC network element during subscription analysis; and making a decision and initiating reinforcement learning-based analysis quality monitoring. This subscription mechanism allows for proactive monitoring and optimization of analysis quality, making it suitable for scenarios requiring long-term stable analysis quality, such as network service level agreement (NSA) guarantees.
[0044] In this embodiment of the disclosure, the aforementioned quality monitoring request carries an identifier ID of the analyzed data and / or a suggestion to utilize reinforcement learning for quality monitoring of the analysis. This request mechanism carrying an identifier ID enables fine-grained management of specific analysis tasks and is suitable for scenarios requiring in-depth analysis of specific users or services, such as VIP user service assurance and critical business performance monitoring.
[0045] In another embodiment, the method further includes: receiving a termination instruction from the 5GC network element indicating termination of analysis quality monitoring or reinforcement learning; and sending the termination instruction to the first entity, wherein the termination instruction contains a reinforcement learning association ID. The termination instruction can promptly stop unnecessary reinforcement learning processes, saving computational resources and is suitable for scenarios with large network load variations and the need for dynamic adjustment of analysis strategies.
[0046] In another embodiment, the method further includes: receiving a termination instruction for reinforcement learning or model quality monitoring sent by a first entity, wherein the termination instruction includes a reinforcement learning association ID. This mechanism enables bidirectional control of the reinforcement learning process, enhances the system's synergy and controllability, and is suitable for multi-entity collaborative analysis scenarios, such as cross-carrier network optimization.
[0047] In this embodiment of the disclosure, step S302 may further include: when the environment interpreter is set in the 5GC network element, performing model inference, evaluating the model inference results to obtain feedback, and generating rewards and / or states based on the feedback. This setup integrates reinforcement learning execution with data analysis into the same entity, improving the real-time performance and efficiency of the analysis. It is suitable for scenarios requiring rapid response to network changes, such as real-time network resource scheduling and dynamic traffic control.
[0048] In one embodiment of this disclosure, the 5GC network element makes decisions and initiates reinforcement learning-based analysis quality monitoring. This decision-making mechanism enables the 5GC network element to proactively adjust its analysis strategy according to network requirements, improving the network's adaptability and making it suitable for automated network operation and maintenance scenarios, such as network fault self-healing and automatic resource optimization.
[0049] In one embodiment of this disclosure, a termination instruction is sent to a first entity, wherein the termination instruction includes a reinforcement learning association ID. The reinforcement learning association ID enables precise control over a specific reinforcement learning process, making it suitable for scenarios requiring the management of multiple independent analysis tasks, such as multi-service parallel optimization and multi-user personalized service assurance.
[0050] In one embodiment of this disclosure, a subscription request for a machine learning model is sent to a first entity, wherein the subscription request includes an analysis ID; the first entity receives the machine learning model as feedback from the first entity, which is used to make decisions and initiate reinforcement learning-based model quality monitoring. This subscription mechanism enables on-demand allocation of models, improves resource utilization efficiency, and is suitable for scenarios with limited network resources and requiring fine-grained management of model usage, such as edge computing resource optimization and network slice management.
[0051] In another embodiment, the method further includes receiving a termination instruction for reinforcement learning or model quality monitoring sent by a first entity, wherein the termination instruction includes a reinforcement learning association ID. This mechanism enables collaborative control of the reinforcement learning process among multiple entities, enhancing the overall coordination and stability of the system. It is suitable for multi-operator network collaboration scenarios, such as cross-operator network optimization and joint resource scheduling.
[0052] In one embodiment of this disclosure, prior to step S302, the method further includes: sending a joining request for reinforcement learning analysis output quality monitoring to a first entity; and receiving a joining response for reinforcement learning analysis output quality monitoring returned by the first entity, wherein the joining response carries indication information indicating agreement to join. This joining request mechanism enables dynamic joining of the reinforcement learning process and is suitable for scenarios where network analysis needs change rapidly, such as handling sudden traffic events and defending against network attacks.
[0053] In this embodiment, the reward is a positive or negative number, and the state is a discrete value with more than one state. By setting positive and negative rewards and multiple states, fine-grained rewards or penalties for model behavior can be achieved, which is suitable for scenarios that require highly precise control of model behavior, such as fine-grained allocation of network resources and deep learning of user behavior.
[0054] In this embodiment of the disclosure, the first entity and the second entity are at least one of the following: MTLF, AnLF, and NWDAF. This entity configuration enables multi-layered optimization of network data analysis and is suitable for scenarios requiring cross-level and cross-functional network element collaborative analysis, such as overall network performance optimization and improvement of user service experience.
[0055] This disclosure also provides a reinforcement learning processing method. Figure 4 is a flowchart of a reinforcement learning processing method according to an embodiment of this disclosure. As shown in Figure 4, the method is applied to a first entity and includes the following steps:
[0056] Step S402: Receive the reward and / or state required for reinforcement learning sent by the environment interpreter, wherein the reward and / or state is generated by the environment interpreter based on the feedback of the model inference results of the machine learning model used for output analysis;
[0057] Step S404: Determine the action feedback based on the reward and / or status;
[0058] Step S406: Send the action feedback to the environment interpreter.
[0059] Through the above steps S402 to S406, the problem that the core network in the related technology does not support reinforcement learning-related technologies, resulting in the lack of a real-time analysis quality feedback mechanism and the inability to perform model enhancement and analysis quality monitoring can be solved. This enables the core network to support reinforcement learning-related technologies, and allows for a real-time analysis quality feedback mechanism, as well as model enhancement and analysis quality monitoring.
[0060] The embodiments disclosed herein enable the first entity to automatically adjust the model according to changes in the network environment, thereby improving the accuracy and efficiency of the analysis. This is applicable to scenarios such as network resource management and user experience optimization.
[0061] In this embodiment of the disclosure, the aforementioned action feedback is applied to the environment interpreter to ensure the quality of the analysis output or the model quality of the machine learning model. Through action feedback, dynamic adjustments to the model can be achieved, improving the adaptability and accuracy of the analysis. This is suitable for scenarios requiring real-time optimization of analysis results, such as rapid network fault location and user behavior prediction.
[0062] In an optional embodiment, step S404 may specifically include: locally calculating a reward value; and determining the action feedback of the machine learning model based on the reward value, reward, and / or state. This reward-based decision-making mechanism enables deeper optimization of model behavior and is suitable for scenarios requiring deep optimization of model behavior, such as deep optimization of network resources and deep learning of user behavior.
[0063] In this embodiment of the disclosure, step S404 may specifically include: inputting the reward, state, and / or return value into a pre-trained policy model to obtain action feedback output by the policy model; or generating action feedback based on the reward, state, and / or return value using locally configured policy information. The use of such a policy model or policy information enables intelligent decision-making regarding model behavior, improves the level of intelligence in analysis, and is suitable for scenarios requiring intelligent decision support, such as intelligent network operation and maintenance, and automated resource scheduling.
[0064] In one embodiment of this disclosure, if the network element lacks reinforcement learning-related policy capabilities, it requests a policy model and / or policy information from a target network element with AI Agent capabilities; or it receives a policy model and / or policy information provided by the target network element. This mechanism enables cross-entity sharing of policy capabilities, improves the overall intelligent decision-making capability of the system, and is suitable for scenarios with limited network resources that require shared intelligent decision-making capabilities, such as edge computing resource optimization and intelligent network slice management.
[0065] In one embodiment of this disclosure, the method further includes: sending a network element discovery request with AI Agent capability to the Network Repository Function (NRF); receiving network element information returned by the NRF; and determining the target network element based on the network element information. This network element discovery mechanism can automatically discover and select network elements with AI Agent capability, improving the efficiency of intelligent decision-making in the system. It is suitable for scenarios with large changes in network topology and the need for dynamic selection of intelligent decision-making network elements, such as network reconstruction and dynamic resource allocation.
[0066] In one embodiment of this disclosure, the network element discovery request with AI Agent capability includes reinforcement learning execution duration information and an analysis ID. Multiple network elements with AI Agent capability are registered in the NRF. By including the duration information and analysis ID in the request, precise control over the reinforcement learning process of a specific analysis task can be achieved. This is suitable for scenarios requiring deep optimization of specific tasks, such as ensuring critical business performance and optimizing VIP user services.
[0067] In this embodiment of the disclosure, step S406 may specifically include: sending action feedback to the environment interpreter through instruction information, wherein the instruction information of the action feedback includes at least one of the following: first instruction information indicating the use of a new machine learning model, wherein the first instruction information includes at least one of the following: a new model file, a link to the model file of the new machine learning model, and the ID of the new machine learning model; second instruction information indicating the updating of the machine learning model, wherein the second instruction information carries the model parameter information to be updated and the corresponding values; third instruction information indicating continued use of the machine learning model; and fourth instruction information indicating the acquisition of a new machine learning model from other network elements, wherein the fourth instruction information includes the identification information or address information of the other network elements. This refined action feedback information enables fine-grained control of the model, improves the accuracy and efficiency of the analysis, and is suitable for scenarios requiring in-depth model optimization, such as fine-grained management of network resources and in-depth analysis of user behavior.
[0068] In one embodiment of this disclosure, the method further includes: obtaining model information of a new machine learning model from a third entity based on action feedback, wherein the model information includes a model file or a link to a model file. This mechanism enables cross-entity model updates, improving the system's flexibility and adaptability. It is suitable for scenarios with widely distributed network resources that require cross-entity model updates, such as cross-regional network optimization and multi-operator network collaboration.
[0069] In one embodiment of this disclosure, the method further includes: receiving a termination instruction from an environment interpreter indicating termination of analysis quality monitoring or reinforcement learning; and sending the termination instruction to a network element with AI Agent capabilities, wherein the termination instruction includes a reinforcement learning association ID. This termination instruction mechanism enables precise control over the reinforcement learning process, improves system controllability and efficiency, and is suitable for scenarios requiring dynamic adjustment of analysis strategies, such as scenarios with large changes in network load requiring rapid adjustment of analysis strategies.
[0070] In one embodiment of this disclosure, the method further includes: initiating a termination command to the second entity and the network element with AI Agent capabilities to terminate reinforcement learning or model quality monitoring, wherein the termination command includes a reinforcement learning association ID. This mechanism enables collaborative control of the reinforcement learning process among multiple entities, enhances the overall coordination and stability of the system, and is suitable for scenarios requiring cross-entity collaborative analysis, such as multi-operator network collaboration and multi-network element joint optimization.
[0071] In one embodiment of this disclosure, the method further includes: receiving a subscription request for a machine learning model from a 5GC network element, wherein the subscription request carries information about whether the 5GC network element has an environment interpreter or supports reinforcement learning capabilities; if the 5GC network element has an environment interpreter or supports reinforcement learning capabilities, sending a request for reward and / or status acquisition to the 5GC network element during model feedback. This subscription request mechanism enables intelligent identification and resource allocation of 5GC network elements with reinforcement learning capabilities, improving the efficiency of intelligent decision-making in the system. It is suitable for scenarios requiring intelligent identification and resource allocation, such as intelligent management of network resources and intelligent optimization of user services.
[0072] In one embodiment of this disclosure, after receiving a subscription request for a machine learning model from a 5GC network element, the method further includes initiating reinforcement learning-based model quality monitoring. This subscription-based model quality monitoring mechanism enables continuous optimization of specific models, improving the accuracy and efficiency of analysis. It is suitable for scenarios requiring continuous model quality optimization, such as continuous optimization of network performance and continuous improvement of user experience.
[0073] In this embodiment of the disclosure, the first entity and the second entity are at least one of the following: MTLF, AnLF, and NWDAF. This entity configuration enables multi-layered optimization of network data analysis and is suitable for scenarios requiring cross-level and cross-functional network element collaborative analysis, such as overall network performance optimization and user service experience improvement, further enhancing the network's intelligent decision-making capabilities and resource utilization efficiency.
[0074] In this embodiment, the first entity and the third entity are at least one of the following: ADRF, MTLF, AnLF, and NWDAF. This entity setup enables cross-entity collaboration in network data analysis and model updates, improving the overall intelligent decision-making capability and resource utilization efficiency of the system. It is suitable for scenarios requiring cross-entity collaborative analysis and model updates, such as cross-operator network optimization and multi-network element joint resource scheduling, further enhancing the network's intelligent decision-making capability and resource utilization efficiency, and providing technical support for building a smarter and more efficient 5G network.
[0075] This disclosed embodiment not only automatically adjusts and optimizes the machine learning model used for output analysis, significantly improving the quality of the analysis output and the adaptability of the model, but also effectively manages and terminates the reinforcement learning process by introducing reinforcement learning association IDs, enhancing the flexibility and controllability of the method. This method reduces human intervention while enhancing the system's intelligent decision-making capabilities, especially in dynamically changing network environments, enabling rapid response to changes in network states and improving network service quality and user experience. Furthermore, by refining action feedback information, more precise control of the model can be achieved, making it suitable for scenarios requiring highly accurate analysis, such as network resource scheduling and traffic prediction, further improving network resource utilization efficiency and analysis accuracy. Through positive and negative rewards and multi-state settings, refined rewards or penalties for model behavior can be achieved, suitable for scenarios requiring highly precise control of model behavior, such as fine-grained allocation of network resources and deep learning of user behavior, further enhancing the network's intelligent decision-making capabilities and resource utilization efficiency. By using strategy models or strategy information, intelligent decision-making on model behavior can be achieved, improving the level of intelligence in analysis. This is applicable to scenarios that require intelligent decision support, such as intelligent network operation and maintenance, and automated resource scheduling, providing technical support for building a smarter and more efficient 5G network.
[0076] The following example illustrates the embodiments of this disclosure, using MTLF as the first entity, AnLFE as the second entity, and ADRF as the third entity.
[0077] Figure 5 is a flowchart of reinforcement learning for per-analytics quality monitoring initiated by AnLF according to an embodiment of the present disclosure, as shown in Figure 5, including:
[0078] 1. When a consumer network element subscribes to analytics, it requests AnLF to perform quality monitoring on the accuracy of the analytics (i.e., monitoring the model inference output). This request must include one or more analytics IDs, and the consumer network element can also specify in this message whether AnLF recommends that AnLF use reinforcement learning methods to perform analytics quality monitoring for each analytics ID.
[0079] 2. After receiving a request from a consumer network element, the AnLF decides whether to use reinforcement learning to monitor the quality of the analytics provided to the consumer network element. If this AnLF decides that reinforcement learning is needed for analytics quality monitoring, or if it receives a suggestion from the consumer network element but cannot act as an environment interpreter, it can select another AnLF as the environment interpreter through NRF.
[0080] 3. AnLF sends a reinforcement learning analytics output quality monitoring join request to the selected MTLF. This request includes information such as the duration of the reinforcement learning execution (e.g., 3 hours), the reinforcement learning association ID, the analytics ID, the model ID currently performing the inference task, and the maximum response time (the time from receiving a reward and / or status to providing action feedback). The selected MTLF can be the MTLF that trained the inference model.
[0081] 4. MTLF returns whether to add reinforcement learning for analysis quality monitoring. If MTLF refuses, AnLF selects another MTLF and returns to step 3.
[0082] 5. AnLF uses model inference to generate analysis, sends the analysis to consumers, and requests feedback from consumers on the quality of the analysis.
[0083] 6. Consumer network elements apply the analysis to core network elements, collect relevant data from core network elements, generate analysis quality feedback, and send it to AnLF. The generated analysis quality feedback may carry time information and information about the applied network elements (e.g., network element type, network element identifier, etc.).
[0084] 7. AnLF converts the received feedback information into rewards and / or states. Rewards can be positive or negative (e.g., +2, -3), and states are discrete values (S0, S1, ...) with more than one value. The conversion of feedback information into rewards and / or states can be assisted by policies provided by the AI Agent entity.
[0085] 8. AnLF sends rewards and / or status to MTLF. This message may also include the reinforcement learning association ID and analytics ID.
[0086] 9a. MTLF determines actions based on received rewards, states, and policies. To determine actions, MTLF may locally compute reward values. MTLF can use a pre-trained policy model as input to rewards, states, and / or reward values, and the model outputs an action. MTLF can also determine actions using locally configured policy information and reward, state, or reward values. Determined actions may include: continuing output analysis using the existing model, retraining the existing model and deploying it to AnLF output analysis, or acquiring a new model from ADRF or other MTLFs and deploying it to AnLF output analysis.
[0087] 9b. If the MTLF does not possess AI Agent capabilities, it requests a policy model and / or policy from a network element that does. This message may also include the reinforcement learning association ID, analytics ID, reward, payout, or state. The MTLF can discover network elements with AI Agent capabilities through NRF.
[0088] 9c, The AI Agent network element sends policies and / or policy models to the MTLF. The AI Agent network element can also proactively update policies to the MTLF without the MTLF requesting policies.
[0089] Among them, 9b / 9c may occur before 9a.
[0090] 10. Based on the action determined by the MTLF through its policy, the MTLF may request a new model from the ADRF. This message may contain the model ID.
[0091] 11. ADRF feeds back model information (model file or link containing model file) to MTLF.
[0092] 12. MTLF sends action feedback to AnLF, which may include instructions to use the new model and new model inference input information for analysis generation, new model information (such as new model files or links containing new model files, new model IDs), model update information (such as model parameter information to be updated and their corresponding values, models to be merged and updated, etc.), model update instructions, instructions to continue using the original model for inference for analysis, or instructions to collect inference data from other network elements (data sources) and new data source address information, etc.
[0093] If MTLF obtains a new model from ADRF, it sends the new model ID in step 12; otherwise, the action that MTLF sends back to AnLF does not include the model ID.
[0094] 13. AnLF applies actions to perform model inference output and repeats steps 5-13.
[0095] 14. The consumer network element initiates a termination command for quality monitoring based on the accuracy of the analysis, or terminates the quality monitoring based on the accuracy of the analysis using reinforcement learning.
[0096] 15. AnLF sends a reinforcement learning termination instruction to MTLF, which includes the reinforcement learning association ID.
[0097] 16. MTLF sends a reinforcement learning termination instruction to the AI Agent subject, which includes the reinforcement learning association ID.
[0098] Figure 6 is a flowchart of reinforcement learning for per-analytics quality monitoring initiated by a consumer network element according to an embodiment of the present disclosure, as shown in Figure 6, including:
[0099] 1. Consumer network element decision-making requires the use of reinforcement learning processes to monitor the quality of analysis.
[0100] 2. The consumer network element sends a reinforcement learning analysis output quality monitoring join request to the selected MTLF. This request includes information such as the duration of the reinforcement learning execution (e.g., 3 hours), the reinforcement learning association ID, the analysis ID, the model ID currently performing the inference task, and the maximum response time (the time from receiving the reward and / or status to providing action feedback). The selected MTLF can be the MTLF from which the inference model has been trained.
[0101] 3. The MTLF returns whether to add reinforcement learning for analysis quality monitoring. If the MTLF refuses, the AnLF selects another MTLF and returns to step 2.
[0102] 4. The consumer network element performs model inference and interacts with other 5GC network elements to evaluate the inference results and generate rewards and / or states. Rewards can be positive or negative (e.g., +2, -3), and states are discrete values with more than one possible value (e.g., S0, S1, ...). This message may also carry time information and application network element information (such as network element type, network element identifier, etc.). The evaluation of the inference results and generation of rewards and / or states requests the AI Agent entity to provide policy support to assist in the transition.
[0103] 5. The environment interpreter sends the calculated reward and / or state to the MTLF. This message may also include the reinforcement learning association ID and the analysis ID.
[0104] 6a. MTLF determines actions based on received rewards, states, and policies. To determine actions, MTLF may compute reward values locally. MTLF can use a pre-trained policy model as input, along with rewards, states, and / or reward values, and the policy model outputs actions. MTLF can also determine actions based on locally configured policy information, considering rewards, states, and / or reward values. Determined actions may include: continuing output analysis using the existing model, retraining the existing model and deploying it to AnLF output analysis, or obtaining a new model from ADRF and deploying it to AnLF output analysis.
[0105] 6b. If the MTLF does not possess AI Agent capabilities, it requests a policy model and / or policy from a network element that does. This message may also include the reinforcement learning association ID, analytics ID, reward, state, and / or payoff value. The MTLF can discover network elements with AI Agent capabilities through NRF.
[0106] 6c, The AI Agent network element sends policies and / or policy models to the MTLF. The AI Agent network element can also proactively update policies to the MTLF without the MTLF requesting policies.
[0107] 7. Based on the action determined by the MTLF through its policy, the MTLF may request a new model from the ADRF. This message may contain the model ID.
[0108] 8. ADRF feeds back model information (model file or link containing model file) to MTLF.
[0109] 9. MTLF sends action feedback to AnLF, which may include instructions to use the new model and new model inference input information for analysis generation, new model information (such as new model files or links containing new model files, new model IDs), model update information (such as model parameter information to be updated and corresponding values, models to be merged and updated, etc.), model update instructions, instructions to continue using the original model for inference for analysis, or instructions to collect inference data from other network elements (data sources) and new data source address information, etc.
[0110] If MTLF obtains a new model from ADRF, it sends the new model ID in step 12; otherwise, the action that MTLF sends back to AnLF does not include the model ID.
[0111] 10. Consumer network element application actions to perform model inference output, and repeat steps 4-10.
[0112] 11. The consumer network element sends an analysis quality monitoring termination or reinforcement learning termination instruction to the MTLF, which includes the reinforcement learning association ID.
[0113] 12. MTLF sends a reinforcement learning termination instruction to the AI Agent, which includes the reinforcement learning association ID.
[0114] Figure 7 is a flowchart of the reinforcement learning architecture construction according to an embodiment of the present disclosure, as shown in Figure 7, including:
[0115] 1. AnLF, MTLF or any 5GC network element registers or updates its reinforcement learning-related capabilities with NRF, including whether it has the ability to act as an environment interpreter and / or AI agent.
[0116] 2. NRF returns a registration result indication.
[0117] 3. AnLF, MTLF or any 5GC network element initiates an AI Agent discovery request, which includes information such as the duration of reinforcement learning execution (e.g., 3 hours) and analysis ID.
[0118] 4. NRF returns one or more AI Agent network element information that meets the requirements.
[0119] 5. The 5GC network element selects one of the network elements with AI Agent capabilities returned by the NRF to send a reinforcement learning joining request or a policy information request. The content of this joining request includes steps 2 and 6b in Figure 6 and steps 3 and 9b in Figure 5.
[0120] 6. Network elements with AI Agent capabilities return a response message or policy information regarding whether to join reinforcement learning.
[0121] 7. If the 5GC network element acting as the environment interpreter receives a request to refuse to join, it selects another 5GC network element with AI Agent capability from the list returned by NRF in step 4, and repeats steps 5 / 6 until it receives a request to 'agree to join reinforcement learning'.
[0122] Figure 8 is a flowchart of reinforcement learning for per-ML Model quality monitoring initiated by MTLF according to an embodiment of the present disclosure, as shown in Figure 8, including:
[0123] 1. After receiving an analysis request from a consumer network element, AnLF subscribes to a model from MTLF. The subscription request includes the analysis ID and whether AnLF supports reinforcement learning-related capabilities (e.g., as an environment interpreter).
[0124] 2. MTLF feedback subscription model initiates a reinforcement learning-based model quality monitoring process, which includes the model ID of the monitored model, the reinforcement learning association ID, and requests AnLF to provide a status reward. This request may also include: the duration of reinforcement learning execution (e.g., 3 hours) and the maximum response time (the time from receiving the reward / status provision request to providing the reward / status).
[0125] 3. AnLF uses the acquired model to perform model inference and generate analysis (prediction).
[0126] 4. AnLF provides analysis results (predictions) to 5GC network elements and requests feedback on analysis quality.
[0127] 5. The 5GC network element provides analysis quality feedback.
[0128] 6. AnLF converts the received feedback information into rewards and / or states. Rewards can be positive or negative (e.g., +2, -3), and states are discrete values with more than one possible value (e.g., S0, S1, ...). The conversion of feedback information into rewards and / or states can be assisted by policies provided by the AI Agent entity.
[0129] 7. AnLF returns the reward and / or status to MTLF. This message may also include the reinforcement learning association ID and the analytics ID.
[0130] 8a. If the MTLF lacks AI Agent capabilities (e.g., does not support policy generation) or needs to request a policy model, the MTLF sends a policy request to a network element that possesses an AI Agent. This message may also include the reinforcement learning association ID, analytics ID, reward, and / or status. It may also include the model ID for quality monitoring, and a reinforcement learning-based model quality monitoring indication.
[0131] 8b. Network elements with AI Agent capabilities return policies and / or policy models to the MTLF. AI Agent network elements can also proactively update policies to the MTLF without the MTLF requesting policies.
[0132] 9. MTLF determines actions based on the received reward, state, and policy. To determine actions, MTLF may locally calculate the reward value. MTLF can use a pre-trained policy model or the policy model received in 8b, with the reward and / or state as input, and the model outputs an action. MTLF can also determine actions based on locally configured policy information. The determined actions may include: continuing output analysis using the existing model (if the existing model has no quality issues), deactivating the existing model and retraining a new model to deploy to AnLF output analysis, or updating the existing model and continuing output analysis.
[0133] 10. MTLF sends the action decided in step 9 to AnLF, which may include instructions to continue outputting analysis using the original model, instructions to deactivate the original model, the retrained model file (e.g., the model file or a link containing the model file), and model update information (e.g., the model parameter information to be updated and its corresponding values, the model to be merged and updated, etc.).
[0134] The model ID remains unchanged after retraining or updating the model.
[0135] 11. AnLF performs relevant processing according to the instructions in step 10 and continues to perform model inference and prediction.
[0136] Steps 3-11 are repeated until the duration of federated learning exceeds the prescribed limit or the MTLF initiates a reinforcement learning termination process.
[0137] 12. MTLF sends a termination instruction for reinforcement learning or model quality monitoring to AnLF and AI Agent Entity, which includes the reinforcement learning association ID.
[0138] Figure 9 is a flowchart of reinforcement learning for per-ML Model quality monitoring initiated by MTLF according to an embodiment of the present disclosure, as shown in Figure 9, including:
[0139] 1. Any 5GC network element subscribes to the MTLF model. The subscription request includes the analysis ID and whether the 5GC network element supports reinforcement learning-related capabilities (e.g., as an environment interpreter).
[0140] 2. MTLF feedback subscribes to the model and initiates a reinforcement learning-based model quality monitoring process, which includes the ML Model ID, the reinforcement learning association ID, and requests AnLF to provide a state reward. This request may also include: the duration of the reinforcement learning execution (e.g., 3 hours) and the maximum response time (the time from receiving the reward / state provision request to providing the reward / state).
[0141] 3. 5GC network elements perform model inference to generate predictions; and interact with other 5GC network elements to evaluate the prediction results, generating rewards and / or states. Rewards can be positive or negative (e.g., +2, -3), and states are discrete values with more than one (e.g., S0, S1, ...). The evaluation of the inference results and generation of rewards and / or states requests the AI Agent entity to provide policies to assist in the transition.
[0142] 4. The 5GC network element sends the calculated reward and / or status to the MTLF. This message may also include the reinforcement learning association ID and the analysis ID.
[0143] 5. If the MTLF does not have AI Agent capabilities, it requests a policy model and / or policy from a network element that does have AI Agent capabilities. This message may also include the reinforcement learning association ID, analysis ID, reward and / or state, quality monitoring model ID, and reinforcement learning-based model quality monitoring indication.
[0144] 6. The AI Agent network element sends policies and / or policy models to the MTLF. The AI Agent network element can also proactively update policies to the MTLF without the MTLF requesting policies.
[0145] 7. MTLF determines actions based on the received reward, state, and policy. To determine actions, MTLF may locally compute the reward value. MTLF can use a pre-trained policy model as input, with the reward and / or state as input, and the model outputs an action. MTLF can also determine actions based on locally configured policy information. The determined actions may include: continuing output analysis using the existing model, retraining the existing model and deploying it to AnLF output analysis, or obtaining a new model from ADRF and deploying it to AnLF output analysis.
[0146] 8. MTLF sends the action decided in step 7 to AnLF, which may include instructions to continue outputting analysis using the original model, instructions to stop using the original model, the retrained model file (e.g., the model file or a link containing the model file), and model update information (e.g., the model parameter information to be updated and the corresponding values, the model to be merged and updated, etc.).
[0147] 9. The consumer network element performs relevant processing according to the instructions received in step 8 and performs model inference.
[0148] Steps 3-9 are repeated until the duration of federated learning exceeds the prescribed limit or the MTLF initiates a reinforcement learning termination process.
[0149] 10. MTLF sends a termination instruction for reinforcement learning or model quality monitoring to AnLF and AI Agent Entity, which includes the reinforcement learning association ID.
[0150] This disclosure also provides a reinforcement learning processing device. FIG10 is a block diagram of a reinforcement learning processing device according to an embodiment of this disclosure. As shown in FIG10, the device is applied to an environment interpreter and includes:
[0151] The generation module 102 is configured to obtain result feedback based on model inference of the machine learning model used for output analysis, and generate the reward and / or state required for reinforcement learning based on the result feedback.
[0152] The acquisition module 104 is configured to acquire action feedback from the MTLF based on the reward and / or the state.
[0153] This disclosure also provides a reinforcement learning processing apparatus. FIG11 is a block diagram of a reinforcement learning processing apparatus according to an embodiment of this disclosure. As shown in FIG11, the apparatus is applied to a first entity and includes:
[0154] The receiving module 112 is configured to receive the reward and / or state required for reinforcement learning sent by the environment interpreter, wherein the reward and / or the state are generated by the environment interpreter based on the model inference results of the machine learning model.
[0155] Decision module 114 is configured to determine action feedback based on the reward and / or the state;
[0156] The sending module 116 is configured to send the action feedback to the environment interpreter.
[0157] This disclosure also provides a computer program product, including computer program instructions, wherein the computer program instructions cause a computer to implement the steps in any of the above method embodiments.
[0158] Embodiments of this disclosure also provide a computer-readable storage medium storing a computer program configured to perform the steps in any of the above method embodiments when executed.
[0159] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0160] Embodiments of this disclosure also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0161] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0162] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0163] It is obvious to those skilled in the art that the modules or steps of this disclosure described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this disclosure is not limited to any particular combination of hardware and software.
[0164] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A reinforcement learning processing method applied to an environment interpreter, the method comprising: The results feedback is obtained based on the model inference of the machine learning model used for output analysis, and the rewards and / or states required for reinforcement learning are generated based on the results feedback. Action feedback is obtained from the first entity based on the reward and / or the state.
2. The method according to claim 1, wherein, The action feedback is applied to ensure the quality of the analysis output or the model quality of the machine learning model.
3. The method according to claim 2, wherein, Applying the action feedback to ensure the quality of the analysis output or the model quality of the machine learning model includes: Based on the instruction information from the action feedback, the machine learning model is processed by at least one of the following: using a new machine learning model for model inference, updating the machine learning model and performing model inference, continuing to use the machine learning model for inference, or obtaining a new machine learning model from other network elements and performing model inference.
4. The method according to claim 3, wherein, The action feedback indication information includes at least one of the following: The first instruction information indicating the use of a new machine learning model includes at least one of the following: a new model file, a link to the model file of the new machine learning model, and the ID of the new machine learning model; The second instruction information indicates that the machine learning model should be updated, wherein the second instruction information carries the model parameter information to be updated and the corresponding values; The third instruction information indicates that the machine learning model should continue to be used; A fourth instruction information is provided to indicate the acquisition of a new machine learning model from other network elements, wherein the fourth instruction information includes the identification information or address information of the other network elements.
5. The method according to claim 1, wherein, Obtaining action feedback from the first entity based on the reward and / or the state includes: Send the reward and / or the status to the first entity; Receive action feedback from the machine learning model determined by the first entity based on the reward and / or the state.
6. The method according to claim 1, wherein, Obtaining result feedback based on model inference from the machine learning model used for output analysis, and generating rewards and / or states required for reinforcement learning based on said result feedback, includes: When the environment interpreter is set in the second entity, the machine learning model is used to perform model inference, and the model inference result is sent to the 5GC network element; Receive the result feedback of the model inference result generated by the 5GC network element; The feedback result is converted into the reward and / or the status.
7. The method according to claim 6, wherein, The method further includes: Receive the analysis quality monitoring request sent by the 5GC network element during subscription analysis; Make decisions and initiate reinforcement learning-based analytics quality monitoring.
8. The method according to claim 7, wherein, The quality monitoring request carries the identifier ID of the analysis and / or a suggestion to use reinforcement learning to monitor the quality of the analysis.
9. The method according to claim 6, wherein, The method further includes: Receive termination commands from 5GC network elements, such as termination of analysis quality monitoring or termination of reinforcement learning. Send the termination instruction to the first entity, wherein the termination instruction contains the reinforcement learning association ID.
10. The method according to claim 6, wherein, The method further includes: After receiving the analysis request from the 5GC network element, a subscription request for the machine learning model is sent to the first entity, wherein the subscription request contains the analysis ID; The machine learning model is received from the feedback of the first entity, wherein the first entity is used to make decisions and initiate model quality monitoring based on reinforcement learning.
11. The method according to claim 6, wherein, The method further includes: Receive a termination instruction for reinforcement learning or model quality monitoring sent by the first entity, wherein the termination instruction includes a reinforcement learning association ID.
12. The method according to claim 1, wherein, Obtaining result feedback based on model inference from the machine learning model used for output analysis, and generating rewards and / or states required for reinforcement learning based on said result feedback, includes: When the environment interpreter is set to a 5GC network element, model inference is performed, and the model inference results are evaluated to obtain the result feedback. The reward and / or the state are generated based on the result feedback.
13. The method according to claim 12, wherein, The method further includes: The 5GC network element makes decisions and initiates reinforcement learning-based analysis quality monitoring.
14. The method according to claim 12, wherein, The method further includes: Send a termination instruction to the first entity, wherein the termination instruction includes the reinforcement learning association ID.
15. The method according to claim 12, wherein, The method further includes: Send a subscription request to the first entity to subscribe to the machine learning model, wherein the subscription request carries whether the 5GC network element has an environment interpreter or supports reinforcement learning capabilities; When the first entity receives a request from the model to obtain the reward and / or state, provided that the 5GC network element has the environment interpreter or supports reinforcement learning capabilities, the request is made during model feedback.
16. The method according to claim 12, wherein, The method further includes: Receive a termination instruction initiated by the first entity for reinforcement learning or model quality monitoring, wherein the termination instruction includes a reinforcement learning association ID.
17. The method according to any one of claims 1 to 16, wherein, Before obtaining result feedback based on model inference from a machine learning model used for output analysis, and generating the reward and / or state required for reinforcement learning based on the result feedback, the method further includes: Send a request to the first entity to join the reinforcement learning analysis output quality monitoring; The system receives a joining response from the first entity regarding the quality monitoring of the reinforcement learning analysis output, wherein the joining response carries an indication of consent to join.
18. The method according to any one of claims 1 to 16, wherein, The reward is a positive or negative number, and the state is a discrete value greater than one.
19. The method according to claim 6, wherein, The first entity and the second entity are at least one of the following: Model Training Logic Function (MTLF), Analysis Logic Function (AnLF), and Network Data Analysis Function (NWDAF).
20. A reinforcement learning processing method applied to a first entity, the method comprising: Receive the reward and / or state required for reinforcement learning sent by the environment interpreter, wherein the reward and / or the state are generated by the environment interpreter based on the feedback of the model inference results of the machine learning model used for output analysis; The action feedback is determined based on the reward and / or the state; The action feedback is sent to the environment interpreter.
21. The method according to claim 20, wherein, The action feedback is applied to the environment interpreter to ensure the quality of the analysis output or the model quality of the machine learning model.
22. The method according to claim 20, wherein, The action feedback determined based on the reward and / or the state includes: Calculate the return value locally; The action feedback of the machine learning model is determined based on the reward value, the reward, and / or the state.
23. The method according to claim 22, wherein, The action feedback determined based on the reward and / or the state includes: The reward, the state, and / or the return value are input into a pre-trained policy model to obtain the action feedback output by the policy model; or The action feedback is generated based on the locally configured policy information, the reward, the status, and / or the reward value.
24. The method according to claim 23, wherein, The method further includes: If the network element lacks reinforcement learning-related policy capabilities, it requests the policy model and / or the policy information from a target network element with AIAgent capabilities; or Receive the policy model and / or policy information provided by the target network element.
25. The method of claim 24, wherein, The method further includes: Send a network element discovery request with the AI Agent capability to the Network Repository Function (NRF); Receive network element information returned by the NRF; The target network element is determined based on the network element information.
26. The method according to claim 25, wherein, The network element discovery request with AI Agent capability includes information on the duration of reinforcement learning execution and an analysis ID. The NRF has multiple network elements with AI Agent capability registered.
27. The method of claim 20, wherein, Sending the action feedback to the environment interpreter includes: The action feedback is sent to the environment interpreter via indication information, wherein the indication information of the action feedback includes at least one of the following: The first instruction information indicating the use of a new machine learning model includes at least one of the following: a new model file, a link to the model file of the new machine learning model, and the ID of the new machine learning model; The second instruction information indicates that the machine learning model should be updated, wherein the second instruction information carries the model parameter information to be updated and the corresponding values; The third instruction information indicates that the machine learning model should continue to be used; A fourth instruction information is provided to indicate the acquisition of a new machine learning model from other network elements, wherein the fourth instruction information includes the identification information or address information of the other network elements.
28. The method according to claim 27, wherein, The method further includes: Based on the action feedback, the model information of the new machine learning model is obtained from the third entity, wherein the model information includes a model file or a link to a model file.
29. The method according to claim 20, wherein, The method further includes: Receive a termination command sent by the environment interpreter to terminate the analysis quality monitoring or the reinforcement learning. The termination instruction is sent to a network element with AI Agent capabilities, wherein the termination instruction includes a reinforcement learning association ID.
30. The method of claim 20, wherein, The method further includes: A termination instruction is sent to the second entity and the network element with AI Agent capability to terminate reinforcement learning or model quality monitoring, wherein the termination instruction contains the reinforcement learning association ID.
31. The method according to claim 20, wherein, The method further includes: Receive a subscription request for a machine learning model sent by a 5GC network element, wherein the subscription request carries whether the 5GC network element has an environment interpreter or supports reinforcement learning capabilities; If the 5GC network element has the environment interpreter or supports reinforcement learning capabilities, a request to obtain the reward and / or state is sent to the 5GC network element during model feedback.
32. The method according to claim 31, wherein, After receiving a subscription request for a machine learning model from a 5GC network element, the method further includes: Initiate model quality monitoring based on reinforcement learning.
33. The method according to claim 30, wherein, The first entity and the second entity are at least one of the following: Model Training Logic Function (MTLF), Analysis Logic Function (AnLF), and Network Data Analysis Function (NWDAF).
34. The method according to claim 28, wherein, The first entity and the third entity are at least one of the following: Analysis Decision Function (ADRF), Model Training Logic Function (MTLF), Analysis Logic Function (AnLF), and Network Data Analysis Function (NWDAF).
35. A computer-readable storage medium storing a computer program, wherein, The computer program is configured to execute the method described in any one of claims 1 to 19, 20 to 34 when it is run.
36. An electronic device comprising a memory and a processor, the memory storing a computer program, the processor being configured to run the computer program to perform the method of any one of claims 1 to 19, 20 to 34.
37. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 19, 20 to 34.