Communication method and communication apparatus

By providing first information in the communication system, the problem that traditional machine learning models cannot be applied to reinforcement learning training is solved, thereby achieving accuracy and device adaptability in reinforcement learning model training and reducing modification and transmission overhead.

WO2026031806A1PCT designated stage Publication Date: 2026-02-12HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102975
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-06-24
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Traditional machine learning model training methods are not applicable to training reinforcement learning-based communication system models, as they lack effective environmental representation and behavioral impact analysis.

Method used

By providing initial information, including a list of initial metrics and the functionality of the initial model, we can help determine the training environment and behavior of reinforcement learning models, reduce device processing complexity, and improve the feasibility and adaptability of model training.

Benefits of technology

It enables accurate determination of environment and behavior during the training process of reinforcement learning-based models, improves the feasibility and device adaptability of model training, and reduces the degree of modification and transmission overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102975_12022026_PF_FP_ABST
    Figure CN2025102975_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of communications, and provides a communication method and a communication apparatus, capable of providing support for application of reinforcement learning in communication systems. The method comprises: determining first information, the first information being used for training a first model, and a training method for the first model being reinforcement learning; and sending the first information to a first device, the first device being a device for training the first model. In embodiments of the present application, a model training user sends information related to reinforcement learning-based model training to a model training provider, thereby helping to provide support for model training based on reinforcement learning methods. For example, the first information may include information related to an environment for model interaction during a reinforcement learning process, thereby helping to determine an environment for reinforcement learning-based model training. For another example, the first information may include information related to model behavior during the reinforcement learning process, thereby helping to determine model behavior during a training process.
Need to check novelty before this filing date? Find Prior Art

Description

Communication method and communication apparatus

[0001] The present application claims priority from the Chinese Patent Application No. 202411103686.7 filed on August 9, 2024, and entitled "Communication method and communication apparatus", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the field of communication, and more particularly, to a communication method and a communication apparatus. BACKGROUND

[0003] Reinforcement learning (RL) is an important branch of machine learning. Due to its numerous advantages, reinforcement learning has been widely applied in games, automatic control, robot control, finance, and medical treatment, etc. Then, how to apply reinforcement learning to a communication system is a problem to be solved. SUMMARY

[0004] The present application provides a communication method and a communication apparatus, which can support the application of reinforcement learning in a communication system, such as the training of a reinforcement learning model.

[0005] In a first aspect, an embodiment of the present application provides a communication method, which can be applied to a second device. The second device can be an entity physical device, a virtualized function module, a software function, a circuit, a chip (such as a system on chip (SoC) chip or a system in package (SIP) chip), etc.

[0006] Exemplarily, the second device is a user of model training.

[0007] Exemplarily, the second device can be a cross-domain management function unit, which can also be referred to as a network management system (NMS).

[0008] The method comprises: determining first information, wherein the first information is used for training a first model, and a training method of the first model is reinforcement learning; and sending the first information to a first device, wherein the first device is a device for training the first model.

[0009] In the embodiments of the present application, the user of the model training sends information related to the reinforcement learning-based model training to the provider of the model training, which helps to support the model training by the reinforcement learning method. For example, the first information can include information related to the environment in which the model interacts in the reinforcement learning process, thereby helping to determine the environment of the reinforcement learning model training. For another example, the first information can include information related to the behavior of the model in the reinforcement learning process, thereby helping to determine the behavior of the model in the training process.

[0010] In some embodiments, in the process of the model training, the first model interacts with a first environment, and the first information includes one or more of the following: a first index list, a state of the first environment is determined by one or more indexes in the first index list; a first parameter list, a behavior of the first model is determined by one or more parameters in the first parameter list; a first training environment, a training environment of the first model is the first training environment, or a training environment of the first model is one of the first training environment; a function of the first model; state information of the first environment; behavior information of the first model; or a performance indicator of the first model.

[0011] In traditional machine learning, a one-to-one mapping relationship can be obtained by model training through single-class training data. Taking supervised learning as an example, single-class training data can refer to single-feature training data, and a mapping relationship between the feature and the label can be obtained by model training.

[0012] Unlike traditional machine learning, reinforcement learning involves the interaction between an agent and an environment, and the change of behavior to the environment will affect the subsequent behavior. However, the training data used in traditional machine learning cannot represent the environment involved in reinforcement learning. That is, the training method of the traditional machine learning model cannot be applied to the reinforcement learning-based model training.

[0013] However, the embodiments of the present application can provide the provider of the model training with a basis for determining the above-mentioned environment through the first index list, the function of the first model, and other information, thereby helping to solve the problem that the traditional training data cannot be applied to the reinforcement learning-based model training.

[0014] In some embodiments, the first information includes the first index list, and the determining the first information includes: determining the first information based on model training requirement information, the model training requirement information including a second index list, indexes in the second index list being indexes expected to be used to determine the first environment; and the first index list being the second index list.

[0015] According to the method for determining the first index list and the model training based on the first index list, the training result of the first model can meet the requirements of the proposing party of the training requirement, that is, the user requirement.

[0016] In some embodiments, the capability information of the first device includes a third index list, the third index list includes indexes supported by the first device when determining the interaction environment of the model, and the first information is determined based on the requirement information of the model training and the capability information of the first device, including: determining the first index list based on the requirement information of the model training and the capability information of the first device, wherein: if the third index list includes the second index list, the first index list is the second index list; if the third index list includes part of the indexes in the second index list, the first index list includes indexes common to the second index list and the third index list.

[0017] The method for determining the first index list can avoid the case that the first device does not support the configured indexes, thereby improving the feasibility of the model training.

[0018] In some embodiments, the first information includes the function of the first model, and the function of the first model is determined based on the requirement information of the model training. This scheme is simple to implement.

[0019] In some embodiments, the first information includes the first parameter list, and the capability information of the first device includes a corresponding relationship between indexes in a third index list and parameters in a second parameter list, wherein the third index list includes indexes supported by the first device when determining the interaction environment of the model, and the second parameter list includes parameters supported by the first device when determining the behavior of the model, and the first information is determined based on the first index list and the corresponding relationship, wherein the first parameter list includes parameters having the corresponding relationship with the indexes in the first index list.

[0020] The corresponding relationship between the indexes and the parameters can be understood as that when the first model is trained in the environment determined by the indexes, the behavior determined by the parameters can be supported, or in other words, the change of the parameters will affect the state of the indexes. This is because when the first model is trained in different environments determined by indexes, which parameters can support the behavior determined by the parameters may be the same or different. Therefore, determining the first parameter list based on the above corresponding relationship can improve the rationality of the determined behavior.

[0021] In addition, different devices or different device manufacturers can support different above-mentioned corresponding relationships, therefore, determining the first parameter list based on the above-mentioned corresponding relationship can configure the supported first parameter list for different devices, which helps to improve the adaptability of the configuration information of model training and the training device.

[0022] In some embodiments, the method further comprises: obtaining the capability information of the first device, the capability information of the first device being related to reinforcement learning, wherein the capability information of the first device comprises one or more of the following: a first function, the first device supporting training of a model implementing the first function; a second environment, an environment supported by the first device, the second environment being used for training of a model; a third index list, the third index list comprising indexes supported by the first device for use in determining an interaction environment of a model; a second parameter list, the second parameter list comprising parameters supported by the first device for use in determining a behavior of a model; or a corresponding relationship between an index in the third index list and a parameter in the second parameter list.

[0023] Exemplarily, obtaining the capability information of the first device can also be replaced by: receiving the capability information of the first device, or subscribing to the capability information of the first device.

[0024] In the embodiments of the present application, by reporting the capability information of the provider of model training (such as the capability information related to reinforcement learning model training), the user of model training can determine the configuration information of model training that matches the capability information of the provider of model training, thereby helping to improve the feasibility of reinforcement learning model training.

[0025] In some embodiments, the method further comprises: receiving the capability information of the first device, the capability information of the first device being related to reinforcement learning, wherein the capability information of the first device comprises one or more of the following: a first function, the first device supporting training of a model implementing the first function; a second environment, an environment supported by the first device, the second environment being used for training of a model; a third index list, the third index list comprising indexes supported by the first device for use in determining an interaction environment of a model; a second parameter list, the second parameter list comprising parameters supported by the first device for use in determining a behavior of a model; or a corresponding relationship between an index in the third index list and a parameter in the second parameter list.

[0026] By carrying the first information in the model training request, the degree of modification of the related technology can be reduced. In addition, this scheme is relatively easy to implement.

[0027] In some embodiments, the method further comprises: receiving a model training report, the model training report comprising a second attribute and / or a third attribute, the second attribute being used to indicate a training environment of the first model, and the third attribute being used to indicate a performance index of the first model.

[0028] By carrying the training environment of the model and / or the performance indicator of the first model in the model training report, the second device can understand the model training situation, and can support the model management of the second device. For example, based on the performance indicator of the first model, a management strategy in the model inference process can be determined, such as determining an acceptable performance indicator in the model inference process.

[0029] In some embodiments, the method further includes: sending a first threshold to the first device, the first threshold being used to indicate an acceptable performance indicator of the first model in the model inference process.

[0030] The acceptable performance indicator may, for example, refer to a minimum value allowed by the performance indicator. Taking the performance indicator as a reward value as an example, the acceptable performance indicator may refer to a minimum reward value allowed in the model inference process based on the first model.

[0031] By setting the first threshold as described above, the reliability of the model inference can be improved.

[0032] Exemplarily, since the reward values in the last few rounds of the model training process are close to the reward values that can be obtained in the model inference process, the first threshold can be determined based on the reward values in the last few rounds of the model training process, which helps to improve the accuracy of the first threshold.

[0033] Exemplarily, the first threshold can be determined based on the performance indicator of the first model in the model training process.

[0034] In some embodiments, the method further includes: sending second information to the first device, the second information being used to indicate that the model inference based on the first model is performed; and receiving a model inference report, the model inference report including a performance indicator of the first model in the model inference process.

[0035] By reporting the performance indicator in the model inference process, the user of the model training can understand the inference state of the model, understand the accuracy of the model inference result, and support the management and monitoring of the model inference process.

[0036] In a second aspect, an embodiment of the present application provides a communication method, which can be applied to a first device. The first device can be an entity physical device, a virtualized function module, a software function, a circuit, a chip (such as a system on chip (SoC) chip or a system in package (SIP) chip), etc.

[0037] Exemplarily, the first device is a provider of the model training, or the model training is deployed in the first device.

[0038] Exemplarily, the first device is a provider of model training and a provider of model inference, or both model training and model inference are deployed in the first device.

[0039] Exemplarily, the first device can be a domain management function unit, which can also be referred to as an element management system (EMS).

[0040] The method comprises: receiving first information, the first information being used for training a first model, a training method of the first model being reinforcement learning; and training the first model based on the first information.

[0041] In the embodiments of the present application, a user of model training sends information related to model training based on reinforcement learning to a provider of model training, which helps to provide support for model training by the reinforcement learning method. For example, the first information can include related information of an environment in which the model interacts in the reinforcement learning process, thereby helping to determine the environment of reinforcement learning model training. For another example, the first information can include related information of model behavior in the reinforcement learning process, thereby helping to determine the behavior of the model in the training process.

[0042] In some embodiments, in the process of model training, the first model interacts with a first environment, and the first information includes one or more of the following: a first index list, a state of the first environment being determined by one or more indexes in the first index list; a first parameter list, a behavior of the first model being determined by one or more parameters in the first parameter list; a first training environment, a training environment of the first model being the first training environment, or a training environment of the first model being one of the first training environment; a function of the first model; state information of the first environment; behavior information of the first model; or a performance index of the first model.

[0043] In traditional machine learning, a one-to-one mapping relationship can be obtained by model training through a single category of training data. Taking supervised learning as an example, a single category of training data can refer to a single feature of training data, and a mapping relationship between the feature and a label can be obtained by model training.

[0044] Unlike traditional machine learning, reinforcement learning involves the interaction between an agent and an environment, and the change of behavior to the environment will affect the subsequent behavior. However, the training data used in traditional machine learning cannot represent the environment involved in reinforcement learning. That is, the training method of the traditional machine learning model cannot be applied to the model training based on reinforcement learning.

[0045] The first index list and the function of the first model can provide a basis for determining the environment for a provider of model training, thereby helping to solve the problem that the traditional training data cannot be applied to the model training based on reinforcement learning.

[0046] In some embodiments, the first information includes a function of the first model, and the training of the first model based on the first information includes: determining third information based on the function of the first model; and training the first model based on the third information, where the third information includes one or more of: one or more indexes for determining a state of the first environment; one or more parameters for determining a behavior of the first model; the first environment; the behavior of the first model; state information of the first environment; behavior information of the first model; or a training environment of the first model.

[0047] In the embodiments of the present application, the reinforcement learning capability is encapsulated as a function, so that the second device does not need to care about the underlying information in the reinforcement learning model training process, which helps to reduce the processing complexity of the second device and reduce the transmission overhead.

[0048] In some embodiments, the method further includes: notifying the capability information of the first device, the capability information of the first device being related to reinforcement learning, where the capability information of the first device includes one or more of: a first function, the first device supporting training of a model implementing the first function; a second environment, an environment supported by the first device, the second environment being used for training of the model; a third index list, the third index list including indexes supported by the first device when determining an interaction environment of the model; a second parameter list, the second parameter list including parameters supported by the first device when determining a behavior of the model; or a correspondence between the indexes in the third index list and the parameters in the second parameter list.

[0049] For example, the notification of the capability information of the first device can also be replaced by reporting the capability information of the first device or sending the capability information of the first device.

[0050] In the embodiments of the present application, by notifying the capability information of the model training provider (such as the capability information related to the reinforcement learning model training), the user of the model training can determine the configuration information of the model training that matches the capability information of the model training provider, thereby helping to improve the feasibility of the reinforcement learning model training.

[0051] In some embodiments, the receiving the first information comprises: receiving a first request, the first request being used to request training of the first model; and wherein the first request comprises a first attribute, the first attribute being used to indicate the first information.

[0052] By carrying the first information in the model training request, the degree of modification to the related technology can be reduced. In addition, the scheme is relatively easy to implement.

[0053] In some embodiments, the method further comprises: sending, to the second device, a model training report, the model training report comprising a second attribute and / or a third attribute, the second attribute being used to indicate a training environment of the first model, and the third attribute being used to indicate a performance indicator of the first model.

[0054] By carrying the training environment of the model and / or the performance indicator of the first model in the model training report, the second device can understand the model training situation, and support the model management of the second device. For example, based on the performance indicator of the first model, a management strategy in the model inference process can be determined, such as determining an acceptable performance indicator in the model inference process.

[0055] In some embodiments, the method further comprises: receiving a first threshold, the first threshold being used to indicate an acceptable performance indicator of the first model in the model inference process.

[0056] The acceptable performance indicator may, for example, refer to the minimum value allowed by the performance indicator. Taking the performance indicator as a reward value as an example, the acceptable performance indicator may, for example, refer to the minimum reward value allowed in the model inference process based on the first model.

[0057] By setting the above-mentioned first threshold, the reliability of the model inference can be improved.

[0058] Exemplarily, since the reward values in the last few rounds of the model training process are close to the reward values that can be obtained in the model inference process, the first threshold can be determined based on the reward values in the last few rounds of the model training process, which helps to improve the accuracy of the first threshold.

[0059] Exemplarily, the first threshold can be determined based on the performance indicator of the first model in the model training process.

[0060] In some embodiments, the method further comprises: receiving second information, the second information being used to indicate performing model inference based on the first model; and sending, to the second device, a model inference report, the model inference report comprising a performance indicator of the first model in the model inference process.

[0061] By reporting the performance indicators in the model inference process, it is helpful for the user of the model training to understand the inference state of the model, helpful for understanding the accuracy of the model inference result, and can provide support for the management and monitoring of the model inference process.

[0062] In some embodiments, if the first device is not the provider of the model inference (i.e., the model inference and the model training are deployed separately), the first device can forward the information related to the model inference issued by the second device to the provider of the model inference (referred to as the third device), and the first device can send the information related to the model inference reported by the third device to the second device.

[0063] The information related to the model inference issued by the second device may, for example, include the second information and / or the first threshold described above. The information related to the model inference reported by the third device may, for example, include the model inference report and / or the notification message described above.

[0064] Exemplarily, the second device can send the second information and / or the first threshold described above to the first device, and the first device can send the second information and / or the first threshold described above to the third device.

[0065] Exemplarily, the third device can send the model inference report and / or the notification message described above to the first device, and the first device can send the model inference report and / or the notification message described above to the second device.

[0066] In a third aspect, a communication apparatus is provided. In one possible design of the third aspect, the communication apparatus can have the functionalities of the first aspect or the second aspect described above. For example, the communication apparatus can include modules or units or means corresponding to the operations of the first aspect or the second aspect described above. The modules or units or means can be implemented in software, hardware, or a combination of software and hardware.

[0067] In a fourth aspect, a communication apparatus is provided. The communication apparatus can include an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory is configured to store part or all of the computer program or instructions necessary to implement the functionalities of the first aspect or the second aspect described above. The one or more processors can execute the computer program or instructions, which when executed cause the communication apparatus to implement the method in any possible design or implementation manner of the first aspect or the second aspect described above. The interface circuit is configured to implement the communication function within the communication apparatus and / or the communication function of the communication apparatus with other apparatuses or components.

[0068] In one possible implementation manner, the processor is configured to communicate with other apparatuses or components through the interface circuit.

[0069] In a possible implementation, the communication apparatus further includes the memory.

[0070] In a fifth aspect, the present application provides a communication system, including a first device and a second device. The first device or the second device can be the communication apparatus provided in the fourth aspect or the third aspect.

[0071] In a possible implementation, the second device can execute the method provided in the first aspect, and the first device can execute the method provided in the second aspect.

[0072] In a sixth aspect, the present application provides a computer readable storage medium, which stores computer readable instructions. When the computer readable instructions are executed, the method in any one of the aspects or any possible implementation of any one of the aspects is executed.

[0073] In a seventh aspect, the present application provides a computer program product, which includes computer program instructions. When the computer program instructions are executed, the method in any one of the aspects or any possible implementation of any one of the aspects is implemented.

[0074] In an eighth aspect, the present application provides a chip, which includes a processor. When the processor executes a program or instructions, the method in any possible implementation of the first aspect and the second aspect is executed. BRIEF DESCRIPTION OF DRAWINGS

[0075] FIG. 1 is a schematic diagram of an autonomous layered operation and maintenance system architecture;

[0076] FIG. 2 is a schematic diagram of an AI / ML workflow;

[0077] FIG. 3A is an example diagram of a deployment mode of model inference and model training;

[0078] FIG. 3B is another example diagram of a deployment mode of model inference and model training;

[0079] FIG. 3C is yet another example diagram of a deployment mode of model inference and model training;

[0080] FIG. 4 is a schematic diagram of the principle of reinforcement learning;

[0081] FIG. 5 is a schematic flowchart of a communication method provided by an embodiment of the present application;

[0082] FIG. 6 is a flowchart of another communication method provided by an embodiment of the present application;

[0083] FIG. 7 is a flowchart of yet another communication method provided by an embodiment of the present application;

[0084] FIG. 8 is a flow diagram of another communication method according to an embodiment of the present application;

[0085] FIG. 9 is a flow diagram of another communication method according to an embodiment of the present application;

[0086] FIG. 10 is a schematic block diagram of a communication device according to an embodiment of the present application;

[0087] FIG. 11 is a schematic block diagram of another communication device according to an embodiment of the present application. DETAILED DESCRIPTION

[0088] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0089] In the description of the present application, unless otherwise specified, " / " represents that the objects before and after the " / " are in an "or" relationship, for example, A / B can represent A or B; "and / or" in the present application is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, in the description of the present application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, "first", "second", and the like are used to distinguish the same items or similar items with basically the same function and effect. Those skilled in the art can understand that "first", "second", and the like do not limit the quantity and execution order, and "first", "second", and the like do not necessarily mean different. In the present application, "in the case of", "if", "when", "if", and the like can be used instead. In addition, these descriptions all mean that the corresponding processing will be done under certain objective conditions, and are not limited by time, and do not require a judgment action when implemented, nor does it mean that there are other limitations. In the present application, "greater than or equal to" can be replaced by "greater than", and correspondingly, "less than" can be replaced by "less than or equal to".

[0090] In the method embodiments of the present application, the size of the serial number does not mean the execution order, and the execution order should be determined by its function and internal logic, and should not construct any limitation on the implementation process of the embodiments of the present application.

[0091] It can be understood that some optional features in the embodiments of the present application can be implemented independently in some scenarios, solve corresponding technical problems, and achieve corresponding effects, without relying on other features, such as the scheme currently based on. In some scenarios, the optional features can be combined with other features according to requirements. Correspondingly, the apparatuses given in the embodiments of the present application can also implement these features or functions, and details are not described herein.

[0092] In the present application, the same or similar parts between various embodiments can be mutually referred to, unless otherwise specified. In the various embodiments of the present application, and the various implementation manners / implementation methods / realization methods in the various embodiments, the terms and / or descriptions between different embodiments, and the various implementation manners / implementation methods / realization methods in the various embodiments have consistency and can be mutually referred to, unless otherwise specified and logically conflicted. The technical features in different embodiments, and the various implementation manners / implementation methods / realization methods in the various embodiments can be combined to form new embodiments, implementation manners, implementation methods, or realization methods according to their inherent logical relationship. The implementation manners of the present application described below do not limit the protection scope of the present application.

[0093] The communication method provided by the present application can be applied to the architecture shown in FIG. 1. Referring to the autonomous layered operation and maintenance system architecture shown in FIG. 1, the autonomous layered operation and maintenance system architecture can include a cross-domain management function unit, a domain management function unit, and a network element. In some embodiments, the autonomous layered operation and maintenance system architecture can also include a business operation unit, and the functions of each unit and network element are briefly described below.

[0094] (1) Business operation unit: also referred to as a communication service management function, can provide functions and management services such as charging, settlement, accounting, customer service, business, network monitoring, communication service life cycle management, and service intent translation. The business operation unit can include an operator's operation system or a vertical industry operation system (vertical operational technology system).

[0095] (2) Cross-domain management function unit: also can be called network management function (NMF). The cross-domain management function unit can provide one or more of the following functions or management services: network lifecycle management, network deployment, network fault management, network performance management, network configuration management, network assurance, network optimization function, translation of intent from communication service provider (intent-CSP), translation of intent from communication service consumer (intent-CSC), etc. The network here can include one or more network elements, sub-networks or network slices. For example, the cross-domain management function unit can be a network slice management function (NSMF), or a management data analytical function (MDAF), or a cross-domain self-organization network function (SON-function), or a cross-domain intent management function unit.

[0096] For example, in some deployment scenarios, the cross-domain management function unit can also provide one or more of the following management functions or management services: sub-network lifecycle management, sub-network deployment, sub-network fault management, sub-network performance management, sub-network configuration management, sub-network assurance, sub-network optimization function, translation of sub-network intent from communication service provider, translation of sub-network intent from communication service consumer, etc. The sub-network can be composed of multiple small sub-networks or multiple network slice sub-networks.

[0097] For example, in implementation, the cross-domain management function unit can be implemented by a management platform and multiple management applications, where the management applications implement the management functions of one or more networks or sub-networks above.

[0098] (3) Domain management function unit: also can be called subnetwork management function (subnetwork NMF) or network element / function management function. The domain management function unit can provide one or more of the following functions or management services: lifecycle management of subnetwork or network element, deployment of subnetwork or network element, fault management of subnetwork or network element, performance management of subnetwork or network element, assurance of subnetwork or network element, optimization management of subnetwork or network element, and intent translation of subnetwork or network element. The subnetwork here includes one or more network elements. Alternatively, the subnetwork here can also include one or more subnetworks, i.e., one or more subnetworks form a larger coverage subnetwork. Alternatively, the subnetwork here can also include one or more network slice subnetworks. The subnetwork can include one of the following descriptions:

[0099] Network of a certain technology domain, such as wireless access network, core network, transmission network, etc.

[0100] Network of a certain standard, such as global system for mobile communications (GSM) network, long term evolution (LTE) network, 5th generation mobile communication technology (5G) network, etc.

[0101] Network provided by a certain equipment vendor, such as network provided by equipment vendor X, etc.

[0102] Network of a certain geographical area, such as network of factory A, network of prefecture-level city B, etc.

[0103] (4) a network element (NE), which is an entity for providing network services, and can include a core network element and / or an access network element, etc. For example, the core network element can include, but is not limited to, an access and mobility management function (AMF) entity, a session management function (SMF) entity, a policy control function (PCF) entity, a network data analysis function (NWDAF) entity, a network repository function (NRF), a gateway, etc. The access network element can include, but is not limited to, various types of base stations (for example, a generation node B (gNB), an evolved Node B (eNB), a central unit control panel (CUCP), a central unit (CU), a distributed unit (DU), a central unit user panel (CUUP), etc.

[0104] It should be understood that the number of units or network elements shown in FIG. 1 is only an example, and the number of units and network elements is not limited in the present application.

[0105] In order to improve the intelligent and automated level of the network, artificial intelligence (AI) and machine learning (ML) technologies are being applied in more and more fields, and the 3rd generation partnership project (3GPP) working group is studying a plurality of related topics for network intelligence. Among them, the related topics study the life cycle management of the model.

[0106] FIG. 2 is a schematic diagram of an AI / ML workflow. The process shown in FIG. 2 mainly includes model training 210, model testing 220, model simulation 230, model deployment 240, and model inference 250.

[0107] Generally, a model training consumer (MLT consumer) can invoke a model training service provided by a model training provider (MLT producer).

[0108] Taking the 3GPP network domain as an example, the network management system (i.e., the cross-domain management function unit mentioned above) can be used as a model training user. The model training provider can be a network element management device (i.e., the single-domain management function unit mentioned above), or can be a network element managed by an EMS, such as a radio access network (RAN), a base station (gNodeB, gNB), or a core network (CN) network element, such as a network data analysis network element (NWDAF), etc.

[0109] Among them, the base station can refer to a device that connects the fixed part and the wireless part in the mobile communication system, and is connected with the mobile terminal through the wireless channel in the air. The NWDAF has AI training, inference and other intelligent computing functions. The network element can generally refer to gNB, NWDAF and other network elements.

[0110] In addition, in the ORAN network domain, the service management and orchestration function (SMO) can be used as a model training user, and the network element directly managed by the SMO can be used as a model training provider. Among them, the role of the SMO in the network architecture is similar to that of the NMS, and it is responsible for the operation, management and maintenance of various network services and orchestration functions; the network element directly managed by it can be heterogeneous, such as EMS, gNB, NWDAF, etc. can be directly managed. Exemplarily, the provider of model training can be referred to as a machine learning training function (MLTF), which generally refers to a network element with ML training function.

[0111] Model inference and model training can be deployed on the same device or on different devices. Taking the 3GPP network domain (RAN domain) as an example, FIGS. 3A-3C show schematic diagrams of different deployment modes of model inference and model training.

[0112] Referring to FIG. 3A, model inference and model training are both deployed on the domain management function unit, i.e., the EMS.

[0113] Referring to FIG. 3B, model inference is deployed on a network element such as a gNB, and model training is deployed on a domain management function unit.

[0114] Referring to FIG. 3C, model inference and model training are both deployed on a network element such as a gNB.

[0115] In some embodiments, a model training request can be initiated by a model training user to a model training provider. In the model training request, the model training user can specify information such as a candidate training data source (candidateTrainingDataSource), performance requirements (performanceRequirements), and the like, as shown in Table 1. Among them, the format of the performance requirement data is ModelPerformance, which can include performance metrics (performanceMetric) and the like, and the performance metrics are used to specify the loss function during training, as shown in Table 2.

[0116] Table 1

[0117] Referring to Table 1, the support qualifier of a certain content indicates whether the content must be included. If it is mandatory (M), it must be included. If it is optional (O), it can not be included. If it is conditional (C), it can be included / not included under certain conditions. If it is CM, it must be included under certain conditions. The readability of a certain content indicates that the content can be read / not read. The writeability of a certain content indicates that the content can be written / not written. Whether it can be read and whether it can be written can be indicated by true (T) and false (F).

[0118] Table 2

[0119] Referring to Table 2, T / F (NOTE) indicates that the qualifier of isWritable is T when the attribute is used in the model training request, and the qualifier of isWritable is F in other cases.

[0120] In 3GPP, the related issues begin to focus on reinforcement learning, and study the full-process management and operation capability of AI / ML in the 5th generation mobile communication technology (5G) system to support various AI / ML technologies including reinforcement learning.

[0121] Reinforcement learning is an important branch of machine learning, which mainly focuses on how to take actions in an environment to maximize a certain cumulative reward, i.e., the training goal of reinforcement learning is to maximize the reward value. The core of reinforcement learning is that an agent learns the optimal behavior policy through interaction with an environment. For example, the agent can select a suitable action according to the current state of the environment (which can be acquired by an interpreter), and observe the feedback of the environment (such as the change of the state of the environment) to the action, so as to learn how to optimize its own decision, as shown in FIG. 4.

[0122] Reinforcement learning has many advantages, such as reinforcement learning can adjust behavior according to changes in the state of the environment, and has strong adaptability and high flexibility. Therefore, reinforcement learning has been widely applied in many fields, such as games, automatic control, robot control, finance, and medical treatment, etc.

[0123] It can be seen that if reinforcement learning is introduced into a communication system, it will help to improve the performance of the communication system. Then, how to apply reinforcement learning to a communication system is a problem to be solved.

[0124] To solve the above problem, the embodiments of the present application provide a communication method. In the method, a user of model training sends information related to reinforcement learning-based model training to a provider of model training, which helps to provide support for model training by a reinforcement learning method. For example, the first information can include information related to the environment of the model interaction in the reinforcement learning process, thereby helping to determine the environment of the reinforcement learning model training. For another example, the first information can include information related to the behavior of the model in the reinforcement learning process, thereby helping to determine the behavior of the model in the training process.

[0125] The method provided by the present application will be further described below in conjunction with the drawings.

[0126] FIG. 5 is a flow diagram of a communication method provided by the embodiments of the present application. The method shown in FIG. 5 involves the interaction of a first device and a second device, the first device can be a provider of model training, and the second device can be a user of model training. For example, the first device can be the EMS mentioned above, and the second device can be the NMS mentioned above.

[0127] It should be understood that the first device or the second device in the method provided by the present application can be an entity physical device, a virtualized function module, a software function, a circuit, a chip (such as a system on chip (SoC) chip or a system in package (SIP) chip), etc.

[0128] It should be noted that the following is an introduction to the method provided by the embodiments of the present application from the perspective of the second device interacting with the first device.

[0129] The method shown in FIG. 5 can include S510 and S520, which are described as follows.

[0130] S510, the second device determines first information.

[0131] The first information described above can be used to train the first model, and the training method of the first model is reinforcement learning, or the first model is trained by using the reinforcement learning method.

[0132] Based on the principle of reinforcement learning mentioned above, the first model can be understood as an agent or an ML model that constitutes an agent. For ease of description, in the following, the object that the first model interacts with in the model training process, i.e., the environment mentioned above, is referred to as the first environment; and the output of the first model, i.e., the behavior mentioned above, is referred to as the first behavior.

[0133] In some embodiments, the first information can include one or more of the following: a first index list; a first parameter list; a first training environment; a function of the first model; state information of the first environment; behavior information of the first model; or a performance indicator of the first model. The above information is described in detail as follows.

[0134] First index list

[0135] Exemplarily, the index mentioned in the embodiments of the present application can be a set composed of network performance measurement indicators or key performance indicators (KPIs).

[0136] Exemplarily, the state of the first environment can be determined by one or more indicators in the first index list. Or in other words, the state parameter of the first environment can include one or more indicators in the first index list. For example, the state information of the first environment can be determined by the state parameter of the first environment and the value of the state parameter of the first environment.

[0137] As an example, the indicators in the first index list can include indicators specified in 3GPP management data analytics (MDA) (refer to 3GPP TS 28.104), indicators specified in 3GPP NWDAF (refer to 3GPP TS 23.288), and indicators in 3GPP RAN.

[0138] For example, the indicators in the first indicator list can include one or more of a reference signal received power (RSRP), a physical resource block (PRB) utilization, or a handover success rate.

[0139] Optionally, the state of the first environment can be determined by all indicators in the first indicator list, that is, the state information of the first environment can be determined by a value set of all indicators in the first indicator list. In this case, the indicators for determining the state of the first environment are provided by the second device.

[0140] For example, the first indicator list includes RSRP, PRB utilization, and handover success rate, the state of the first environment can be determined by RSRP, PRB utilization, and handover success rate, and the state information of the first environment can be determined by a value set of RSRP, PRB utilization, and handover success rate.

[0141] Optionally, the state of the first environment can be determined by some indicators in the first indicator list, that is, the state information of the first environment can be determined by a value set of some indicators in the first indicator list. In this case, the second device provides candidate indicators for determining the first environment, and the first device can select indicators for determining the state of the first environment from the candidate indicators.

[0142] For example, the first indicator list includes RSRP, PRB utilization, and handover success rate, the state of the first environment can be determined by RSRP and PRB utilization, or the state of the first environment can be determined by RSRP and handover success rate. The indicators for determining the first environment can be determined by the first device.

[0143] If the state of the first environment is determined by RSRP and PRB utilization, the value set of RSRP is X1, X2, X3, X4, X5, X6, and the value set of PRB utilization is 20%, 30%, 40%, 60%, 80%, and 90%, then the state information of the first environment is composed of the above values of RSRP and PRB utilization.

[0144] It should be noted that the values of the above indicators are only exemplary, and the values of the indicators can be discrete or continuous, which is not limited in the present application.

[0145] The first parameter list

[0146] Exemplarily, the behavior of the first model can be determined by one or more parameters in the first parameter list. In other words, the behavior (action) parameters of the first model can include one or more parameters in the first parameter list. For example, the parameters in the first parameter list can include antenna settings, cell settings, beam settings, etc. For another example, the parameters in the first parameter list can include parameters specified in 3GPP management data analytics (MDA) (refer to 3GPP TS 28.104), parameters specified in 3GPP NWDAF (refer to 3GPP TS 23.288), and performance parameters in 3GPP RAN.

[0147] Optionally, the behavior of the first model can be determined by all parameters in the first parameter list, that is, the behavior information of the first model can be determined by the value set of all parameters in the first parameter list. In this case, the parameters for determining the first behavior are provided by the second device.

[0148] Taking an example in which the first parameter list includes antenna settings, cell settings, and beam settings, the first behavior can be determined by the antenna settings, the cell settings, and the beam settings, and the behavior information of the first model can be determined by the value set of the antenna settings, the cell settings, and the beam settings.

[0149] Optionally, the behavior of the first model can be determined by part of the parameters in the first parameter list, that is, the behavior information of the first model can be determined by the value set of part of the parameters in the first parameter list. In this case, the candidate parameters for determining the first behavior are provided by the second device, and the parameters for determining the first behavior can be selected by the first device from the candidate parameters.

[0150] Still taking an example in which the first parameter list includes antenna settings, cell settings, and beam settings, the first behavior may, for example, be determined by the antenna settings and the beam settings, or the first behavior may, for example, be determined by the cell settings and the beam settings. The parameters for determining the first behavior can be determined by the first device.

[0151] The first training environment

[0152] The training environment of the first model can refer to an environment platform capable of providing the first environment, or in other words, the training environment of the first model can provide the first environment for model training. Specifically, the environment platform can provide the state of the first environment, and the state of the first environment provided by the environment platform can respond to different behaviors.

[0153] Exemplarily, the first training environment can include one or more training environments of the first model, such as a simulation environment or a real environment. For example, the first training environment can include a simulation function, such as a network digital twin function (NDT Function), an available live network environment, and the like. Optionally, the first information can indicate the first training environment by an identifier of the simulation function or an identifier of the available live network environment, so as to save indication overhead.

[0154] Optionally, the training environment of the first model can be the first training environment, or the training environment of the first model can be one of the first training environment. If the first training environment includes only one training environment, the training environment of the first model can be the first training environment; if the first training environment includes multiple training environments, the first device can select the training environment of the first model from the multiple training environments. This scheme helps to improve the flexibility of model training.

[0155] Exemplarily, the second device can determine the available live network environment based on the function of the first model and / or the indicator of the state of the first environment. For example, if the indicator of the state of the first environment includes RSRP, the live network environment that supports RSRP collection or has RSRP collection information can be used as the first training environment. For another example, the live network environment that supports model training of the function can be used as the first training environment.

[0156] For example, the first device can select the training environment of the first model based on its capability information. As an example, for the first device that supports training a model in a live network environment, the live network environment can be selected as the training environment of the first model. As another example, for the first device that supports training a model in a simulation function, the simulation function can be selected as the training environment of the first model.

[0157] It is considered that it is relatively easy for the second device to obtain information of a live network environment, and the information processing capability of the second device is usually strong. Therefore, the first training environment is provided by the second device, which is relatively simple to implement and helps to reduce the processing complexity of the first device.

[0158] Function of the first model

[0159] Exemplarily, the first device can train the first model based on the function of the first model issued by the second device. In this way, the information related to training the first model can be determined by the first device, thereby helping to improve the freedom and flexibility of the first device in training the first model.

[0160] For example, the first device can determine information (may be referred to as third information) related to the training of the first model based on the function of the first model, and then train the first model based on the third information. Wherein, the third information can include one or more of the following: one or more indicators for determining the state of the first environment; one or more parameters for determining the behavior of the first model; the first environment; the behavior of the first model; the state information of the first environment; the behavior information of the first model; or the training environment of the first model.

[0161] Optionally, the third information can further include the first indicator list and / or the first parameter list.

[0162] Optionally, the third information can include the performance indicator of the first model. The meaning of the performance indicator of the first model can be referred to in the following description, and will not be described here for brevity.

[0163] The method of determining the third information based on the function of the first model is described below.

[0164] As an example, the indicators for determining the state of the first environment can be determined based on the association between the function of the first model and the indicators (i.e., the indicators for determining the state of the environment).

[0165] As another example, the first indicator list can be determined based on the association between the function of the first model and the indicator list. Wherein, the indicators for determining the state of the environment can be considered as KPIs for model training using reinforcement learning method.

[0166] Further, the first device can construct the first environment based on the first indicator list, or the above-mentioned indicators for determining the state of the first environment.

[0167] Optionally, the association between the function of the model and the indicators, and / or the association between the function of the model and the indicator list can be pre-stored in the first device.

[0168] As another example, the parameters for determining the first behavior can be determined based on the association between the function of the first model and the parameters (i.e., the parameters for determining the behavior). Optionally, the association between the function and the parameters can be pre-stored in the first device.

[0169] As another example, during the model training process, the first device can adjust the behavior strategy of the first model based on the state of the current environment and the function of the first model, which helps to speed up the training speed of the model.

[0170] It should be understood that the third information can also include other information for the training of the first model, which is not limited by the present application.

[0171] State information of the first environment / behavior information of the first model

[0172] In some embodiments, the first information can further include the state information of the first environment and / or the behavior information of the first model. Compared with determining the information by the first device, providing the state information of the first environment and / or the behavior information of the first model by the second device can help reduce the processing complexity of the first device.

[0173] Performance indicator of the first model

[0174] In some embodiments, the performance indicator of the first model can include, for example, a reward indicator, a reward function. The reward indicator can include, for example, an immediate reward and a cumulative reward. The cumulative reward can include, for example, a cumulative reward within a certain time granularity and / or a cumulative reward in the entire training process, etc. The reward function defines the feedback signal that the agent or the ML model constituting the agent can obtain after performing the behavior.

[0175] For example, the reward function of the first model is an initial reward function, which can be adjusted and iterated according to the training of the model.

[0176] For example, the performance indicator of the first model can represent the mapping relationship between the reward value and the value of the indicator of the first environment (or the change amount of the value of the indicator). The reward value mentioned here can include an immediate reward and / or a cumulative reward.

[0177] It should be understood that the first information can include one or more of the above information, and the first information can also include other information for reinforcement learning not shown in the present application, which is not limited in the present application.

[0178] S520, the second device sends the first information to the first device, and correspondingly, the first device receives the first information sent by the second device. The first device is a device for training the first model, that is, the provider of the model training mentioned above.

[0179] In some embodiments, the first information can be carried in the model training request. That is, the second device can send a first request to the first device, and the first request can be used to request to train the first model, wherein the first information is included in the first request. For example, the first request can include a first attribute, which can be used to indicate the first information, such as the first attribute can be used to indicate the first training environment.

[0180] Table 3 is an example of the model training request provided by the embodiments of the present application.

[0181] Table 3

[0182] Referring to Table 3, the support qualifier of the first attribute is M / O. For example, when reinforcement learning is requested, the support qualifier of the first attribute can be M, i.e., the first attribute is mandatory; in other cases, the first attribute can be optional.

[0183] Table 3 gives an example of the first attribute related qualifier, and it should be understood that the first attribute related qualifier can also be of other types, which are not limited by the present application.

[0184] By carrying the first information in the model training request, the degree of modification to related technologies can be reduced. In addition, this scheme is relatively easy to implement.

[0185] In some embodiments, the first information can be carried in the model performance. For example, the model performance can include a fourth attribute, which can be used to indicate the first information, such as the first indicator list and / or the first parameter list.

[0186] Table 4 gives an example of the model performance provided by the embodiments of the present application.

[0187] Table 4

[0188] Table 4 gives an example of the fourth attribute related qualifier, and it should be understood that the fourth attribute related qualifier can also be of other types, which are not limited by the present application.

[0189] It should be understood that multiple information in the first information can be carried in the same message or in different messages. For example, the first training environment can be carried in the model training request, and the first indicator list can be carried in the model performance.

[0190] It should be understood that the first information can also be carried in other types of information, which are not limited by the present application.

[0191] In the embodiments of the present application, the user of the model training sends information related to reinforcement learning based model training to the provider of the model training, which helps to support the application of reinforcement learning technology.

[0192] In traditional machine learning, through single category training data, model training can obtain one-to-one mapping relationship. For example, in supervised learning, single category training data can refer to single feature training data, and through model training, the mapping relationship between the feature and the label can be obtained.

[0193] Different from traditional machine learning, reinforcement learning involves the interaction between an agent and an environment, and the change of the environment by the behavior will affect the subsequent behavior. The training data used by traditional machine learning usually cannot represent the environment interacted by reinforcement learning. That is, the training method of the traditional machine learning model cannot be applied to the training of the reinforcement learning-based model.

[0194] The embodiments of the present application can provide the provider of model training with a basis for determining the above-mentioned environment through the first index list and the function of the first model, thereby helping to solve the problem that the traditional training data cannot be applied to the model training based on reinforcement learning.

[0195] The way of determining the first information can include various ways, and the way of determining different information in the first information can be the same or different. The way of determining the first information is described in detail below.

[0196] In some embodiments, the first information can be determined based on the requirement information of model training. For example, the second device can receive the requirement information of model training, and then determine the first information based on the requirement information of model training.

[0197] For example, the requirement information of model training can include the function of the first model, or the requirement information of model training can include the function of the first model and the second index list. The index in the second index list is an index expected to be used to determine the state of the first environment.

[0198] For example, the function of the first model in the first information is the function of the first model included in the requirement information of model training.

[0199] For another example, the first information includes the first index list, and the first device can determine the first index list based on the function of the first model.

[0200] As an example, the first index list can be determined based on the association between the function of the first model and the index.

[0201] Table 5 is an example of the association between the function of the model and the index.

[0202] Table 5

[0203] Referring to Table 5, function 1 has an association with index 1 and index 2, function 2 has an association with index 3 to index 5, and function 3 has an association with index 6 to index 8. If the function of the first model is function 2, the first index list can include index 3 to index 5.

[0204] As another example, the first indicator list can be determined based on the association between the function of the first model and the indicator list. Table 6 is an example of the association between the function of the model and the indicator list.

[0205] Table 6

[0206] Referring to Table 6, function 1 has an association with indicator list 1, function 2 has an association with indicator list 2, and function 3 has an association with indicator list 3. If the function of the first model is function 3, the first indicator list is indicator list 3.

[0207] Further, the indicators in the first indicator list can be determined based on the correspondence between the first indicator list and the indicators. As an example, the first information can include an identifier of the first indicator list, and the first device can determine the indicators in the first indicator list based on the correspondence between the first indicator list and the indicators. Alternatively, the second device can determine the indicators in the first indicator list based on the correspondence between the first indicator list and the indicators. In this case, the first information can include each indicator in the first indicator list.

[0208] Optionally, the association between the function of the model and the indicators, and / or the association between the function of the model and the indicator list can be pre-stored in the second device and / or the first device.

[0209] For another example, the first information includes the first indicator list, and if the demand information for model training includes a second indicator list, the first indicator list can be determined based on the second indicator list, such as the first indicator list being the second indicator list.

[0210] For another example, the first information can include a first parameter list, and the first device can determine the first parameter list based on the function of the first model. As an example, the first parameter list can be determined based on the association between the function of the first model and the parameters. As another example, the first parameter list can be determined based on the association between the function of the first model and the parameter list.

[0211] Optionally, the first information can include an identifier of the parameter list, and the first device can determine the parameters in the first parameter list based on the identifier and the correspondence between the first parameter list and the parameters. Alternatively, the second device can determine the parameters in the first parameter list based on the correspondence between the first parameter list and the parameters, and in this case, the first information includes each parameter in the first parameter list.

[0212] Optionally, the association between the function of the model and the parameters, and / or the association between the function of the model and the parameter list can be pre-stored in the second device and / or the first device.

[0213] It should be understood that the method of determining the first parameter list based on the function of the first model is similar to the method of determining the first index list based on the function of the first model mentioned above, and for brevity, will not be repeated here.

[0214] It should be understood that the method of determining the first parameter list based on the function of the first model can also include: first determining the first index list based on the function of the first model, and then determining the first parameter list according to the correspondence between the indexes of the first index list and the parameters. The method will be described in detail below, and will not be repeated here.

[0215] For another example, the first information includes the state information of the first environment. As an example, the state information of the first environment can be determined based on the first index list, and the first index list can be determined based on the method described above. As another example, the state information of the first environment can be determined directly based on the function of the first model, such as based on the association between the function of the model and the state information of the first environment.

[0216] For another example, the first information includes the behavior information of the first model. As an example, the behavior information of the first model can be determined based on the first parameter list, and the first parameter list can be determined based on the method described above. As another example, the behavior information of the first model can be determined directly based on the function of the first model, such as based on the association between the function of the model and the behavior information of the first model.

[0217] For another example, the first information includes the first training environment, and the performance of the first model can be determined based on the function of the first model.

[0218] For another example, the first information includes the performance index of the first model, and the performance of the first model can be determined based on the function of the first model.

[0219] In some embodiments, the first information can be determined based on the capability information of the first device, such as the second device can obtain the capability information of the first device, and then determine the first information based on the capability information of the first device. The capability information of the first device can refer to the related description below, and will not be repeated here.

[0220] The capability information of the first device may, for example, include a third index list, which can include indexes that the first device supports to use when determining the interaction environment of the model. That is, the capability information of the first device is used to indicate which indexes (i.e., indexes in the third index list) the first device supports to use to determine the interaction object of the first model, i.e., the state of the environment.

[0221] Optionally, the third index list supported by the first device is associated with the training environment, that is, the third index list supported by the first device can be the same or different in different training environments.

[0222] For example, the first information includes a function of the first model, and the capability information of the first device includes a third index list. The first index list can be determined based on a correspondence between an index in the third index list and the function of the first model. For example, the first index list can include an index in the third index list that corresponds to the function of the first model.

[0223] In some embodiments, the first information can be determined based on the capability information of the first device and the requirement information of the model training. For example, the first index list can be determined based on the second index list and the third index list, e.g., the first index list includes part of the indexes in the third index list.

[0224] As an example, if the third index list includes the second index list, the first index list can be the second index list.

[0225] According to the above method of determining the first index list and performing model training based on the first index list, the training result of the first model can meet the requirements of the proposer of the training requirement, i.e., meet the user requirements.

[0226] As another example, if the third index list includes part of the indexes in the second index list, the first index list includes indexes common to the second index list and the third index list. In other words, the indexes in the first index list belong to both the second index list and the third index list.

[0227] The above method of determining the first index list can avoid the case where the first device does not support the configured indexes, thereby helping to improve the feasibility of model training.

[0228] If the third index list does not include the second index list, the first index list can be determined based on the third index list, e.g., the first index list is a subset of the third index list. Further, the first device can send a response message to the proposer of the model training requirement. The response message can be used to indicate, for example, the indexes not adopted by the proposer of the model training requirement and / or the first index list adopted.

[0229] The capability information of the first device may, for example, include a second parameter list, which can include parameters supported by the first device when determining the behavior of the model. That is, the capability information of the first device is used to indicate which parameters (i.e., parameters in the second parameter list) supported by the first device are used to determine the behavior of the model.

[0230] If the first information includes a first parameter list, the first parameter list can be determined based on a correspondence between parameters in the second parameter list and functions of the first model, or the first parameter list can be determined based on a parameter list provided by a proposer of the model training requirement and the second parameter list, in a case where the capability information of the first device includes the second parameter list. For example, the first parameter list can include parameters common to the second parameter list and the parameter list provided by the proposer of the training requirement.

[0231] In some embodiments, the capability information of the first device includes a third indicator list, a second parameter list, and a correspondence (which can be referred to as a first correspondence for ease of description) between indicators in the third indicator list and parameters in the second parameter list. In this case, the first parameter list can be determined based on indicators in the first indicator list and the first correspondence, where the first parameter list includes parameters that have the correspondence with the indicators in the first indicator list.

[0232] The correspondence between an indicator and a parameter can be understood as that, when training the first model in an environment determined by the indicator, the parameter can support a determined behavior, or in other words, a change in the parameter will affect a state of the indicator. This is because, when training the first model in different environments determined by indicators, which parameters can support determined behaviors can be the same or different. Therefore, determining the first parameter list based on the above correspondence helps to improve the rationality of the determined behavior.

[0233] As an example, the first correspondence can be stored in the form of a list, as shown in Table 7. Table 7 gives an example of the first correspondence.

[0234] Table 7

[0235] Referring to Table 7, indicator 1 has a correspondence with parameter 1 and parameter 2; indicator 2 has a correspondence with parameter 3; and indicator 3 has a correspondence with parameter 4, parameter 5, and parameter 6.

[0236] Taking the first correspondence shown in Table 7 as an example, if the third indicator list includes indicator 1, indicator 2, and indicator 3, and the second parameter list includes parameters 1 to 6. If the first indicator list includes indicator 1 and indicator 2, then the first parameter list can include parameter 1, parameter 2, and parameter 3.

[0237] In addition, different devices or different device manufacturers can support different above-mentioned correspondences, therefore, determining the first parameter list based on the above-mentioned correspondences can configure the supported first parameter list for different devices, which helps to improve the adaptability of the configuration information of the model training and the training device.

[0238] In some embodiments, the plurality of first correspondence relations can be associated with a plurality of devices one by one, or the plurality of first correspondence relations can be associated with a plurality of manufacturers one by one, that is, the first correspondence relations associated with the devices belonging to a unified manufacturer are the same.

[0239] For example, the first correspondence relation can be pre-configured (or stored) in the second device, or reported by the first device.

[0240] For example, for the case that the plurality of first correspondence relations can be associated with a plurality of devices one by one, the first correspondence relation can be reported by the first device, such as when the first device needs to perform model training. In this case, the second device can not need to maintain the plurality of first correspondence relations associated with the plurality of first devices, which helps to reduce the memory overhead of the second device.

[0241] For example, for the case that the plurality of first correspondence relations can be associated with a plurality of devices one by one, the first correspondence relation can be reported by the first device, such as when the first device needs to perform model training. In this case, the second device can not need to maintain the plurality of first correspondence relations associated with the plurality of first devices, which helps to reduce the memory overhead of the second device.

[0242] In some embodiments, the first information can include index information, such as any one of the index information in Table 5, Table 6, or Table 7, to indicate the function of the first model and the corresponding indicators and / or parameters, thereby helping to save the indication overhead of the first information. For example, the first information can include index 1, and through the index 1, it can be indicated that the function of the first model is function 1, and the first indicator list includes indicators 1 and 2. For another example, the first information can include index 5, and through the index 5, it can be indicated that the function of the first model is function 2, and the first indicator list is indicator list 2.

[0243] To solve the problem that the communication system cannot support the training of the reinforcement learning model in the related art, the embodiments of the present application provide a communication method, which reports the capability information of the provider of the model training (such as the capability information related to the training of the reinforcement learning model), helps the user of the model training to determine the configuration information of the model training that matches the capability information of the provider of the model training, and thereby helps to improve the feasibility of the training of the reinforcement learning model.

[0244] FIG. 6 is a flow diagram of another communication method provided by the embodiments of the present application. The method shown in FIG. 6 involves the interaction between the first device and the second device. The first device can be a provider of model training, such as the EMS mentioned above, and the second device can be a user of model training, such as the NMS mentioned above.

[0245] The method shown in FIG. 6 can include S610. The method provided by the embodiments of the present application is introduced below from the perspective of the interaction between the first device and the second device.

[0246] In S610, the second device acquires the capability information of the first device. Accordingly, the first device can inform the second device of the capability information of the first device.

[0247] In some embodiments, the capability information of the first device is related to reinforcement learning, or the capability information of the first device can be referred to as information of the reinforcement learning capability of the first device. For example, the capability information of the first device can include one or more of the following: a first function, the first device supports training of a model implementing the first function; a second environment, the first device supports an environment for training of the model; a third index list, the third index list includes indexes supported by the first device for use in determining an interaction environment of the model; a second parameter list, the second parameter list includes parameters supported by the first device for use in determining a behavior of the model; or a correspondence between the indexes in the third index list and the parameters in the second parameter list (which can be referred to as a first correspondence).

[0248] As an example, the first function in the above capability information can be stored in association with the indexes in the third index list; and / or the first function can be stored in association with the parameters in the second parameter list. In this way, when determining the function of the first model, the indexes used to determine the state of the first environment can be determined based on the correspondence between the first function and the indexes in the third index list, or the parameters used to determine the behavior of the first model can be determined based on the correspondence between the first function and the parameters in the second parameter list.

[0249] For example, the capability information of the first device can include a value range of the indexes in the third index list. For example, the third index list includes PRB utilization, the variation range of the PRB utilization supported by different first devices can be the same or different; or the variation range of the PRB utilization supported by the same device in different environments can be the same or different. Based on the value range of the indexes in the third index list, the state information of the first environment is determined, which helps to improve the feasibility and reliability of the training of the first model.

[0250] For example, the capability information of the first device can include a value range of the parameters in the second parameter list. For example, the second parameter list includes antenna settings, the adjustment range of the antenna angle supported by different first devices can be the same or different. Based on the value range of the parameters in the second parameter list, the behavior information of the first model is determined, which helps to improve the feasibility and reliability of the training of the first model.

[0251] For example, the second environment can refer to an available training environment supported by the first device, such as NDT or an available live network environment.

[0252] In some embodiments, the capability information of the first device can be used to determine one or more of the following: an indicator for determining the state of the first environment; a parameter for determining the behavior information of the first model; the state information of the first environment; or the behavior information of the first model, etc.

[0253] For example, the third indicator list can be used to determine the state and / or the state information of the first environment, such as determining the state and / or the state information of the first environment according to part or all of the indicators in the third indicator list. As an example, an indicator for determining the state of the first environment can be determined first, and then the state information of the first environment is determined based on the value range of the indicator.

[0254] For another example, the second parameter list can be used to determine the behavior and / or the behavior information of the first model, such as determining the behavior and / or the behavior information of the first model through part or all of the parameters in the second parameter list. As an example, a parameter for determining the behavior of the first model can be determined first, and then the behavior information of the first model is determined based on the value range of the parameter.

[0255] The interaction mode of the capability information of the first device (the second device acquires the capability information of the first device, or the method for the first device to notify the capability information of the first device) can include multiple modes, which will be introduced respectively.

[0256] In some embodiments, the capability information of the first device can be reported periodically. When a preset period is reached, the first device sends the capability information of the first device to the second device.

[0257] In some embodiments, the capability information of the first device is reported in an updated manner, that is, when the capability information of the first device changes, the first device sends the capability information of the first device to the second device, which helps to update the capability information of the first device on the second device side in a timely manner, thereby helping to avoid the impact of the update of the capability information of the first device on the training performance of the first model. Optionally, the first device can send the updated capability information of the first device to the second device, or the first device can send the updated capability information to the second device to save transmission overhead.

[0258] In some embodiments, the capability information of the first device is reported in a registered manner. That is, when the first device is registered, such as when it is registered in the management domain of the second device, the first device sends the capability information of the first device to the second device.

[0259] In some embodiments, the second device can request the capability information of the first device from the first device, and in response to the request, the first device can send the capability information of the first device to the second device.

[0260] In some embodiments, the second device can subscribe to the capability information of the first device. Exemplarily, different subscription manners, the second device obtains the capability information of the first device at different time. Optionally, after the second device subscribes to the capability information of the first device, the second device can periodically obtain the capability information of the first device, i.e., the subscription message is periodically triggered. Optionally, after the second device subscribes to the capability information of the first device, if the capability information of the first device changes, the second device can obtain the updated capability information of the first device, i.e., the subscription message is event triggered. It should be understood that the obtaining of the capability information of the first device mentioned herein can include receiving a notification including the capability information of the first device.

[0261] It should be understood that the two communication methods mentioned above can also be used in combination. For example, the scheme can include steps 1 to 3. The parts not described in detail can refer to the description above.

[0262] Step 1, the second device obtains the capability information of the first device. Correspondingly, the first device notifies the second device of the capability information of the first device.

[0263] Step 2, the second device determines the first information.

[0264] In some embodiments, the second device can determine the first information based on the requirement information of the model training and / or the capability information of the first device. The scheme can refer to the description above, and for brevity, will not be described here.

[0265] Step 3, the second device sends the first information to the first device, and correspondingly, the first device receives the first information sent by the second device.

[0266] The first information can be used to train the first model. In some embodiments, the first device can determine one or more of the first environment, the state information of the first environment, the behavior information of the first model, the training environment or the performance index based on the first information. Further, the first device can train the first model based on these information.

[0267] In some embodiments, the capability information of the first device can be encapsulated as a function, i.e., the capability information of the first device includes a first function. The first function may, for example, include one or more of various RAN domain functions (see use cases in 3GPP TS 38.300), such as but not limited to a mobile load balancing function, a coverage optimization function, and the like; one or more of various management data analysis functions (MDAF) in a domain management function (see use cases in 3GPP TS 28.104), such as but not limited to a coverage problem analysis function, a slice coverage analysis function, and the like; or one or more of various functions defined in a NWDAF of a CN domain (see use cases in 3GPP TS 23.288).

[0268] In this case, the first information can include the function of the first model (but not the first list of indicators, the first list of parameters, the state information of the first environment, and the behavior information of the first model, etc.), and the first device can determine the information related to the training of the first model based on the function of the first model, such as the first list of indicators, the first list of parameters, the state information of the first environment, and the behavior information of the first model, etc.

[0269] In this way, the second device only needs to focus on the information at the model function level, and does not need to focus on the information at the indicator and parameter levels, which helps to reduce the processing complexity of the second device and the transmission overhead of the second device. In addition, the first device can store information associated with the first function for training the model using the reinforcement learning method to support the above scheme.

[0270] In some embodiments, the capability information of the first device is not encapsulated as a function, i.e., the capability information of the first device does not include the first function, but includes the third list of indicators, the second list of parameters, and the first correspondence, etc. In this case, the first information can include the first list of indicators, the first list of parameters, the state information of the first environment, and the behavior information of the first model, etc.

[0271] It should be noted that if the capability information of the first device includes the second list of parameters, the first information can include the first list of parameters, or the first list of parameters and the behavior information of the first model, or can not include the first list of parameters and the behavior information of the first model. If the capability information of the first device does not include the second list of parameters, the first information can not include the first list of parameters and the behavior information of the first model. This scheme helps to avoid the case that the first device does not support determining the behavior information of the first model based on the parameters in the first list of parameters, and improves the reliability of training the first model.

[0272] Alternatively, if the first information includes the first parameter list, i.e., the second device provides parameters or candidate parameters for determining the behavior of the first model, the first device can report the second parameter list to the second device; if the first information does not include the first parameter list, the first device can not report the second parameter list to the second device.

[0273] In the embodiments of the present application, by considering the capability information of the first device in the process of configuring the first information, the feasibility of training the first model is improved.

[0274] In some embodiments, the first device can send the environment used for training the first model to the second device. For example, the first device can select a second training environment from the first training environment, and train the first model in the second training environment. As an example, the first device can send the second training environment to the second device. Alternatively, the second training environment can be indicated by the identifier of the training environment, so as to save the indication overhead.

[0275] In some embodiments, if the first device determines the first environment by using part of the indicators in the first indicator list, the first device can report the indicators for determining the state of the first environment to the second device. If the first device determines the behavior of the first model by using part of the parameters in the first parameter list, the first device can report the parameters for determining the behavior of the first model to the second device. For example, the model training report can include the parameters for determining the behavior of the first model and / or the indicators for determining the state of the first environment.

[0276] In some embodiments, the first device can send the performance indicators of the first model to the second device, such as the performance indicators in the training process. For example, the performance indicators can include reward indicators and reward functions. The reward indicators can include single reward and cumulative reward, for example. The cumulative reward can include cumulative reward within a certain time granularity and / or cumulative reward in the entire training process, etc. The reward function defines the feedback signal that the agent or the ML model constituting the agent can obtain after performing the behavior.

[0277] For example, the performance indicators within a period of time, such as the performance indicators in the last several rounds of training, can be reported to the second device, or all the performance indicators in the training process can be reported to the second device. The number of rounds of performance indicators to be reported can be determined according to a preset value. For example, if the preset value is 5, the first device can send the performance indicators of the last 5 rounds in the training process of the first model to the second device.

[0278] Exemplarily, the performance indicator of the first model can include an interaction record of the environment, the behavior, and the reward value. The environment, the behavior, and the reward value in the interaction record have a time sequence relationship or a chronological relationship. This scheme helps the second device understand the training process of the first model. Further, based on the interaction record, the second device can determine a model management strategy, such as updating the model training parameters, thereby helping to improve the system performance.

[0279] In some embodiments, the first device can send a model training report to the second device, and accordingly, the second device can receive the model training report sent by the first device. The model training report can include one or more of the following: the environment (such as the second training environment) used for training the first model, the parameter for determining the behavior of the first model, the indicator for determining the state of the first environment, or the performance indicator of the first model.

[0280] Exemplarily, the model training report can include a second attribute and / or a third attribute. The second attribute is used to indicate the training environment of the first model, such as the second training environment. The third attribute is used to indicate the performance indicator of the first model, such as the reward value in the process of training the first model.

[0281] By carrying the training environment of the model and the performance of the first model in the model training report, it helps the second device to understand the model training situation and provides support for the model management of the second device.

[0282] The foregoing describes the training process of the first model. The inference process of the first model will be described below.

[0283] In some embodiments, the first device can send second information to the second device, and accordingly, the second device can receive the second information of the first device. The second information can be used to indicate the execution of model inference based on the first model. Alternatively, the second information can also be understood as the activation information of the model inference.

[0284] In some embodiments, the first device can send a model inference report to the second device, and accordingly, the second device can receive the model inference report sent by the first device. Alternatively, the model inference report includes the performance indicator of the first model in the model inference process. The performance indicator of the first model included in the inference report can include one or more of the following: the reward value in the entire model inference process, the reward value in the model inference process within a period of time, the size relationship between the reward value and the first threshold value (such as whether the reward value is less than the first threshold value), the reward value less than the first threshold value, or the interaction record of the environment-behavior-reward value.

[0285] In some embodiments, the first device can receive the second device sending the first threshold value, and accordingly, the second device can send the first threshold value to the first device. The first threshold value can be used to indicate an acceptable performance indicator (or a threshold value of an acceptable performance indicator) of the first model in the model inference process, such as a minimum value allowed by the performance indicator. Taking the reward value as an example, the acceptable performance indicator can refer to the minimum reward value allowed in the model inference process based on the first model.

[0286] Exemplarily, since the reward values in the last few rounds of the model training process are close to the reward values that can be obtained in the model inference process, the first threshold value can be determined based on the reward values in the last few rounds of the model training process, which helps to improve the accuracy of the first threshold value.

[0287] Exemplarily, the first threshold value can be determined based on the performance indicator of the first model in the model training process.

[0288] Exemplarily, in the inference process of the first model, if the reward value is less than the first threshold value, the first device can send a notification message to the second device, which can be used to notify the second device that the reward value is less than the first threshold value, or the notification message can be used to notify the second device that the reward value is less than the first threshold value and the current reward value. Based on the notification message, the second device can timely manage the inference process of the first model, such as optimizing the first model, replacing the inference model, etc., thereby helping to improve the system performance.

[0289] In order to increase the robustness of the system, when the number of times that the reward value is less than the first threshold value reaches a preset number of times, the first device can send a notification message to the second device, which can be used to notify the second device that the number of times that the reward value is less than the first threshold value reaches the preset value, or the notification message can be used to notify the second device that the number of times that the reward value is less than the first threshold value reaches the preset number of times, and the reward value less than the first threshold value in the inference process of the first model.

[0290] Exemplarily, the notification message can be periodically sent. For example, the notification message is used to notify the second device that the reward value in the inference process of the first model within a certain time granularity, and the relationship between the reward value and the first threshold value, such as the reward value is greater than the first threshold value, the reward value is partially greater than the first threshold value, etc. In order to save overhead, the notification message can be used to notify the second device that the reward value less than the first threshold value in the inference process of the first model within a certain time granularity.

[0291] As an example, the model training report can include the above notification message.

[0292] In some embodiments, if the first device is not the provider of model inference (i.e., the scenario where model inference and model training are deployed separately), the first device can forward the information related to model inference issued by the second device to the provider of model inference (referred to as the third device), and the first device can send the information related to model inference reported by the third device to the second device.

[0293] The information related to model inference issued by the second device may, for example, include the second information and / or the first threshold described above. The information related to model inference reported by the third device may, for example, include the model inference report and / or the notification message described above.

[0294] For example, the second device can send the second information and / or the first threshold described above to the first device, and the first device can send the second information and / or the first threshold described above to the third device.

[0295] For example, the third device can send the model inference report and / or the notification message described above to the first device, and the first device can send the model inference report and / or the notification message described above to the second device.

[0296] By reporting the performance indicators in the model inference process, it is helpful for the user of model training to understand the inference state of the model, to understand the accuracy of the model inference result, and to provide support for the management and monitoring of the model inference process.

[0297] It should be noted that the reward value mentioned in the embodiments of the present application can include a single reward value and / or a cumulative reward value, wherein the cumulative reward value can include one or more of the cumulative reward value within a certain time, the cumulative reward value in the entire model training process, or the cumulative reward value in the entire model inference process.

[0298] It should be understood that the reward value mentioned in the embodiments of the present application can be a reward value in the model training process, or a reward value in the model inference process.

[0299] It should be understood that the list mentioned in the embodiments of the present application is only an example, and the indicators in the indicator list and / or the parameters in the parameter list can also be stored in other forms, such as an array, which is not limited in the present application.

[0300] It should be noted that the method provided by the embodiments of the present application can also be applied to the ORAN architecture. In the ORAN architecture, the second device can be an SMO, the first device can be an MLTF, and the third device can be a network element mentioned above, such as a gNB or a NWDAF. Under the ORAN architecture, the actions performed by the user of model training (SMO), the provider of model training (MLTF), and the provider of model inference (network element) can be referred to the introduction in the foregoing. For the sake of brevity, it will not be repeated here.

[0301] For ease of understanding, the method provided by the embodiments of the present application will be introduced below with the first model as an example in combination with FIG. 7 to FIG. 9.

[0302] Scenario one of training and inference co-deployment

[0303] FIG. 7 is a flow diagram of another communication method provided by the embodiments of the present application. The method shown in FIG. 7 can involve the interaction between the first device and the second device, and belongs to the scenario of co-deployment of the provider of model training and the user of model training. Among them, the second device is the NMS, and the first device is the EMS.

[0304] The method shown in FIG. 7 can include steps 1 to 10.

[0305] In step 1, the NMS sends a capability query message to the EMS. The message is used to query the capability information related to the reinforcement learning model training.

[0306] In step 2, the EMS sends the capability information to the NMS.

[0307] Exemplarily, the capability information can include a list of adjustable KPIs, i.e., the third index list mentioned above. Alternatively, the capability information can also include a list of adjustable parameters, i.e., the second parameter list mentioned above. In this case, the capability information can also include the correspondence between the indexes in the third index list and the parameters in the second parameter list.

[0308] For example, the list of adjustable KPIs can include wireless network performance indicators such as RSRP, PRB utilization, handover success rate, etc., which can be used to determine the environment of reinforcement learning. For example, the adjustable parameters can be wireless network parameters such as antenna settings, cell settings, etc. that affect the above wireless network performance indicators, which can be used to select the behavior of reinforcement learning.

[0309] Exemplarily, the capability information can include the available training environment A supported by the EMS, such as NDT or the available live network environment.

[0310] In step 3, the NMS receives new function requirement information.

[0311] The function mentioned here is the function of the model, that is, the NMS receives new model training requirement information, and the model training requirement information includes the function of the model. Alternatively, the function requirement information can also include the second index list corresponding to the function, and the indexes in the list are the indexes expected to be used to determine the model training environment.

[0312] In step 4, the NMS determines the first parameter list and the first index list of the EMS.

[0313] Exemplarily, the NMS can determine the first indicator list of the EMS according to the third indicator list in step 2 (i.e., carried in the capability information) and the second indicator list in step 3 (i.e., carried in the function requirement information).

[0314] If the third indicator list in step 2 includes all indicators in the second indicator list in step 3, the first indicator list of the EMS can be the second indicator list in step 3.

[0315] If the third indicator list in step 2 includes part of the indicators in the second indicator list in step 3, the first indicator list of the EMS can include the part of the indicators.

[0316] Exemplarily, the NMS can determine the first parameter list and the first indicator list according to the correspondence between the parameters in the second parameter list in step 2 and the indicators in the third indicator list. For example, the first parameter list of the EMS can include the parameters that have the correspondence with the indicators in the first indicator list of the EMS.

[0317] The EMS can perform model training by using the reinforcement learning method according to the first parameter list and / or the first indicator list. For example, the indicators in the first indicator list are the indicators for determining the state of the environment for model training; and the parameters in the first parameter list are the parameters for selecting the behavior for model training.

[0318] If the NMS manages the model training of multiple EMSs, the NMS can determine the first parameter list and the first indicator list of the multiple EMSs respectively.

[0319] In step 5, the NMS initiates a model training request to the EMS.

[0320] Exemplarily, the model training uses the reinforcement learning method.

[0321] Exemplarily, the model training request can include the first information mentioned above, such as the function of the first model, the available training environment (i.e., the subset of the available training environment A mentioned above), the first indicator list, the first parameter list, and the performance indicator of the first model.

[0322] The available training environment can support the construction of the environment in the reinforcement learning model training process, such as the available training environment can be NDT or an available live network environment identifier.

[0323] The indicators in the first indicator list can be used to determine the state of the environment for model training, and / or the state information of the environment.

[0324] The parameters in the first parameter list can be used to determine the behavior of the first model, and / or the first parameter list can be used to determine the behavior information of the first model.

[0325] The first parameter list is optional. If the first parameter list is not included in the training request, the capability information reported in step 2 can not include the second parameter list.

[0326] In step 6, the EMS performs model training of the first model based on the model training request, such as the first information in the model training request.

[0327] In step 7, the EMS reports a model training report to the NMS.

[0328] Exemplarily, the training report can include a training environment used in the model training, such as an available live network environment or NDT used.

[0329] Exemplarily, the training report can include a performance indicator of the first model. The performance indicator of the first model can include, for example, a reward indicator and / or a reward function. As an example, the performance indicator of the first model can include one or more of a single reward value, a cumulative reward value, or an interaction record of action-environment-reward value. Optionally, the training report can include an interaction record of action-environment-reward value of the last few rounds.

[0330] In step 8, the NMS issues a threshold of an acceptable performance indicator, i.e., the first threshold mentioned above, to the EMS.

[0331] The first threshold can be, for example, a threshold of an acceptable reward value when the agent or the ML model constituting the agent makes an action. Exemplarily, the NMS can configure the first threshold according to the performance indicator of the first model obtained in step 6, such as the interaction record of action-environment-reward value. For example, the first threshold can be an average value of the reward values obtained in the last few rounds, or a certain proportion of the average value, such as 90% of the average value, etc.

[0332] In step 9, the NMS sends second information, i.e., activation information for performing model inference, to the EMS.

[0333] If the second information is received, the EMS can perform model inference of the first model based on the first threshold.

[0334] Exemplarily, in the model inference process, it can be determined whether the performance indicator of the model, such as the reward value, exceeds the first threshold. Optionally, it can be determined whether the reward value exceeds the threshold within a certain time granularity.

[0335] Exemplarily, if the second information includes the first threshold, step 8 and step 9 can be combined as follows: the NMS sends second information (i.e., activation information for performing model inference) to the EMS, and the second information includes the first threshold.

[0336] In step 10, the EMS sends a model inference report to the NMS.

[0337] Exemplarily, the inference report can include the performance indicators of the first model in the model inference process. For example, the inference report can include one or more of the reward value of the model in the model inference process, the reward value of the model within a certain time granularity, the size relationship (the reward value is greater than, less than or equal to the first threshold) between the reward value and the first threshold, or the reward value less than the first threshold.

[0338] In the embodiments of the present application, the NMS can query the reinforcement learning capability of the EMS, and configure a reinforcement learning training request according to the reinforcement learning capability of the EMS and the required function, which helps to improve the feasibility and reliability of the reinforcement learning model training, and helps to enable the reinforcement learning function in the 3GPP management domain.

[0339] Scenario two of training and inference co-deployment

[0340] FIG. 8 is a flow diagram of another communication method provided by the embodiments of the present application. The method shown in FIG. 8 can involve the interaction between the first device and the second device, and belongs to the scenario of co-deployment of the provider of model training and the user of model training. Among them, the second device is the NMS, and the first device is the EMS.

[0341] The main difference between the method shown in FIG. 8 and the method shown in FIG. 7 is that the method shown in FIG. 7 does not encapsulate the reinforcement learning capability as a function, and the method shown in FIG. 8 encapsulates the reinforcement learning capability as a function.

[0342] The method shown in FIG. 8 can include steps 1 to 9.

[0343] In step 1, the NMS sends a capability query message to the EMS. The message is used to query the capability information related to the reinforcement learning model training.

[0344] In step 2, the EMS sends the capability information to the NMS.

[0345] Exemplarily, the capability information can include the functions supported by the EMS, such as mobile load balancing function, coverage optimization function, etc.

[0346] Exemplarily, the capability information can include the available training environment A supported by the EMS, such as NDT or available live network environment.

[0347] In step 3, the NMS receives new function requirement information.

[0348] The function mentioned here is the function of the model, that is, the NMS receives new model training requirement information, which includes the function of the model. Optionally, the function requirement information can also include a second index list corresponding to the function, and the indexes in the list are indexes expected to be used to determine the model training environment.

[0349] In step 4, the NMS initiates a model training request to the EMS.

[0350] Exemplarily, the model training adopts a reinforcement learning method.

[0351] Exemplarily, the model training request can include the first information mentioned above, such as the function of the first model, the available training environment B (which is a subset of the available training environment A), and the performance index of the first model.

[0352] Optionally, the EMS can determine third information according to the function of the first model, and the third information includes one or more of the following: one or more indexes used to determine the state of the first environment; one or more parameters used to determine the behavior of the first model; the first environment; the behavior of the first model; state information of the first environment; behavior information of the first model; or the training environment of the first model.

[0353] In step 5, the EMS performs model training of the first model based on the model training request.

[0354] Exemplarily, the EMS determines the third information based on the first information in the model training request, and then performs model training of the first model based on the third information.

[0355] In step 6, the EMS reports a model training report to the NMS.

[0356] Exemplarily, the training report can include the training environment used in the model training (which is a subset of the available training environment B or the available training environment B), such as the available live network environment or the NDT.

[0357] Exemplarily, the training report can include the performance index of the first model. The performance index of the first model can include, for example, a reward index and / or a reward function. As an example, the performance index of the first model can include one or more of a single reward value, a cumulative reward value, or an interaction record of behavior-environment-reward value. Optionally, the training report can include an interaction record of behavior-environment-reward value of the last few rounds.

[0358] In step 7, the NMS issues a threshold of acceptable reward value, that is, the first threshold mentioned above, to the EMS.

[0359] The first threshold value may be, for example, a threshold value of a reward value that can be accepted when the agent or the ML model constituting the agent makes a behavior. Illustratively, the NMS can configure the first threshold value according to the performance indicator of the first model obtained in step 6, such as the interaction record of behavior-environment-reward value. For example, the first threshold value can be the average value of the reward values obtained in the last few rounds, or a certain proportion of the average value, such as 90% of the average value, etc.

[0360] In step 8, the NMS sends the second information, i.e., the activation information of performing model inference, to the EMS.

[0361] If the second information is received, the EMS can perform model inference of the first model based on the first threshold value.

[0362] Illustratively, in the model inference process, it can be judged whether the performance indicator of the model, such as the reward value, exceeds the first threshold value. Alternatively, it can be judged whether the reward value exceeds the threshold value within a certain time granularity.

[0363] Illustratively, if the first threshold value is included in the second information, steps 7 and 8 can be combined into: the NMS sends the second information (i.e., the activation information of performing model inference) to the EMS, and the second information includes the first threshold value.

[0364] In step 9, the EMS sends the model inference report to the NMS.

[0365] Illustratively, the inference report can include the performance indicator of the first model in the model inference process. For example, the inference report can include one or more of the reward value of the model in the model inference process, the reward value of the model within a certain time granularity, the size relationship between the reward value and the first threshold value (the reward value is greater than, less than or equal to the first threshold value), or the reward value less than the first threshold value.

[0366] In the embodiments of the present application, by encapsulating the reinforcement learning capability as a function, the second device does not need to care about the underlying information in the reinforcement learning model training process, which helps to reduce the processing complexity of the second device and helps to reduce the transmission overhead.

[0367] Scenario of separate deployment of training and inference (ORAN architecture)

[0368] FIG. 9 is a flow diagram of another communication method provided by the embodiments of the present application. The method shown in FIG. 9 can involve the interaction between the second device and the first device and the third device, and belongs to the scenario of separate deployment of the provider of model training and the user of model training. Among them, the second device is SMO, the first device is MLTF, and the third device is NE.

[0369] The method shown in FIG. 9 can include steps 1 to 10.

[0370] In step 1, the SMO sends a capability query message to the MLTF. The message is used to query capability information related to reinforcement learning model training.

[0371] In step 2, the MLTF sends the capability information to the SMO.

[0372] Exemplarily, the capability information can include a list of adjustable KPIs, i.e., the third list of indicators mentioned above. Optionally, the capability information can also include a list of adjustable parameters, i.e., the second list of parameters mentioned above. In this case, the capability information can also include a correspondence between the indicators in the third list of indicators and the parameters in the second list of parameters.

[0373] For example, the list of adjustable KPIs can include wireless network performance indicators such as RSRP, PRB utilization, handover success rate, etc., which can be used to determine the environment of reinforcement learning. For example, the adjustable parameters can be wireless network parameters such as antenna settings, cell settings, etc., which affect the above-mentioned wireless network performance indicators, and can be used to select the behavior of reinforcement learning.

[0374] Exemplarily, the capability information can include available training environments C supported by the MLTF, such as NDT or available live network environments.

[0375] In step 3, the SMO receives new function requirement information.

[0376] The function mentioned here is the function of the model, that is, the SMO receives new model training requirement information, which includes the function of the model. Optionally, the function requirement information can also include a second list of indicators corresponding to the function, the indicators in the list being indicators expected to be used to determine the model training environment.

[0377] In step 4, the SMO determines a first list of parameters and a first list of indicators of the MLTF.

[0378] Exemplarily, the SMO can determine the first list of indicators of the MLTF according to the third list of indicators in step 2 (i.e., carried in the capability information) and the second list of indicators in step 3 (i.e., carried in the function requirement information).

[0379] If the third list of indicators in step 2 includes all the indicators in the second list of indicators in step 3, the first list of indicators of the MLTF can be the second list of indicators in step 3.

[0380] If the third list of indicators in step 2 includes part of the indicators in the second list of indicators in step 3, the first list of indicators of the MLTF can include the part of the indicators.

[0381] Exemplarily, the SMO can determine the first parameter list and the first index list according to the correspondence between the parameters in the second parameter list in step 2 and the indexes in the third index list. For example, the first parameter list of the MLTF can include parameters that have a correspondence with the indexes in the first index list of the MLTF.

[0382] The MLTF can perform model training using a reinforcement learning method according to the first parameter list and / or the first index list. For example, the indexes in the first index list are indexes for determining the state of the environment for model training; and the parameters in the first parameter list are parameters for selecting the behavior of model training.

[0383] If the SMO manages the model training of multiple MLTFs, the SMO can determine the first parameter list and the first index list of the multiple MLTFs respectively.

[0384] In step 5, the SMO initiates a model training request to the MLTF.

[0385] Exemplarily, the model training uses a reinforcement learning method.

[0386] Exemplarily, the model training request can include the first information mentioned above, such as the function of the first model, the available training environment D (i.e., the first training environment mentioned above, which is a subset of the available training environment C), the first index list, the first parameter list, and the performance index of the first model, etc.

[0387] The available training environment can support the construction of the environment in the reinforcement learning model training process, such as the available training environment can be NDT or an available live network environment identifier.

[0388] The indexes in the first index list can be used to determine the state of the environment for model training, and / or the state information of the environment.

[0389] The parameters in the first parameter list can be used to determine the behavior of the first model, and / or the first parameter list can be used to determine the behavior information of the first model.

[0390] The first parameter list is optional. If the first parameter list is not included in the training request, the capability information reported in step 2 can not include the second parameter list.

[0391] In step 6, the MLTF performs model training of the first model based on the model training request, such as the first information in the model training request.

[0392] In step 7, the MLTF reports a model training report to the SMO.

[0393] Exemplarily, the training report can include a training environment adopted by the model training, such as an available online environment or NDT adopted.

[0394] Exemplarily, the training report can include a performance indicator of the first model. The performance indicator of the first model can include, for example, a reward indicator and / or a reward function. As an example, the performance indicator of the first model can include one or more of a single reward value, a cumulative reward value, or an interaction record of behavior-environment-reward value. Optionally, the training report can include an interaction record of behavior-environment-reward value of the last few rounds.

[0395] In step 8, the SMO sends, through the MLTF, a threshold of an acceptable performance indicator, i.e., the first threshold mentioned above, to the NE.

[0396] Exemplarily, the SMO can send the first threshold to the MLTF, and the MLTF can send the first threshold to the NE.

[0397] The first threshold can be, for example, a threshold of an acceptable reward value when the agent or the ML model constituting the agent makes a behavior. Exemplarily, the NMS can configure the first threshold according to the performance indicator of the first model obtained in step 6, such as the interaction record of behavior-environment-reward value. For example, the first threshold can be an average value of the reward values obtained in the last few rounds, or a certain proportion of the average value, such as 80% of the average value, etc.

[0398] In step 9, the SMO sends, through the MLTF, second information, i.e., activation information for performing model inference, to the NE.

[0399] Exemplarily, the SMO can send the second information to the MLTF, and the MLTF can send the second information to the NE.

[0400] Upon receiving the second information, the NE can perform model inference of the first model based on the first threshold.

[0401] Exemplarily, in the process of model inference, it can be determined whether the performance indicator of the model, such as the reward value, exceeds the first threshold. Optionally, it can be determined whether the reward value exceeds the threshold within a certain time granularity.

[0402] Exemplarily, if the second information includes the first threshold, step 8 and step 9 can be combined into: the SMO sends, through the MLTF, second information (i.e., activation information for performing model inference) to the NE, and the second information includes the first threshold.

[0403] In step 10, the NE sends a model inference report to the SMO.

[0404] Exemplarily, the NE sends a model inference report to the MLTF, and the MLTF can send the model inference report to the SMO.

[0405] Exemplarily, the inference report can include a performance indicator of the first model in the model inference process. For example, the inference report can include one or more of a reward value of the model in the model inference process, a reward value of the model within a certain time granularity, a size relationship (the reward value is greater than, less than or equal to the first threshold) between the reward value and the first threshold, or a reward value less than the first threshold.

[0406] In the embodiments of the present application, the interface between the SMO and the MLTF adds the functions of querying and reporting reinforcement learning capability; the training request supports reinforcement learning information configuration, including available training environment, index list, etc.; the training report adds an indication of the used training environment, which is helpful to enable the reinforcement learning function in the ORAN architecture. In addition, the performance indicator requirement is added in the NE reinforcement learning model inference configuration, which is helpful to improve the reliability of model inference.

[0407] It should be understood that the methods shown in FIGS. 7 to 9 are only specific examples in specific scenarios and should not determine any limitation on the present application.

[0408] The method embodiments provided by the present application are described above, and the device embodiments provided by the present application will be described below. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments, and therefore, the content not described in detail can be referred to the method embodiments described above, and will not be described here for brevity.

[0409] FIG. 10 is a schematic block diagram of a communication device according to an embodiment of the present application. As shown in FIG. 10, the communication device 1000 can include a transceiver unit 1010 and / or a processing unit 1020. The transceiver unit 1010 can implement corresponding communication functions, and the processing unit 1020 is configured to perform data processing. The transceiver unit 1010 can also be referred to as a communication interface or a communication unit. Optionally, the device 1000 can also include a storage unit, which can be used to store instructions and / or data, and the processing unit 1020 can read the instructions and / or data in the storage unit to enable the device to implement the foregoing method embodiments.

[0410] In a possible design, the device 1000 can be the second device in the method embodiments above, or can be a chip, processor or chip system for implementing the function of the second device.

[0411] Specifically, the processing unit 1020 can be configured to determine first information, the first information being used for training a first model, and a training method of the first model being reinforcement learning. The transceiver unit 1010 can be configured to send the first information to a first device, the first device being a device for training the first model.

[0412] In some embodiments, during the model training, the first model interacts with a first environment, and the first information comprises one or more of the following: a first indicator list, a state of the first environment is determined by one or more indicators in the first indicator list; a first parameter list, a behavior of the first model is determined by one or more parameters in the first parameter list; a first training environment, a training environment of the first model is the first training environment, or a training environment of the first model is one of the first training environment; a function of the first model; state information of the first environment; behavior information of the first model; or a performance indicator of the first model.

[0413] In some embodiments, the first information comprises the first indicator list, and the determining the first information comprises determining the first information based on model training requirement information, the model training requirement information comprising a second indicator list, indicators in the second indicator list being expected to be used to determine the first environment; and wherein the first indicator list is the second indicator list.

[0414] In some embodiments, the capability information of the first device comprises a third indicator list, the third indicator list comprising indicators supported by the first device when determining an interaction environment of a model, and the determining the first information based on the model training requirement information comprises determining the first indicator list based on the model training requirement information and the capability information of the first device, wherein: if the third indicator list comprises the second indicator list, the first indicator list is the second indicator list; and if the third indicator list comprises part of the indicators in the second indicator list, the first indicator list comprises indicators common to the second indicator list and the third indicator list.

[0415] In some embodiments, the first information comprises a function of the first model, and the function of the first model is determined based on model training requirement information.

[0416] In some embodiments, the first information comprises the first parameter list, and the capability information of the first device comprises a correspondence between indicators in a third indicator list and parameters in a second parameter list, wherein the third indicator list comprises indicators supported by the first device when determining an interaction environment of a model, and the second parameter list comprises parameters supported by the first device when determining a behavior of a model, and the determining the first information comprises determining the first parameter list based on indicators in the first indicator list and the correspondence, wherein the first parameter list comprises parameters having the correspondence with the indicators in the first indicator list.

[0417] In some embodiments, the processing unit 1020 can be configured to obtain capability information of the first device, the capability information of the first device being related to reinforcement learning, wherein the capability information of the first device comprises one or more of: a first function, the first device supporting training of a model implementing the first function; a second environment, the second environment being an environment supported by the first device, the second environment being used for training of the model; a third index list, the third index list comprising indexes supported by the first device for use in determining an interaction environment of the model; a second parameter list, the second parameter list comprising parameters supported by the first device for use in determining a behavior of the model; or a correspondence between an index in the third index list and a parameter in the second parameter list.

[0418] In some embodiments, the sending of the first information to the first device comprises sending a first request to the first device, the first request being used to request training of the first model, wherein the first request comprises a first attribute, the first attribute being used to indicate the first information.

[0419] In some embodiments, the transceiver 1010 can be configured to receive a model training report, the model training report comprising a second attribute and / or a third attribute, the second attribute being used to indicate a training environment of the first model, and the third attribute being used to indicate a performance indicator of the first model.

[0420] In some embodiments, the transceiver 1010 can be configured to send a first threshold to the first device, the first threshold being used to indicate an acceptable performance indicator of the first model in a model inference process.

[0421] In some embodiments, the transceiver 1010 can be configured to send second information to the first device, the second information being used to indicate execution of model inference based on the first model, and receive a model inference report, the model inference report comprising a performance indicator of the first model in a model inference process.

[0422] In a possible design, the apparatus 1000 can be the first device in the above method embodiments, or can be a chip, processor, or chip system that implements functions of the first device.

[0423] Specifically, the transceiver 1010 can be configured to receive first information, the first information being used to train a first model, a training method of the first model being reinforcement learning. The processing unit 1020 can be configured to train the first model based on the first information.

[0424] In some embodiments, during the model training, the first model interacts with a first environment, and the first information includes one or more of the following: a first indicator list, a state of the first environment being determined by one or more indicators in the first indicator list; a first parameter list, a behavior of the first model being determined by one or more parameters in the first parameter list; a first training environment, a training environment of the first model being the first training environment or being one of the first training environment; a function of the first model; state information of the first environment; behavior information of the first model; or a performance indicator of the first model.

[0425] In some embodiments, the first information includes a function of the first model, and the training the first model based on the first information includes: determining third information based on the function of the first model; and training the first model based on the third information, wherein the third information includes one or more of the following: one or more indicators for determining a state of the first environment; one or more parameters for determining a behavior of the first model; the first environment; the behavior of the first model; state information of the first environment; behavior information of the first model; or a training environment of the first model.

[0426] In some embodiments, the transceiver 1010 can be configured to: notify capability information of the first device, the capability information of the first device being related to reinforcement learning, wherein the capability information of the first device includes one or more of the following: a first function, the first device supporting training of a model implementing the first function; a second environment, the second environment being supported by the first device and being used for training of a model; a third indicator list, the third indicator list including indicators supported by the first device when determining an interaction environment of a model; a second parameter list, the second parameter list including parameters supported by the first device when determining a behavior of a model; or a correspondence between indicators in the third indicator list and parameters in the second parameter list.

[0427] In some embodiments, the receiving the first information includes: receiving a first request, the first request being used to request training of the first model; and wherein the first request includes a first attribute, the first attribute being used to indicate the first information.

[0428] In some embodiments, the transceiver 1010 can be configured to: send a model training report to a second device, the model training report including a second attribute and / or a third attribute, the second attribute being used to indicate a training environment of the first model, and the third attribute being used to indicate a performance indicator of the first model.

[0429] In some embodiments, the transceiver 1010 can be configured to receive a first threshold, the first threshold being used to indicate an acceptable performance indicator of the first model in a model inference process.

[0430] In some embodiments, the transceiver 1010 can be configured to receive second information, the second information being used to indicate that a model inference based on the first model is performed; and send a model inference report to a second device, the model inference report comprising a performance indicator of the first model in a model inference process.

[0431] It should be understood that the division of units in the above apparatus is only a logical function division, one function unit can correspond to one function, or two or more functions can be integrated into one function unit. In actual implementation, all or part of the units can be integrated into one physical entity, or distributed in different physical entities. In addition, the above function units can be realized in the form of hardware, or in the form of software, or in the form of hardware combined with software. Whether a certain function is executed in hardware or software depends on the specific application and design constraints of the technical scheme. Professional technicians can use different methods to implement the described functions for specific applications, but such implementation should not be considered beyond the scope of the present application.

[0432] It should be understood that the "units" in the apparatus 1000 can be implemented by hardware, or by software, or by hardware executing corresponding software. For example, the "units" can refer to application specific integrated circuits (ASIC), electronic circuits, processors (such as shared processors, dedicated processors, or group processors, etc.) and memories for executing one or more software or firmware programs, integrated logic circuits, and / or other suitable components supporting the described functions. For another example, the transceiver 1010 can be replaced by a transceiver circuit (which can include a receiving circuit and a transmitting circuit), and the processing unit 1020 can be replaced by a processor or a processing circuit.

[0433] FIG. 11 shows a schematic block diagram of another communication apparatus provided by the embodiments of the present application. The communication apparatus 1100 can be a first device or a second device, or a chip, a chip system, or a processor, etc. implemented in the first device or the second device to implement the above method. The apparatus can be used to implement the method described in the above method embodiments, and specific implementation can be referred to the description in the above method embodiments.

[0434] The communication device 1100 can include one or more processors 1110, which can also be referred to as processing units, to implement certain control functions. The processor 1110 can be a general processor or a special purpose processor, etc. For example, it can be a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, and the central processing unit can be used to control the communication device, execute software programs, and process data of the software programs.

[0435] In an alternative design, the processor 1110 can also store instructions and / or data, which can be executed by the processor 1110, so that the communication device 1100 performs the methods described in the above method embodiments.

[0436] In another alternative design, the communication device 1100 can include a communication interface 1120 for implementing receiving and transmitting functions. For example, the communication interface 1120 can be a transceiver circuit, an interface, an interface circuit, or a transceiver, etc. The transceiver circuit, the interface, the interface circuit, or the transceiver for implementing receiving and transmitting functions can be separate or integrated together. The above transceiver circuit, the interface, the interface circuit, or the transceiver can be used for reading and writing of codes / data, or the above transceiver circuit, the interface, the interface circuit, or the transceiver can be used for transmission or transfer of signals.

[0437] Optionally, the communication device 1100 can include one or more memories 1130, which can store instructions executable by the processor 1110, so that the communication device 1100 performs the methods described in the above method embodiments. Optionally, the memory 1130 can also store data. Optionally, the processor 1110 can also store instructions and / or data. The processor 1110 and the memory 1130 can be separately provided or integrated together.

[0438] It should be understood that, in a possible design, the steps in the method embodiments provided by the embodiments of the present application can be completed by integrated logic circuits of hardware in the processor or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being completed by a hardware processor, or completed by a combination of hardware and software modules in the processor. The software modules can be located in random access memories, flash memories, read-only memories, programmable read-only memories, or electrically erasable programmable memories, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads information in the memory and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0439] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with a signal processing capability. In the implementation process, the steps of the above method embodiments can be completed by an integrated logic circuit or an instruction in the form of software in the processor. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the storage, and the processor reads the information in the storage, and combines the hardware to complete the steps of the above method.

[0440] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM). It should be noted that the memory of the system and method described herein is intended to include but not limited to these and any other suitable types of memory.

[0441] The various device embodiments and method embodiments described above are fully corresponding, and the corresponding steps are performed by corresponding modules or units, for example, the communication unit or communication interface performs the steps of receiving or sending in the method embodiments, and other steps except sending and receiving can be performed by the processing unit or processor.

[0442] The present application also provides a computer program product, comprising computer program instructions, which, when executed, cause the steps or processes performed by the first device or the second device in any of the method embodiments described above to be performed.

[0443] The present application also provides a computer-readable storage medium, which stores a computer program or instructions, which, when executed, cause the steps or processes performed by the first device or the second device in any of the method embodiments described above to be performed.

[0444] The present application also provides a chip, comprising a processor, for calling and running a computer program or instructions from a memory, which, when executed, cause the steps or processes performed by the first device or the second device in any of the method embodiments described above to be performed.

[0445] The present application also provides a communication system, comprising at least one of the first device or the second device.

[0446] Those skilled in the art should understand that embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.

[0447] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the computer or other programmable data processing apparatus produce the functions specified in one or more flows in the flowcharts and / or one or more blocks in the block diagrams.

[0448] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart or flowsheets and / or block or blocks of the block diagrams.

[0449] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart or flowsheets and / or block or blocks of the block diagrams.

[0450] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A communication method characterized by comprising: The method comprises: determining first information, the first information being used for training a first model, the training method of the first model being reinforcement learning; sending the first information to a first device, the first device being a device for training the first model.

2. The method of claim 1, wherein, In the process of model training, the first model interacts with a first environment, and the first information comprises one or more of the following: a first index list, a state of the first environment being determined by one or more indexes in the first index list; a first parameter list, a behavior of the first model being determined by one or more parameters in the first parameter list; a first training environment, the training environment of the first model being the first training environment, or the training environment of the first model being one of the first training environment; a function of the first model; state information of the first environment; behavior information of the first model; or a performance index of the first model.

3. The method of claim 2, wherein, The first information comprises the first index list, and the determining of the first information comprises: determining the first information based on model training requirement information, the model training requirement information comprising a second index list, indexes in the second index list being indexes expected to be used for determining the first environment; wherein the first index list is the second index list.

4. The method of claim 3, wherein, The capability information of the first device comprises a third index list, the third index list comprising indexes supported by the first device when determining an interaction environment of a model, the determining of the first information based on the model training requirement information comprises: determining the first index list based on the model training requirement information and the capability information of the first device, wherein: if the third index list comprises the second index list, the first index list is the second index list; if the third index list comprises part of the indexes in the second index list, the first index list comprises indexes common to the second index list and the third index list.

5. The method according to any one of claims 2-4, characterized in that, The first information indicates a function of the first model, and the function of the first model is determined based on model training requirement information.

6. The method according to any one of claims 2-5, characterized in that, The first information comprises the first parameter list, and the capability information of the first device comprises a correspondence between indexes in a third index list and parameters in a second parameter list, wherein the third index list comprises indexes supported by the first device when determining an interaction environment of a model, and the second parameter list comprises parameters supported by the first device when determining a behavior of a model, the determining of the first information comprises: the first parameter list being determined based on indexes in the first index list and the correspondence, wherein the first parameter list comprises parameters having the correspondence with the indexes in the first index list.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: obtaining capability information of the first device, the capability information of the first device being related to reinforcement learning, wherein the capability information of the first device comprises one or more of the following: a first function, the first device supporting training of a model implementing the first function; a second environment, the environment supported by the first device, the second environment being used for training of the model; a third index list, the third index list comprising indexes used by the first device to determine an interaction environment of the model; a second parameter list, the second parameter list comprising parameters supported by the first device to determine a behavior of the model; or a correspondence between the indexes in the third index list and the parameters in the second parameter list.

8. The method according to any one of claims 1-7, characterized in that, The sending of the first information to the first device comprises: sending a first request to the first device, the first request being used to request training of the first model; wherein the first request comprises a first attribute, the first attribute being used to indicate the first information.

9. The method according to any one of claims 1-8, characterized in that, The method further comprises: receiving a model training report, the model training report comprising a second attribute and / or a third attribute, the second attribute being used to indicate a training environment of the first model, and the third attribute being used to indicate a performance index of the first model.

10. The method according to any one of claims 1-9, characterized in that, The method further comprises: sending a first threshold to the first device, the first threshold being used to indicate an acceptable performance index of the first model in a model inference process.

11. The method according to any one of claims 1-9, characterized in that, The method further comprises: sending second information to the first device, the second information being used to indicate execution of a model inference based on the first model; receiving a model inference report, the model inference report comprising a performance index of the first model in a model inference process.

12. A communication method characterized by comprising: comprises: receiving first information, the first information being used to train a first model, a training method of the first model being reinforcement learning; training the first model based on the first information.

13. The method of claim 12, wherein, In a process of model training, the first model interacts with a first environment, and the first information comprises one or more of the following: a first index list, a state of the first environment being determined by one or more indexes in the first index list; a first parameter list, a behavior of the first model being determined by one or more parameters in the first parameter list; a first training environment, a training environment of the first model being the first training environment, or a training environment of the first model being one of the first training environment; a function of the first model; state information of the first environment; behavior information of the first model; or a performance index of the first model.

14. The method of claim 13, wherein, The first information comprises a function of the first model, and the training of the first model based on the first information comprises: determining third information based on the function of the first model; training the first model based on the third information, wherein the third information comprises one or more of the following: one or more indexes used to determine a state of the first environment; one or more parameters used to determine a behavior of the first model; the first environment; the behavior of the first model; state information of the first environment; behavior information of the first model; or a training environment of the first model.

15. The method according to any one of claims 12-14, characterized in that, The method further comprises: informing the first device of capability information related to reinforcement learning, wherein the capability information of the first device comprises one or more of the following: a first function, the first device supporting training of a model implementing the first function; a second environment, the first device supporting an environment for training of the model; a third list of indicators, the third list of indicators comprising indicators used by the first device in determining an interaction environment of the model; a second list of parameters, the second list of parameters comprising parameters supported by the first device in determining a behavior of the model; or a correspondence between the indicators in the third list of indicators and the parameters in the second list of parameters.

16. The method according to any one of claims 12-15, characterized in that, The receiving the first information comprises: receiving a first request, the first request being for requesting training of the first model; wherein the first request comprises a first attribute, the first attribute being for indicating the first information.

17. The method according to any one of claims 12-16, characterized by, The method further comprises: sending, to a second device, a model training report, the model training report comprising a second attribute and / or a third attribute, the second attribute being for indicating a training environment of the first model, the third attribute being for indicating a performance indicator of the first model.

18. The method according to any one of claims 12-17, characterized by, The method further comprises: receiving a first threshold, the first threshold being for indicating an acceptable performance indicator of the first model in a model inference process.

19. The method according to any one of claims 12-18, characterized by, The method further comprises: receiving second information, the second information being for indicating execution of a model inference based on the first model; sending, to a second device, a model inference report, the model inference report comprising a performance indicator of the first model in a model inference process.

20. A communications device, characterized by comprising means for performing the method of any one of claims 1-11.

21. A communications device, characterized by comprising means for performing the method of any one of claims 12-19.

22. A communications device, characterized by comprising a processor that, when executing a program or instructions, causes the method of any one of claims 1-11 to be performed, or causes the method of any one of claims 12-19 to be performed.

23. A readable storage medium, on which a computer program or instructions are stored, characterized in that, The computer program or instructions, when executed, cause the method of any one of claims 1-11 to be performed, or cause the method of any one of claims 12-19 to be performed.

24. A computer program product, characterised in that, comprising computer program instructions that, when executed, cause the method of any one of claims 1-11 to be performed, or cause the method of any one of claims 12-19 to be performed.

25. A chip, characterized by comprising a processor for invoking and running a computer program from a memory, such that the method of any one of claims 1-11 is performed, or such that the method of any one of claims 12-19 is performed.

Citation Information

Patent Citations

  • Communication method and communication device

    CN116233857A

  • Reinforcement learning method and device

    CN117474121A

  • Information transmission method, device and system

    CN117764146A

  • Fluid recovery in semiconductor processing

    KR1020220131395A