Communication method and apparatus

By determining performance and time requirements in the communication network, selecting appropriate training parameters and network elements, and monitoring training progress, the impact of machine learning training on network performance was resolved, improving the efficiency and rationality of training and ensuring the quality of user service.

WO2026067119A1PCT designated stage Publication Date: 2026-04-02HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

When machine learning training is performed in a real communication network, it can affect the network performance and lead to a decline in the quality of service for users. It is difficult to balance the impact of training efficiency and network performance.

Method used

By determining the initial information, the performance and time requirements for machine learning training are indicated, and machine learning training is performed to meet the requirements of the communication network, including selecting appropriate training parameters and network elements, monitoring training progress, and adjusting strategies.

Benefits of technology

It achieves a balance between the efficiency of machine learning training and its impact on network performance while ensuring the quality of service for users of the communication network, thereby improving the rationality and flexibility of training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025121468_02042026_PF_FP_ABST
    Figure CN2025121468_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A communication method and a communication apparatus. According to the present application, the method comprises: determining first information, and on the basis of the first information, executing machine learning training, wherein the first information is used for indicating a first performance requirement, and the first performance requirement indicates requirements for communication network performance for executing the machine learning training. Thus, the efficiency of machine learning training and the degree of impact on the communication network performance of a communication network can be balanced, thereby ensuring the quality of service for communication network users.
Need to check novelty before this filing date? Find Prior Art

Description

Communication method and apparatus

[0001] The present application claims priority to the Chinese patent application No. 202411392437.4, filed on September 30, 2024, with the State Intellectual Property Office of China, and the Chinese patent application No. 202411392437.4 has the title of “Communication method and apparatus”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of communication, and more particularly, to a communication method and apparatus. BACKGROUND

[0003] In order to improve the intelligent and automated level of the network, artificial intelligence (AI) and machine learning (ML) technologies are being applied in more and more fields.

[0004] In order to ensure the training efficiency and the performance of the trained model, the machine learning process can perform experiments and learn from mistakes in the real communication network. However, this will have an impact on the communication network performance of the real communication network, resulting in a decline in the service quality of the communication network users. SUMMARY

[0005] The embodiments of the present application provide a communication method and apparatus, which can balance the machine learning training efficiency and the degree of impact on the communication network performance of the communication network, and ensure the service quality of the communication network users.

[0006] The communication network to which the embodiments of the present application are applicable can include a real communication network (for example, a network specified by the 3rd Generation Partnership Project (3GPP) or open RAN (ORAN)), or the communication network to which the embodiments of the present application are applicable can include a simulation network simulating a real network. The embodiments of the present application do not limit this.

[0007] In the present application, “communication network performance” and “network performance” can be used interchangeably, and they represent the same meaning unless otherwise specified.

[0008] In the scheme provided by the embodiments of the present application, the inference function can refer to a function associated with an AI / ML inference name (aimlinferencename) (for example, the inference function can include a function corresponding to the values of the management data analytics type (the values of the MDA type), or a function corresponding to the analytics identity(s) of network data analytics function (Analytics ID(s) of NWDAF), etc.).

[0009] For example, the aimlinferencename can include allowedValues, the allowedValues can include vendor's specific extensions, and any one of the following three: the values of the management data analytics (MDA) type (for example, see the protocol 3GPP TS 28.104); the analytics identity (identity, ID) of the network data analytics function (NWDAF) (Analytics ID(s) of NWDAF) (for example, see the protocol 3GPP TS 23.288); and the types of inference for the radio access network (RAN).

[0010] In a first aspect, a communication method is provided for machine learning training. The method can be performed by a first apparatus capable of performing machine learning training. The first apparatus can refer to a device on a network device side (e.g., a network device), a component in the device (e.g., a communication module, a processor, a circuit, a chip, or a chip system, etc.), or a logic module or software capable of implementing all or part of the functions of the communication device. The network device side can include at least one of a network device or an AI entity on the network device side. The AI entity on the network device side can be the network device itself or an AI entity serving the network device, such as a RAN intelligent controller (RIC), an operation administration and maintenance (OAM), or a server, such as an OTT server or a cloud server, etc. Communication between servers can be achieved through a communication link between a terminal device and a network device, or through other communication devices outside the servers, or through a wired link.

[0011] For ease of description, the method performed by the first apparatus is described below as an example.

[0012] The method includes determining first information, the first information being used to indicate a first performance requirement, the first performance requirement indicating a requirement of performing machine learning training on a communication network performance; and performing machine learning training based on the first information.

[0013] Specifically, after determining the first information used to indicate the first performance requirement, the first apparatus can perform machine learning training according to the requirement of performing machine learning training on the communication network performance indicated by the first performance requirement.

[0014] For example, the network element performing the machine learning training satisfies the first performance requirement.

[0015] For example, the first apparatus can be a device where a management service producer (MnS producer) is located. For example, the first apparatus can be a device where an element management system (EMS) is located, or a device where a base station gNB is located, or a RIC.

[0016] In some possible implementations, the first apparatus described above can be referred to as: a production entity; an EMS; a domain management node / system / function; an MnS producer; a single-domain management system (domain management system, for short); or a single-domain management function (domain management function, for short). Correspondingly, an apparatus where a management service consumer / user / caller (management service consumer, MnS consumer, for short) that jointly implements machine learning training with the first apparatus described above can be referred to as: a consumption entity; a network management system (network management system, NMS, for short); a cross-domain management node / system / function; an MnS consumer; a cross-domain management system (Cross-domain management system, CD-MnS, for short); or a cross-domain management function (Cross-domain management function, CD-MnF, for short); a service management and orchestration (function) (service management and orchestration, SMO, for short).

[0017] For example, determining the first information can be determining the first information according to the received indication information, for example, the indication information received by the first apparatus indicates the first information, and the first apparatus can determine the first information according to the received indication information; or determining the first information can also be determining the first information according to a requirement to be met by a key performance indicator (key performance indicator, KPI, for short) or a performance measurement (performance measurement, PM, for short) of the related communication network, and / or a current performance status of the related communication network, for example, the first apparatus can determine the requirement for the communication network performance of the machine learning training by calculation according to the requirement to be met by the KPI or the PM of the related communication network, and / or the current performance status of the related communication network. The embodiments of the present application do not make any limitation in this regard.

[0018] The first information can be received.

[0019] For example, the method comprises: receiving first information, the first information being used to indicate a first performance requirement, the first performance requirement indicating a requirement for the communication network performance of the machine learning training; and performing the machine learning training based on the first information.

[0020] The determining the first information can be understood as that the MnS producer receives information configured for the MnS consumer to create a machine learning training request management object instance (MLTrainingRequest MOI).

[0021] For example, a property related to the performance requirement of the existing network can be added in a machine learning training request information object class (MLTrainingRequest IOC).

[0022] For example, the property related to the performance requirement of the existing network can include a reinforcement learning network performance requirement (rlnetworkperformancerequirement).

[0023] The performing the machine learning training based on the first information can be understood as that the MnS producer receives a create management object instance (MOI) request of the MLTrianingRequest and creates the MLTrianingRequest MOI.

[0024] For example, the first device can calculate a value capable of meeting the minimum KPI or PM requirement and the current communication network condition based on the minimum KPI or PM requirement and the current communication network condition, and the calculated value is used to represent the requirement of the machine learning training on the communication network performance. The minimum KPI or PM requirement and the current communication network condition are related to the inference function of the machine learning corresponding to the machine learning training.

[0025] Since the efficiency of the machine learning training and the degree of influence on the communication network performance of the real communication network are mutually restricted, by performing the machine learning training based on the requirement of the machine learning training on the communication network performance, on the one hand, the machine learning training efficiency is too low or the model obtained by the machine learning training is out of date due to the too long training time of the machine learning training can be avoided; on the other hand, the service quality of the communication network users does not meet the requirement due to the too great influence of the machine learning training on the communication network performance of the real communication network can be avoided.

[0026] Based on the scheme provided in the embodiments of the present application, by performing machine learning training based on the first information indicating the first performance requirement, the machine learning training efficiency and the degree of influence on the communication network performance can be balanced, and the service quality of the communication network users can be ensured.

[0027] In some possible implementation manners, the requirement of the machine learning training on the communication network performance can include at least one of the following: a lower threshold of an allowed KPI or PM; or, a loss value of an allowed KPI or PM loss; or, a ratio of an allowed KPI or PM loss; or, a range of an allowed network fluctuation.

[0028] For example, the requirement of the machine learning training on the communication network performance can include: a plurality of KPIs and a KPI lower threshold / range of deviation / maximum loss ratio corresponding to each KPI (at this time, the requirement of the machine learning training on the communication network performance can also be information in a list form). Alternatively, the requirement of the machine learning training on the communication network performance can include: a single KPI and a KPI lower threshold / range of deviation / maximum loss ratio corresponding to the KPI; or, a plurality of PMs and a PM lower threshold / range of deviation / maximum loss ratio corresponding to each PM (at this time, the requirement of the machine learning training on the communication network performance can also be information in a list form); or, the requirement of the machine learning training on the communication network performance can include: a single PM and a PM lower threshold / range of deviation / maximum loss ratio corresponding to the PM.

[0029] The influence of the machine learning training based on the first information on the communication network performance can meet: an allowed KPI lower threshold; or, an allowed KPI loss value; or, an allowed KPI loss ratio; or, an allowed network fluctuation range; or, an allowed PM lower threshold; or, an allowed PM loss value; or, an allowed PM loss ratio.

[0030] In some possible implementation manners, the first information is further used to indicate a first training time requirement, and the first training time requirement indicates a requirement of the machine learning training on a training time.

[0031] Specifically, after the first information used to indicate the first performance requirement and the first training time requirement is determined, the first device can perform the machine learning training according to the requirement of the machine learning training on the communication network performance and the requirement of the machine learning training on the training time indicated by the first performance requirement.

[0032] Exemplarily, the first training time requirement includes information from a training request for a model corresponding to a reinforcement learning function received by a management service consumer (MnS consumer) in a network, and a training time requirement corresponding to the training request received by the MnS consumer (for example, the training time requirement corresponding to the training request can be the fourth information).

[0033] For example, when the network management system NMS is the MnS consumer, it can receive a training request for a model corresponding to a reinforcement learning function in the network and a training time requirement corresponding to the training request. When the MnS consumer initiates a training request to the MnS producer, the first training time requirement sent by the MnS consumer can include information of the training time requirement corresponding to the training request.

[0034] Based on the scheme provided in the embodiments of the present application, by performing machine learning training based on the first information indicating the first training time requirement, the training time of the machine learning training can meet the expected time requirement, so that the implementation of the machine learning training is more reasonable.

[0035] In some possible implementation manners, the machine learning training is reinforcement learning training, and the first training time requirement includes a training time point at which the reinforcement learning training is expected to be completed and / or a training duration.

[0036] In some possible implementation manners, the network element performing the machine learning training meets the first performance requirement and the first training time requirement.

[0037] In some possible implementation manners, the requirement for the training time of the machine learning training can include at least one of the following: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0038] The training time of the machine learning training performed based on the first information can meet the following requirements: a time point at which the machine learning training is expected to be completed; or an expected training duration, and the like.

[0039] In some possible implementation manners, performing the machine learning training based on the first information includes performing the machine learning training based on a first parameter, where the first parameter is determined according to the first performance requirement and a corresponding relationship, the corresponding relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and a first value corresponding to the first parameter meets the first performance requirement.

[0040] Exemplarily, the training parameter includes a learning rate, or the learning rate and a decay ratio of the learning rate.

[0041] For example, the training parameter includes an exploration rate, or the exploration rate and a decay ratio of the exploration rate.

[0042] For example, when the machine learning training is reinforcement learning (RL) training, the machine learning is performed according to a learning rate (or the learning rate and a decay ratio of the learning rate), or according to an exploration rate (or the exploration rate and a decay ratio of the exploration rate), and when the learning rate or the exploration rate decays to 0, the machine learning training no longer obtains a new action.

[0043] For example, when the EMS is the MnS producer, the EMS can select a training parameter for performing the machine learning training as a first parameter in the training parameters according to the correspondence relationship between the existing first value and the training parameter, and the first value satisfying the correspondence relationship with the first parameter satisfies the first performance requirement. For example, the first value includes a KPI / PM loss value of the network when the machine learning training is performed based on the training parameter, a ratio of the KPI / PM loss, or a range of KPI / PM fluctuation.

[0044] For example, when the gNB is the MnS producer, the EMS can select a training parameter for performing the machine learning training as a first parameter in the training parameters according to the correspondence relationship between the existing first value and the training parameter, and the first value satisfying the correspondence relationship with the first parameter satisfies the first performance requirement. The EMS can send indication information indicating the first parameter determined by the EMS to the gNB, and the gNB can determine the first parameter according to the received indication information and perform the machine learning training according to the first parameter. At this time, the device where the EMS and the gNB are located can be one possible implementation of the first device.

[0045] For example, the first performance requirement is that the handover success rate loss does not exceed 5%, and in the correspondence relationship between the existing first value and the training parameter, when the exploration rate is 10%, 20% or 30%, the handover success rate loss corresponding to the exploration rate is 3%, 5% or 7% respectively. Then the exploration rate of 20% can be selected as the first parameter.

[0046] In some possible implementations, performing the machine learning training based on the first information includes: performing the machine learning training based on the first parameter, wherein the first parameter is determined according to the first performance requirement and the correspondence relationship, the correspondence relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first performance requirement and the first training time requirement.

[0047] For example, when the EMS is the MnS producer, the EMS can select a training parameter for performing the machine learning training as the first parameter according to the correspondence between the existing first values and the training parameters, and the first value corresponding to the first parameter satisfies the first performance requirement and the first training time requirement. For example, the first value includes the KPI / PM loss value, the ratio of the KPI / PM loss, or the range of the KPI / PM fluctuation of the network based on the training parameter performing the machine learning training, and the time point at which the machine learning training is expected to be completed or the expected training duration.

[0048] For example, when the gNB is the MnS producer, the EMS can select a training parameter for performing the machine learning training as the first parameter according to the correspondence between the existing first values and the training parameters, and the first value corresponding to the first parameter satisfies the first performance requirement and the first training time requirement. The EMS can send the indication information indicating the determined first parameter to the gNB, and the gNB can determine the first parameter according to the received indication information and perform the machine learning training according to the first parameter. At this time, the device where the EMS and the gNB are located can be one possible implementation of the first device.

[0049] For example, the first performance requirement is that the handover success rate loss does not exceed 5%, the first training time requirement is that the training duration does not exceed 200ms, and in the correspondence between the existing first values and the training parameters, when the exploration rate is 10%, 20% or 30%, the handover success rate loss corresponding to the exploration rate is 3%, 5% or 7%, and the training duration corresponding to the exploration rate is 300ms, 200ms or 100ms. Then the exploration rate of 20% can be selected as the first parameter.

[0050] In some possible implementations, the first information is further used to indicate a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

[0051] For example, the network element performing the machine learning training can belong to the plurality of first network elements indicated by the first information, and the network element performing the machine learning training satisfies the first performance requirement.

[0052] When the first information indicates the first performance requirement, the network element performing the machine learning training satisfies the first performance requirement.

[0053] For example, the multiple network elements indicated by the first information include network element #1, network element #2 and network element #3, wherein the communication network performance corresponding to the network element #1 is 80%, the communication network performance corresponding to the network element #2 is 85%, and the communication network performance corresponding to the network element #3 is 90% (as a possible example, the communication network performance corresponding to the network element #1, the network element #2 or the network element #3 can be a handover success rate). When the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 90%, the network element #3 can be selected as the network element performing the machine learning training; when the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 85%, any one of the network element #3 or the network element #2 can be randomly selected as the network element performing the machine learning training.

[0054] When the first information indicates the first performance requirement and the first training time requirement, the network element performing the machine learning training meets the first performance requirement and the first training time requirement.

[0055] For example, the multiple network elements indicated by the first information include network element #4, network element #5 and network element #6, wherein the communication network performance corresponding to the network element #4 is 80%, the training duration corresponding to the network element #4 is 300 ms, the communication network performance corresponding to the network element #5 is 85%, the training duration corresponding to the network element #5 is 350 ms, the communication network performance corresponding to the network element #6 is 90%, and the training duration corresponding to the network element #6 is 400 ms. When the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 85% and the upper limit of the training duration is 350 ms, the network element #5 can be selected as the network element performing the machine learning training; when the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 85% and the upper limit of the training duration is 400 ms, any one of the network element #5 or the network element #6 can be randomly selected as the network element performing the machine learning training.

[0056] Based on the scheme provided in the embodiments of the present application, by indicating multiple network elements through the first information, the network element performing the machine learning training belongs to the preselected multiple network elements, which can make the network element performing the machine learning training more reasonable.

[0057] In some possible implementation manners, the network element performing the machine learning training belongs to multiple second network elements; and the first information is further used to indicate at least one of the following: an RL environment of the multiple second network elements; a training position of the multiple second network elements; and a training function of the multiple second network elements.

[0058] For example, the network element performing the machine learning training is determined according to the related information of the second network element indicated by the first information, and the network element performing the machine learning training meets the first information.

[0059] For example, the first information can indicate an RL environment of the plurality of second network elements, a training position of the plurality of second network elements, or a training function of the plurality of second network elements, and the first device can select one RL environment from the RL environments of the plurality of second network elements, or select one training position from the training positions of the plurality of second network elements, or select one training function from the training functions of the plurality of second network elements, according to the first performance requirement (or according to the first performance requirement and the first training time requirement). The one RL environment, the one training position, or the one training function meets the first performance requirement (or meets the first performance requirement and the first training time requirement). After the first device selects the one RL environment, the one training position, or the one training function, the first device can determine the network element for performing the machine learning training according to the network element corresponding to the one RL environment, the one training position, or the one training function.

[0060] Based on the scheme provided in the embodiments of the present application, the network element for performing the machine learning training meets the information of the plurality of network elements by indicating the information of the plurality of network elements through the first information, which can make the network element for performing the machine learning training more reasonable.

[0061] In some possible implementation manners, the method further includes determining the network element for performing the machine learning training according to the communication network performance of the plurality of first network elements.

[0062] For example, the EMS can determine the network element for performing the machine learning training from the plurality of first network elements according to the current communication network performance of the plurality of first network elements and the first performance requirement (or the first performance requirement and the first training time requirement).

[0063] In some possible implementation manners, the method further includes determining the network element for performing the machine learning training according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0064] For example, the EMS can determine one RL environment, one training position, or one training function that meets the first performance requirement (or meets the first performance requirement and the first training time requirement) according to any one of the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements, and the first performance requirement (or the first performance requirement and the first training time requirement), and determine the network element for performing the machine learning training from the plurality of second network elements according to the network element corresponding to the one RL environment, the one training position, or the one training function.

[0065] For example, the method further includes determining the network element for performing the machine learning training according to the RL environment of the plurality of second network elements and / or the training position of the plurality of second network elements.

[0066] Based on the scheme provided in the embodiments of the present application, the network element for performing machine learning training is enabled to meet the first performance requirement (or meet the first performance requirement and the first training time requirement) by determining the network element for performing machine learning training according to the network performance of the plurality of network elements, which is helpful to the implementation of machine learning training.

[0067] In some possible implementation manners, the method further includes: receiving second information, the second information being used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of the communication network performance for performing the machine learning training; or a second training time requirement, the second training time requirement indicating a requirement of the training time for performing the machine learning training; and after performing the machine learning training based on the first information, the method further includes: performing the machine learning training based on the second information.

[0068] In some possible implementation manners, the method further includes: receiving second information, the second information being used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of the communication network performance for performing the machine learning training, the first performance requirement meeting the second performance requirement; or a second training time requirement, the second training time requirement indicating a requirement of the training time for performing the machine learning training, the first training time requirement meeting the second training time requirement, the first training time requirement being indicated by the first information, the first training time requirement indicating a requirement of the training time for performing the machine learning training; and after performing the machine learning training based on the first information, the method further includes: performing the machine learning training based on the second information.

[0069] For example, the second information includes the second performance requirement and / or the second training time requirement.

[0070] For example, the first device can further receive second information, the second information can indicate a second performance requirement and / or a second training time requirement, the first performance requirement meets the second performance requirement, and the first training time requirement meets the second training time requirement. When the machine learning training based on the first information does not meet the requirement (for example, the machine learning training is not completed within a time point at which the machine learning training is expected to be completed or an expected training time length), the first device can perform the machine learning training according to the second performance requirement and / or the second training time requirement indicated by the second information.

[0071] When the machine learning training based on the first information does not meet the requirement, the requirement of the communication network performance for performing the machine learning training and / or the requirement of the training time for performing the machine learning training can be relaxed, and the machine learning training can be performed again based on the relaxed requirement of the communication network performance and / or the relaxed requirement of the training time.

[0072] Based on the scheme provided in the embodiments of the present application, by performing machine learning training again based on the second performance requirement and / or the second training time requirement which are more relaxed than the first performance requirement and / or the first training time requirement, the flexibility and rationality of performing machine learning training can be improved, and the efficiency of performing machine learning training can be improved.

[0073] In some possible implementation manners, the method further includes: sending third information, where the third information is used to indicate at least one of the following: a training time of performing machine learning training; or a value of an impact of performing machine learning training on communication network performance; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform machine learning training again.

[0074] For example, when the machine learning training performed based on the first information meets the requirement (for example, the machine learning training is completed within the expected completion time of the machine learning training or the expected training duration), the MnS producer can indicate: a training time of completing the machine learning training; or a value of an impact of completing the machine learning training on communication network performance; or a requirement (for example, a third performance requirement) for determining whether to perform machine learning training again based on the first information performing machine learning training and / or a requirement (for example, a third training time requirement) for a training time.

[0075] Based on the scheme provided in the embodiments of the present application, by feeding back the related information after completing the machine learning training, information can be provided for adjusting the execution of machine learning training according to actual conditions, and the strategy of performing machine learning training again is more reasonable after adjusting the execution strategy of machine learning training according to actual conditions.

[0076] In some possible implementation manners, the third performance requirement is determined according to the value of the impact of performing machine learning training on communication network performance; and the third training time requirement is determined according to the training time of performing machine learning training.

[0077] For example, the value of the impact of performing machine learning training on communication network performance can include an average value of the value of the impact of the training process on communication network performance, and the third performance requirement can include a maximum value of the loss of communication network performance in the training process; or the training time of performing machine learning training can include the overall time of performing machine learning training, and the third training time requirement can include 1 / n of the overall time of the training, where n is a preset value, and n can be a positive integer. (The early stop in the training process is usually n times of the overall training process, and the specific value depends on the trained model.)

[0078] Based on the scheme provided in the embodiments of the present application, the third performance requirement or the third training time requirement is determined by completing the machine learning training to affect the communication network performance or the training time, which can provide an information basis for adjusting the execution of machine learning training according to actual conditions, so that the execution of machine learning training is more reasonable.

[0079] In some possible implementation ways, before the first information is determined, the method further includes: determining the correspondence relationship according to historical network performance and historical training parameters of executing machine learning training; and / or determining the correspondence relationship according to historical training time and historical training parameters of executing machine learning training.

[0080] Specifically, the first device can determine the correspondence relationship according to historical data of executing machine learning training by each network element (for example, training parameters of the completed machine learning training and corresponding communication network performance, and / or training parameters of the completed machine learning training and corresponding training time).

[0081] For example, the network element #A executes a certain machine learning training, and the training time of a certain training is 300 ms, and the loss of a certain KPI / PM of the communication network caused by executing the machine learning training is 10%; the network element #A executes the machine learning training for another time, and the training time of the completed training is 350 ms, and the loss of the certain KPI / PM of the communication network caused by executing the machine learning training is 5%. For the machine learning training, the training time of 300 ms corresponds to the loss value of the certain KPI / PM of 10%, and the training time of 350 ms corresponds to the loss value of the certain KPI / PM of 5%. The training time and the loss value of the certain KPI / PM that meet the correspondence (for example, the training time and the loss value of the RSRP distribution that meet the correspondence) can be regarded as meeting the above-mentioned correspondence relationship, and the correspondence between different training time and different KPI / PM can be regarded as the above-mentioned determination of the correspondence relationship.

[0082] Based on the scheme provided in the embodiments of the present application, the correspondence relationship is determined based on the historical data of executing machine learning training, which can provide reliable and real basis for the determination of the first parameter, and can improve the rationality of executing machine learning training and the efficiency of executing machine learning training.

[0083] In some possible implementation ways, the first information is determined according to the first performance index, and the first performance index is associated with an inference function to which a model corresponding to the machine learning training belongs.

[0084] For example, the first information can be determined according to a communication network performance index associated with an inference function to which a model corresponding to the machine learning training belongs.

[0085] For example, the first information can be determined according to acceptable KPIs or PMs of the reinforcement learning in the existing network. The device for calculating the first information can calculate a value that can meet the requirements of the minimum KPIs or PMs and the current communication network status by the requirements of the minimum KPIs or PMs and the current communication network status, and the calculated value is used to represent the requirements of the machine learning training on the communication network performance. The requirements of the minimum KPIs or PMs and the current communication network status are related to the inference function of the machine learning corresponding to the machine learning training.

[0086] For example, the network function corresponding to the model is the MDA use case coverage problem analysis, when the RL is applied to the coverage problem analysis use case, the KPI / PM can include the RSRP distribution; for example, the network function corresponding to the model is the load balancing optimization, when the RL is applied to the base station cell load balancing, mobility optimization or mobile robustness optimization use case, the KPI / PM can include the cell load and the handover success rate.

[0087] Based on the scheme provided in the embodiments of the present application, by determining the first information according to the communication network performance index associated with the inference function corresponding to the model of the machine learning training, the communication network performance in the training process can be maintained, and the service quality of the communication network users can be ensured.

[0088] In some possible implementation manners, the method further includes: updating the correspondence according to the third information.

[0089] For example, the first device can update the correspondence according to the data of completing the machine learning training (for example, the training parameters of the completed machine learning training and the corresponding communication network performance, and / or the training parameters of the completed machine learning training and the corresponding training time).

[0090] In a second aspect, a communication method applied to machine learning training is provided, which can be performed by a second device capable of performing machine learning training. The second device can refer to a device (for example, a network device) on a network device side, a component (for example, a communication module, a processor, a circuit, a chip, or a chip system) in the device, or a logic module or software capable of realizing all or part of the functions of the communication device. The network device side can include at least one of a network device or an AI entity on the network device side. The AI entity on the network device side can be the network device itself or an AI entity serving the network device, for example, a radio access network (RAN) intelligent controller (RIC), an operation administration and maintenance (OAM), or a server, such as an OTT server or a cloud server. The communication between servers can be realized through a communication link between a terminal device and a network device, or through other communication devices outside the servers, or through a wired link. For ease of description, the method performed by the second device is described below as an example.

[0091] The method includes determining first information and sending the first information. The first information is used to indicate a first performance requirement, and the first performance requirement indicates a requirement of performing machine learning training on a communication network performance.

[0092] For example, the first information is used to perform machine learning training.

[0093] For example, the second device can be a device where the MnS consumer is located. For example, the second device can be a device where the NMS is located.

[0094] For example, the first information can be determined according to received indication information. For example, the second device receives indication information indicating the first information, and the second device can determine the first information according to the received indication information. Alternatively, the first information can be determined according to a requirement to be met by a key performance indicator (KPI) or a performance management (PM) of a related communication network and / or a current performance status of the related communication network. For example, the second device can determine the requirement of performing machine learning training on the communication network performance by calculation according to the requirement to be met by the KPI or the PM of the related communication network and / or the current performance status of the related communication network. The embodiments of the present application do not limit this.

[0095] Based on the scheme provided in the embodiments of the present application, by indicating the first information for performing the machine learning training, the machine learning training efficiency and the degree of influence on the communication network performance of the communication network can be balanced, and the service quality of the communication network users is ensured.

[0096] In some possible implementation manners, the network element performing the machine learning training satisfies the first performance requirement.

[0097] In some possible implementation manners, the requirement of the communication network performance for performing the machine learning training can include at least one of the following: a lower threshold of an allowed KPI or PM; or, a loss value of an allowed KPI or PM loss; or, a ratio of an allowed KPI or PM loss; or, a range of an allowed network fluctuation.

[0098] In some possible implementation manners, the first information is further used to indicate a first training time requirement, and the first training time requirement indicates a requirement of the training time for performing the machine learning training.

[0099] In some possible implementation manners, the requirement of the training time for performing the machine learning training can include at least one of the following: a time point at which the machine learning training is expected to be completed; or, an expected training duration.

[0100] In some possible implementation manners, the machine learning training is reinforcement learning training, and the first training time requirement includes a training time point and / or a training duration at which the reinforcement learning training is expected to be completed.

[0101] In some possible implementation manners, the network element performing the machine learning training satisfies the first performance requirement and the first training time requirement.

[0102] In some possible implementation manners, the first parameter is used to perform the machine learning training, wherein the first parameter is determined according to the first performance requirement and a corresponding relationship, the corresponding relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first performance requirement.

[0103] For example, the training parameter includes a learning rate, or the learning rate and a decay ratio of the learning rate.

[0104] For example, the training parameter includes an exploration rate, or the exploration rate and a decay ratio of the exploration rate.

[0105] In some possible implementation manners, the first parameter is used to perform the machine learning training, wherein the first parameter is determined according to the first performance requirement and a corresponding relationship, the corresponding relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first performance requirement and the first training time requirement.

[0106] In some possible implementation, the first information is further used to indicate a plurality of first network elements, the network element performing the machine learning training belongs to the plurality of first network elements.

[0107] In some possible implementation, the network element performing the machine learning training is determined according to a communication network performance of the plurality of first network elements.

[0108] In some possible implementation, the network element performing the machine learning training belongs to a plurality of second network elements; the first information is further used to indicate at least one of: an RL environment of the plurality of second network elements; a training location of the plurality of second network elements; a training function of the plurality of second network elements.

[0109] For example, the first information is further used to indicate at least one of: the RL environment of the plurality of second network elements; the training location of the plurality of second network elements.

[0110] In some possible implementation, the network element performing the machine learning training is determined according to the RL environment of the plurality of second network elements, the training location of the plurality of second network elements, or the training function of the plurality of second network elements.

[0111] For example, the RL environment of the plurality of second network elements and / or the training location of the plurality of second network elements are used to determine the network element performing the machine learning training.

[0112] In some possible implementation, the method further includes: determining a first time, the first time is used to monitor whether the machine learning training is completed; based on a case that the machine learning training is not completed according to the first time, sending second information, the second information is used to indicate at least one of: a second performance requirement, the second performance requirement indicates a requirement of the communication network performance for performing the machine learning training, the first performance requirement meets the second performance requirement; or, a second training time requirement, the second training time requirement indicates a requirement of the training time for performing the machine learning training, the first training time requirement meets the second training time requirement, the first training time requirement is indicated by the first information, the first training time requirement indicates the requirement of the training time for performing the machine learning training.

[0113] For example, the second performance requirement indicates the requirement of the communication network performance for performing the machine learning training; the second training time requirement indicates the requirement of the training time for performing the machine learning training.

[0114] For example, the second information can include the second performance requirement and / or the second training time requirement.

[0115] For example, the second device can monitor whether the machine learning training is completed within a time period corresponding to the first time or before a time node according to the preset or input first time. In a case where the machine learning training is not completed based on the first time, the second device can send second information, which is more relaxed than the first information (i.e., the first performance requirement meets the second performance requirement, or the first training time requirement meets the second training time requirement).

[0116] Based on the scheme provided in the embodiments of the present application, on the one hand, by monitoring the time of machine learning training, the time of machine learning training can be avoided to be too long or unresponsive, which helps to improve the rationality of machine learning training; on the other hand, by indicating the second performance requirement and / or the second training time requirement which is more relaxed than the first performance requirement and / or the first training time requirement for executing the machine learning training again, the flexibility and rationality of executing the machine learning training can be improved, and the efficiency of executing the machine learning training can be improved.

[0117] In some possible implementation ways, the method further includes: receiving third information; and updating the first information according to the third information, wherein the third information is used to indicate at least one of the following: a training time of executing the machine learning training; or a value of an influence of executing the machine learning training on a communication network performance; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement is used to execute the machine learning training again.

[0118] In some possible implementation ways, the third performance requirement is determined according to the value of the influence of executing the machine learning training on the communication network performance; and the third training time requirement is determined according to the training time of executing the machine learning training.

[0119] In some possible implementation ways, the first information is determined according to a first performance index, and the first performance index is associated with an inference function to which a model corresponding to the machine learning training belongs.

[0120] For example, the first information is determined according to a first performance index, and the first performance index is associated with an inference function to which a model corresponding to the machine learning training belongs.

[0121] In some possible implementation ways, the method further includes: receiving fourth information, and the fourth information is used to determine the first training time requirement.

[0122] For example, the first information is determined according to the fourth information, and the first training time requirement is determined according to the fourth information.

[0123] In a third aspect, a device is provided, which is applied to machine learning training. The device has the functions of the first aspect. For example, the device includes modules or units or means corresponding to the operations of the first aspect. The modules or units or means can be implemented in software, hardware, or a combination of software and hardware.

[0124] For example, the device can be the first device. For example, the device can be a module or unit (for example, a chip, a chip system, or a circuit) corresponding to the method or operation or step or action of the first aspect.

[0125] In some possible implementation ways, the device includes a processing unit (or a processing module). The processing unit is configured to: determine first information, the first information being used to indicate a first performance requirement, the first performance requirement indicating a requirement of performing machine learning training on a communication network performance; and perform the machine learning training based on the first information.

[0126] In some possible implementation ways, the requirement of performing the machine learning training on the communication network performance can include at least one of: a lower threshold of an allowed KPI or PM; or a loss value of an allowed KPI or PM loss; or a ratio of an allowed KPI or PM loss; or a range of an allowed network fluctuation.

[0127] In some possible implementation ways, the first information is further used to indicate a first training time requirement, the first training time requirement indicating a requirement of performing the machine learning training on a training time.

[0128] In some possible implementation ways, the requirement of performing the machine learning training on the training time can include at least one of: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0129] In some possible implementation ways, the processing unit is specifically configured to: perform the machine learning training based on the first parameter, wherein the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicating a relationship between a first value and a training parameter, the first parameter belonging to the training parameter, and the first value corresponding to the first parameter satisfying the first performance requirement.

[0130] In some possible implementation ways, the processing unit is specifically configured to: perform the machine learning training based on the first parameter, wherein the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicating a relationship between a first value and a training parameter, the first parameter belonging to the training parameter, and the first value corresponding to the first parameter satisfying the first performance requirement and the first training time requirement.

[0131] In some possible implementation manners, the first information is further used to indicate a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

[0132] In some possible implementation manners, the processing unit is further used to determine the network element performing the machine learning training according to the communication network performance of the plurality of first network elements.

[0133] In some possible implementation manners, the network element performing the machine learning training belongs to a plurality of second network elements; and the first information is further used to indicate at least one of the following: an RL environment of the plurality of second network elements; a training position of the plurality of second network elements; and a training function of the plurality of second network elements.

[0134] In some possible implementation manners, the processing unit is further used to determine the network element performing the machine learning training according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0135] In some possible implementation manners, the apparatus further includes a transceiver, which is used to receive second information, and the second information is used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of the machine learning training on the communication network performance, the first performance requirement meeting the second performance requirement; or a second training time requirement, the second training time requirement indicating a requirement of the machine learning training on a training time, the first training time requirement meeting the second training time requirement, the first training time requirement being indicated by the first information, the first training time requirement indicating the requirement of the machine learning training on the training time; and after performing the machine learning training based on the first information, the processing unit is further used to perform the machine learning training based on the second information.

[0136] For example, the second information includes the second performance requirement and / or the second training time requirement.

[0137] In some possible implementation manners, the transceiver is further used to send third information, and the third information is used to indicate at least one of the following: a training time of the machine learning training; or a value of an influence of the machine learning training on the communication network performance; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

[0138] In some possible implementation manners, the third performance requirement is determined according to the value of the influence of the machine learning training on the communication network performance; and the third training time requirement is determined according to the training time of the machine learning training.

[0139] In some possible implementation, the processing unit is further configured to determine the correspondence relationship according to historical network performance and historical training parameters of performing the machine learning training; and / or determine the correspondence relationship according to historical training time and historical training parameters of performing the machine learning training.

[0140] In some possible implementation, the first information is determined according to a first performance index, the first performance index being associated with an inference function to which a model corresponding to the machine learning training belongs.

[0141] In some possible implementation, the processing unit is further configured to update the correspondence relationship according to the third information.

[0142] In some possible implementation, the processing unit includes a processor.

[0143] In some possible implementation, the transceiving unit includes a transceiver.

[0144] In some possible implementation, the apparatus is a chip.

[0145] In the fourth aspect, a communication apparatus is provided, and the communication apparatus has the functions of the second aspect, for example, the communication apparatus includes modules or units or means corresponding to the operations of the second aspect, and the modules or units or means can be implemented in software, hardware or a combination of software and hardware.

[0146] For example, the communication apparatus can be the second apparatus, for example, a module or unit (for example, a chip or a chip system or a circuit) corresponding to the method or operation or step or action described in the second aspect.

[0147] In some possible implementation, the communication apparatus includes a transceiving unit (or a communication module) and a processing unit (or a processing module) connected to the transceiving unit. The processing unit is configured to determine the first information. The transceiving unit is configured to send the first information, and the first information is used to indicate a first performance requirement, and the first performance requirement indicates a requirement of performing the machine learning training on the communication network performance.

[0148] For example, the first information is used to perform the machine learning training.

[0149] In some possible implementation, the requirement of performing the machine learning training on the communication network performance can include at least one of the following: a lower threshold of an allowed KPI or PM; or a loss value of an allowed KPI or PM loss; or a ratio of an allowed KPI or PM loss; or a range of an allowed network fluctuation.

[0150] In some possible implementation manners, the first information is further used to indicate a first training time requirement, the first training time requirement indicating a requirement of the performing the machine learning training on training time.

[0151] In some possible implementation manners, the requirement of the performing the machine learning training on training time can include at least one of the following: a time point at which the machine learning training is expected to be completed; or, an expected training duration.

[0152] In some possible implementation manners, the first parameter is used to perform the machine learning training, where the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicating a relationship between the first value and a training parameter, the first parameter belonging to the training parameter, and the first value corresponding to the first parameter satisfying the first performance requirement.

[0153] In some possible implementation manners, the first parameter is used to perform the machine learning training, where the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicating a relationship between the first value and a training parameter, the first parameter belonging to the training parameter, and the first value corresponding to the first parameter satisfying the first performance requirement and the first training time requirement.

[0154] In some possible implementation manners, the first information is further used to indicate a plurality of first network elements, the network element performing the machine learning training belonging to the plurality of first network elements.

[0155] In some possible implementation manners, the network element performing the machine learning training is determined according to a communication network performance of the plurality of first network elements.

[0156] In some possible implementation manners, the network element performing the machine learning training belongs to a plurality of second network elements; and the first information is further used to indicate at least one of the following: an RL environment of the plurality of second network elements; a training position of the plurality of second network elements; and a training function of the plurality of second network elements.

[0157] In some possible implementation manners, the network element performing the machine learning training is determined according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0158] In some possible implementation, the processing unit is further configured to determine a first time, the first time being used to monitor whether the machine learning training is completed; and based on a case that the first time monitors that the machine learning training is not completed, the transceiver is further configured to send second information, the second information being used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of performing the machine learning training on a performance of the communication network, the first performance requirement meeting the second performance requirement; or a second training time requirement, the second training time requirement indicating a requirement of performing the machine learning training on a training time, the first training time requirement meeting the second training time requirement, the first training time requirement being indicated by the first information, the first training time requirement indicating the requirement of performing the machine learning training on the training time.

[0159] In some possible implementation, the transceiver is further configured to receive third information; and update the first information according to the third information, wherein the third information is used to indicate at least one of the following: a training time of performing the machine learning training; or a value of an impact of performing the machine learning training on a performance of the communication network; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

[0160] In some possible implementation, the third performance requirement is determined according to the value of the impact of performing the machine learning training on the performance of the communication network; and the third training time requirement is determined according to the training time of performing the machine learning training.

[0161] In some possible implementation, the first information is determined according to a first performance index, the first performance index being associated with an inference function to which a model corresponding to the machine learning training belongs.

[0162] In some possible implementation, the transceiver is further configured to receive fourth information, the fourth information being used to determine the first training time requirement.

[0163] In some possible implementation, the processing unit includes a processor.

[0164] In some possible implementation, the transceiver includes a transceiver.

[0165] In some possible implementation, the apparatus is a chip.

[0166] In a fifth aspect, a communication method is provided for machine learning training. The method can be performed by a third device capable of performing machine learning training. The third device can refer to a device on a network device side (e.g., a network device), a component in the device (e.g., a communication module, a processor, a circuit, a chip, or a chip system, etc.), or a logic module or software capable of realizing all or part of the functions of the communication device. The network device side can include at least one of a network device or an AI entity on the network device side. The AI entity on the network device side can be the network device itself or an AI entity serving the network device, such as a radio access network (RAN) intelligent controller (RIC), an operation administration and maintenance (OAM), or a server, such as an OTT server or a cloud server, etc. The communication between servers can be realized through a communication link between a terminal device and a network device, or through other communication devices outside the servers, or through a wired link. For ease of description, the method performed by the third device is described below as an example.

[0167] The method includes determining first information and performing machine learning training based on the first information, where the first information is used to indicate at least one of the following: a first performance requirement, which indicates the requirement of the machine learning training on the performance of a communication network; or a first training time requirement, which indicates the requirement of the machine learning training on the training time.

[0168] For example, the determination of the first information can be based on received indication information, e.g., the third device receives indication information indicating the first information, and the third device can determine the first information based on the received indication information. Alternatively, the determination of the first information can also be based on the requirement to be met by the KPI or PM of a related communication network and / or the current performance status of the related communication network, e.g., the third device can determine the requirement of the machine learning training on the performance of the communication network through calculation based on the requirement to be met by the KPI or PM of the related communication network and / or the current performance status of the related communication network. The present application does not limit the determination of the first information.

[0169] The determination of the first information can be receiving the first information.

[0170] The determining the first information can be understood as that the MnS producer receives information configured by the MnS consumer for creating a machine learning training request management object instance (MLTrainingRequest MOI).

[0171] Exemplarily, a related attribute of the performance requirement of the existing network can be added in a machine learning training request information object class (MLTrainingRequest IOC).

[0172] Exemplarily, the related attribute of the performance requirement of the existing network can include a reinforcement learning network performance requirement (rlnetworkperformancerequirement).

[0173] The performing machine learning training based on the first information can be understood as that the MnS producer receives a create management object instance (MOI) request of the MLTrianingRequest and creates the MLTrianingRequest MOI.

[0174] Exemplarily, the third device can be a device where a management service producer (MnS producer) is located. For example, the third device can be a device where an element management system (EMS) is located, or a device where a base station gNB is located, or a RIC.

[0175] In some possible implementation manners, the third apparatus can be referred to as: a production entity; an EMS; a domain management node / system / function; a MnS producer; a single-domain management system (domain management system, MnS); or a single-domain management function (domain management function, MnF). Correspondingly, an apparatus where a management service consumer / user / caller (MnS consumer) that jointly implements machine learning training with the third apparatus can be referred to as: a consumption entity; a network management system (network management system, NMS); a cross-domain management node / system / function; a MnS consumer; a cross-domain management system (Cross-domain management system, CD-MnS); or a cross-domain management function (Cross-domain management function, CD-MnF); a service management and orchestration (function) (service management and orchestration, SMO).

[0176] Based on the scheme provided in the embodiments of the present application, by performing machine learning training based on the first information indicating the first performance requirement and / or the first training time requirement, the machine learning training efficiency and the degree of influence on the communication network performance of the real communication network can be balanced, and the service quality of the communication network user can be ensured.

[0177] In some possible implementation manners, the requirement of performing machine learning training on the communication network performance can include at least one of: a lower threshold of an allowed KPI or PM; or a loss value of an allowed KPI or PM loss; or a ratio of an allowed KPI or PM loss; or a range of an allowed network fluctuation.

[0178] The influence of performing machine learning training based on the first information on the communication network performance can meet: a lower threshold of an allowed KPI; or a loss value of an allowed KPI loss; or a ratio of an allowed KPI loss; or a range of an allowed network fluctuation; or a lower threshold of an allowed PM; or a loss value of an allowed PM loss; or a ratio of an allowed PM loss, and the like.

[0179] In some possible implementation manners, the requirement of performing machine learning training on the training time can include at least one of: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0180] The training time of performing the machine learning training based on the first information can meet: a time point at which the machine learning training is expected to be completed, or an expected training duration, and the like.

[0181] In some possible implementation manners, performing the machine learning training based on the first information includes: performing the machine learning training based on a first parameter, where the first parameter is determined according to the first performance requirement and a corresponding relationship, the corresponding relationship indicates a relationship between a first value and a training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter meets the first performance requirement.

[0182] In some possible implementation manners, performing the machine learning training based on the first information includes: performing the machine learning training based on a first parameter, where the first parameter is determined according to the first performance requirement and a corresponding relationship, the corresponding relationship indicates a relationship between a first value and a training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter meets the first training time requirement.

[0183] In some possible implementation manners, performing the machine learning training based on the first information includes: performing the machine learning training based on a first parameter, where the first parameter is determined according to the first performance requirement and a corresponding relationship, the corresponding relationship indicates a relationship between a first value and a training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter meets the first performance requirement and the first training time requirement.

[0184] For example, the training parameter includes a learning rate, or the learning rate and a decay ratio of the learning rate.

[0185] For example, the training parameter includes an exploration rate, or the exploration rate and a decay ratio of the exploration rate.

[0186] In some possible implementation manners, the first information is further used to indicate a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

[0187] In some possible implementation manners, the method further includes: determining the network element performing the machine learning training according to a communication network performance of the plurality of first network elements.

[0188] In some possible implementation manners, the network element performing the machine learning training belongs to a plurality of second network elements, and the first information is further used to indicate at least one of the following: an RL environment of the plurality of second network elements, a training position of the plurality of second network elements, and a training function of the plurality of second network elements.

[0189] In some possible implementation manners, the method further includes: determining the network element performing the machine learning training according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0190] In some possible implementation manners, the method further includes: receiving second information, the second information being used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of performing the machine learning training on a performance of the communication network, the first performance requirement satisfying the second performance requirement; or a second training time requirement, the second training time requirement indicating a requirement of performing the machine learning training on a training time, the first training time requirement satisfying the second training time requirement; and after performing the machine learning training based on the first information, the method further includes: performing the machine learning training based on the second information.

[0191] For example, the second information includes the second performance requirement and / or the second training time requirement.

[0192] In some possible implementation manners, the method further includes: sending third information, wherein the third information is used to indicate at least one of the following: a training time of performing the machine learning training; or a value of an influence of performing the machine learning training on the performance of the communication network; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

[0193] In some possible implementation manners, the third performance requirement is determined according to the value of the influence of performing the machine learning training on the performance of the communication network; and the third training time requirement is determined according to the training time of performing the machine learning training.

[0194] In some possible implementation manners, before determining the first information, the method further includes: determining the correspondence relationship according to historical network performance and historical training parameters of performing the machine learning training; and / or determining the correspondence relationship according to historical training time and historical training parameters of performing the machine learning training.

[0195] In some possible implementation manners, the first information is determined according to a first performance index, the first performance index being associated with an inference function to which a model corresponding to the machine learning training belongs.

[0196] In some possible implementation manners, the method further includes: updating the correspondence relationship according to the third information.

[0197] In a sixth aspect, a communication method applied to machine learning training is provided, which can be performed by a fourth apparatus capable of performing machine learning training. The fourth apparatus can refer to a device (for example, a network device) on a network device side, or a component (for example, a communication module, a processor, a circuit, a chip, or a chip system, etc.) in the device, or a logic module or software capable of realizing all or part of the functions of the communication device. The network device side can include at least one of a network device or an AI entity on the network device side. The AI entity on the network device side can be the network device itself, or an AI entity serving the network device, for example, a radio access network (RAN) intelligent controller (RIC), an operation administration and maintenance (OAM), or a server, such as an OTT server or a cloud server. The communication between servers can be realized through a communication link between a terminal device and a network device, or through other communication devices other than the servers, or through a wired link. For ease of description, the fourth apparatus performing the method is taken as an example in the following description.

[0198] The method includes determining first information, and transmitting the first information, where the first information is used to indicate at least one of the following: a first performance requirement, the first performance requirement indicating a requirement of the machine learning training on a communication network performance; or a first training time requirement, the first training time requirement indicating a requirement of the machine learning training on a training time.

[0199] For example, the first information is used to perform the machine learning training.

[0200] In some possible implementation manners, the requirement of the machine learning training on the communication network performance can include at least one of the following: a lower threshold of an allowed KPI or PM; or a loss value of an allowed KPI or PM loss; or a ratio of an allowed KPI or PM loss; or a range of an allowed network fluctuation.

[0201] In some possible implementation manners, the requirement of the machine learning training on the training time can include at least one of the following: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0202] In some possible implementation manners, the first parameter is used to perform the machine learning training, where the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicating a relationship between the first value and a training parameter, the first parameter belonging to the training parameter, and the first value corresponding to the first parameter satisfying the first performance requirement.

[0203] In some possible implementation, the first parameter is used for performing the machine learning training, wherein the first parameter is determined according to the first training time requirement and a correspondence relationship, the correspondence relationship indicates a relationship between the first value and a training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first training time requirement.

[0204] In some possible implementation, the first parameter is used for performing the machine learning training, wherein the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicates a relationship between the first value and a training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first performance requirement and the first training time requirement.

[0205] For example, the training parameter includes a learning rate, or the learning rate and a decay ratio of the learning rate.

[0206] For example, the training parameter includes an exploration rate, or the exploration rate and a decay ratio of the exploration rate.

[0207] In some possible implementation, the first information is further used for indicating a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

[0208] In some possible implementation, the network element performing the machine learning training is determined according to a communication network performance of the plurality of first network elements.

[0209] In some possible implementation, the network element performing the machine learning training belongs to a plurality of second network elements, and the first information is further used for indicating at least one of the following: an RL environment of the plurality of second network elements, a training position of the plurality of second network elements, and a training function of the plurality of second network elements.

[0210] In some possible implementation, the network element performing the machine learning training is determined according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0211] In some possible implementation, the method further includes: determining a first time, the first time is used for monitoring whether the machine learning training is completed; and based on a case that the machine learning training is not completed according to the first time, sending second information, the second information is used for indicating at least one of the following: a second performance requirement, the second performance requirement indicates a requirement of the communication network performance for performing the machine learning training, and the first performance requirement satisfies the second performance requirement; or a second training time requirement, the second training time requirement indicates a requirement of the training time for performing the machine learning training, and the first training time requirement satisfies the second training time requirement.

[0212] In some possible implementation manners, the method further includes: receiving third information; and updating the first information according to the third information, wherein the third information is used to indicate at least one of the following: a training time of performing the machine learning training; or a value of an influence of performing the machine learning training on the communication network performance; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

[0213] In some possible implementation manners, the third performance requirement is determined according to the value of the influence of performing the machine learning training on the communication network performance; and the third training time requirement is determined according to the training time of performing the machine learning training.

[0214] In some possible implementation manners, the first information is determined according to a first performance index, the first performance index being associated with an inference function to which a model corresponding to the machine learning training belongs.

[0215] In some possible implementation manners, the method further includes: receiving fourth information, the fourth information being used to determine the first training time requirement.

[0216] In a seventh aspect, a device is provided, which is applied to machine learning training, and the communication device has the functions of implementing the fifth aspect, for example, the communication device includes a module or unit or means corresponding to the operations described in the fifth aspect, and the module or unit or means can be implemented by software, or by hardware, or by a combination of software and hardware.

[0217] For example, the communication device can be the third device described above, for example, a module or unit (for example, a chip, a chip system, or a circuit) corresponding to the method or operation or step or action described in the fifth aspect.

[0218] In some possible implementation manners, the communication device includes a processing unit (or a processing module). The processing unit is configured to: determine the first information; and perform the machine learning training based on the first information, wherein the first information is used to indicate at least one of the following: a first performance requirement, the first performance requirement indicating a requirement of performing the machine learning training on the communication network performance; or a first training time requirement, the first training time requirement indicating a requirement of performing the machine learning training on a training time.

[0219] In some possible implementation manners, the requirement of performing the machine learning training on the communication network performance can include at least one of the following: a lower threshold of an allowed KPI or PM; or a loss value of an allowed KPI or PM loss; or a ratio of an allowed KPI or PM loss; or a range of an allowed network fluctuation.

[0220] In some possible implementation manners, the requirement of the machine learning training on training time can include at least one of the following: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0221] In some possible implementation manners, the processing unit is specifically configured to perform the machine learning training based on the first parameter, where the first parameter is determined according to the first performance requirement and the correspondence relationship, the correspondence relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first performance requirement.

[0222] In some possible implementation manners, the processing unit is specifically configured to perform the machine learning training based on the first parameter, where the first parameter is determined according to the first training time requirement and the correspondence relationship, the correspondence relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first training time requirement.

[0223] In some possible implementation manners, the processing unit is specifically configured to perform the machine learning training based on the first parameter, where the first parameter is determined according to the first performance requirement and the correspondence relationship, the correspondence relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter satisfies the first performance requirement and the first training time requirement.

[0224] In some possible implementation manners, the first information is further used to indicate a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

[0225] In some possible implementation manners, the processing unit is further configured to determine the network element performing the machine learning training according to the communication network performance of the plurality of first network elements.

[0226] In some possible implementation manners, the network element performing the machine learning training belongs to a plurality of second network elements; and the first information is further used to indicate at least one of the following: an RL environment of the plurality of second network elements; a training position of the plurality of second network elements; and a training function of the plurality of second network elements.

[0227] In some possible implementation manners, the processing unit is further configured to determine the network element performing the machine learning training according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0228] In some possible implementation, the apparatus further includes a transceiver, configured to: receive second information, the second information being used to indicate at least one of: a second performance requirement, the second performance requirement indicating a requirement of performing the machine learning training on a performance of the communication network, the first performance requirement satisfying the second performance requirement; or a second training time requirement, the second training time requirement indicating a requirement of performing the machine learning training on a training time, the first training time requirement satisfying the second training time requirement, the first training time requirement being indicated by the first information, the first training time requirement indicating a requirement of performing the machine learning training on a training time; and the processing unit is further configured to perform the machine learning training based on the second information after performing the machine learning training based on the first information.

[0229] For example, the second information includes the second performance requirement and / or the second training time requirement.

[0230] In some possible implementation, the transceiver is further configured to: transmit third information, wherein the third information is used to indicate at least one of: a training time of performing the machine learning training; or a value of an impact of performing the machine learning training on a performance of the communication network; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

[0231] In some possible implementation, the third performance requirement is determined according to the value of the impact of performing the machine learning training on the performance of the communication network; and the third training time requirement is determined according to the training time of performing the machine learning training.

[0232] In some possible implementation, the processing unit is further configured to: determine the correspondence relationship according to historical network performance and historical training parameters of performing the machine learning training; and / or determine the correspondence relationship according to historical training time and historical training parameters of performing the machine learning training.

[0233] In some possible implementation, the first information is determined according to a first performance indicator, the first performance indicator being associated with an inference function to which a model corresponding to the machine learning training belongs.

[0234] In some possible implementation, the processing unit is further configured to: update the correspondence relationship according to the third information.

[0235] In some possible implementation, the processing unit includes a processor.

[0236] In some possible implementation, the transceiver includes a transceiver.

[0237] In some possible implementation, the apparatus is a chip.

[0238] In an eighth aspect, a communication apparatus is provided with the function of implementing the sixth aspect, for example, the communication apparatus includes a module or unit or means corresponding to the operation of the sixth aspect, which can be implemented by software, or by hardware, or by a combination of software and hardware.

[0239] For example, the communication apparatus can be the fourth apparatus, for example, a module or unit (for example, a chip, a chip system, or a circuit) corresponding to the method or operation or step or action described in the sixth aspect.

[0240] In some possible implementation manners, the communication apparatus includes a transceiver (or a communication module) and a processing unit (or a processing module) connected to the transceiver. The processing unit is configured to determine the first information. The transceiver is configured to send the first information, where the first information is used to indicate at least one of the following: a first performance requirement, the first performance requirement indicating a requirement of performing machine learning training on a communication network performance; or a first training time requirement, the first training time requirement indicating a requirement of performing machine learning training on a training time.

[0241] For example, the first information is used to perform the machine learning training.

[0242] In some possible implementation manners, the requirement of performing machine learning training on the communication network performance can include at least one of the following: a lower threshold of an allowed KPI or PM; or a loss value of an allowed KPI or PM loss; or a ratio of an allowed KPI or PM loss; or a range of an allowed network fluctuation.

[0243] In some possible implementation manners, the requirement of performing machine learning training on the training time can include at least one of the following: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0244] In some possible implementation manners, the first parameter is used to perform the machine learning training, where the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicating a relationship between a first value and a training parameter, the first parameter belonging to the training parameter, and the first value corresponding to the first parameter satisfying the first performance requirement.

[0245] In some possible implementation manners, the first parameter is used to perform the machine learning training, where the first parameter is determined according to the first performance requirement and a correspondence relationship, the correspondence relationship indicating a relationship between a first value and a training parameter, the first parameter belonging to the training parameter, and the first value corresponding to the first parameter satisfying the first performance requirement and the first training time requirement.

[0246] In some possible implementation, the first information is further used to indicate a plurality of first network elements, the network element performing the machine learning training belongs to the plurality of first network elements.

[0247] In some possible implementation, the network element performing the machine learning training is determined according to the communication network performance of the plurality of first network elements.

[0248] In some possible implementation, the network element performing the machine learning training belongs to a plurality of second network elements; the first information is further used to indicate at least one of the following: an RL environment of the plurality of second network elements; a training position of the plurality of second network elements; a training function of the plurality of second network elements.

[0249] In some possible implementation, the network element performing the machine learning training is determined according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0250] In some possible implementation, the processing unit is further used to: determine a first time, the first time is used to monitor whether the machine learning training is completed; based on a case that the machine learning training is not completed based on the first time, the transceiver is further used to send second information, the second information is used to indicate at least one of the following: a second performance requirement, the second performance requirement indicates a requirement of the communication network performance for performing the machine learning training, the first performance requirement meets the second performance requirement; or a second training time requirement, the second training time requirement indicates a requirement of the training time for performing the machine learning training.

[0251] In some possible implementation, the transceiver is further used to: receive third information; and update the first information according to the third information, wherein the third information is used to indicate at least one of the following: a training time of performing the machine learning training; or a value of an influence of performing the machine learning training on the communication network performance; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement is used to perform the machine learning training again.

[0252] In some possible implementation, the third performance requirement is determined according to the value of the influence of performing the machine learning training on the communication network performance; and the third training time requirement is determined according to the training time of performing the machine learning training.

[0253] In some possible implementation, the first information is determined according to a first performance index, the first performance index is associated with an inference function to which a model corresponding to the machine learning training belongs.

[0254] In some possible implementation, the transceiver is further used to: receive fourth information, the fourth information is used to determine the first training time requirement.

[0255] In some possible implementation manners, the processing unit comprises a processor.

[0256] In some possible implementation manners, the transceiving unit comprises a transceiver.

[0257] In some possible implementation manners, the apparatus is a chip.

[0258] In a ninth aspect, a communication apparatus is provided, which comprises: a processor, configured to execute computer instructions stored in a memory, so that the apparatus performs the method in the first aspect or any possible implementation manner of the first aspect; or the processor is configured to execute the method in the second aspect or any possible implementation manner of the second aspect; or the communication apparatus further comprises a transceiver, and the processor and the transceiver are configured to perform the method in the fifth aspect or any possible implementation manner of the fifth aspect; or the processor and the transceiver are configured to perform the method in the sixth aspect or any possible implementation manner of the sixth aspect.

[0259] In some possible implementation manners, the processor can be a general processor, which can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit or the like; when implemented by software, the processor can be a general processor, which is implemented by reading software codes stored in a memory. The memory can be integrated in the processor or exist independently outside the processor.

[0260] In some possible implementation manners, the apparatus further comprises a memory.

[0261] In some possible implementation manners, the apparatus further comprises a communication interface, which is coupled with the processor, and is configured to input and / or output information.

[0262] In some possible implementation manners, the apparatus is a chip.

[0263] In a tenth aspect, a chip or a chip system is provided, which comprises: a circuit, configured to perform the method in the first aspect or any possible implementation manner of the first aspect; or perform the method in the second aspect or any possible implementation manner of the second aspect; or perform the method in the fifth aspect or any possible implementation manner of the fifth aspect; or perform the method in the sixth aspect or any possible implementation manner of the sixth aspect.

[0264] In an eleventh aspect, a computer program product is provided. When a computer program in the computer program product is executed by a communication device, the method in the first aspect or any possible implementation of the first aspect is implemented; or the method in the second aspect or any possible implementation of the second aspect is implemented; or the method in the fifth aspect or any possible implementation of the fifth aspect is implemented; or the method in the sixth aspect or any possible implementation of the sixth aspect is implemented.

[0265] In a twelfth aspect, a computer readable storage medium is provided. The computer readable storage medium stores a computer program or instructions. When the computer program or instructions is executed by a processor, the method in the first aspect or any possible implementation of the first aspect is implemented; or the method in the second aspect or any possible implementation of the second aspect is implemented; or the method in the fifth aspect or any possible implementation of the fifth aspect is implemented; or the method in the sixth aspect or any possible implementation of the sixth aspect is implemented.

[0266] As an example, the computer readable storage includes, but is not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), Flash memory, electrically EPROM (EEPROM), and hard drive.

[0267] In some possible implementations, the storage medium can be a non-volatile storage medium.

[0268] In a thirteenth aspect, a system is provided. The system includes the apparatus in the ninth aspect or any possible implementation of the ninth aspect.

[0269] The implementation of the solution and the beneficial effects brought by the above-mentioned second aspect can refer to the specific description of the first aspect; the implementation of the solution and the beneficial effects brought by the above-mentioned third aspect or fourth aspect can refer to the specific description of the first aspect and the second aspect respectively; the implementation of the solution and the beneficial effects brought by the above-mentioned fifth aspect to thirteenth aspect can refer to the first aspect to the fourth aspect respectively. For the sake of brevity, no longer be described here. BRIEF DESCRIPTION OF DRAWINGS

[0270] FIG. 1 is a schematic diagram of a wireless communication system suitable for embodiments of the present application.

[0271] FIG. 2 is a schematic diagram of an AI / ML workflow.

[0272] FIG. 3 is a schematic diagram of a training process of reinforcement learning.

[0273] FIG. 4 is a possible use case of reinforcement learning.

[0274] FIG. 5 is a schematic diagram of a system architecture suitable for embodiments of the application.

[0275] FIG. 6 is a schematic diagram of a communication method provided by embodiments of the application.

[0276] FIG. 7 is a schematic diagram of a communication method provided by embodiments of the application.

[0277] FIG. 8 is a schematic diagram of a communication method provided by embodiments of the application.

[0278] FIG. 9 is a schematic diagram of a communication method provided by embodiments of the application.

[0279] FIG. 10 is a schematic block diagram of a communication apparatus provided by embodiments of the application.

[0280] FIG. 11 is a schematic block diagram of another communication apparatus provided by embodiments of the application. DETAILED DESCRIPTION

[0281] The technical solutions in the application will be described below with reference to the drawings.

[0282] Before introducing the solutions of the application, the following points are explained.

[0283] (1) The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and the appended claims of the application, the meaning of "a plurality" or "a plurality of" is two or more; the singular expression "a", "an", "said", "the", "above", "preceding", "this" and "that" are intended to include, for example, the expression "one or more" unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of the present application, "at least one", "at least one", "one or more" means one, two or more. "And / or", which describes the relationship between the associated objects, means that there can be three relationships, for example, A and / or B, which can represent the following cases: A exists alone, A and B exist together, B exists alone, where A and B can be singular or plural. In the textual description of the present application, the character " / " generally represents a "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b and c, can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c. Where a, b and c can be single or multiple.

[0284] (2) The ordinal numbers "first", "second", "#1", "#2" and the like mentioned in the embodiments of the present application are used to distinguish a plurality of objects, and are not used to limit the size, content, order, time sequence, priority or importance of the plurality of objects. For example, the first information and the second information can be the same information or different information, and such names do not indicate that the contents, sizes, application scenarios, sending / receiving ends, priorities or importance of the two pieces of information are different. In addition, the numbering of steps in each embodiment introduced in the present application is only for the purpose of distinguishing different steps, and the numbering of steps is not used to limit the order between steps unless otherwise stated.

[0285] (3) In the present application, "sending" and "receiving" represent the direction of signal transmission, and "transmitting" can include at least one of sending and / or receiving. For example, "sending information to XX" can be understood as that the destination of the information is XX, which can include direct sending through the air interface, or indirect sending through the air interface by other units or modules. "Receiving information from YY" can be understood as that the source of the information is YY, which can include direct receiving from YY through the air interface, or indirect receiving from YY through the air interface by other units or modules. "Sending" can also be understood as "output" of a chip interface, and "receiving" can also be understood as "input" of a chip interface. In other words, sending and receiving can be carried out between devices, such as between network devices and terminal devices, or can be carried out within a device, such as between components, modules, chips, software modules or hardware modules within a device through a bus, wire or interface.

[0286] (4) In the present application, the terms "comprising", "having", "including" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device containing a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices; or, information #A includes B, which can be equivalent to carrying the entire content of B in information #A.

[0287] (5) In the present application, "indication" can include direct indication and indirect indication. When describing that certain indication information indicates A, it can include that the indication information directly indicates A or indirectly indicates A, unless otherwise stated, which does not mean that A is necessarily carried in the indication information. Wherein, direct indication information A means including the information A; implicit indication information A means indicating information A by the corresponding relationship between information A and information B and directly indicating information B. Wherein, the corresponding relationship between information A and information B can be pre-defined, pre-stored, pre-burned or pre-configured.

[0288] (6) In the present application, the reference "one embodiment" or "some embodiments" and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in some possible implementations", "in other possible implementations" and the like appearing in different places in the present application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variations mean "include but are not limited to", unless otherwise specifically emphasized.

[0289] (7) In various embodiments of the present application, the terms and / or descriptions of different embodiments are consistent and can be referred to each other if there is no special description and logical conflict. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0290] (8) In the present application, the words such as "exemplary", "for example", etc. are used to represent examples, illustrations or descriptions. Any embodiment or design scheme described as "exemplary" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "exemplary" is intended to present the concept in a specific way. In the embodiments of the present application, "of", "corresponding" and "corresponding" are sometimes used interchangeably. It should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent.

[0291] (9) In the present application, the descriptions such as "when", "based on the case" and the like all refer to the objective situation in which the device will make corresponding processing, not limited to time, and also does not require the device to have a judgment action when implemented, nor does it mean that there are other limitations.

[0292] The technical solutions of the embodiments of the present application can be applied to various communication systems, such as: long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD), 5G system, or new radio (NR) and future communication systems. The technical solutions provided in the present application can also be applied to device to device (D2D) communication, vehicle-to-everything (V2X) communication, machine to machine (M2M) communication, machine type communication (MTC), and internet of things (IoT) communication system.

[0293] In addition, the embodiments of the present application are applicable to homogeneous network and heterogeneous network scenarios, and there is no limitation on the transmission point, which can be multi-point cooperative transmission systems between macro base stations and macro base stations, micro base stations and micro base stations, and macro base stations and micro base stations. The embodiments of the present application are applicable to low frequency scenarios and high frequency scenarios, terahertz, optical communication, etc.

[0294] A device in a communication system can transmit or receive signals to or from another device. The signals can include reference signals, information, signaling, or data, etc. In this application, a device can be replaced with an entity, a network entity, a communication device, a communication module, a node, a communication node, etc.

[0295] FIG. 1 is a schematic diagram of a wireless communication system 100 suitable for embodiments of the present application. As shown in FIG. 1, the wireless communication system includes a radio access network 100. The radio access network 100 can be a next generation (e.g., a future communication network or a higher version) radio access network, or a legacy (e.g., 5G, 4G, 3G or 2G) radio access network. One or more terminal devices (120a-120j, collectively referred to as 120) can be connected to each other or connected to one or more network devices (110a, 110b, collectively referred to as 110) in the radio access network 100. Network elements in the wireless communication system are connected through interfaces (e.g., NG, Xn) or air interfaces.

[0296] FIG. 1 is only a schematic diagram. The wireless communication system can further include other devices, such as a core network (CN) device, a wireless relay device, and / or a wireless backhaul device, etc., which are not shown in FIG. 1.

[0297] In practical applications, the wireless communication system can include multiple network devices at the same time, or can include multiple terminal devices at the same time, which is not limited. One network device can serve one or more terminal devices at the same time. One terminal device can access one or more network devices at the same time. Embodiments of the present application do not limit the number of terminal devices and network devices included in the wireless communication system.

[0298] The above-mentioned communication system suitable for embodiments of the present application is only an example. The communication system suitable for embodiments of the present application is not limited to this. Any communication system capable of realizing the functions of the above-mentioned devices is suitable for embodiments of the present application.

[0299] The terminal device in the embodiments of the present application can refer to a user equipment (user equipment, UE), an access terminal, a user unit, a user station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent or a user device. The terminal device can also be a cellular phone, a cordless phone, a session initiation protocol (session initiation protocol, SIP) phone, a wireless local loop (wireless local loop, WLL) station, a personal digital assistant (personal digital assistant, PDA), a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a wearable device, a terminal device in a 5G network or a terminal device in a future evolved public land mobile network (public land mobile network, PLMN) and the like. The embodiments of the present application are not limited thereto.

[0300] Among them, the wearable device can also be called a wearable smart device, which is a general term for devices that can be designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing and shoes, etc. The wearable device is a portable device that can be directly worn on the body or integrated into the user's clothes or accessories. The wearable device is not only a hardware device, but also a powerful function realized through software support and data interaction, cloud interaction. The general wearable smart device includes full function, large size, and can realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, etc., and only focuses on a certain application function, and needs to cooperate with other devices such as a smart phone, such as various smart wristbands, smart jewelry and the like for monitoring body signs.

[0301] In addition, the terminal device can also be a terminal device in an internet of things (internet of things, IoT) system. IoT is an important part of future information technology development, and its main technical feature is to connect objects through communication technology and network, so as to realize the intelligent network of man-machine interconnection and interconnection.

[0302] It should be understood that the specific form of the terminal device is not limited in the present application.

[0303] The network device can be a device in a wireless network. For example, the network device can be a device deployed in a wireless network to provide wireless communication functions for terminal devices. For example, the network device can be a radio access network (RAN) node that accesses terminal devices to a wireless network. Wherein, the RAN can be connected with a core network (for example, a core network of long term evolution (LTE), or a core network of 5G, etc.).

[0304] The network device in the embodiments of the present application can be an access network device, including but not limited to various types of base stations, such as next generation node B (gNodeB, gNB), evolved node B (eNB), or base station device in future evolved communication system, can also be an enabling server, wearable device, vehicle-mounted device, wireless relay node, wireless backhaul node, transmission point (TP), or transmission and reception point (TRP), etc., can also be one or a group of antenna panels (including multiple antenna panels) of a base station, or can also be a network node constituting a base station, such as a bandwidth based unit (BBU), or a distributed unit (DU), etc. Wherein, the base station can be a macro base station, a micro base station, a pico base station, a small station, a relay station or a balloon station, etc.

[0305] The communication between the access network device and the terminal device follows a certain protocol layer structure. The protocol layer can include a control plane protocol layer and a user plane protocol layer. The control plane protocol layer can include at least one of the following: radio resource control (RRC) layer, packet data convergence protocol (PDCP) layer, radio link control (RLC) layer, media access control (MAC) layer, or physical (PHY) layer, etc. The user plane protocol layer can include at least one of the following: service data adaptation protocol (SDAP) layer, PDCP layer, RLC layer, MAC layer, or physical layer, etc.

[0306] The network device in the embodiments of the present application can also be a core network device, including but not limited to: an access and mobility management function (AMF) network element, a session management function (SMF) network element, a user plane function (UPF) network element, a policy control function (PCF) network element, or a unified data management function (UDM) network element, etc.

[0307] The application layer network element refers to a network device responsible for processing application layer protocols in a computer network, including but not limited to: a data collection application function (DCAF) network element, a provisioning application function (PAF) network element, an event consumer application function (ECAF) network element, etc.

[0308] It can be understood that all or part of the functions of the network device or the terminal device in the present application can also be implemented by software functions running on hardware, or by virtualized functions instantiated on a platform (such as a cloud platform).

[0309] In some deployments, the network device mentioned in the embodiments of the present application can be a device including a centralized unit (CU), a DU, or a device including a CU and a DU, or a control plane CU node (central unit-control plane (CU-CP)) and a user plane CU node (central unit-user plane (CU-UP)) and a DU node. For example, the network device can include a gNB-CU-CP, a gNB-CU-UP, and a gNB-DU.

[0310] In some deployments, wireless access is facilitated by a terminal being cooperatively assisted by multiple RAN nodes that respectively implement part of the functionalities of a base station. For example, a RAN node can be a CU, a DU, a CU-CP, a CU-UP, or a radio unit (RU), etc. A CU and a DU can be separately configured or can be included in the same network element, e.g., a BBU. A RU can be included in a radio frequency device or a radio frequency unit, e.g., a remote radio unit (RRU), a massive active antenna system (AAU), or a remote radio head (RRH).

[0311] In different systems, a CU (or a CU-CP and a CU-UP), a DU, or a RU can also have different names, but those skilled in the art can understand their meanings. For example, in an ORAN system, a CU can also be referred to as an O-CU (open CU), a DU can also be referred to as an O-DU, a CU-CP can also be referred to as an O-CU-CP, a CU-UP can also be referred to as an O-CU-UP, and a RU can also be referred to as an O-RU. Any of the CUs (or CU-CPs, CU-UPs), DUs, and RUs in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.

[0312] For the correspondence between a network element in an open radio access network (ORAN) system and the protocol layer functions that can be implemented by the network element, refer to Table 1 below. In Table 1, O- can represent “open”.

[0313] Table 1

[0314] In order to support machine learning functions in a wireless communication system, an artificial intelligence (AI) node can also be introduced in the wireless communication system.

[0315] Optionally, the communication system also includes at least one AI node.

[0316] Optionally, the AI node is deployed in one or more of the following: a network device, a terminal device, a core network, or a positioning device; or the AI node can also be separately deployed, such as being deployed at a location other than any of the above devices. The AI node can communicate with other devices in the communication system, which can be one or more of the following: a network device, a terminal device, a core network element, or a sensing device.

[0317] Optionally, the AI node is configured to perform AI-related operations. As an example, the AI-related operations can include one or more of model failure testing, model performance testing, model training testing, or data collection, etc.

[0318] For example, the network device can forward the data related to the AI model reported by the terminal device to the AI node, and the AI node performs the AI-related operations. As another example, the network device or the terminal device can forward the data related to the AI model to the AI node, and the AI node performs the AI-related operations. As another example, the AI node can send one or more of the outputs of the AI-related operations, such as a trained neural network model, model evaluation, or test results, etc., to the network device and / or the terminal device. For example, the AI node can directly send the outputs of the AI-related operations to the network device and the terminal device. As another example, the AI node can send the outputs of the AI-related operations to the terminal device through the network device. As another example, the AI node can send the outputs of the AI-related operations to the network device through the terminal device.

[0319] It can be understood that the number of AI nodes is not limited in the present application. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on functions, such as different AI nodes being responsible for different functions.

[0320] It can also be understood that the AI node can be a separate device, or can be integrated into the same device to implement different functions, or can be a network element in a hardware device, or can be a software function running on a dedicated hardware, or a virtualized function instantiated on a platform (e.g., a cloud platform), and the specific form of the AI node is not limited in the present application.

[0321] For example, the AI node can be an AI network element or an AI module.

[0322] It should be understood that FIG. 1 is an example of a communication system applicable to the embodiments of the present application, and is a simplified schematic diagram for ease of understanding. The above communication system can also include other network devices or can also include other terminal devices, which are not shown in FIG. 1. The communication system to which the embodiments of the present application are applied is not limited to this.

[0323] It should also be understood that FIG. 1 is only an example of the application scenario of the embodiments of the present application, and the application scenario of the method is not limited in the present application. The present application can be applied to network device and network device communication, network device and terminal device communication, terminal device and terminal device communication, etc., and the embodiments of the present application are not limited to this.

[0324] The basic principle of AI is to combine massive data with super strong operation processing capability and intelligent algorithms to establish an AI model for solving specific problems, so that the AI model can automatically induce and learn potential patterns or features from data, thereby realizing a thinking mode close to that of humans.

[0325] An AI model, i.e., an AI algorithm (or AI operator), is a general term for a mathematical algorithm constructed based on the principle of artificial intelligence, and is also the basis for solving specific problems using AI. According to different specific methods and / or technologies for realizing artificial intelligence, the AI model can also be referred to as a machine learning model, a deep learning model, or a reinforcement learning model.

[0326] For the convenience of understanding the embodiments of the present application, the basic concepts involved in the present application are described.

[0327] The following first specifically describes machine learning, machine learning model, deep learning, deep learning model, neural network, reinforcement learning, etc.

[0328] Machine learning (ML) is a method for realizing artificial intelligence, and the goal of this method is to design and analyze some algorithms (i.e., models) that allow computers to automatically "learn". The designed algorithm is called a machine learning model.

[0329] A machine learning model is a type of algorithm that automatically analyzes rules from data and uses the rules to predict unknown data. Machine learning models are diverse, and according to whether the model needs to rely on the labels corresponding to the training data during training, the machine learning model can be divided into: 1, supervised learning model; 2, unsupervised learning model.

[0330] 1. Supervised learning model: a model obtained after determining the parameters of the initial AI model according to the data in the given training data set and the labels corresponding to the data in the training data set. The process of determining the parameters of the initial AI model using the data in the training data set and the labels corresponding to the data is also called supervised learning (or supervised training). The labels of the data in the training data set are usually manually annotated and used to identify the correct answer of the data on a specific task. Typical supervised learning models include support vector machines, neural network models, logistic regression models, decision trees, naive Bayes models, Gaussian discriminant models, etc. Supervised learning models are usually used for classification or regression.

[0331] 2. Unsupervised learning model: a model obtained after determining the parameters of an initial AI model according to unlabeled data in a given training data set. The process of determining the parameters of the initial AI model using unlabeled training data is also called unsupervised learning (or unsupervised training). Through unsupervised learning, the model can discover meaningful information and associations in the data and then make predictions of the results of the data. There are many kinds of unsupervised learning models, and the commonly used ones include clustering, principal component analysis (PCA), anomaly detection model, autoencoder, generative adversarial network (GAN), etc.

[0332] Any AI model needs to be trained before being used to solve a specific technical problem. The training of an AI model refers to the process of using a specified initial model to calculate training data, adjusting the parameters in the initial model according to the calculation results, so that the model gradually learns certain rules and has specific functions. The AI model with stable functions after training can be used for inference. The inference of an AI model is the process of using a trained AI model to calculate input data and obtain a predicted inference result.

[0333] The most common way is to train an AI model in a supervised manner. For example: most deep learning models are trained in a supervised manner.

[0334] The following describes the most widely used supervised training method for deep learning models.

[0335] In the training phase, a training set for the deep learning model needs to be constructed based on the target. The training set includes multiple training data, and each training data is provided with a label. The label of the training data is the correct answer of the training data on a specific problem, and the label can represent the target of training the deep learning model using the training data. For example: for a deep learning model to be trained to recognize different animals, the training set can include multiple images of different animals (i.e. training data), and each image can have a label identifying the type of animal contained therein, such as cat, dog. In this example, the type of animal corresponding to each image is the label of the training data.

[0336] When training a deep learning model, training data can be inputted to the deep learning model after parameter initialization in batches, and the deep learning model calculates (i.e., inference) the training data to obtain a prediction result for the training data. The prediction result obtained through inference and the label corresponding to the training data are used as data for calculating the loss according to the loss function. The loss function is a function used to calculate the gap (i.e., loss value) between the prediction result of the model for the training data and the label of the training data in the model training stage. The loss function can be implemented by using different mathematical functions, and the expressions of commonly used loss functions are: mean square error loss function, logarithmic loss function, least squares method, etc.

[0337] The loss value calculated based on the loss function can be used to update the parameters of the deep learning model. The gradient descent method is commonly used for parameter updating. The training of the model is a repeated iterative process. In each iteration, different training data is inferred, and the loss value is calculated. The goal of multiple iterations is to continuously update the parameters of the deep learning model to find the parameter configuration that makes the loss value of the loss function the lowest or tends to be stable.

[0338] In the training stage, in order to make the training efficiency of the model and the performance of the model after training more optimal, some reasonable hyperparameters need to be set for training. The hyperparameters of the deep learning model refer to a class of parameters that cannot be obtained by learning the training data or cannot be changed due to the training data in the training process. It is a concept relative to the parameters in the model. The hyperparameters of the deep learning model are usually set by humans according to experience or experiments. The hyperparameters include: learning rate, batch size, network structure hyperparameters (such as: network layer number (also known as depth), interaction mode between network layers, number of convolution kernels and size of convolution kernels, activation function), etc. Among them, the learning rate as a hyperparameter is used to control the amplitude of the parameter weight update of the model in the training process, which greatly affects the speed and accuracy of the training.

[0339] The trained deep learning model can be used for inference of input data. In the inference stage, the data of the actual application scenario is usually used as input data, and the inference of the trained deep learning model can obtain an inference result. The inference stage is the actual application of the trained deep learning model, which can quickly use the AI capability to solve specific technical problems. Nowadays, there are many AI application scenarios, and the inference of the deep learning model can also be used in various application scenarios, such as personnel identification scenarios for access control security systems, video pornography and violence detection, express delivery order number detection and identification, etc.

[0340] The above only takes the training of the most typical deep learning model as an example, and the training of other types of models has slight differences, but the principle is similar, which is to infer the training data, adjust the parameters in the model according to the inference result, and obtain the parameter combination that stabilizes the performance of the model as the goal.

[0341] Training is also mainly divided into supervised training and unsupervised training, and the training process of the foregoing deep learning model belongs to supervised training. Taking images as an example, if an AI model is trained unsupervisedly, the training images in the training image set used for training are not labeled, and the training images in the training image set are input into the AI model in turn, and the AI model gradually identifies the association and potential rules between the training images in the training image set until the AI model can be used to judge or identify the type or characteristics of the input image. For example, clustering, after the AI model used for clustering receives a large number of training images, it can learn the characteristics of each training image and the association and difference between the training images, and automatically divide the training images into multiple types. Different task types can use different AI models, some AI models can only be trained in a supervised learning manner, some AI models can only be trained in an unsupervised learning manner, and some AI models can be trained in both a supervised learning manner and an unsupervised learning manner.

[0342] Deep learning is a new technical field in the research process of machine learning. Specifically, deep learning is a method of learning based on deep representation learning of data in machine learning. Deep learning explains data by establishing a neural network that simulates the analysis and learning of the human brain.

[0343] In the field of AI, deep learning is a learning technology based on a deep neural network algorithm. A deep learning model includes an input layer, a hidden layer, and an output layer, and uses multiple nonlinear transformations to process data.

[0344] In machine learning methods, almost all features need to be determined by industry experts, and then the features are encoded. However, deep learning algorithms attempt to learn features from data themselves, and algorithms designed according to the deep learning idea are called deep learning models.

[0345] A typical structure of a current deep learning model is a deep neural network. A neural network is a mathematical or computational model that mimics the structure and function of a biological neural network (the central nervous system of an animal, especially the brain). A neural network performs computation by a large number of neuron connections. A neural network can include multiple neural network layers with different functions, each layer including parameters and computation rules. Different layers in a neural network have different names according to different computation formulas or functions, for example, a layer performing convolution computation is called a convolution layer, which is often used for feature extraction of an input signal (e.g., an image). A neural network can also be composed of multiple sub-neural networks. Neural networks with different structures can be suitable for different scenarios (e.g., classification, recognition) or provide different effects when used for the same scenario. The structure of a neural network can be different in one or more of the following aspects: the number of network layers in the neural network, the order of the network layers, the weights, parameters, or computation formulas in each network layer. There are many different neural networks with high accuracy for different application scenarios such as recognition or classification in the industry. Some neural networks can be trained by a specific data set and then used alone or combined with other neural networks (or other functional modules) to complete a task.

[0346] In other words, a deep learning model is actually a machine learning model with a complex structure of a neural network. According to whether a label corresponding to training data is needed during training of a deep learning model, a deep learning model can also be divided into a supervised learning model and an unsupervised learning model, which will not be described here. Classical deep learning models include convolutional neural network (CNN), recurrent neural network (RNN), recursive neural network (RNN), etc.

[0347] In order to improve the intelligent and automated level of the network, AI and ML technologies are being applied in more and more fields, and there are also multiple related topics being studied for network intelligence (e.g., research on the life cycle management of models by the 3rd Generation Partnership Project (3GPP) working group).

[0348] As shown in FIG. 2, FIG. 2 is a schematic diagram of an AI / ML workflow. The AI / ML workflow mainly includes ML model training, ML model testing, AI / ML inference emulation, ML model deployment, AI / ML inference, etc. The use cases and management functions involved in each of the above processes can be discussed separately.

[0349] Reinforcement learning (RL): also known as re-education, evaluation learning or enhancement learning, used to describe and solve the problem of learning a strategy by an agent in the process of interacting with the environment to maximize the reward or achieve a specific goal. With the research of network intelligence, reinforcement learning has begun to attract attention, and the full-process management and operation capability of AI / ML in the 5G system (5G system, 5GS) can also be studied to support various AI / ML technologies including reinforcement learning.

[0350] Reinforcement learning is a learning method of an agent in a trial-and-error manner, which determines the change of state and the corresponding reward according to the interaction between each action and the environment, so as to guide the behavior according to the reward. The training goal of reinforcement learning is to make the agent obtain the maximum reward. Reinforcement learning does not require a training data set. The reinforcement signal (i.e. reward) provided by the environment in reinforcement learning evaluates the good and bad of the action, rather than telling the reinforcement learning system how to produce the correct action. Since the information provided by the external environment is little, the agent must learn by itself. In this way, the agent obtains knowledge in the action-evaluation (i.e. reward) environment and improves the action plan to adapt to the environment.

[0351] FIG. 3 is a schematic diagram of a training process of reinforcement learning. As shown in FIG. 3, reinforcement learning mainly includes four elements: agent, environment state, action and reward, wherein the input of the agent is the state and the output is the action.

[0352] In the prior art, the training process of reinforcement learning is: through multiple interactions between the agent and the environment, the action, state and reward of each interaction are obtained; the multiple sets of information (action, state and reward) are used as training data to train the agent once. The above process is used to train the agent in the next round until the convergence condition is met.

[0353] The process of obtaining the action, state and reward of an interaction is shown in FIG. 3. The current state s(t) of the environment is input to the agent to obtain the action a(t) output by the agent. The reward r(t) of this interaction is calculated according to the relevant performance indicators of the environment under the action a(t). At this point, the action a(t), state s(t) and reward r(t) of this interaction are obtained. The action a(t), state s(t) and reward r(t) of this interaction are recorded for subsequent use to train the agent. The next state s(t+1) of the environment under the action a(t) is also recorded to enable the next interaction between the agent and the environment.

[0354] The agent refers to an entity that can think and interact with the environment. For example, the agent can be a computer system or a part of a computer system in a certain environment. The agent can perceive the environment according to its own perception, follow existing instructions or learn autonomously, and communicate and cooperate with other agents to autonomously complete the set goals in the environment. The agent can be a software or a combination of software and hardware entity.

[0355] Reinforcement learning can be applied to different fields. Taking the management data analysis (MDA) use case coverage problem analysis as an example, when RL is applied to the coverage problem analysis use case, the RL agent can be an ML model for coverage problem analysis, and the RL environment can be a simulation environment. The action can be the value of the network adjustable parameter, such as the recommended action (e.g., changing the transmission power of the NR sector carrier frequency, etc.). The state can be the performance measurement (PM) or key performance indicator (KPI) of the network, such as the distribution of the reference signal received power (RSRP), etc. The reward can be the score of the RL performance indicator, which is used to evaluate the PM or KPI.

[0356] As shown in FIG. 4, FIG. 4 is a possible reinforcement learning use case. In this case, the RL environment can be a network digital twin (NDT) function network element or a gNB.

[0357] For example, a model training provider / model training service provider (e.g., machine learning training producer, MLT producer) can allow a model training provider / model training service caller (e.g., machine learning training consumer, MLT consumer) to query available RL capabilities; the MLT producer can allow the MLT consumer to specify an RL environment; and the MLT producer can report performance of an ML model that is finally trained by RL.

[0358] The above description of the terms is only for the convenience of understanding and does not limit the protection scope of the embodiments of the present application.

[0359] In order to ensure the training efficiency and the performance of the model obtained by training, machine learning training (e.g., performing an RL process in a real communication network) can be selected to be performed in a real communication network. In a real communication network, the RL process performs experiments and learns from mistakes, which can have a negative impact on the real communication network and can cause the quality of service of users of the communication network to decrease.

[0360] Therefore, embodiments of the present application provide a communication method and device, which can balance the machine learning training efficiency and the degree of influence on the performance of the communication network, and ensure the quality of service of users of the communication network.

[0361] FIG. 5 is a schematic diagram of a system architecture suitable for embodiments of the present application. As shown in FIG. 5, in the 3GPP network domain, a network management system (NMS) can serve as a model training consumer (MLT consumer), and the NMS is responsible for the operation, management and maintenance functions of the network, and can also be referred to as a cross-domain management system. A model training provider (MLT producer) can be an element management system (EMS), which is used to manage one or more network elements of a certain category, and can also be referred to as a domain management system or a single-domain management system; the MLT producer can also be an EMS-managed network element (such as a base station gNB of a RAN, in a mobile communication system, the gNB can be a device that connects the fixed part and the wireless part, and is connected to the mobile terminal through the air channel) or a CN function (such as a network data analytics function (NWDAF) network element, the NWDAF network element has AI training, inference and other intelligent computing functions) and the like. The MLT consumer can invoke the model training service provided by the MLT producer. In addition, in the ORAN network domain, a service management and orchestration function (SMO) can serve as an MLT consumer, and the SMO directly managed network elements can serve as an MLT producer. The role of the SMO in the network architecture is similar to that of the NMS, and the SMO can be responsible for the operation, management and maintenance of various network services and orchestration functions, and the directly managed network elements can be heterogeneous, such as directly managing EMS, gNB, NWDAF and the like.

[0362] The AI / ML management service (MnS) interface in operations, administration, and maintenance (OAM) is a key component for managing and coordinating network resources. The MnS interface allows different network management and orchestration entities to interact for functions such as network resource configuration, performance monitoring, fault management, and the like. The providing entity of the management service can be referred to as a management service producer, an AI / ML management service producer (AI / ML MnS producer), or an AI / ML update management service producer. The calling entity of the management service can be referred to as a management service consumer, an AI / ML management service consumer (AI / ML MnS consumer), or an AI / ML update management service consumer. Model training (MLT) is a possible implementation of the MnS.

[0363] Model inference refers to a process of performing inference or prediction based on an ML model to generate an inference result or a prediction result. This process can also be referred to as a process of using an ML model.

[0364] The capability or function of performing model inference can be referred to as a model inference function or simply an inference function, an ML inference function. Illustratively, the model inference function can be deployed in the management service producer described above; the model inference function can also be deployed in the EMS; the model inference function can also be deployed in a network device; the model inference function can also be deployed in a core network element such as a NWDAF network element; the model inference function can also be deployed in other devices, for which the present application does not make special limitations. In the present application, gNB, EMS, or NMS can refer to a combination of hardware and software, for which the embodiments of the present application do not make limitations.

[0365] The communication network to which the embodiments of the present application are applicable can include a real communication network (for example, a network specified by the 3rd Generation Partnership Project (3GPP) or open RAN (ORAN)), or the communication network to which the embodiments of the present application are applicable can include a simulation network that simulates a real network. The embodiments of the present application do not make limitations thereto.

[0366] The communication method provided by the embodiments of the present application will be described in detail below with reference to FIGS. 6-9.

[0367] In the present application, “communication network performance” and “network performance” can be used interchangeably, and in the absence of special explanations, they represent the same meaning.

[0368] In the scheme provided by the embodiments of the present application, the inference function can refer to a function associated with an AI / ML inference name (aimlinferencename) (for example, the inference function can include a function corresponding to the values of the management data analytics type (the values of the MDA type) or a function corresponding to the analytics identity(s) of network data analytics function (Analytics ID(s) of NWDAF), etc.).

[0369] For example, the aimlinferencename can include allowedValues, the allowedValues can include vendor's specific extensions, and any one of the following three: the values of the management data analytics (MDA) type (for example, refer to the protocol 3GPP TS 28.104); the analytics identity (identity, ID) of the network data analytics function (NWDAF) (Analytics ID(s) of NWDAF) (for example, refer to the protocol 3GPP TS 23.288); and the types of inference for the radio access network (RAN).

[0370] FIG. 6 shows a schematic diagram of a communication method 600 provided by the embodiments of the present application. The implementation of the communication method 600 has two possible situations (situation #1 and situation #2).

[0371] Situation #1: In some possible implementation manners, the communication method can be applied between the first device and the second device, and the method 600 includes:

[0372] S610a, determining first information, the first information being used to indicate a first performance requirement, the first performance requirement indicating a requirement of performing machine learning training on the performance of the communication network.

[0373] Determining the first information can be receiving the first information.

[0374] Determining the first information can be understood as: the MnS producer receives information configured for the MnS consumer to create a machine learning training request management object instance (MLTrainingRequest MOI).

[0375] Exemplarily, a related attribute of the in-network performance requirement can be added in a machine learning training information object class (MLTrainingRequest IOC).

[0376] Exemplarily, the related attribute of the in-network performance requirement can include a reinforcement learning network performance requirement (rlnetworkperformancerequirement).

[0377] S620c, performing machine learning training based on the first information.

[0378] Performing machine learning training based on the first information can be understood as: the MnS producer receives a create management object instance (MOI) request of the MLTrianingRequest and creates the MLTrianingRequest MOI.

[0379] Specifically, after determining the first information indicating the first performance requirement, the first device can perform machine learning training according to the requirement of the machine learning training on the communication network performance indicated by the first performance requirement.

[0380] The communication method provided by the embodiments of the present application can be applied to machine learning (for example, can be applied to reinforcement learning).

[0381] Exemplarily, the communication method provided by the embodiments of the present application can optimize reinforcement learning. The NMS configures a reinforcement learning strategy for model training functions in the network to perform reinforcement learning parameter configuration, so as to balance the training efficiency of reinforcement learning training in the network, the model performance of the model obtained by reinforcement learning training and the communication network performance fluctuation caused by reinforcement learning training, and can balance the training efficiency and the model performance while avoiding excessive communication network performance fluctuation.

[0382] Since the efficiency of machine learning training and the degree of impact on the communication network performance of the real communication network are mutually restricted, by performing machine learning training based on the requirements of the communication network performance for performing machine learning training, on the one hand, the low efficiency of machine learning training caused by too long training time of performing machine learning training or the poor timeliness of the model obtained by machine learning training can be avoided; on the other hand, the service quality of the communication network users cannot meet the demand caused by too great impact of performing machine learning training on the communication network performance of the real communication network.

[0383] Based on the scheme provided in the embodiments of the present application, by performing machine learning training based on the first information indicating the first performance requirement, the efficiency of machine learning training and the degree of impact on the communication network performance of the communication network can be balanced, and the service quality of the communication network users can be guaranteed.

[0384] Case #2: In some possible implementation manners, the communication method can be applied between the third device and the fourth device, and the method 600 includes:

[0385] S610b, determining the first information, wherein the first information is used to indicate at least one of the following: the first performance requirement, the first performance requirement indicating the requirement of the communication network performance for performing machine learning training; or the first training time requirement, the first training time requirement indicating the requirement of the training time for performing machine learning training.

[0386] The determination of the first information can be receiving the first information.

[0387] The determination of the first information can be understood as that the MnS producer receives the information configured by the MnS consumer for creating the machine learning training request management object instance (MLTrainingRequest MOI).

[0388] For example, the related attribute of the performance requirement of the real communication network can be added in the machine learning training request information object class (MLTrainingRequest IOC).

[0389] For example, the related attribute of the performance requirement of the real communication network can include the reinforcement learning network performance requirement (rlnetworkperformancerequirement).

[0390] S620c, performing the machine learning training based on the first information.

[0391] The performing the machine learning training based on the first information can be understood as that the MnS producer receives a create management object instance (MOI) request of the MLTrianingRequest and creates the MLTrianingRequest MOI.

[0392] Specifically, after determining the first information indicating the first performance requirement and / or the first training time requirement, the third apparatus can perform the machine learning training according to a requirement of the machine learning training on the communication network performance and / or the training time indicated by the first performance requirement and / or the first training time requirement.

[0393] Since the efficiency of the machine learning training and the degree of influence of the machine learning training on the communication network performance of the real communication network are mutually restricted, by performing the machine learning training based on the requirement of the machine learning training on the communication network performance and / or the training time, on the one hand, it can avoid the situation that the efficiency of the machine learning training is too low or the model obtained by the machine learning training does not meet the timeliness due to too long training time of the machine learning training; on the other hand, it can avoid the situation that the service quality of the communication network users does not meet the requirement due to too great influence of the machine learning training on the communication network performance of the real communication network.

[0394] Based on the scheme provided in the embodiments of the present application, by performing the machine learning training based on the first information indicating the first performance requirement and / or the first training time requirement, the efficiency of the machine learning training and the degree of influence of the machine learning training on the communication network performance can be balanced, and the service quality of the communication network users can be guaranteed.

[0395] It should be understood that, the specific description of the communication method 600 hereinafter is applicable to both the case #1 and the case #2, unless otherwise specified or contradictory.

[0396] For example, the first apparatus / third apparatus can be a device where an element management system (EMS) is located or a device where a base station (gNB) is located, or the first apparatus / third apparatus can be a smart controller (RIC).

[0397] In some possible implementation manners, the first device / third device can be referred to as a production entity, an EMS, a domain management node / system / function, a MnS producer, a single-domain management system (domain management system, for short), or a single-domain management function (domain management function, for short). Correspondingly, a device in which a management service consumer / user / caller (MnS consumer) that jointly implements machine learning training with the first device / third device is located can be referred to as a consumption entity, a network management system (NMS), a cross-domain management node / system / function, a MnS consumer, a cross-domain management system (CD-MnS), or a cross-domain management function (CD-MnF); and a service management and orchestration (function) (SMO).

[0398] For example, the first device / third device can determine the first information according to received indication information, for example, the indication information received by the first device / third device indicates the first information, and the first device / third device can determine the first information according to the received indication information. Alternatively, the first device / third device can determine the first information according to a requirement to be met by a key performance indicator (KPI) or performance measurement (PM) of a related communication network and / or a current performance status of the related communication network, for example, the first device / third device can determine, by calculation, a requirement of the machine learning training to the performance of the communication network according to the requirement to be met by the KPI or PM of the related communication network and / or the current performance status of the related communication network. The embodiments of the present application do not make any limitation in this regard.

[0399] For example, the first device / third device can calculate, by using a minimum KPI / PM requirement and a current communication network status, a value that can meet the minimum KPI / PM requirement and the current communication network status, and the calculated value is used to represent the requirement of the machine learning training to the performance of the communication network. The minimum KPI / PM requirement and the current communication network status are related to an inference function of machine learning corresponding to the machine learning training.

[0400] In some possible implementation manners, the requirement of the machine learning training on the communication network performance can include at least one of the following: a lower threshold of an allowed KPI or PM; or, a loss value of an allowed KPI or PM loss; or, a ratio of an allowed KPI or PM loss; or, a range of an allowed network fluctuation.

[0401] For example, the requirement of the machine learning training on the communication network performance can include: a plurality of KPIs and a KPI lower threshold / range of deviation / maximum loss ratio corresponding to each KPI (at this time, the requirement of the machine learning training on the communication network performance can also be information in a list form). Alternatively, the requirement of the machine learning training on the communication network performance can include: a single KPI and a KPI lower threshold / range of deviation / maximum loss ratio corresponding to the KPI; or, a plurality of PMs and a PM lower threshold / range of deviation / maximum loss ratio corresponding to each PM (at this time, the requirement of the machine learning training on the communication network performance can also be information in a list form); or, the requirement of the machine learning training on the communication network performance can include: a single PM and a PM lower threshold / range of deviation / maximum loss ratio corresponding to the PM.

[0402] The influence of the machine learning training based on the first information on the communication network performance can meet: an allowed KPI lower threshold; or, a loss value of an allowed KPI loss; or, a ratio of an allowed KPI loss; or, a range of an allowed network fluctuation; or, an allowed PM lower threshold; or, a loss value of an allowed PM loss; or, a ratio of an allowed PM loss.

[0403] In some possible implementation manners, in the case #1, the first information is further used to indicate a first training time requirement, and the first training time requirement indicates a requirement of the machine learning training on a training time.

[0404] Specifically, after the first information used to indicate the first performance requirement and the first training time requirement is determined, the first apparatus can perform the machine learning training according to the requirement of the machine learning training on the communication network performance and the requirement of the machine learning training on the training time indicated by the first performance requirement.

[0405] Based on the scheme provided in the embodiments of the present application, by performing the machine learning training based on the first information indicating the first training time requirement, the training time of the machine learning training can meet the expected time requirement, so that the implementation of the machine learning training is more reasonable.

[0406] Exemplarily, in the case #1 or the case #2, the information included in the first training time requirement can be obtained from that the management service consumer (MnS consumer) receives a training request for a model corresponding to a reinforcement learning function in a network, and the MnS consumer determines the information included in the first training time requirement according to a training time requirement corresponding to the training request received by the MnS consumer (for example, the training time requirement corresponding to the training request can be the fourth information).

[0407] For example, when the network management system (NMS) is the MnS consumer, the NMS can receive a training request for a model corresponding to a reinforcement learning function in a network and a training time requirement corresponding to the training request. When the MnS consumer initiates a training request to the MnS producer, the first training time requirement sent by the MnS consumer can include the information of the training time requirement corresponding to the training request.

[0408] In some possible implementation manners, the requirement for a training time for performing the machine learning training can include at least one of the following: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0409] The training time for performing the machine learning training based on the first information can meet the following requirements: a time point at which the machine learning training is expected to be completed; or an expected training duration.

[0410] In some possible implementation manners, in the case #1, the S620c includes: performing the machine learning training based on the first parameter, where the first parameter is determined according to the first performance requirement and a corresponding relationship, the corresponding relationship indicates a relationship between the first value and the training parameter, the first parameter belongs to the training parameter, and the first value corresponding to the first parameter meets the first performance requirement.

[0411] Exemplarily, the training parameter includes a learning rate, or a learning rate and a decay ratio of the learning rate.

[0412] Exemplarily, the training parameter includes an exploration rate, or an exploration rate and a decay ratio of the exploration rate.

[0413] Exemplarily, when the machine learning training is reinforcement learning (RL) training, the RL training performs machine learning according to a learning rate (or a learning rate and a decay ratio of the learning rate), or according to an exploration rate (or an exploration rate and a decay ratio of the exploration rate), and the machine learning training no longer obtains a new action when the learning rate or the exploration rate decays to 0.

[0414] For example, when the EMS is the MnS producer, the EMS can select a training parameter as the first parameter from the training parameters according to the correspondence between the existing first values and the training parameters, and the first value corresponding to the first parameter satisfies the first performance requirement according to the correspondence. For example, the first value includes the KPI / PM loss value, the ratio of KPI / PM loss, or the range of KPI / PM fluctuation of the communication network when the machine learning training is performed based on the training parameter.

[0415] For example, when the gNB is the MnS producer, the EMS can select a training parameter as the first parameter from the training parameters according to the correspondence between the existing first values and the training parameters, and the first value corresponding to the first parameter satisfies the first performance requirement according to the correspondence. The EMS can send indication information indicating the first parameter determined by the EMS to the gNB, and the gNB can determine the first parameter according to the received indication information and perform machine learning training according to the first parameter. At this time, the device where the EMS and the gNB are located can be one possible implementation of the first device.

[0416] For example, the first performance requirement is that the handover success rate loss does not exceed 5%, and in the correspondence between the existing first values and the training parameters, when the exploration rate is 10%, 20% or 30%, the handover success rate loss corresponding to the exploration rate is 3%, 5% or 7% respectively. Then the exploration rate of 20% can be selected as the first parameter.

[0417] In some possible implementations, in the case #2, S620c includes: performing the machine learning training based on the first parameter, wherein the first parameter is determined according to the first performance requirement and the correspondence, the correspondence indicates the relationship between the first values and the training parameters, the first parameter belongs to the training parameters, and the first value corresponding to the first parameter satisfies the first performance requirement and / or the first training time requirement.

[0418] For example, when the EMS is the MnS producer, the EMS can select a training parameter as the first parameter from the training parameters according to the correspondence between the existing first values and the training parameters, and the first value corresponding to the first parameter satisfies the first performance requirement according to the correspondence. For example, the first value includes the KPI / PM loss value, the ratio of KPI / PM loss, or the range of KPI / PM fluctuation of the communication network when the machine learning training is performed based on the training parameter.

[0419] Exemplarily, when the gNB is the MnS producer, the EMS can select a training parameter in the training parameters as the first parameter for performing the machine learning training according to the correspondence between the existing first values and the training parameters, and the first value corresponding to the first parameter satisfies the first performance requirement and / or the first training time requirement. The EMS can send the indication information indicating the determined first parameter to the gNB, and the gNB can determine the first parameter according to the received indication information and perform the machine learning training according to the first parameter. At this time, the apparatuses where the EMS and the gNB are located can be a possible implementation of the first apparatus.

[0420] In some possible implementation manners, the first information is further used to indicate a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

[0421] Exemplarily, the network element performing the machine learning training can belong to the plurality of first network elements indicated by the first information, and the network element performing the machine learning training satisfies the first performance requirement.

[0422] When the first information indicates the first performance requirement, the network element performing the machine learning training satisfies the first performance requirement.

[0423] For example, the plurality of network elements indicated by the first information include a network element #1, a network element #2 and a network element #3, wherein the communication network performance corresponding to the network element #1 is 80%, the communication network performance corresponding to the network element #2 is 85%, and the communication network performance corresponding to the network element #3 is 90% (as a possible example, the communication network performance corresponding to the network element #1, the network element #2 or the network element #3 can be the handover success rate). When the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 90%, the network element #3 can be selected as the network element performing the machine learning training; when the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 85%, any one of the network element #3 or the network element #2 can be randomly selected as the network element performing the machine learning training.

[0424] When the first information indicates the first performance requirement and the first training time requirement, the network element performing the machine learning training satisfies the first performance requirement and the first training time requirement.

[0425] For example, the multiple network elements indicated by the first information include network element #4, network element #5 and network element #6, wherein the communication network performance corresponding to the network element #4 is 80%, the training duration corresponding to the network element #4 is 300 ms, the communication network performance corresponding to the network element #5 is 85%, the training duration corresponding to the network element #5 is 350 ms, the communication network performance corresponding to the network element #6 is 90%, and the training duration corresponding to the network element #6 is 400 ms. When the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 85% and the upper limit of the training duration is 350 ms, the network element #5 can be selected as the network element for performing the machine learning training; when the first performance requirement indicated by the first information is that the lower limit of the communication network performance is 85% and the upper limit of the training duration is 400 ms, any one of the network element #5 or the network element #6 can be randomly selected as the network element for performing the machine learning training.

[0426] Based on the scheme provided in the embodiments of the present application, the network element for performing the machine learning training is selected from the multiple network elements indicated by the first information, so that the network element for performing the machine learning training is more reasonable.

[0427] In some possible implementation manners, the network element for performing the machine learning training belongs to multiple second network elements; and the first information is further used to indicate at least one of the following: an RL environment of the multiple second network elements; a training position of the multiple second network elements; and a training function of the multiple second network elements.

[0428] For example, the network element for performing the machine learning training is determined according to the related information of the second network element indicated by the first information, and the network element for performing the machine learning training satisfies the first information.

[0429] For example, the first information can indicate at least one of the RL environment of the multiple second network elements, the training position of the multiple second network elements, or the training function of the multiple second network elements, and the first device can select one RL environment from the RL environments of the multiple second network elements, or select one training position from the training positions of the multiple second network elements, or select one training function from the training functions of the multiple second network elements according to the first performance requirement (or according to the first performance requirement and the first training time requirement). The one RL environment, the one training position or the one training function satisfies the first performance requirement (or satisfies the first performance requirement and the first training time requirement). After the first device selects the one RL environment, the one training position or the one training function, the network element for performing the machine learning training can be determined according to the network element corresponding to the one RL environment, the one training position or the one training function.

[0430] Based on the scheme provided in the embodiments of the present application, the information of the plurality of network elements is indicated by the first information, and the network element performing the machine learning training meets the information of the plurality of network elements, so that the network element performing the machine learning training is more reasonable.

[0431] In some possible implementation manners, the method 600 further includes:

[0432] S630a, determining the network element performing the machine learning training according to the communication network performance of the plurality of first network elements.

[0433] For example, the EMS can determine the network element performing the machine learning training from the plurality of first network elements according to the communication network performance of the plurality of first network elements and the first performance requirement (or the first performance requirement and the first training time requirement).

[0434] In some possible implementation manners, the method 600 further includes:

[0435] S630b, determining the network element performing the machine learning training according to the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements.

[0436] For example, the EMS can determine one RL environment, one training position, or one training function meeting the first performance requirement / first training time requirement (or meeting the first performance requirement and the first training time requirement) according to any one of the RL environment of the plurality of second network elements, the training position of the plurality of second network elements, or the training function of the plurality of second network elements and the first performance requirement / first training time requirement (or the first performance requirement and the first training time requirement), and determine the network element performing the machine learning training from the plurality of second network elements according to the network element corresponding to the one RL environment, one training position, or one training function.

[0437] Based on the scheme provided in the embodiments of the present application, the network element performing the machine learning training is determined according to the network performance of the plurality of network elements, which can make the network element performing the machine learning training meet the first performance requirement / first training time requirement (or meet the first performance requirement and the first training time requirement), and is helpful for the implementation of the machine learning training.

[0438] In some possible implementation manners, the method 600 further includes:

[0439] S640, receiving second information sent by the second device / the fourth device, the second information being used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of performing the machine learning training on a communication network performance, the first performance requirement meeting the second performance requirement; or a second training time requirement, the second training time requirement indicating a requirement of performing the machine learning training on a training time, the first training time requirement meeting the second training time requirement, the first training time requirement being indicated by the first information, the first training time requirement indicating the requirement of performing the machine learning training on the training time.

[0440] After performing the machine learning training based on the first information, the method further includes performing the machine learning training based on the second information.

[0441] For example, the second information includes the second performance requirement and / or the second training time requirement.

[0442] For example, the first device can further receive the second information, the second information can indicate the second performance requirement and / or the second training time requirement, the first performance requirement meeting the second performance requirement, and the first training time requirement meeting the second training time requirement. When the machine learning training based on the first information does not meet the requirement (for example, the machine learning training is not completed within a time point at which the machine learning training is expected to be completed or a training time length expected to be completed), the first device can perform the machine learning training according to the second performance requirement and / or the second training time requirement indicated by the second information.

[0443] When the machine learning training based on the first information does not meet the requirement, the requirement of performing the machine learning training on the communication network performance and / or the requirement of performing the machine learning training on the training time can be relaxed, and the machine learning training can be performed again based on the relaxed requirement of performing the machine learning training on the communication network performance and / or the relaxed requirement of performing the machine learning training on the training time.

[0444] Based on the scheme provided in the embodiments of the present application, by performing the machine learning training again based on the second performance requirement and / or the second training time requirement which are more relaxed than the first performance requirement and / or the first training time requirement, the flexibility and rationality of performing the machine learning training can be improved, and the efficiency of performing the machine learning training can be improved.

[0445] In some possible implementation manners, the method further includes:

[0446] S650, sending third information to the second device / the fourth device, wherein the third information is used to indicate at least one of the following: a training time of performing the machine learning training; or a value of an influence of performing the machine learning training on a communication network performance; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

[0447] Exemplarily, when the machine learning training based on the first information meets the requirement (for example, the machine learning training is completed at a time point expected to complete the machine learning training or within an expected training duration), the MnS producer can indicate: a training time at which the machine learning training is completed; or a value of an impact of the machine learning training on the performance of the communication network; or a requirement (for example, a third performance requirement) for the performance of the communication network and / or a requirement (for example, a third training time requirement) for a training time for determining to perform the machine learning training again after the machine learning training based on the first information is performed.

[0448] Based on the scheme provided in the embodiments of the present application, by feeding back the related information after the machine learning training is completed, information can be provided for adjusting the execution of the machine learning training according to the actual situation, and the strategy for performing the machine learning training again is more reasonable after the execution strategy of the machine learning training is adjusted according to the actual situation.

[0449] In some possible implementation manners, the second device / the fourth device can update the first information according to the received third information.

[0450] Exemplarily, the second device / the fourth device can adjust the requirement for the performance of the communication network and / or the requirement for the training time for performing the machine learning training again according to the related information after the machine learning training is performed.

[0451] In some possible implementation manners, the third performance requirement is determined according to the value of the impact of the machine learning training on the performance of the communication network, and the third training time requirement is determined according to the training time of the machine learning training.

[0452] Exemplarily, the value of the impact of the machine learning training on the performance of the communication network can include an average value of the values of the impact of the machine learning training on the performance of the communication network in a training process, and the third performance requirement can include a maximum value of the loss of the performance of the communication network in the training process; or the training time of the machine learning training can include an overall time of the machine learning training, and the third training time requirement can include 1 / n of the overall time of the training, n is a preset value, and n can be a positive integer. (Early stop in the training process is usually n times of the overall training process, and a specific value depends on the trained model.)

[0453] Based on the scheme provided in the embodiments of the present application, the third performance requirement or the third training time requirement is determined according to the impact of the machine learning training on the performance of the communication network or the training time, which can provide a basis for adjusting the execution of the machine learning training according to the actual situation, so that the machine learning training is more reasonable.

[0454] In some possible implementation manners, before determining the first information, the method 600 further includes:

[0455] S660, determining the corresponding relationship according to historical network performance and historical training parameters of performing machine learning training; and / or, determining the corresponding relationship according to historical training time and historical training parameters of performing machine learning training.

[0456] Specifically, the first device / third device can determine the corresponding relationship according to historical data of performing machine learning training by each network element (for example, training parameters of completed machine learning training and corresponding communication network performance, and / or training parameters of completed machine learning training and corresponding training time).

[0457] For example, the network element #A performs a certain machine learning training, and the training time of a certain training is 300 ms, and the loss of a certain KPI / PM of the communication network caused by performing the machine learning training is 10%; the network element #A performs the machine learning training for another time, and the training time of the training is 350 ms, and the loss of the certain KPI / PM of the communication network caused by performing the machine learning training is 5%. For the machine learning training, the training time of 300 ms and the loss value of the certain KPI / PM of 10% correspond, and the training time of 350 ms and the loss value of the certain KPI / PM of 5% correspond. Satisfying the corresponding training time and the loss value of the certain KPI / PM (for example, the corresponding training time and the loss value of RSRP distribution) can be regarded as satisfying the above corresponding relationship, and determining the correspondence of different training time and different KPI / PM can be regarded as determining the above corresponding relationship.

[0458] Based on the scheme provided in the embodiments of the present application, the corresponding relationship is determined through the historical data of performing machine learning training, which can provide reliable and real basis for determining the first parameter, and can improve the rationality and efficiency of performing machine learning training.

[0459] In some possible implementation manners, the first information is determined according to the first performance index, and the first performance index is associated with an inference function to which a model corresponding to the machine learning training belongs.

[0460] For example, the first information can be determined according to a communication network performance index associated with an inference function to which a model corresponding to the machine learning training belongs.

[0461] For example, the first information can be determined according to acceptable KPI / PM of reinforcement learning in the existing network. The device for calculating the first information can calculate a value capable of meeting the requirement of the minimum KPI / PM and the current communication network condition by the requirement of the minimum KPI / PM and the current communication network condition, and the calculated value is used to represent the requirement of performing machine learning training on the communication network performance. The minimum KPI / PM and the current communication network condition are related to the inference function of machine learning corresponding to the machine learning training.

[0462] For example, when the RL is applied to the coverage problem analysis use case, the KPI / PM can include the RSRP distribution, and the network function corresponding to the model is the MDA coverage problem analysis; for example, when the RL is applied to the base station cell load balancing, mobility optimization or mobile robustness optimization use case, the KPI / PM can include the cell load and the handover success rate, and the network function corresponding to the model is the load balancing optimization.

[0463] Based on the scheme provided in the embodiments of the present application, the first information is determined according to the communication network performance index associated with the inference function corresponding to the model to which the machine learning training belongs, so that the communication network performance in the training process can be maintained and the service quality of the communication network users can be ensured.

[0464] In some possible implementation manners, the method 600 further includes:

[0465] S670, updating the correspondence according to the third information.

[0466] For example, the first device / third device can update the correspondence according to the data of completing the machine learning training (for example, the training parameter of the completed machine learning training and the corresponding communication network performance, and / or the training parameter of the completed machine learning training and the corresponding training time).

[0467] It can be understood that S670 can be performed before S650, or can be performed after S650, or can be performed simultaneously with S650. The embodiments of the present application do not limit this.

[0468] In some possible implementation manners, before S640, the second device / fourth device can further determine a first time, and the first time is used to monitor whether the machine learning training is completed.

[0469] Based on the case that the first time monitoring indicates that the machine learning training is not completed, the second device / the fourth device can further perform S640: sending second information, the second information being used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of performing the machine learning training on a communication network performance, the first performance requirement meeting the second performance requirement; or a second training time requirement, the second training time requirement indicating a requirement of performing the machine learning training on a training time, the first training time requirement meeting the second training time requirement, the first training time requirement being indicated by the first information, the first training time requirement indicating the requirement of performing the machine learning training on the training time.

[0470] For example, the second device / the fourth device can monitor whether the machine learning training is completed within a time length corresponding to the first time or before a time node according to the preset or input first time. Based on the case that the first time monitoring indicates that the machine learning training is not completed, the second device / the fourth device can send second information, the second information being more relaxed than the requirement of the first information (i.e., the first performance requirement meets the second performance requirement; or the first training time requirement meets the second training time requirement).

[0471] Based on the scheme provided in the embodiments of the present application, on the one hand, by monitoring the time of the machine learning training, it can be avoided that the machine learning training is too long or unresponsive, which helps to improve the rationality of the machine learning training; on the other hand, by indicating the second performance requirement and / or the second training time requirement that is more relaxed than the first performance requirement and / or the first training time requirement for performing the machine learning training again, the flexibility and rationality of performing the machine learning training can be improved, and the efficiency of performing the machine learning training can be improved.

[0472] In some possible implementation manners, before S660, the second device / the fourth device can further receive fourth information, the fourth information being used to determine the first training time requirement.

[0473] For example, the information contained in the first training time requirement can come from: the MnS consumer receiving a training request of a model corresponding to a certain reinforcement learning function in the network, and the MnS consumer determining the information contained in the first training time requirement according to the training time requirement corresponding to the training request received by the MnS consumer (for example, the training time requirement corresponding to the training request can be the fourth information).

[0474] For example, when the network management system NMS is a possible implementation of the second device / fourth device, the NMS can receive a training request for a model corresponding to a certain reinforcement learning function in the network and fourth information, which can include a training time requirement corresponding to the training request. When the NMS initiates a training request to the MnS producer, the first training time requirement sent by the NMS can include the information of the training time requirement corresponding to the training request.

[0475] The above method 600 will be described in detail below in combination with FIG. 7 to FIG. 9.

[0476] FIG. 7 shows a schematic diagram of a communication method 700 provided by an embodiment of the present application. The method 700 can be a possible implementation of the method 600. The method 700 can be applicable to: in a 3GPP management domain, an NMS as a MnS consumer sends a training request to an EMS as a MnS producer according to the received training request, the training request sent by the NMS includes a communication network performance fluctuation requirement, the EMS can determine configuration parameters for reinforcement learning training according to the communication network performance fluctuation requirement, and a machine learning training function (MLTF) network element performs reinforcement learning training.

[0477] In the embodiment shown in the method 700, the MLTF network element is deployed in the EMS. The MLTF network element generally refers to a network element with ML training function.

[0478] For example, the MLTF network element can be a network digital twin (NDT) module. The NDT is a technology that creates a virtual copy of a physical entity using digital technology, which can simulate, analyze and optimize the performance of the physical entity.

[0479] The method 700 will be described in detail below in combination with FIG. 7. The method 700 can include:

[0480] S701, receiving a training request.

[0481] The NMS can receive a training request for a certain RL function in the network.

[0482] The NMS can also receive a training time requirement corresponding to the training request.

[0483] For example, the training time requirement corresponding to the training request can be preset or externally input. The embodiments of the present application do not limit this.

[0484] For example, the training time requirement corresponding to the training request can be a possible implementation of the fourth information.

[0485] S705, record the correspondence between the configuration parameter and the change of the KPI / PM of the communication network in the historical RL training process of the different MLTF network elements, and / or record the correspondence between the configuration parameter and the change of the training duration in the historical RL training process of the different MLTF.

[0486] The EMS can record the correspondence between the configuration parameter and the change of the KPI / PM of the communication network in the historical RL training process of the different MLTF network elements, and / or record the correspondence between the configuration parameter and the change of the training duration in the historical RL training process of the different MLTF.

[0487] For example, one configuration parameter can correspond to the change of one KPI of the communication network and one training duration, or one configuration parameter can correspond to the change of one PM of the communication network and one training duration. That is, according to the determined configuration parameter, one KPI of the communication network corresponding to the configuration parameter and one training duration can be uniquely determined, or according to the determined configuration parameter, one PM of the communication network corresponding to the configuration parameter and one training duration can be uniquely determined.

[0488] For example, the configuration parameter can include the learning rate, or the learning rate and the decay rate of the learning rate.

[0489] For example, the configuration parameter can include the exploration rate, or the exploration rate and the decay rate of the exploration rate.

[0490] For example, the recorded correspondence can be one possible implementation of the correspondence in the above method 600. S705 can be one possible implementation of S660 in method 600.

[0491] S710, query the function-related KPI / PM.

[0492] The NMS can request / query the related KPI / PM corresponding to the aIMLInferenceType of the training request received in S701 from the EMS.

[0493] It should be understood that which EMS the NMS communicates with can be determined according to the related art, and will not be described here.

[0494] S720, send / receive a response (Response).

[0495] The EMS can feed back the KPI / PM or the list of KPI / PM to the NMS in response to the request / query of the NMS.

[0496] S730, determine the performance fluctuation requirement of the communication network.

[0497] The NMS can initialize the acceptable KPI / PM fluctuation range for RL training in the live network, and the result of the initialization can be referred to as a fluctuation range configuration strategy.

[0498] For example, the NMS can multiply the average value of all KPI / PMs by a loss coefficient, and the product is the result of the initialization of the acceptable KPI / PM fluctuation range. Alternatively, the communication network fluctuation requirement should be able to guarantee the minimum service requirement of the corresponding inference function, the MnS consumer can monitor the communication network performance, and query the minimum KPI / PM requirement of the corresponding inference function according to the training request, and then determine the requirement for communication network performance (such as KPI / PM and the corresponding KPI / PM lower threshold / deviation range / maximum loss ratio, etc.) for executing machine learning training according to the current communication network performance.

[0499] The acceptable KPI / PM fluctuation range is used to maintain the performance of the network during the RL process, and balance the RL training efficiency and the performance degradation impact on the communication network caused by the RL training in the communication network.

[0500] The communication network performance fluctuation requirement in S730 can be a possible implementation of the first information in method 600.

[0501] It should be understood that the network performance fluctuation requirement is determined by multiplying the average value of all KPI / PMs by a loss coefficient, or the network performance fluctuation requirement is determined by querying the minimum KPI / PM requirement of the corresponding inference function and the current network performance, which are only exemplary solutions and do not limit the present application.

[0502] S740, send / receive the communication network performance fluctuation requirement (send / receive MLTraniningRequest including information of the communication network performance fluctuation requirement).

[0503] The NMS can initiate a training request to the EMS and configure the related reinforcement learning strategy according to the training request received in S701. The related reinforcement learning strategy can include the communication network performance fluctuation requirement (the communication network performance fluctuation requirement can be in the form of a list: a list of multiple KPI / PMs, and the KPI / PM lower threshold / deviation range / maximum loss ratio corresponding to each KPI / PM; or the communication network performance fluctuation requirement can be a single KPI / PM value and the corresponding KPI / PM lower threshold / deviation range / maximum loss ratio).

[0504] The related reinforcement learning policy can further include a training time requirement and / or an MLTF selection range. The training time requirement can include a training time point and / or a training duration at which the RL training is expected to be completed, and the MLTF selection range can include a plurality of MLTF network elements capable of performing the RL training, or the MLTF selection range can include an RL environment, a training location, or a training function of the plurality of MLTF network elements capable of performing the RL training.

[0505] At S750, the training parameters are determined according to the communication network performance fluctuation requirement and the machine learning training is performed.

[0506] The EMS can determine the RL configuration parameters / training parameters (e.g., the configuration parameters / training parameters can include a maximum exploration rate) according to historical experience (e.g., a correspondence between the configuration parameters and the changes in the communication network KPI / PM in the historical RL training processes of different MLTFs recorded in S705, and / or a correspondence between the configuration parameters and the training duration in the historical RL training processes), the communication network performance fluctuation requirement in S740, and / or the training time requirement in S740, and perform the RL training.

[0507] For example, the network performance fluctuation requirement is ≤X, the training time requirement is ≤Y, and the exploration rates corresponding to the communication network performance fluctuation requirement ≤X and / or the training time requirement ≤Y in the correspondence recorded in S705 are A, B, and C, respectively, the maximum of A, B, and C is selected as the training parameter for performing the machine learning training.

[0508] An excessively low exploration rate can result in excessively slow training convergence, and an excessively high exploration rate can result in excessively large communication network performance degradation, and the training speed and the communication network performance degradation are a set of trade offs.

[0509] For example, S750 can be a possible implementation of S620c in scenario #1 or scenario #2 in method 600.

[0510] At S780, a training report (MLTrainingReport) is sent / received.

[0511] The training report can include at least one of an actual training time of the RL training, an actual fluctuation of the communication network, or a policy adjustment suggestion.

[0512] After the RL training is completed, the EMS can feed back the actual training time of the RL training and / or the actual fluctuation of the communication network to the NMS.

[0513] The EMS can also provide a policy adjustment suggestion according to the actual training time of the RL training and / or the actual fluctuation of the communication network. The policy adjustment suggestion can be used to determine the RL configuration parameters / training parameters for the next RL training.

[0514] For example, the policy adjustment suggestion can include an adjustment suggestion for the discount factor.

[0515] For example, the policy adjustment suggestion can be a possible implementation of the third performance requirement and / or the third training time requirement.

[0516] For example, S780 can be a possible implementation of S650 in method 600. The training report can be a possible implementation of the third information.

[0517] S790, updating the communication network performance fluctuation requirement.

[0518] For example, the NMS can adjust the fluctuation range configuration policy determined in S730 according to the actual fluctuation of the communication network and / or the actual training time in S780, or according to the policy adjustment suggestion (the adjustment basis can be that different users have different sensitivities to different KPIs, such as the average human reaction time of 200-300 ms, so the transmission-related communication network KPI / PM configuration in continuous services (such as large model inference) only needs to meet the basic requirements).

[0519] S790 can be a possible implementation of updating the first information in method 600.

[0520] Optionally, the method 700 can further include:

[0521] S760, updating the communication network performance fluctuation requirement.

[0522] S770, re-determining the training parameters according to the updated communication network performance fluctuation requirement and performing machine learning training.

[0523] In S740, if the relevant reinforcement learning policy does not include a training time requirement, the NMS can determine a monitoring time for monitoring whether the RL training is timed out; in S740, if the relevant reinforcement learning policy includes a training time requirement, the training time requirement can be regarded as the monitoring time for monitoring whether the RL training is timed out. If the RL training is not monitored to be completed based on the monitoring time, it can be considered that the training is timed out. At this time, S760 and S770 can be performed: the NMS sends an updated communication network performance fluctuation requirement, and the EMS performs RL training again according to the updated communication network performance fluctuation requirement.

[0524] The monitoring time can be a possible implementation of the first time described above. The updated communication network performance fluctuation requirement in S760 can be a possible implementation of the second information.

[0525] For example, after the machine learning training is timed out according to the maximum exploration rate, the updated communication network performance fluctuation requirement and / or the training time requirement can be indicated, X < updated communication network performance fluctuation requirement ≤ X', and / or, Y < updated training time requirement ≤ Y'. The EMS can re-determine the training parameters and perform machine learning training according to the updated communication network performance fluctuation requirement and / or the training time requirement.

[0526] For example, in the correspondence recorded in S805, the exploration rates corresponding to the communication network performance fluctuation requirement ≤ X' and / or the training time requirement ≤ Y' are A', B', and C', respectively. The maximum of A', B', and C' is selected as the re-determined training parameters.

[0527] FIG. 8 shows a schematic diagram of a communication method 800 provided by an embodiment of the present application. The method 800 can be a possible implementation of the method 600. The method 800 can be applicable to: in a 3GPP management domain, an NMS as a MnS consumer sends a training request to an EMS according to the received training request, the training request sent by the NMS includes a communication network performance fluctuation requirement, the EMS can determine configuration parameters for reinforcement learning training according to the communication network performance fluctuation requirement, and an MLTF network element performs reinforcement learning training.

[0528] In the embodiment shown in the method 800, the MLTF network element is deployed in a gNB.

[0529] The method 800 will be described in detail below in combination with FIG. 8. The method 800 can include:

[0530] S801, receiving a training request.

[0531] S805, recording the correspondence between the configuration parameters and the changes of the communication network KPI / PM in the historical RL training process of different MLTF network elements, and / or recording the correspondence between the configuration parameters and the changes of the training duration in the historical RL training process of different MLTFs.

[0532] S810, querying function-related KPI / PM.

[0533] S820, sending / receiving a response (Response).

[0534] S830, determining a communication network performance fluctuation requirement.

[0535] S840, sending / receiving a communication network performance fluctuation requirement (MLTrainingRequest including information of the communication network performance fluctuation requirement).

[0536] The implementation of the above steps S801-S840 can refer to the description of the above steps S701-S740, and will not be repeated here.

[0537] S850, determining the training parameter according to the communication network performance fluctuation requirement.

[0538] The EMS can determine the RL configuration parameter / training parameter (for example, the configuration parameter / training parameter can include the maximum exploration rate) according to the historical experience (for example, the correspondence between the change of the configuration parameter and the communication network KPI / PM in the historical RL training process of different MLTF recorded in S805, and / or the correspondence between the change of the configuration parameter and the training time in the historical RL training process), the communication network performance fluctuation requirement in S840, and / or the training time requirement in S840.

[0539] For example, the communication network performance fluctuation requirement is ≤X, the training time requirement is ≤Y, and the exploration rate corresponding to the communication network performance fluctuation requirement ≤X and / or the training time requirement ≤Y in the correspondence recorded in S805 is A, B, and C respectively, then the maximum of A, B, and C is selected as the training parameter for performing machine learning training.

[0540] If the exploration rate is too low, the training may converge too slowly, and if the exploration rate is too high, the communication network performance may decay too much, and the training speed and the communication network performance decay are a set of trade-offs.

[0541] S853, sending / receiving the training parameter.

[0542] The EMS can send the training parameter determined by the EMS to the network element in the gNB performing RL training through the communication interface.

[0543] S856, performing machine learning training.

[0544] The network element in the gNB can perform machine learning training according to the training parameter received in S853.

[0545] S879 and S880, sending / receiving a training report (MLTraninigReport).

[0546] In S879, the training report sent by the gNB can include the actual training time of the RL training or the actual fluctuation of the communication network, and the performance of the trained model.

[0547] In S880, the training report sent by the EMS can include at least one of the actual training time of the RL training, the actual fluctuation of the communication network, or a policy adjustment suggestion, which can be determined by the EMS according to the training report in S879.

[0548] After the RL training is completed, the gNB / EMS can feed back the actual training time of the RL training and / or the actual fluctuation of the communication network.

[0549] The EMS can also provide a policy adjustment suggestion according to the actual training time of the RL training and the actual fluctuation of the communication network. The policy adjustment suggestion can be used to determine the RL configuration parameters / training parameters for the next RL training.

[0550] In S890, the fluctuation requirement of the performance of the communication network is updated.

[0551] The implementation of the above steps S880-S890 can refer to the textual description of the above steps S780-S790, and will not be repeated here.

[0552] Optionally, the method 800 can further include:

[0553] In S860, the fluctuation requirement of the performance of the communication network is updated.

[0554] In S870, the training parameters are re-determined according to the updated fluctuation requirement of the performance of the communication network.

[0555] In S873, the training parameters are updated.

[0556] In S876, the machine learning training is performed according to the updated training parameters.

[0557] In S840, if the relevant reinforcement learning policy does not include training time, the NMS can determine a monitoring time for monitoring whether the RL training is timed out; in S840, if the relevant reinforcement learning policy includes training time requirement, the training time requirement can be regarded as the monitoring time for monitoring whether the RL training is timed out. If the RL training is not monitored to be completed based on the monitoring time, it can be considered that the training is timed out. At this time, S860-S876 can be performed: the NMS sends the updated fluctuation requirement of the performance of the communication network; the EMS updates the training parameters according to the updated fluctuation requirement of the performance of the communication network; the EMS sends the updated training parameters; and the gNB performs the machine learning training according to the updated training parameters.

[0558] The monitoring time can be a possible implementation of the first time described above. The updated communication network performance fluctuation requirement in S860 can be a possible implementation of the second information.

[0559] For example, after the machine learning training is performed according to the maximum exploration rate, the updated communication network performance fluctuation requirement and / or the training time requirement can be indicated, X < updated communication network performance fluctuation requirement ≤ X', and / or, Y < updated training time requirement ≤ Y'. The EMS can re-determine the training parameters and perform machine learning training according to the updated communication network performance fluctuation requirement and / or the training time requirement.

[0560] For example, in the correspondence recorded in S805, the exploration rates corresponding to the communication network performance fluctuation requirement ≤ X' and / or the training time requirement ≤ Y' are A', B', and C', respectively. The maximum of A', B', and C' is selected as the re-determined training parameters.

[0561] FIG. 9 shows a schematic diagram of a communication method 900 provided by an embodiment of the present application. The method 900 can be a possible implementation of the method 600. The method 900 can be applicable to: in a 3GPP management domain, an NMS as a MnS consumer sends a training request to an EMS according to the received training request, the training request sent by the NMS includes a communication network performance fluctuation requirement, the EMS can determine configuration parameters for reinforcement learning training and a training environment according to the communication network performance fluctuation requirement, and an MLTF network element performs reinforcement learning training.

[0562] In the embodiment shown in the method 900, the MLTF network element can be deployed in the EMS or outside the EMS, and the MLTF network element can communicate with the EMS through a communication interface.

[0563] The method 900 will be described in detail below in conjunction with FIG. 9. The method 900 can include:

[0564] S901, receiving a training request.

[0565] S905, recording the correspondence between the configuration parameters and the changes in the communication network KPI / PM in the historical RL training process of different MLTF network elements, and / or recording the correspondence between the configuration parameters and the changes in the training duration in the historical RL training process of different MLTFs.

[0566] S910, querying function-related KPI / PM.

[0567] S920, sending / receiving a response (Response).

[0568] S930, determining a communication network performance fluctuation requirement.

[0569] S940, sending / receiving the communication network performance fluctuation requirement (MLTrainingRequest including information of the communication network performance fluctuation requirement).

[0570] The MLTF selection range can also be carried in S940.

[0571] The implementation of the above steps S901-S940 can refer to the description of the above steps S701-S740, and will not be repeated here.

[0572] S950, determining the training parameter according to the communication network performance fluctuation requirement; determining the network element for performing the RL training according to the MLTF selection range.

[0573] The EMS can determine the RL configuration parameter / training parameter (for example, the configuration parameter / training parameter can include the maximum exploration rate) according to the historical experience (for example, the correspondence relationship between the configuration parameter and the change of the communication network KPI / PM in the historical RL training process of different MLTFs recorded in S905, and / or the correspondence relationship between the configuration parameter and the training time in the historical RL training process), the communication network performance fluctuation requirement in S940, and / or the training time requirement in S940.

[0574] For example, the communication network performance fluctuation requirement is ≤X, the training time requirement is ≤Y, and the exploration rates corresponding to the communication network performance fluctuation requirement ≤X and / or the training time requirement ≤Y in the correspondence relationship recorded in S905 are A, B, and C respectively, then the maximum of A, B, and C is selected as the training parameter for performing the machine learning training.

[0575] If the exploration rate is too low, the training may converge too slowly, and if the exploration rate is too high, the communication network performance may decay too much, and the training speed and the communication network performance decay are a set of trade off.

[0576] The EMS can also determine the network element for performing the RL training according to the MLTF selection range carried in S940.

[0577] As a possible implementation manner, the EMS can determine the network element for performing the RL training according to the identification ID of different network elements, and the EMS can also determine the network element for performing the RL training according to the selection condition (for example, the geographical location corresponding to the network element).

[0578] Different parameters have different requirements for performance under different network performances. The EMS can determine the network that can meet the communication network performance fluctuation requirement according to the communication network performance fluctuation requirement in S940 and the network performance. The EMS can determine the network element for performing RL training according to the network that meets the communication network performance fluctuation requirement. If the indicated MLTF selection range does not meet the communication network performance fluctuation requirement, since the communication network performance fluctuation is fixed under the fixed training duration, the EMS can also select the relatively best one in the MLTF selection range as the network element for performing RL training (for example, the one with the shortest training duration under the communication network performance fluctuation requirement, or the one with the smallest communication network performance fluctuation under the training duration requirement).

[0579] It can be understood that when the EMS does not determine the network element for performing RL training, the network element for performing RL training can be preset or determined by the MnS producer.

[0580] The MLTF selection range can be a plurality of network elements with the smallest impact on the network selected by the NMS according to the communication network performance fluctuation requirement. This helps to avoid conflicts between multiple training when multiple RL training exists at the same time.

[0581] S953, sending / receiving training parameters.

[0582] The EMS can send the training parameters to the network element for performing RL training determined in S950 through the communication interface.

[0583] S956, performing machine learning training.

[0584] S979 and S980, sending / receiving training report (MLTraninigReport).

[0585] In S979, the training report sent by the MLTF can include the actual training time of the RL training or the actual fluctuation of the communication network, and the performance of the model obtained by training.

[0586] In S980, the training report sent by the EMS can include at least one of the actual training time of the RL training, the actual fluctuation of the communication network, or the policy adjustment suggestion, which can be determined by the EMS according to the training report in S979.

[0587] After the RL training is completed, the MLTF / EMS can feed back the actual training time of the RL training and the actual fluctuation of the communication network.

[0588] The EMS can also provide a policy adjustment suggestion according to the actual training time of the RL training and the actual fluctuation of the communication network. The policy adjustment suggestion can be used to determine the RL configuration parameter / training parameter at the next RL training.

[0589] S990, updating the communication network performance fluctuation requirement.

[0590] Optionally, the method 900 can further include:

[0591] S960, updating the communication network performance fluctuation requirement.

[0592] S970, re-determining the training parameter according to the updated communication network performance fluctuation requirement.

[0593] S973, updating the training parameter.

[0594] S976, performing machine learning training according to the updated training parameter.

[0595] In S940, if the relevant reinforcement learning policy does not include training time, the NMS can determine a monitoring time for monitoring whether the RL training is timed out; in S940, if the relevant reinforcement learning policy includes training time requirement, the training time requirement can be regarded as the monitoring time for monitoring whether the RL training is timed out. If no RL training completion is monitored based on the monitoring time, it can be considered that the training is timed out. At this time, S960-S976 can be performed: the NMS sends the updated communication network performance fluctuation requirement; the EMS updates the training parameter according to the updated communication network performance fluctuation requirement; the EMS sends the updated training parameter; and the MLTF performs machine learning training according to the updated training parameter.

[0596] The implementation of the above steps S953-S990 can refer to the textual description of the above steps S853-S890, and will not be repeated here.

[0597] For example, the embodiments of the present application are applicable to the MnS interface of AIML in enhanced OAM.

[0598] For example, a new attribute rlNetworkPerformanceReqirement is added in MLTrainingRequest; and / or, new attributes timeConsumption and / or rlNetworkPerformance are added in MLTrainingReport.

[0599] rlNetworkPerformanceReqirement can be a possible implementation of the above communication network performance fluctuation requirement.

[0600] The timeConsumption can be a possible implementation of the actual training time described above.

[0601] The rlNetworkPerformance can be a possible implementation of the actual fluctuation of the communication network described above.

[0602] It should be noted that, for the embodiments of FIGS. 6-9 described above:

[0603] (1) The numbering of each step described in the embodiments is only an example and does not constitute a limitation on the present application. Some steps can be added or deleted according to actual needs in the embodiments of the present application.

[0604] (2) The embodiments of FIGS. 6-9 described above can be implemented independently or in combination with each other. For example, the embodiment shown in FIG. 6 and the embodiments shown in FIGS. 7-9 can be combined with each other, the embodiments shown in FIGS. 7-9 can be combined with each other, and so on.

[0605] FIGS. 10 and 11 are schematic block diagrams of communication apparatuses provided by embodiments of the present application. These communication apparatuses can be used to implement the functions of the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) in the method embodiments described above, and thus can also achieve the beneficial effects possessed by the method embodiments described above.

[0606] FIG. 10 is a schematic block diagram of a communication apparatus 1000 provided by an embodiment of the present application. As shown in FIG. 10, the communication apparatus 1000 can include a processing unit 1010, and optionally, the communication apparatus 1000 can also include a transceiver unit (or communication unit) 1020. The communication apparatus 1000 can be used to implement the functions of the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) in the method embodiments shown in FIGS. 6-9 described above.

[0607] When the communication apparatus 1000 is used to implement the functions of the first device / third device in the method embodiments shown in FIG. 6, the transceiver unit 1020 can be used to receive indication information, receive second information, or send third information, and the processing unit 1010 can be used to determine first information according to the indication information, determine the first information, perform machine learning training, determine a network element that performs machine learning training, determine a correspondence relationship, update the correspondence relationship, and so on.

[0608] When the communication apparatus 1000 is configured to implement the function of the EMS in the method embodiment shown in FIG. 7, the transceiver unit 1020 can be configured to receive a request for querying function-related KPI / PM, send a response, receive a communication network performance fluctuation requirement, receive an updated communication network performance fluctuation requirement, send a training report, etc., and the processing unit 1010 can be configured to determine a training parameter according to the communication network performance fluctuation requirement, redetermine the training parameter according to the updated communication network performance fluctuation requirement, perform machine learning training, etc.

[0609] When the communication apparatus 1000 is configured to implement the function of the NMS in the method embodiment shown in FIG. 7, the transceiver unit 1020 can be configured to receive a training request, send a request for querying function-related KPI / PM, receive a response, send a communication network performance fluctuation requirement, send an updated communication network performance fluctuation requirement, receive a training report, etc., and the processing unit 1010 can be configured to determine a communication network performance fluctuation requirement, update the communication network performance fluctuation requirement, etc.

[0610] When the communication apparatus 1000 is configured to implement the function of the EMS in the method embodiment shown in FIG. 8, the transceiver unit 1020 can be configured to receive a request for querying function-related KPI / PM, send a response, receive a communication network performance fluctuation requirement, send a training parameter, receive an updated communication network performance fluctuation requirement, send an updated training parameter, receive a training report, send a training report, etc., and the processing unit 1010 can be configured to determine a training parameter according to the communication network performance fluctuation requirement, redetermine the training parameter according to the updated communication network performance fluctuation requirement, determine a policy adjustment suggestion according to the training report, etc.

[0611] When the communication apparatus 1000 is configured to implement the function of the NMS in the method embodiment shown in FIG. 8, the transceiver unit 1020 can be configured to receive a training request, send a request for querying function-related KPI / PM, receive a response, send a communication network performance fluctuation requirement, send an updated communication network performance fluctuation requirement, receive a training report, etc., and the processing unit 1010 can be configured to determine a communication network performance fluctuation requirement, update the communication network performance fluctuation requirement, etc.

[0612] When the communication apparatus 1000 is configured to implement the function of the gNB in the method embodiment shown in FIG. 8, the transceiver unit 1020 can be configured to receive a training parameter, receive an updated training parameter, send a training report, etc., and the processing unit 1010 can be configured to perform machine learning training, determine a training report, etc.

[0613] When the communication apparatus 1000 is configured to implement the function of the EMS in the method embodiment shown in FIG. 9, the transceiver unit 1020 can be configured to receive a request for querying function-related KPI / PM, send a response, receive a communication network performance fluctuation requirement, send a training parameter, receive an updated communication network performance fluctuation requirement, send an updated training parameter, receive a training report, send a training report, etc., and the processing unit 1010 can be configured to determine a training parameter according to the communication network performance fluctuation requirement, redetermine the training parameter according to the updated communication network performance fluctuation requirement, determine a network element for performing RL training according to a MLTF selection range, determine a policy adjustment suggestion according to the training report, etc.

[0614] When the communication apparatus 1000 is configured to implement the function of the NMS in the method embodiment shown in FIG. 9, the transceiver unit 1020 can be configured to receive a training request, send a request for querying function-related KPI / PM, receive a response, send a communication network performance fluctuation requirement, send an updated communication network performance fluctuation requirement, receive a training report, etc., and the processing unit 1010 can be configured to determine a communication network performance fluctuation requirement, update a communication network performance fluctuation requirement, etc.

[0615] When the communication apparatus 1000 is configured to implement the function of the MLTF in the method embodiment shown in FIG. 9, the transceiver unit 1020 can be configured to receive a training parameter, receive an updated training parameter, send a training report, etc., and the processing unit 1010 can be configured to perform machine learning training, determine a training report, etc.

[0616] Optionally, the apparatus 1000 can further include a storage unit (not shown in FIG. 10), which can be configured to store instructions and / or data, and the processing unit 1010 can read the instructions and / or data in the storage unit to enable the apparatus to implement the foregoing method embodiments.

[0617] For example, the storage unit can record the correspondence between historical RL configuration parameters of different MLTFs and changes in communication network KPI / PM, the correspondence between configuration parameters and training duration, etc.

[0618] For brevity, the foregoing will not be repeated here.

[0619] The apparatus 1000 of each of the above-mentioned solutions has the function of implementing the corresponding steps performed by the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) in the above-mentioned methods. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions; for example, the transceiver unit can be replaced by a transceiver (for example, the transmitting unit in the transceiver unit can be replaced by a transmitter, and the receiving unit in the transceiver unit can be replaced by a receiver), and other units such as the processing unit can be replaced by a processor, which respectively performs the transceiving operation and the related processing operation in each method embodiment.

[0620] In addition, the transceiver unit mentioned above can also be a transceiver circuit (for example, which can include a receiving circuit and a transmitting circuit), and the processing unit can be a processing circuit. The processing circuit can be one or more processors, or all or part of the circuit in the one or more processors for control or processing functions. In the embodiments of the present application, the apparatus in FIG. 10 can be the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) in the foregoing embodiments, or a chip or a chip system, for example, a system on chip (SoC). Wherein, the transceiver unit can be an input / output circuit, a communication interface; and the processing unit is a processor or a microprocessor integrated on the chip or an integrated circuit. Herein, no limitation is made.

[0621] It should also be understood that the apparatus 1000 herein is embodied in the form of functional units. The term “unit” herein can refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (for example, a shared processor, a dedicated processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combination logic circuit, and / or other suitable components that support the described functions.

[0622] FIG. 11 is a schematic block diagram of a communication apparatus 1100 provided by the embodiments of the present application. The apparatus 1100 includes a processing circuit. The apparatus can also include a communication circuit. Wherein, the processing circuit and the communication circuit communicate with each other through an internal connection path, and the processing circuit is used to execute instructions to control the communication circuit to transmit and / or receive signals.

[0623] When the communication apparatus 1100 is used to implement the method shown in FIG. 11, the processor 1110 is configured to implement the functions of the processing unit 1010, and the transceiver 1120 is configured to implement the functions of the transceiving unit 1020.

[0624] In a possible implementation, the apparatus 1100 is configured to implement the respective procedures and steps of the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) in the method embodiments.

[0625] It can be understood that the apparatus 1100 can be specifically the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) in the above embodiments, and can also be a chip or a chip system. Correspondingly, the communication circuit can be an interface circuit of the chip, or an input / output circuit, which is not limited herein. Specifically, the apparatus 1100 can be configured to execute the respective steps and / or procedures of the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) in the method embodiments.

[0626] When the communication apparatus 1100 is used to implement the method shown in FIG. 11, the processor 1110 is configured to implement the functions of the processing unit 1010, and the transceiver 1120 is configured to implement the functions of the transceiving unit 1020.

[0627] It can be understood that, in order to implement the functions in the above embodiments, the first device, the second device, the third device, the fourth device, the EMS (or the network device where the EMS is located), or the NMS (or the network device where the NMS is located) includes a hardware structure and / or a software module for executing each function. It can be easily understood by those skilled in the art that the units and method steps of the examples described in combination with the embodiments disclosed in the present application can be implemented in the form of hardware or hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application scenario and design constraints of the technical solution.

[0628] It can be appreciated that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), image processors, artificial intelligence processors, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor can be a microprocessor, or any conventional processor.

[0629] The method steps in the embodiments of the present application can be implemented in hardware or in software instructions executable by a processor. The software instructions can be composed of corresponding software modules, which can be stored in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from and write information to the storage medium. The storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in the first device, the second device, the third device, the fourth device, the EMS (or a network device where the EMS is located), or the NMS (or a network device where the NMS is located). The processor and the storage medium can also exist as discrete components in the first device, the second device, the third device, the fourth device, the EMS (or a network device where the EMS is located), or the NMS (or a network device where the NMS is located).

[0630] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are performed. The computer can be a general purpose computer, a special purpose computer, a computer network, a network device, a user equipment or other programmable apparatus. The computer programs or instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer programs or instructions can be transferred from one website site, computer, enabling server or data center to another website site, computer, enabling server or data center through wired or wireless manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as an enabling server, data center and the like integrated with one or more available media. The available media can be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape; an optical medium, such as a digital video disc; and a semiconductor medium, such as a solid state disk. The computer readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile storage media.

[0631] In the above various embodiments, the terms and / or descriptions of different embodiments are consistent and can be referred to each other if there is no special description and logical conflict. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0632] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0633] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0634] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The division of the units is merely logical function division. There can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0635] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0636] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0637] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server computer, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0638] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A communication method characterized by comprising: Applicable to machine learning training, comprising: receiving first information, the first information being used to indicate a first performance requirement, the first performance requirement indicating a requirement of performing the machine learning training on communication network performance; performing the machine learning training based on the first information.

2. The method of claim 1, wherein, The first information is also used to indicate a first training time requirement, the first training time requirement indicating a requirement of performing the machine learning training on training time.

3. The method of claim 2, wherein, The machine learning training is reinforcement learning training, and the first training time requirement includes a training time point and / or a training duration at which the reinforcement learning training is expected to be completed.

4. The method according to claim 2 or 3, characterized in that, The network element performing the machine learning training meets the first performance requirement and the first training time requirement.

5. The method according to any one of claims 1 to 3, characterized in that, The network element performing the machine learning training meets the first performance requirement.

6. The method according to any one of claims 1 to 5, characterized in that, The first information is also used to indicate a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

7. The method of claim 6, wherein, The method further comprises: determining the network element performing the machine learning training according to communication network performance of the plurality of first network elements.

8. The method according to any one of claims 1 to 5, characterized in that, The network element performing the machine learning training belongs to a plurality of second network elements, and the first information is also used to indicate at least one of: RL environment of the plurality of second network elements; training location of the plurality of second network elements.

9. The method of claim 8, wherein, The method further comprises: determining the network element performing the machine learning training according to the RL environment of the plurality of second network elements and / or the training location of the plurality of second network elements.

10. The method according to any one of claims 1 to 9, characterized in that, The method further comprises: receiving second information, the second information being used to indicate at least one of: a second performance requirement, the second performance requirement indicating a requirement of performing the machine learning training on communication network performance; or a second training time requirement, the second training time requirement indicating a requirement of performing the machine learning training on training time; After performing the machine learning training based on the first information, the method further comprises: performing the machine learning training based on the second information.

11. The method according to any one of claims 1 to 10, characterized in that, The method further comprises: sending third information, wherein the third information is used to indicate at least one of: training time of performing the machine learning training; or a value of an impact of performing the machine learning training on communication network performance; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

12. The method of claim 11, wherein, The third performance requirement is determined according to the value of the impact of performing the machine learning training on communication network performance; and the third training time requirement is determined according to the training time of performing the machine learning training.

13. A method of communication, comprising: Applicable to machine learning training, comprising: determining first information; sending the first information, the first information being used to indicate a first performance requirement, the first performance requirement indicating a requirement of performing the machine learning training on communication network performance.

14. The method of claim 13, wherein, The first information is also used to indicate a first training time requirement, the first training time requirement indicating a requirement of performing the machine learning training on training time.

15. The method of claim 14, wherein, The machine learning training is reinforcement learning training, and the first training time requirement comprises a training time point and / or a training time length at which the reinforcement learning training is expected to be completed.

16. The method according to claim 14 or 15, characterized in that The network element performing the machine learning training satisfies the first performance requirement and the first training time requirement.

17. The method according to any one of claims 13 to 15, characterized in that, The network element performing the machine learning training satisfies the first performance requirement.

18. The method according to any one of claims 13 to 17, characterized in that, The first information further indicates a plurality of first network elements, and the network element performing the machine learning training belongs to the plurality of first network elements.

19. The method according to any one of claims 13 to 17, characterized in that, The network element performing the machine learning training belongs to a plurality of second network elements, and the first information further indicates at least one of the following: an RL environment of the plurality of second network elements; and / or a training location of the plurality of second network elements.

20. The method of claim 19, wherein, The RL environment of the plurality of second network elements and / or the training location of the plurality of second network elements are used to determine the network element performing the machine learning training.

21. The method according to any one of claims 13 to 20, characterized in that, The method further comprises: determining a first time, the first time being used to monitor whether the machine learning training is completed; and based on a case that the machine learning training is not completed according to the first time, sending second information, the second information being used to indicate at least one of the following: a second performance requirement, the second performance requirement indicating a requirement of the machine learning training on a performance of a communication network; or a second training time requirement, the second training time requirement indicating a requirement of the machine learning training on a training time.

22. The method of any one of claims 13-21, wherein, The method further comprises: receiving third information; and updating the first information according to the third information, wherein the third information is used to indicate at least one of the following: a training time of the machine learning training; or a value of an influence of the machine learning training on a performance of a communication network; or a third performance requirement and / or a third training time requirement, the third performance requirement or the third training time requirement being used to perform the machine learning training again.

23. The method of any one of claims 14-16, wherein, The method further comprises: receiving fourth information, the fourth information being used to determine the first training time requirement; and The determining the first information comprises: determining the first training time requirement according to the fourth information.

24. The method of any one of claims 14 to 23, wherein, The determining the first information comprises: determining the first information according to a first performance index, the first performance index being associated with an inference function to which a model corresponding to the machine learning training belongs.

25. A communications device, characterized by application to machine learning training, comprising a processor and a memory, the memory is configured to store computer instructions; and the processor is configured to execute the computer instructions stored in the memory to implement the method according to any one of claims 1 to 12.

26. A communications device, characterized by application to machine learning training, comprising a processor and a memory, the memory is configured to store computer instructions; and the processor is configured to execute the computer instructions stored in the memory to implement the method according to any one of claims 13 to 24.

27. A chip or chip system, characterized by application to machine learning training, comprising a processor and a memory, the memory is configured to store computer instructions; and the processor is configured to execute the computer instructions stored in the memory to implement the method according to any one of claims 1 to 12.

28. A computer program product, characterised in that, When a computer program in the computer program product is executed by a processor, a method as claimed in any one of claims 1 to 12, or a method as claimed in any one of claims 13 to 24 is implemented.

29. A computer-readable storage medium, characterized in that, The storage medium has stored therein a computer program or instructions, and when the computer program or instructions are executed by a processor, a method as claimed in any one of claims 1 to 12, or a method as claimed in any one of claims 13 to 24 is implemented.

30. A system, comprising: Comprising: The apparatus as claimed in claim 25; and the apparatus as claimed in claim 26.

Citation Information

Patent Citations

  • Handling Training of a Machine Learning Model

    US20230297884A1

  • Communication method and apparatus used for training machine learning model

    WO2023185711A1

  • Communication method and apparatus

    WO2024169515A1

  • Communication method and apparatus

    WO2024169522A1