Model training method and device

By receiving model training requirements information in the wireless access network and conducting targeted model training, the problem of model training failure and resource waste caused by the limited data volume and resources of the access network is solved, and the effect of improving the success rate of model training and reducing resource waste is achieved.

CN120017533APending Publication Date: 2025-05-16HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311531431.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the field of wireless access networks, due to the limitations of equipment costs, wireless environment and application requirements, the amount of data that access network devices can collect and obtainable computing resources and storage resources are limited, resulting in failure in model training and waste of resources.

Method used

Provide a model training method and device, by receiving model training requirements information, targeted model training, improve the success rate of model training, and reduce resource waste. The specific steps include receiving model training requirements information, performing model training based on this information, and sending training results or failure reasons.

Benefits of technology

This improves the success rate of model training, reduces the waste of resources during model training, and enhances the collaboration between the first functional entity and the second functional entity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017533A_ABST
    Figure CN120017533A_ABST
Patent Text Reader

Abstract

The model training method can be applied to the field of wireless access networks, and comprises the following steps: receiving model training demand information sent by a second functional entity, the model training demand information comprises at least one of the following information: layer number information of model training, the maximum value of time used by the model training, the data volume used by the model training, the maximum value of the data volume used by the model training, the minimum value of the data volume used by the model training or the performance threshold value of the model training; and performing model training according to the model training demand information. Through the method, the first functional entity can perform model training in a targeted manner according to the model training demand information, so that the success rate of model training can be improved, and the phenomenon of resource waste during model training can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless access networks, and more specifically, to a model training method and device. Background Art

[0002] Artificial intelligence (AI) is a technology that imitates human cognition, learning and reasoning abilities. Machine learning (ML) is an important technical means in the field of artificial intelligence. Its core idea is to learn a large amount of known data to obtain the relationship between the data, and finally predict and analyze the unknown data. Common AI or ML technologies can include supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, etc. In the field of radio access network (RAN), AI or ML also has a wide range of applications, such as business prediction, resource scheduling and network management based on AI or ML technology.

[0003] However, when applying AI or ML technology in the RAN field, due to objective reasons such as the cost of access network equipment, wireless environment, and application requirements, the amount of data that can be collected by access network equipment and the available computing resources and storage resources are very limited, which may lead to model training failure (for example, model performance does not meet standards and model training timeouts) and waste of computing resources. Summary of the invention

[0004] The present application provides a model training method and device, which can improve the success rate of model training and reduce the occurrence of resource waste during model training.

[0005] In a first aspect, a model training method is provided, which is applied to a first functional entity, and the method includes: receiving model training requirement information sent by a second functional entity, the model training requirement information including at least one of the following information: information on the number of layers of model training, the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or a performance threshold value for model training; and performing model training according to the model training requirement information.

[0006] Optionally, the first functional entity may be a mobile intelligent function (MIF), and further, the first functional entity may be a centralized MIF.

[0007] Optionally, the second functional entity may be an access network device or a MIF. Further, when the second functional entity is a MIF, it may be a distributed MIF.

[0008] Optionally, the first functional entity and / or the second functional entity may be located in a radio access network.

[0009] In an embodiment of the present application, the first functional entity can receive model training requirement information sent by the second functional entity, and perform model training according to the model training requirement information. In this way, the first functional entity can perform model training in a targeted manner, improve the success rate of model training, and reduce the occurrence of resource waste during model training.

[0010] In combination with the first aspect, in some implementations of the first aspect, the method further includes: sending first information to the second functional entity, wherein the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0011] In an embodiment of the present application, after the first functional entity performs model training based on model training requirement information, the first functional entity can inform the second functional entity of the result of the model training through the first information. In this way, it can ensure that the second functional entity obtains the training results of the model in a timely manner, thereby enhancing the collaboration between the first functional entity and the second functional entity.

[0012] In combination with the first aspect, in certain implementations of the first aspect, the first information is used to indicate trained model information, and the first information includes: a first model.

[0013] The first model may be understood as a model trained by the first functional entity based on model training requirement information.

[0014] In combination with the first aspect, in certain implementations of the first aspect, the performing model training according to the model training requirement information includes: receiving a first data set sent by the second functional entity; training a second model according to the first data set and the model training requirement information to obtain the first model.

[0015] Optionally, the second model may also be referred to as an initial model. The second model may be a model whose parameters are initial values, or may be a model trained based on other data sets.

[0016] In the embodiment of the present application, the first functional entity can train the second model according to the first data set and model training requirement information sent by the second functional entity, and obtain the first model. In this way, the first model can be obtained more efficiently and the success rate of the second model training can be improved.

[0017] In combination with the first aspect, in certain implementations of the first aspect, the first information is used to indicate that model training has failed, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0018] Optionally, the failure reason value may indicate the reason for the failure in the form of a number. For example, when the failure reason value is 01, it may indicate that the computing power of the first functional entity is insufficient when performing model training. When the failure reason value is 10, it may indicate that the training samples of the first functional entity are insufficient when performing model training. When the failure reason value is 11, it may indicate that the performance of the first model does not meet the requirements.

[0019] Optionally, the performance of the trained model does not meet the requirements, for example, the accuracy is lower than a threshold value, or the mean square error is higher than a threshold value, etc.

[0020] In an embodiment of the present application, when the first functional entity fails in model training, the first information may include a failure reason value, so that the second functional entity can quickly determine the cause of the model training failure based on the failure reason value and adjust the model training strategy.

[0021] In combination with the first aspect, in certain implementations of the first aspect, before performing model training according to the model training requirement information, the method also includes: sending model training requirement information confirmation information to the second functional entity, and the model training requirement confirmation information is used to indicate that the first functional entity confirms the model training requirement information.

[0022] Alternatively, the first functional entity may send model training requirement change information to the second functional entity, and use the changed model training requirement information to perform model training.

[0023] In a second aspect, a model training method is provided, which is applied to a second functional entity, and the method includes: determining model training requirement information, the model training requirement information including at least one of the following information: the number of layers of model training, the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; and sending the model training requirement information to the first functional entity.

[0024] In an embodiment of the present application, the second functional entity can confirm the model training requirement information and send the model training requirement information to the first functional entity, so that the first functional entity can perform model training based on the model training requirement information. In this way, the first functional entity can perform model training in a targeted manner, improve the success rate of model training, and reduce the occurrence of resource waste during model training.

[0025] In combination with the second aspect, in some implementations of the second aspect, the method further includes: receiving first information sent by the first functional entity, the first information being used to indicate trained model information, or the first information being used to indicate model training failure.

[0026] In an embodiment of the present application, the second functional entity can obtain the result of model training based on the first information. In this way, it can ensure that the second functional entity can obtain the training result of the model in a timely manner, thereby enhancing the collaboration between the first functional entity and the second functional entity.

[0027] In combination with the second aspect, in certain implementations of the second aspect, the first information is used to indicate trained model information, and the first information includes a first model; before receiving the first information sent by the first functional entity, the method also includes: acquiring a first data set; sending the first data set to the first functional entity; the method also includes: training the first model based on the first data set to obtain a third model.

[0028] Alternatively, the second functional entity may collect a second data set, and train the first model according to the second data set to obtain a third model, wherein the data contents of the first data set and the second data set may be different.

[0029] Optionally, in this implementation, the second functional entity may be a distributed MIF.

[0030] In an embodiment of the present application, the second functional entity may collect a first data set and send the first data set to the first functional entity so that the first functional entity can perform model training to obtain a first model. Furthermore, after the first functional entity sends the trained first model to the second functional entity, the second functional entity may train the first model again based on the first data set to obtain a third model. In this way, the collaboration between the first functional entity and the second functional entity in the model training process is enhanced, thereby further improving the reasoning performance of the trained model.

[0031] In combination with the second aspect, in certain implementations of the second aspect, the first information is used to indicate that the model training has failed, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0032] In an embodiment of the present application, when the first functional entity fails in model training, the first information may include a failure reason value, so that the second functional entity can quickly determine the cause of the model training failure based on the failure reason value and adjust the model training strategy.

[0033] In combination with the second aspect, in some implementations of the second aspect, the method further includes: receiving model training requirement information confirmation information sent by the first functional entity, wherein the model training requirement confirmation information is used to indicate that the first functional entity confirms the model training requirement information.

[0034] In a third aspect, a model training method is provided, characterized in that the method is applied to a first functional entity, and the method includes: determining model training requirement information, the model training requirement information including at least one of the following information: model identification, model structure information, model training layer information, maximum time used for model training, data volume used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value for model training; and sending the model training requirement information to a second functional entity.

[0035] Optionally, the first functional entity may be a MIF, and further, the first functional entity may be a centralized MIF.

[0036] Optionally, the second functional entity may be a MIF, and further, the second functional entity may be a distributed MIF.

[0037] Optionally, the first functional entity and / or the second functional entity may be located in a radio access network.

[0038] In an embodiment of the present application, the first functional entity can determine model training requirement information and send the model training requirement information to the second functional entity, so that the second functional entity can perform model training based on the model training requirement information. In this way, the second functional entity can perform model training in a targeted manner, improve the success rate of model training, and reduce the occurrence of resource waste during model training.

[0039] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: receiving first information sent by the second functional entity, the first information being used to indicate trained model information, or the first information being used to indicate model training failure.

[0040] In an embodiment of the present application, the first functional entity can receive the first information sent by the second functional entity, thereby obtaining the result of the model training performed by the second functional entity. In this way, it can ensure that the first functional entity obtains the training result of the model in a timely manner, thereby enhancing the collaboration between the first functional entity and the second functional entity.

[0041] In combination with the third aspect, in certain implementations of the third aspect, the first information is used to indicate trained model information, and the first information includes a first model.

[0042] The first model may be understood as a model trained by the second functional entity based on model training requirement information.

[0043] In combination with the third aspect, in certain implementations of the third aspect, before receiving the first information sent by the second functional entity, the method also includes: receiving a first data set sent by the second functional entity; performing model training based on the first data set to obtain a second model, wherein the first model is obtained by training the second model based on the model training requirement information; and sending the second model to the second functional entity.

[0044] Specifically, the first data set can be collected by the second functional entity, and the data in the first data set can be labeled data or unlabeled data. Among them, the labeled data can include the data itself and the corresponding label. For example, the first data set for AI positioning is {(CIR1, TOD1), (CIR2, TOA2)}, then in each data, the channel impulse response (CIR) is the data itself, and the arrival time (TOA) is the label corresponding to the data. Unlabeled data may include only the data itself, excluding the label corresponding to the data. For example, the data set for AI channel compression is {CSI1, CSI2}, then each data only includes the data itself, that is, the channel state information (CSI).

[0045] In the embodiment of the present application, the first functional entity can perform model training based on the first data set to obtain the second model, and send the second model to the second functional entity, so that the second functional entity can train the second model based on the model training requirement information. In this way, the collaboration between the first functional entity and the second functional entity in model training is enhanced, and the accuracy of model training can be improved.

[0046] In combination with the third aspect, in certain implementations of the third aspect, the first information is used to indicate that the model training has failed, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0047] In an embodiment of the present application, when the second functional entity fails in model training, the first information may include a failure reason value, so that the first functional entity can quickly determine the cause of the model training failure based on the failure reason value and adjust the model training strategy.

[0048] In combination with the third aspect, in certain implementations of the third aspect, before determining the model training requirement information, the method also includes: receiving model training capability information sent by the second functional entity, the model training capability information being used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin or video memory bandwidth; determining the model training requirement information includes: determining the model training requirement information based on the model training capability information.

[0049] In an embodiment of the present application, the first functional entity can determine the model training requirement information based on the model training capability information sent by the second functional entity. In this way, the model training requirement information determined by the first functional entity can be closer to the actual situation of the model training performed by the second functional entity, thereby further improving the efficiency of the model training performed by the second functional entity.

[0050] In a fourth aspect, a model training method is provided, which is applied to a second functional entity, and the method includes: receiving model training requirement information sent by a first functional entity, the model training requirement information including at least one of the following information: a model identifier, model structure information, model training layer information, a maximum time used for model training, a data volume used for model training, a maximum amount of data used for model training, a minimum amount of data used for model training, or a performance threshold value for model training; and performing model training according to the model training requirement information.

[0051] In an embodiment of the present application, the second functional entity can perform model training based on the model training requirement information sent by the first functional entity. In this way, the second functional entity can perform model training in a targeted manner, improve the success rate of model training, and reduce the occurrence of resource waste during model training.

[0052] In combination with the fourth aspect, in certain implementations of the fourth aspect, the method further includes: sending first information to the first functional entity, the first information being used to indicate trained model information, or the first information being used to indicate model training failure.

[0053] In an embodiment of the present application, the second functional entity may send the first information to the first functional entity so that the first functional entity can promptly obtain the model training results of the second functional entity, thereby enhancing the collaboration between the first functional entity and the second functional entity.

[0054] In combination with the fourth aspect, in certain implementations of the fourth aspect, the first information is used to indicate trained model information, and the first information includes: a first model.

[0055] The first model may be understood as a model trained by the second functional entity based on model training requirement information.

[0056] In combination with the fourth aspect, in certain implementations of the fourth aspect, the performing model training according to the model training requirement information includes: acquiring a first data set; sending the first data set to the first functional entity; receiving a second model sent by the first functional entity, wherein the second model is trained based on the first data set; training the second model according to the first data set and the model training requirement information to obtain the first model.

[0057] In an embodiment of the present application, the second functional entity trains the second model according to the first data set to obtain the first model. In this way, the collaboration between the first functional entity and the second functional entity in model training is enhanced, which can improve the accuracy of model training.

[0058] In combination with the fourth aspect, in certain implementations of the fourth aspect, the first information is used to indicate that the model training has failed, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0059] In an embodiment of the present application, when the second functional entity fails in model training, the first information may include a failure reason value, so that the first functional entity can quickly determine the cause of the model training failure based on the failure reason value and adjust the model training strategy.

[0060] In combination with the fourth aspect, in certain implementations of the fourth aspect, before receiving the model training requirement information, the method includes: sending model training capability information to the first functional entity, and the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin or video memory bandwidth.

[0061] In an embodiment of the present application, the second functional entity can send model training capability information to the first functional entity, so that the first functional entity can determine the model training requirement information based on the model training capability information. In this way, the model training requirement information determined by the first functional entity can be closer to the actual situation of the model training performed by the second functional entity, and can further improve the efficiency of the model training performed by the second functional entity.

[0062] In a fifth aspect, a model training method is provided, which is applied to a first functional entity, and the method includes: sending model training requirement information to multiple second functional entities, the model training requirement information including at least one of the following: the maximum time for a single round of model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or a performance threshold value for model training; receiving multiple first model training parameters sent by the multiple second functional entities, the multiple first model training parameters are obtained based on multiple first model trainings; determining global model training parameters based on the multiple model training parameters; and sending the global model training parameters to the multiple second functional entities.

[0063] The model training method may be a distributed training method, that is, multiple second functional entities may perform model training according to the model training requirements sent by the first functional entity, and obtain multiple model training results. Finally, the multiple model training results may be aggregated to obtain a trained model.

[0064] Optionally, the first functional entity may be a MIF, and further, the first functional entity may be a centralized MIF.

[0065] Optionally, the second functional entity may be a MIF, and further, the second functional entity may be a distributed MIF.

[0066] Optionally, the first functional entity and / or the second functional entity may be located in a radio access network.

[0067] Optionally, the first model training parameter may be a gradient value, and the global model training parameter may be understood as: a parameter used for training on the entire data set during the model training process, which parameters may affect the learning process of the entire model.

[0068] In an embodiment of the present application, the first functional entity can send model training requirement information to multiple second functional entities. After receiving the first model training parameters sent by the multiple second functional entities, the first functional entity can determine the global training parameters based on the multiple first model training parameters, and send the global training parameters to the multiple second functional entities, so that the multiple second functional entities can perform model training based on the global training parameters. In this way, the collaboration between the first functional entity and the multiple second functional entities can be improved, thereby improving the success rate of model training. In addition, the efficiency of model training can also be improved through the above-mentioned distributed training method.

[0069] In combination with the fifth aspect, in certain implementations of the fifth aspect, before sending the model training requirement information to multiple second functional entities, the method also includes: receiving multiple model training capability information sent by the multiple second functional entities, the multiple model training capability information being used to indicate at least one of the following contents corresponding to the multiple second functional entities: single-round training time estimate, computing power margin, memory margin, video memory margin or video memory bandwidth; sending the model training requirement information to the multiple second functional entities includes: sending the model training requirement information to the multiple second functional entities according to the multiple model training capability information.

[0070] In an embodiment of the present application, the first functional entity can determine the model training requirement information based on multiple model training capability information sent by multiple second functional entities. In this way, the model training requirement information determined by the first functional entity can be closer to the actual situation of the model training of the multiple second functional entities, thereby further improving the efficiency of the model training of the second functional entities.

[0071] In combination with the fifth aspect, in certain implementations of the fifth aspect, before sending the model training requirement information to the multiple second functional entities, the method includes: sending the multiple first models to the multiple second functional entities.

[0072] In a sixth aspect, a model training method is provided, which is applied to a second functional entity, and the method includes: receiving model training requirement information sent by a first functional entity, the model training requirement information including at least one of the following: the maximum time for a single round of model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or a performance threshold value for model training; training a first model according to the model training requirement information to obtain first model training parameters; and sending the first model training parameters to the first functional entity.

[0073] In an embodiment of the present application, the second functional entity can train the first model based on the model training requirement information sent by the first functional entity to obtain the first model training parameters, and send the first model training parameters to the second functional entity, so that the first functional entity can determine the global training parameters based on multiple first model training parameters. In this way, the collaboration between the first functional entity and the second functional entity can be improved, thereby improving the accuracy of the final model for reasoning. In addition, the efficiency of model training can also be improved through the above-mentioned distributed training method.

[0074] In combination with the sixth aspect, in certain implementations of the sixth aspect, the method further includes: sending model training capability information to the first functional entity, the model training capability information being used to indicate at least one of the following contents corresponding to the second functional entity: single-round training time estimate, computing power margin, memory margin, video memory margin, or video memory bandwidth.

[0075] In an embodiment of the present application, the second functional entity can send model training requirement information to the first functional entity, so that the model training requirement information determined by the first functional entity can be closer to the actual situation of the model training performed by the second functional entity, thereby further improving the efficiency of the model training performed by the second functional entity.

[0076] In combination with the sixth aspect, in some implementations of the sixth aspect, before receiving the model training requirement information sent by the first functional entity, the method also includes: receiving the first model sent by the first functional entity.

[0077] In the seventh aspect, a model training device is provided, which is applied to a first functional entity, and the device includes: a transceiver unit and a processing unit, the transceiver unit is used to: receive model training requirement information sent by the second functional entity, the model training requirement information includes at least one of the following information: the number of layers of model training, the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; the processing unit is used to perform model training according to the model training requirement information.

[0078] In combination with the seventh aspect, in certain implementations of the seventh aspect, the transceiver unit is further used to send first information to the second functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0079] In combination with the seventh aspect, in certain implementations of the seventh aspect, the first information is used to indicate trained model information, and the first information includes: a first model.

[0080] In combination with the seventh aspect, in certain implementations of the seventh aspect, the transceiver unit is specifically used to receive a first data set sent by the second functional entity; the processing unit is also used to train the second model based on the first data set and the model training requirement information to obtain the first model.

[0081] In combination with the seventh aspect, in certain implementations of the seventh aspect, the first information is used to indicate that the model training has failed, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0082] In combination with the seventh aspect, in certain implementations of the seventh aspect, the transceiver unit is further used to send model training requirement information confirmation information to the second functional entity, and the model training requirement confirmation information is used to instruct the first functional entity to confirm the model training requirement information.

[0083] In an eighth aspect, a model training device is provided, which is applied to a second functional entity, and the device includes: a transceiver unit and a processing unit, the processing unit is used to determine model training requirement information, and the model training requirement information includes at least one of the following information: the number of layers of model training, the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; the transceiver unit is used to send the model training requirement information to the first functional entity.

[0084] In combination with the eighth aspect, in certain implementations of the eighth aspect, the transceiver unit is further used to receive first information sent by the first functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0085] In combination with the eighth aspect, in certain implementations of the eighth aspect, the first information is used to indicate the trained model information, and the first information includes a first model; the processing unit is specifically used to: obtain a first data set; send the first data set to the first functional entity; train the first model based on the first data set to obtain a third model.

[0086] In combination with the eighth aspect, in certain implementations of the eighth aspect, the first information is used to indicate that the model training has failed, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0087] In combination with the eighth aspect, in certain implementations of the eighth aspect, the transceiver unit is further used to receive model training requirement information confirmation information sent by the first functional entity, and the model training requirement confirmation information is used to indicate that the first functional entity confirms the model training requirement information.

[0088] In a ninth aspect, a model training device is provided, which is applied to a first functional entity, and the device includes: a transceiver unit and a processing unit, the processing unit is used to: determine model training requirement information, the model training requirement information includes at least one of the following information: model identification, model structure information, model training layer information, maximum time used for model training, data volume used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value for model training; the transceiver unit is used to send the model training requirement information to the second functional entity.

[0089] In combination with the ninth aspect, in certain implementations of the ninth aspect, the transceiver unit is further used to receive first information sent by the second functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0090] In combination with the ninth aspect, in certain implementations of the ninth aspect, the first information is used to indicate trained model information, and the first information includes a first model.

[0091] In combination with the ninth aspect, in certain implementations of the ninth aspect, the transceiver unit is further used to receive a first data set sent by a second functional entity; the processing unit is further used to perform model training based on the first data set to obtain a second model, and the first model is obtained by training the second model based on the model training requirement information; the transceiver unit is also used to send the second model to the second functional entity.

[0092] In combination with the ninth aspect, in certain implementations of the ninth aspect, the first information is used to indicate a failure of model training, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0093] In combination with the ninth aspect, in certain implementations of the ninth aspect, the transceiver unit is further used to receive model training capability information sent by the second functional entity, and the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin or video memory bandwidth; the processing unit is specifically used to determine the model training requirement information based on the model training capability information.

[0094] In the tenth aspect, a model training device is provided, which is applied to a second functional entity, and the device includes: a transceiver unit and a processing unit, the transceiver unit is used to receive model training requirement information sent by the first functional entity, and the model training requirement information includes at least one of the following information: model identification, model structure information, model training layer number information, maximum time used for model training, data volume used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value for model training; the processing unit is used to perform model training according to the model training requirement information.

[0095] In combination with the tenth aspect, in certain implementations of the tenth aspect, the transceiver unit is further used to send first information to the first functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0096] In combination with the tenth aspect, in certain implementations of the tenth aspect, the first information is used to indicate trained model information, and the first information includes: a first model.

[0097] In combination with the tenth aspect, in certain implementations of the tenth aspect, the processing unit is further used to obtain a first data set; the transceiver unit is further used to: send the first data set to the first functional entity; receive a second model sent by the first functional entity, wherein the second model is trained based on the first data set; the processing unit is further used to train the second model according to the first data set and the model training requirement information to obtain the first model.

[0098] In combination with the tenth aspect, in certain implementations of the tenth aspect, the first information is used to indicate a failure of model training, and the first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0099] In combination with the tenth aspect, in certain implementations of the tenth aspect, the transceiver unit is also used to send model training capability information to the first functional entity, and the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin or video memory bandwidth.

[0100] In combination with the eleventh aspect, a model training device is provided, which is applied to a first functional entity, and the device includes: a transceiver unit and a processing unit, the transceiver unit is used to: send model training requirement information to multiple second functional entities, the model training requirement information includes at least one of the following: the maximum time of a single round of model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; receive multiple first model training parameters sent by the multiple second functional entities, the multiple first model training parameters are obtained based on multiple first model training; the processing unit is used to determine the global model training parameters based on the multiple first model training parameters; the transceiver unit is also used to send the global model training parameters to the multiple second functional entities.

[0101] In combination with the eleventh aspect, in certain implementations of the eleventh aspect, the transceiver unit is further used to receive multiple model training capability information sent by the multiple second functional entities, and the multiple model training capability information is used to indicate at least one of the following contents corresponding to the multiple second functional entities: single-round training time estimate, computing power margin, memory margin, video memory margin or video memory bandwidth; the processing unit is also used to send model training requirement information to the multiple second functional entities based on the multiple model training capability information.

[0102] In combination with the eleventh aspect, in certain implementations of the eleventh aspect, the transceiver unit is further used to send the multiple first models to the multiple second functional entities.

[0103] In a twelfth aspect, a model training device is provided, which is applied to a second functional entity, and the device includes: a transceiver unit and a processing unit, the transceiver unit is used to receive model training requirement information sent by the first functional entity, and the model training requirement information includes at least one of the following: the maximum time of a single round of model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; the processing unit is used to train a first model according to the model training requirement information to obtain first model training parameters; the transceiver unit is also used to send the first model training parameters to the first functional entity.

[0104] In combination with the twelfth aspect, in certain implementations of the twelfth aspect, the transceiver unit is also used to send model training capability information to the first functional entity, and the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: single-round training time estimate, computing power margin, memory margin, video memory margin or video memory bandwidth.

[0105] In combination with the twelfth aspect, in some implementations of the twelfth aspect, the transceiver unit is further used to receive the first model sent by the first functional entity.

[0106] In the thirteenth aspect, a model training device is provided, which can be a first device, or a module or unit (for example, a chip, a chip system, or a circuit) in the first device that corresponds one-to-one to executing any method / operation / action from the first to the sixth aspects.

[0107] In combination with the thirteenth aspect, in certain implementations of the thirteenth aspect, the first device may be an access network device, a centralized MIF, a distributed MIF or other forms.

[0108] In the fourteenth aspect, a model training device is provided, comprising: a memory and at least one processor, wherein the at least one processor is coupled to the memory and is used to read and execute instructions in the memory so that the device implements the method in any one of the implementation modes of the first to sixth aspects above.

[0109] In a fifteenth aspect, a chip is provided, the chip comprising a circuit, the circuit being used to execute the method in any one of the implementations of the first to sixth aspects above.

[0110] In the sixteenth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a program code, and when the computer program code runs on a computer, the computer executes a method in any one of the implementation modes of the first to sixth aspects above.

[0111] In the seventeenth aspect, a computer program product is provided, which includes a computer program. When the computer program is run, the computer executes the method in any one of the implementation modes of the first to sixth aspects above.

[0112] In the eighteenth aspect, a system is provided, which includes a first functional entity and a second functional entity, the first functional entity is used to execute the method in any one of the implementations of the first aspect, the third aspect or the fifth aspect, and the second functional entity is used to execute the method in any one of the implementations of the second aspect, the fourth aspect or the sixth aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0113] Figure 1 This is the system architecture applicable to the model training method provided in the embodiment of the present application;

[0114] Figure 2 is a functional schematic diagram of the mobile intelligent function provided by the embodiment of the present application;

[0115] Figure 3It is a model training method provided in an embodiment of the present application;

[0116] Figure 4 This is another model training method provided in the embodiment of the present application;

[0117] Figure 5 This is another model training method provided in the embodiment of the present application;

[0118] Figure 6 This is another model training method provided in the embodiment of the present application;

[0119] Figure 7 This is another model training method provided in the embodiment of the present application;

[0120] Figure 8 This is another model training method provided in the embodiment of the present application;

[0121] Fig. 9 This is another model training method provided in the embodiment of the present application;

[0122] Fig.10 It is a model training device provided in an embodiment of the present application;

[0123] Fig.11 This is another model training device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0124] The technical solution in this application will be described below in conjunction with the accompanying drawings.

[0125] In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0126] The prefixes such as "first" and "second" used in the embodiments of the present application are only used to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers used to distinguish description objects in the embodiments of the present application does not constitute a limitation on the described objects. For the statement of the described objects, please refer to the description in the context of the claims or embodiments, and the use of such prefixes should not constitute an unnecessary limitation.

[0127] The technical solutions of the embodiments of the present application can be applied to various communication systems, for example: global system of mobile communication (GSM) system, code division multiple access (CDMA) system, wideband code division multiple access (WCDMA) system, general packet radio service (GPRS), long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD), universal mobile telecommunication system (UMTS), worldwide interoperability for microwave access (WiMAX) communication system, fifth generation (5G) system or new radio (NR) and future sixth generation (6G) system, etc.

[0128] The terminal device in the embodiments of the present application may refer to user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device. The terminal device may also be a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a wearable device, a terminal device in a future 5G network or a terminal device in a future evolved public land mobile communication network (PLMN), etc., and the embodiments of the present application are not limited to this.

[0129] The network device in the embodiment of the present application can be a device for communicating with a terminal device. The network device can be a base station (base transceiver station, BTS) in a global system of mobile communication (GSM) system or code division multiple access (CDMA), or a base station (NodeB, NB) in a wideband code division multiple access (WCDMA) system, or an evolutionary base station (eNB or eNodeB) in an LTE system, or a wireless controller in a cloud radio access network (CRAN) scenario, or the network device can be a relay station, an access point, a vehicle-mounted device, a wearable device, a network device in a 5G or 6G network, or a network device in a future evolved PLMN network, etc., and the embodiments of the present application are not limited.

[0130] The following introduces the technical problems to be solved by this application and the technical solutions adopted.

[0131] AI is a technology that imitates human cognition, learning and reasoning abilities. ML is an important technical means in the field of artificial intelligence. Its core idea is to learn a large amount of known data to obtain the relationship between data, and finally predict and analyze unknown data. Common AI or ML technologies can include supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, etc. In the RAN field, AI or ML also has a wide range of applications, such as business prediction, resource scheduling and network management based on AI or ML technology.

[0132] However, when applying AI or ML technology in the RAN field, due to objective reasons such as the cost of access network equipment, wireless environment, and application requirements, the amount of data that can be collected by access network equipment and the available computing resources and storage resources are very limited, which may lead to model training failure (for example, model performance does not meet standards and model training timeouts) and waste of computing resources.

[0133] The present application provides a model training method and device, which can improve the success rate of model training and avoid the waste of computing resources.

[0134] Figure 1 This is the system architecture applicable to the model training method provided in the embodiment of the present application.

[0135] like Figure 1 As shown in (a) to (f) of FIG. 1 , the system architecture may include: a base station, a UE, and a mobile intelligent function (MIF). Figure 2 As shown, MIF can be responsible for the AI ​​or ML functions of the base station, including: data management function, computing power management function and model management function, among which the data management function may include: data collection, data storage and data analysis; the computing power management function may include: computing power perception, computing power scheduling, and computing power and transmission coordination; the model management function may include: model training, model reasoning and model lifecycle management.

[0136] Optionally, the MIF may be in the form of: a base station, a network element independent of the base station, a network function independent of the base station, a submodule within the base station, or a subfunction within the base station.

[0137] For example, MIF can be a separate unit such as Figure 1 As shown in (a); MIF can also be divided into two units: centralized MIF and distributed MIF, such as Figure 1As shown in (b) and (c) in . Among them, for centralized AI, the centralized MIF can train the initial model, and the distributed MIF can update the model. For distributed AI, the centralized MIF can be responsible for the division and global training of the model, and the distributed MIF can be responsible for local model training. It should be understood that the above-mentioned centralized MIF and distributed MIF can jointly complete the functions of MIF, and the way in which the above-mentioned MIF is divided into centralized MIF and distributed MIF can be divided from a logical perspective or from a physical structure perspective. As for how the functions are divided, this application does not limit it.

[0138] against Figure 1 In (a) to (c), if the base station adopts a split architecture, the corresponding system architecture can be as follows: Figure 1 As shown in (d) to (e), the separated architecture can be understood as the base station being divided into two parts: a centralized unit (CU) and a distributed unit (DU). The CU is responsible for the wireless high-level protocol functions, and the DU is responsible for the wireless low-level protocol functions.

[0139] It should be understood that in the embodiments of the present application, the base station, the network device and the access network device may be the same concept and may be used interchangeably.

[0140] Figure 3 It is a model training method provided in an embodiment of the present application. Method 300 may include steps S301 to S302.

[0141] S301, the first functional entity receives model training requirement information sent by the second functional entity.

[0142] The model requirement information may include at least one of the following information: the number of layers of model training, the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training. The specific meanings and examples of these information will be introduced in detail in method 500 and method 600 and will not be repeated here.

[0143] Optionally, the first functional entity may be a mobile intelligent function, and further, the first functional entity may be a centralized MIF.

[0144] Optionally, the second functional entity may be an access network device or a MIF. Further, when the second functional entity is a MIF, it may be a distributed MIF.

[0145] Optionally, the first functional entity and / or the second functional entity may be located in a radio access network.

[0146] In one embodiment, before step S301, method 300 may further include: the second functional entity determines model training requirement information.

[0147] S302: The first functional entity performs model training according to model training requirement information.

[0148] In an embodiment of the present application, the first functional entity can receive model training requirement information sent by the second functional entity, and perform model training according to the model training requirement information. In this way, the first functional entity can perform model training in a targeted manner, improve the success rate of model training, and reduce the occurrence of resource waste during model training.

[0149] It should be understood that in this application, model training requirement information may also be referred to as model training configuration information, joint training configuration information or configuration information.

[0150] After performing model training, the first functional entity may inform the second functional entity of the result of the model training.

[0151] In one embodiment, after step S302, the first functional entity may send first information to the second functional entity, where the first information is used to indicate the trained model information, or the first information is used to indicate the failure of model training. In this way, it is possible to ensure that the second functional entity is informed of the training result of the model in a timely manner, thereby enhancing the collaboration between the first functional entity and the second functional entity.

[0152] Optionally, when the first information is used to indicate trained model information, the first information may include a first model, and the first model may be understood as a model trained by the first functional entity based on model training requirement information.

[0153] Optionally, when the first information is used to indicate trained model information, the first information may also include a first model identifier, which may be an integer (eg, 0001) or a character string (eg, model A).

[0154] In one embodiment, before step S302, the first functional entity may receive a first data set sent by the second functional entity. Then, in step S302, the first functional entity may train the second model based on the first data set and the model training requirement information to obtain the first model. In this way, the first model can be obtained more efficiently and the success rate of the second model training can be improved.

[0155] The second model may also be referred to as an initial model. The second model may be a model whose parameters are initial values, or a model trained based on other data sets.

[0156] Optionally, before sending the first data set, the second functional entity may first collect the first data set.

[0157] Specifically, the data in the first data set may be labeled data or unlabeled data. Among them, the labeled data may include the data itself and the corresponding label. For example, the first data set for AI positioning is {(CIR1, TOD1), (CIR2, TOA2)}, then in each data, CIR is the data itself, and TOA is the label corresponding to the data. Unlabeled data may include only the data itself, excluding the label corresponding to the data. For example, the data set for AI channel compression is {CSI1, CSI2}, then each data includes only the data itself, that is, CSI.

[0158] In one embodiment, when the first information is used to indicate that the model training has failed, the first information may include a failure reason value, and the failure reason value may include at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements. In this way, the second functional entity can quickly determine the cause of the model training failure based on the failure reason value and adjust the model training strategy.

[0159] Among them, insufficient computing power can be understood as insufficient computing power when the first functional entity performs model training, insufficient training samples can be understood as insufficient data volume or samples when the first functional entity performs model training, and the performance of the model after training does not meet the requirements. For example, the accuracy rate is lower than the threshold value, or the mean square error is higher than the threshold value, etc.

[0160] Optionally, the failure reason value may indicate the reason for the failure in the form of a number. For example, when the failure reason value is 01, it may indicate that the computing power of the first functional entity is insufficient. When the failure reason value is 10, it may indicate that the first functional entity has insufficient training samples for model training. When the failure reason value is 11, it may indicate that the performance of the first model does not meet the requirements.

[0161] Optionally, the correspondence between the failure cause value and the failure cause may be indicated using a mapping relationship, and the second functional entity may determine the failure cause based on the failure cause value and the mapping relationship.

[0162] For example, the above mapping relationship can be reflected by Table 1, that is, after receiving the failure reason value, the second functional entity can determine the reason for the failure of the first functional entity model training according to the mapping relationship in Table 1.

[0163] Table 1

[0164] Failure reason value Cause of failure 01 Insufficient computing power 10 Insufficient training samples 11 Model performance does not meet requirements

[0165] Before performing model training based on the model training requirement information, the first functional entity may send model training requirement information confirmation information or model training requirement information change information to the second functional entity, so as to better complete the training.

[0166] For example, when the model training requirement information includes the amount of data used for model training, and the value of the amount of data used for model training is 1000, if the first functional entity believes that the model training can be successfully completed according to the model training quantity requirement, the model training requirement information confirmation information is sent to the second functional entity. If the first functional entity believes that the model training cannot be successfully completed according to the model training quantity requirement, and the above-mentioned data amount should be adjusted to 2000, the first functional entity can send the training requirement information change information to the second functional entity, indicating that the above-mentioned data amount is adjusted to 2000.

[0167] The first functional entity and the second functional entity may also complete the training of the model through collaboration.

[0168] In one embodiment, after receiving the first model, the second functional entity can train the first model according to the first data set to obtain a third model, and apply the third model for reasoning. In this way, the collaboration between the first functional entity and the second functional entity in the model training process is enhanced, thereby further improving the reasoning performance of the trained model.

[0169] Alternatively, the second functional entity may collect a second data set, and train the first model according to the second data set to obtain a third model, wherein the data contents of the first data set and the second data set may be different.

[0170] Figure 4 It is a model training method provided by an embodiment of the present application. Method 400 may include steps S401 to S403.

[0171] S401, the first functional entity determines model training requirement information.

[0172] The model training requirement information includes at least one of the following information: model identification, model structure information, model training layer information, maximum time used for model training, amount of data used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value of model training. The specific meaning and examples of these information will be described in detail in method 800, and will not be repeated here.

[0173] Optionally, the first functional entity may be a MIF, and further, the first functional entity may be a centralized MIF.

[0174] Optionally, the second functional entity may be a MIF, and further, the second functional entity may be a distributed MIF.

[0175] Optionally, the first functional entity and / or the second functional entity may be located in a radio access network.

[0176] In one embodiment, before step S401, the first functional entity may receive model training capability information sent by the second functional entity, and the model training capability information may be used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin or video memory bandwidth. The specific meanings and examples of these information will be introduced in detail in method 800 and will not be repeated here. Then in step S401, the first functional entity may determine the model training requirement information based on the above-mentioned model training capability information. In this way, the model training requirement information determined by the first functional entity can be closer to the actual situation of the model training performed by the second functional entity, thereby further improving the efficiency of the model training performed by the second functional entity.

[0177] S402: The first functional entity sends model training requirement information to the second functional entity.

[0178] S403: The second functional entity performs model training according to the model training requirement information.

[0179] In an embodiment of the present application, the first functional entity can determine model training requirement information and send the model training requirement information to the second functional entity, so that the second functional entity can perform model training based on the model training requirement information. In this way, the second functional entity can perform model training in a targeted manner, improve the success rate of model training, and reduce the occurrence of resource waste during model training.

[0180] After performing model training, the second functional entity may inform the first functional entity of the result of the model training.

[0181] In one embodiment, after step S403, the second functional entity may send first information to the first functional entity, where the first information is used to indicate the trained model information, or the first information is used to indicate the failure of model training. In this way, it is possible to ensure that the first functional entity is informed of the training result of the model in a timely manner, thereby enhancing the collaboration between the first functional entity and the second functional entity.

[0182] Optionally, when the first information is used to indicate trained model information, the first information may include a first model, and the first model may be understood as a model trained by the second functional entity based on model training requirement information.

[0183] Optionally, when the first information is used to indicate trained model information, the first information may also include a first model identifier, which may be an integer (eg, 0001) or a character string (eg, model A).

[0184] In one embodiment, before step S403, the first functional entity may receive the first data set sent by the second functional entity, and the first functional entity may also perform model training based on the first data set to obtain a second model, and send the second model to the second functional entity. Then, in step S403, the second functional entity may train the second model based on the first data set and the model training requirement information to obtain the first model. In this way, the collaboration between the first functional entity and the second functional entity in model training is enhanced, and the accuracy of model training can be improved.

[0185] In one embodiment, when the first information is used to indicate that the model training has failed, the first information may include a failure reason value, and the failure reason value may include at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements. In this way, the first functional entity can quickly determine the cause of the model training failure based on the failure reason value and adjust the model training strategy.

[0186] Figure 5 It is a model training method provided by an embodiment of the present application. Method 500 may include steps S501 to S505.

[0187] S501, a first functional entity sends model training requirement information to multiple second functional entities.

[0188] The model training requirement information includes at least one of the following information: the maximum time for a single round of model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training. The specific meanings and examples of these information will be introduced in detail in method 900 and will not be repeated here.

[0189] Optionally, the first functional entity may be a MIF, and further, the first functional entity may be a centralized MIF.

[0190] Optionally, the second functional entity may be a MIF, and further, the second functional entity may be a distributed MIF.

[0191] Optionally, the first functional entity and / or the second functional entity may be located in a radio access network.

[0192] In one embodiment, before step S501, method 500 further includes: the first functional entity receives multiple model training capability information sent by multiple second functional entities, and the multiple model training capability information is used to indicate at least one of the following contents corresponding to the multiple second functional entities: single round training time estimate, computing power margin, memory margin, video memory margin or video memory bandwidth; then in step S501, the first functional entity can send model training requirement information to the multiple second functional entities based on the multiple model training capability information. In this way, the model training requirement information determined by the first functional entity can be closer to the actual situation of the multiple second functional entities performing model training, thereby further improving the efficiency of the multiple second functional entities performing model training.

[0193] Optionally, the corresponding parameters in the above-mentioned multiple model training capability information may be the same or different, and this application does not limit this.

[0194] In one embodiment, before step S501, the method 500 further includes: the first functional entity sends a plurality of first models to a plurality of second functional entities.

[0195] Optionally, the above-mentioned multiple first models can be the same model or different models, which is not limited in the present application.

[0196] S502, multiple second functional entities perform model training according to model training requirement information to obtain multiple first model training parameters.

[0197] The first model training parameters may be training parameters obtained when the first model is trained by multiple second functional entities.

[0198] Optionally, the first model training parameter may be a gradient value.

[0199] S503: Multiple second functional entities send multiple first model training parameters to the first functional entity.

[0200] S504: The first functional entity determines a global model training parameter according to multiple first model training parameters.

[0201] Optionally, global training parameters can be understood as parameters used to train on the entire data set during model training, which will affect the learning process of the entire model.

[0202] S505: The first functional entity sends global model training parameters to multiple second functional entities.

[0203] Among them, multiple second functional entities can perform model training according to global model training parameters to obtain multiple model training results.

[0204] Optionally, the multiple model training results may be multiple trained sub-models.

[0205] It should be understood that method 500 can be a distributed training method, that is, multiple second functional entities perform model training according to the model training requirements sent by the first functional entity, and obtain multiple model training results. Finally, the multiple model training results can be summarized to obtain a trained model.

[0206] In an embodiment of the present application, the first functional entity can send model training requirement information to multiple second functional entities. After receiving the first model training parameters sent by the multiple second functional entities, the first functional entity can determine the global training parameters based on the multiple first model training parameters, and send the global training parameters to the multiple second functional entities, so that the multiple second functional entities can perform model training based on the global training parameters. In this way, the collaboration between the first functional entity and the multiple second functional entities can be improved, thereby improving the success rate of model training. In addition, the efficiency of model training can also be improved through the above-mentioned distributed training method.

[0207] Figure 6 Another model training method provided in the embodiment of the present application is provided. The method 600 can be applied to Figure 1 In the system architecture shown in (a) and (d), method 600 may be a specific introduction to step S301 to step S302 in method 300, and method 600 may include the following steps.

[0208] S601, the access network device sends data set #1 to the MIF.

[0209] The data set #1 can be used for model training, verification and testing by MIF. The access network device can be the second functional entity in the method 300, and the MIF can be the first functional entity in the method 300.

[0210] Optionally, the access network device may obtain data set #1 by collecting data online to obtain the data set, the network management, core network or external server may generate the data set and send it to the access network device, or generate the data set through laboratory simulation and send it to the access network device.

[0211] S602, MIF performs initial model training based on data set #1.

[0212] The initial model may be the second model in method 300 .

[0213] S603, MIF feeds back the initial model to the access network device.

[0214] Optionally, after receiving the initial model, the access network device may use the model for reasoning.

[0215] Optionally, method 600 may not execute steps S601 to S603 and directly execute step S604.

[0216] S604, the access network device sends model training requirement information to the MIF.

[0217] Exemplarily, the model training requirement information may include at least one of the following: the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or a model performance threshold.

[0218] Among them, the maximum value of the time used for model training can be used to indicate the maximum value of the time used by MIF for model training (corresponding to step S608), or the maximum interval between MIF receiving data set #2 (corresponding to step S607) and MIF sending the model to the access network device (corresponding to step S609); the model performance threshold can be used to indicate the threshold setting of the performance of the trained model.

[0219] Optionally, the time used for the MIF training model (corresponding to step S608) may be less than or equal to the maximum time used for model training; the amount of data used for the MIF training model may be less than or equal to the amount of data used for model training; the amount of data used for the MIF training model may be greater than the minimum amount of data used for model training; the performance of the MIF training model may meet the model performance threshold, if the model performance threshold is the minimum value, then the performance of the training model may be greater than or equal to the threshold, if the model performance threshold is the maximum value, then the training performance of the model may be less than or equal to the threshold.

[0220] Optionally, the model performance threshold may include at least one of the following: the maximum and / or minimum value of accuracy, the maximum and / or minimum value of precision, the maximum and / or minimum value of recall, the maximum and / or minimum value of F1 score, the maximum and / or minimum value of recall, the maximum and / or minimum value of mean absolute error (MAE), the maximum and / or minimum value of mean squared error (MSE), the maximum and / or minimum value of root mean square error (RMSE), the maximum and / or minimum value of mean absolute percentage error (MAPE), and the maximum and / or minimum value of cross-entropy.

[0221] S605, MIF sends demand confirmation or demand change information to the access network device.

[0222] Specifically, the MIF may send demand confirmation information to the access network device to confirm the model training requirement in step S605, or the MIF may send demand change information to the access network device to indicate the training capabilities that the MIF can provide.

[0223] Optionally, the requirement change information may include at least one of the following: the maximum time available for model training, the amount of data available for model training, the maximum amount of data available for model training, the minimum amount of data available for model training, or an available model performance threshold.

[0224] Among them, the maximum value of the time used for model training that can be provided can correspond to the maximum value of the time used for model training in step S604. For example, the maximum value of the time used for model training in step S604 is 1 second. MIF believes that this time requirement cannot be met and the above maximum value should be adjusted to 2 seconds. Then in the requirement change information, the maximum value of the time used for model training can be 2 seconds.

[0225] The amount of data available for model training may correspond to the amount of data used for model training in step S604. For example, the amount of data used for model training in step S604 is 1000. MIF believes that this quantity requirement cannot be met and the above data amount should be adjusted to 2000. In this case, the maximum value of the amount of data required for model training in the requirement change information can be 2000.

[0226] The maximum amount of data required for model training that can be provided may correspond to the maximum amount of data used for model training in step S604. For example, the maximum amount of data used for model training in step S604 is 1000. MIF believes that this quantity requirement cannot be met and that the above amount of data needs to be adjusted to 2000. Then the maximum amount of data required for model training in step S605 is 2000.

[0227] The minimum amount of data required for model training that can be provided may correspond to the minimum amount of data used for model training in step S604. For example, the minimum amount of data used for model training in step S604 is 1,000. MIF believes that this quantity requirement cannot be met and that the above amount of data needs to be adjusted to 600. Then the minimum amount of data required for model training in step S605 is 600.

[0228] The available model performance threshold may correspond to the model performance threshold in step S604. For example, the "model performance threshold" in step S604 indicates that the minimum model accuracy is 95%. MIF believes that this performance requirement cannot be met and the above performance index needs to be adjusted to 90%. Then the available model performance threshold in step S605 indicates that the minimum model accuracy is 90%.

[0229] S606, the access network device performs data collection.

[0230] Specifically, the data collected by access network equipment can be used for MIF to conduct model training, verification and testing.

[0231] Optionally, the access network device may collect data in the following ways: the access network device collects data online, the network management, core network or external server generates data and sends it to the access network device, and the laboratory simulation generates data and sends it to the access network device.

[0232] S607, the access network device sends data set #2 to the MIF.

[0233] The data set #2 may be data collected by the access network device in step S606 , and the data set #2 may be the first data set in method 300 .

[0234] S608, MIF performs model training.

[0235] When MIF performs model training, it is necessary to comply with the indicators determined in step S604 and step S605.

[0236] For example, in step S604, the access network device requires that the maximum time used for model training is 1 second, and (step S605) MIF does not indicate the maximum time used for model training that can be provided, then MIF can complete the training within 1 second.

[0237] For another example, in step S604, the access network device requires that the maximum time for model training is 1 second, and (step S605) the MIF indicates that the maximum time that can be provided for model training is 2 seconds, then the MIF should complete the training within 2 seconds.

[0238] S609, MIF sends the model to the access network device.

[0239] Specifically, if in step S608 , the MIF completes the training of the model, the MIF may send the trained model to the access network device, wherein the trained model may be the first model in method 300 .

[0240] Optionally, if the MIF fails to complete the training of the model in step S608, for example, the MIF fails to complete the training within the specified time, or the model trained by the MIF fails to meet the specified model training threshold, the MIF may send a training failure indication message to the access network device to indicate that the model training has failed, and the failure indication message may be the first information in method 300.

[0241] Optionally, the above training failure indication information may include a failure reason value, which can be used to indicate at least one of the following: training timeout, insufficient training samples, and model performance not meeting requirements.

[0242] The access network device can perform different processing after receiving the failure cause value.

[0243] For example, if the failure reason value is training timeout, the access network device can relax the constraint on the model training time when sending a model training request to the MIF in a subsequent process.

[0244] For another example, if the failure reason value is that the model training performance does not meet the requirements, the subsequent access network device can collect more data for model training, or the access network device can abandon the use of the model and switch to other algorithms (for example, optimization algorithms, etc.).

[0245] In an embodiment of the present application, the access network device and MIF can negotiate to determine the indicators of model training, so that MIF can determine the training requirements of the model in a timely manner, thereby better allocating resources for model training to ensure that model training meets user needs.

[0246] Figure 7 Another model training method provided in the embodiment of the present application is provided. The method 700 can be applied to Figure 1 For the system architectures shown in (b) and (e) in , method 700 may be a detailed introduction to steps S301 to S302 of method 300, and method 700 may include the following steps.

[0247] S701, The centralized MIF sends model information to the distributed MIF.

[0248] Among them, the centralized MIF may be the first functional entity in method 300, and the distributed MIF may be the second functional entity in method 300.

[0249] Among them, the model information may include information about the model structure, and the information about the model structure may include at least one of the following: the number of model layers, the model input dimension, or the model output dimension.

[0250] Optionally, the model information may include a model identifier, and the model identifier may be an integer (for example, 0001) or a string (for example, model A).

[0251] S702, The distributed MIF sends joint model training requirement information to the centralized MIF.

[0252] Exemplarily, the joint model training requirement information may be used to indicate the requirement for training the centralized MIF, and may specifically include at least one of the following: information about the number of model layers to be trained, the amount of data used for model training, the maximum value of the amount of data used for model training, the minimum value of the amount of data used for model training, or the model performance threshold. The joint model training requirement information may be the model training requirement information in method 300.

[0253] Among them, the information about the number of model layers to be trained may be used to indicate which layers of the model the centralized MIF needs to train. This information may be an integer N, indicating the first N layers or the last N layers of the model to be trained; or, this information may be two integers M and N (M < N), indicating the Mth layer to the Nth layer of the model to be trained; or, this information may be an index, and this index corresponds to an item in a predefined or preconfigured table (predefined may mean predefined in the protocol, and preconfigured may mean preconfigured by the network management system or a certain network device for the centralized MIF and the distributed MIF), and the content of this item indicates which layers of the model need to be trained.

[0254] The amount of data used for model training can be used to indicate the amount of data used by the centralized MIF for model training; the maximum time used for model training can be used to indicate the maximum amount of data used by the MIF for model training; the minimum amount of data used for model training can be used to indicate the minimum amount of data used by the centralized MIF for model training; the model performance threshold can be used to indicate the threshold setting of the performance of the model trained by the centralized MIF.

[0255] Optionally, the model performance threshold may include at least one of the following: maximum and / or minimum value of accuracy, maximum and / or minimum value of precision, maximum and / or minimum value of recall, maximum and / or minimum value of F1 score, maximum and / or minimum value of recall, maximum and / or minimum value of mean absolute error, maximum and / or minimum value of mean square error, maximum and / or minimum value of root mean square error, maximum and / or minimum value of mean absolute percentage error, maximum and / or minimum value of cross entropy.

[0256] S703, the centralized MIF sends a request confirmation or request change message to the distributed MIF.

[0257] Among them, step S703 is an optional step, and method 700 may not execute step S703 and directly execute step S704.

[0258] Specifically, the centralized MIF may send confirmation information to the distributed MIF to confirm the requirement in step S702, or the centralized MIF may send change request information to the distributed MIF to indicate the need to change the model training or to indicate the training capability that the centralized MIF can provide.

[0259] Optionally, the requirement change information may include at least one of the following: information on the number of trainable model layers, the amount of data required for model training that can be provided, the maximum amount of data required for model training that can be provided, the minimum amount of data required for model training that can be provided, or an available model performance threshold.

[0260] The information on the number of model layers that can be trained may correspond to the information on the number of model layers that need to be trained in step S702. For example, the information on the number of model layers that need to be trained in step S702 is an integer of 5 (indicating that the centralized MIF needs to train the first 5 layers of the model), and the centralized MIF believes that this requirement cannot be met, but needs to adjust the above value to 3 (indicating that the centralized MIF can train the first 3 layers of the model), then the value of the information on the number of model layers that can be trained in step S703 may be 3.

[0261] The amount of data available for model training may correspond to the amount of data used for model training in step S702. For example, the value of the amount of data used for model training in step S702 is 1000. The centralized MIF believes that this quantity requirement cannot be met and the above data amount should be adjusted to 2000. Then the value of the amount of data required for model training in step S703 is 2000.

[0262] The maximum amount of data required for model training that can be provided may correspond to the maximum amount of data used for model training in step S702. For example, the maximum amount of data used for model training in step S702 is 1000. The centralized MIF believes that this quantity requirement cannot be met and needs to adjust the above data amount to 2000. Then the maximum amount of data required for model training in step S703 is 2000.

[0263] The minimum amount of data required for model training that can be provided may correspond to the minimum amount of data used for model training in step S702. For example, the minimum amount of data used for model training in step S702 is 1000. The centralized MIF believes that this quantity requirement cannot be met and needs to adjust the above data amount to 700. Then the minimum amount of data required for model training in step S703 is 700.

[0264] The available model performance threshold may correspond to the model performance threshold in step S702. For example, the model performance threshold in step S702 indicates that the minimum model accuracy is 95%. The centralized MIF believes that this performance requirement cannot be met and needs to adjust the above performance indicator to 90%. Then the available model performance threshold in step S703 indicates that the minimum model accuracy is 90%.

[0265] S704, the distributed MIF performs data collection.

[0266] Specifically, the data is used in the centralized MIF for model training, validation, and testing.

[0267] Optionally, the distributed MIF may obtain data in at least one of the following ways: the distributed MIF collects data online, a network device or a network management or an external server generates data and sends it to the distributed MIF, or a laboratory simulation generates data and sends it to the distributed MIF.

[0268] S705, the distributed MIF sends the data set to the centralized MIF.

[0269] Specifically, the data set may be a data set obtained by the distributed MIF through data collection in step S704 , and the data set may be the first data set in method 300 .

[0270] S707, centralized MIF performs model training.

[0271] Optionally, when the centralized MIF performs model training, the indicators determined in step S702 and step S703 may be followed.

[0272] For example, in step S702, the distributed MIF requires 5 layers of model to be trained, and in step S703, the centralized MIF does not indicate the number of trainable model layers, then in step S706, the centralized MIF can train the first 5 layers of the model.

[0273] For another example, in step S702, the distributed MIF requires that the number of model layers to be trained is 5, and in step S703 the centralized MIF indicates that the number of model layers that can be trained is 3. Then in step S706, the centralized MIF can train the first 3 layers of the model.

[0274] S707, the centralized MIF sends the model to the distributed MIF.

[0275] Specifically, in step S706 , the centralized MIF completes the model training. Then, in step S707 , the centralized MIF may send the trained model to the distributed MIF. The trained model may be the first model in method 300 .

[0276] Alternatively, if the model training fails in step S706, in step S707, the centralized MIF may send training failure indication information to the distributed MIF to indicate the model training failure.

[0277] Optionally, the above-mentioned model training failure indication information may include a failure reason value, which can be used to indicate at least one of the following: training timeout, insufficient training samples, or model performance does not meet requirements.

[0278] The distributed MIF can perform different processing after receiving the failure reason value.

[0279] For example, if the reason value for failure is insufficient computing power, the distributed MIF can reduce the number of layers that require the centralized MIF to train the model when subsequently sending the joint model training requirements to the centralized MIF.

[0280] For another example, if the failure reason value is that the model performance does not meet the requirements, the distributed MIF can collect more data for model training.

[0281] S708, distributed MIF performs model retraining.

[0282] Specifically, after the distributed MIF receives the model sent by the centralized MIF, it can retrain the model.

[0283] Optionally, model retraining may include: training the remaining layers of the model, or training the model with a new data set, and the distributed MIF retraining the model can obtain the third model in method 300.

[0284] In an embodiment of the present application, the distributed MIF and the centralized MIF can negotiate to determine the training indicators of the model joint, so that the centralized MIF can determine its own needs for model training, thereby better allocating resources for model training to ensure that the model training meets the needs.

[0285] Figure 8 Another model training method provided in the embodiment of the present application is provided. The method 800 can be applied to Figure 1 In the system architecture shown in (b) and (e), method 800 may be a specific introduction to steps S401 to S403 in method 400, and method 800 may include the following steps.

[0286] S801, the centralized MIF sends model training capability request information to the distributed MIF.

[0287] Among them, step S801 is an optional step, that is, method 800 may not execute step S801 and directly execute step S802.

[0288] Specifically, the model training capability request information can be used to request the distributed MIF to report the model training capability information, wherein the centralized MIF can be the first functional entity in method 400, and the distributed MIF can be the second functional entity in method 400.

[0289] Optionally, the model training capability request information may include at least one of the following: a computing power margin indication, a memory margin indication, a video memory margin indication, or a video memory bandwidth indication.

[0290] The computing power margin indication can be used to request reporting of computing power margin, and the indication can be for a certain data type, for example, the indication is used to request computing power margin for 32-bit floating-point data (32bit floating-point, FP32), or the indication is used to request computing power margin for 16-bit integer data (16bit integer, INT16). The units of the above-mentioned "computing power" and "computing power margin" can be: the number of floating-point operations per second (floating-point operations persecond, FLOPS) or the total number of operations per second (trillions of operations per second, TOPS), etc.

[0291] The memory margin indication may be used to indicate a request to report the memory margin. The units of the above “memory” and “memory margin” may be gigabytes (GB) and the like.

[0292] The video memory remaining amount indication may be used to request reporting of the video memory remaining amount. The units of the above “video memory” and “video memory remaining amount” may be gigabyte (GB) and the like.

[0293] The memory bandwidth indication may be used to request reporting of the memory bandwidth. The unit of the “memory bandwidth” may be terabyte per second (TB / s), etc.

[0294] S802, the distributed MIF sends a model training capability report to the centralized MIF.

[0295] Specifically, the model training capability report can be used to indicate the model training capability information of the distributed MIF, that is, the model training capability report can be the model training capability information in method 400.

[0296] Optionally, the model training capability information may include at least one of the following items for indicating the distributed MIF: computing power margin, memory margin, video memory margin, or video memory bandwidth.

[0297] The computing power margin can be used to indicate the current computing power margin of the distributed MIF or the computing power that the distributed MIF can currently provide. The unit can be: FLOPS or TOPS, etc. The computing power margin can be for a given data type, for example, the computing power margin for FP32, or the computing power margin for INT16, etc.

[0298] The memory margin may be used to indicate the memory size currently available for the distributed MIF, and the unit may be GB or the like.

[0299] The video memory margin may be used to indicate the current video memory margin of the distributed MIF, or the video memory size that the distributed MIF can currently provide, and the unit may be GB or the like.

[0300] The memory bandwidth may be used to indicate the memory bandwidth currently available for the distributed MIF, and the unit may be TB / s or the like.

[0301] S803, the centralized MIF sends joint training configuration information to the distributed MIF.

[0302] Specifically, the joint model training configuration information can be used to indicate the requirements for the distributed MIF to perform model training, that is, the joint training configuration information can be the model training requirement information in method 400.

[0303] Exemplarily, the joint training configuration information may include at least one of the following: model identifier, model structure information, information on the number of model layers to be trained, the amount of data used for model training, the maximum value of the amount of data used for model training, the minimum value of the amount of data used for model training, or the model performance threshold.

[0304] Among them, the model identifier may be an integer (e.g., 0001) or a string (e.g., model A).

[0305] The information on the model structure may include at least one of the following: the number of model layers, the model input dimension, or the model output dimension.

[0306] The information on the number of model layers to be trained can be used to indicate which layers of the model the distributed MIF needs to train. This information can be an integer N, indicating that the first N layers or the last N layers of the model need to be trained; or, this information can be two integers M and N (M < N), indicating that the Mth layer to the Nth layer of the model need to be trained; or, this information can be an index that corresponds to an item in a predefined or preconfigured table (predefined can mean predefined in the protocol, and preconfigured can mean preconfigured by the network management system or a certain network device for the centralized MIF and the distributed MIF), and the content of this item indicates which layers of the model need to be trained.

[0307] The amount of data used for model training can be used to indicate the amount of data used by the centralized MIF for model training; the maximum value of the time used for model training can be used to indicate the maximum amount of data used by the MIF for model training; the minimum value of the amount of data used for model training can be used to indicate the minimum amount of data used by the centralized MIF for model training; the model performance threshold can be used to indicate the threshold setting of the performance of the model trained by the centralized MIF.

[0308] Optionally, the model performance threshold may include at least one of the following: the maximum and / or minimum value of accuracy, the maximum and / or minimum value of precision, the maximum and / or minimum value of recall, the maximum and / or minimum value of F1 score, the maximum and / or minimum value of recall, the maximum and / or minimum value of mean absolute error, the maximum and / or minimum value of mean square error, the maximum and / or minimum value of root mean square error, the maximum and / or minimum value of mean absolute percentage error, the maximum and / or minimum value of cross entropy.

[0309] S804, the distributed MIF performs data collection.

[0310] Optionally, the distributed MIF may obtain data in at least one of the following ways: the distributed MIF collects data online, a network device or a network management or an external server generates data and sends it to the distributed MIF, and a laboratory simulation generates data and sends it to the distributed MIF.

[0311] S805, the distributed MIF sends the data set to the centralized MIF.

[0312] Specifically, the data set may be a data set obtained by the distributed MIF through data collection in step S805 , and the data set may be the first data set in method 400 .

[0313] S806, the centralized MIF performs model training according to the received data set.

[0314] S807, the centralized MIF sends the trained model to the distributed MIF.

[0315] Among them, the trained model can be the second model in method 400.

[0316] S808, distributed MIF performs model retraining.

[0317] Optionally, when the distributed MIF performs model retraining, the indicators determined in step S803 may be followed, and the model retraining result of the distributed MIF may be the first model in method 400 .

[0318] S809, the distributed MIF sends training result feedback information to the centralized MIF.

[0319] Among them, step S803 may be an optional step, that is, method 800 may not execute step S809.

[0320] Specifically, if the distributed MIF completes the model retraining in step S808, the training result feedback information may include: a model training completion indication, which is used to indicate that the model training is completed. If the distributed MIF fails to complete the model training in step S808, the training result feedback information may include: a failure reason value, which may be used to indicate at least one of the following: training timeout, insufficient training samples, or model performance does not meet the requirements.

[0321] The centralized MIF can perform different processing after receiving the failure reason value.

[0322] For example, if the reason for failure is insufficient computing power, the centralized MIF can reduce the number of model layers that require distributed training when subsequently allocating training tasks.

[0323] In an embodiment of the present application, the centralized MIF can obtain the model training capabilities of the distributed MIF, thereby more accurately allocating joint training tasks to meet the user's model training needs.

[0324] Fig. 9 Another model training method provided in the embodiment of the present application is provided. Method 900 can be applied to Figure 1 In the system architecture shown in (c) and (f), method 900 may be a specific introduction to steps S501 to S505 of method 500, and method 900 may include the following steps.

[0325] S901, the centralized MIF sends sub-model information to the distributed MIF1 and MIF2.

[0326] Herein, step S901 may include two sub-steps, namely, S901a and S901b.

[0327] The centralized MIF may be the first functional entity in the method 500, and the multiple second functional entities in the method 500 may include MIF1 and MIF2.

[0328] Exemplarily, the above-mentioned sub-model information may include: information on the model structure, and the information on the model structure may include at least one of the following: the number of model layers, the model input dimension, or the model output dimension. The above-mentioned sub-model information may correspond to the first model in method 500.

[0329] Optionally, the above sub-model information may include a model identifier, which may be an integer (eg, 0001) or a character string (eg, model A).

[0330] Optionally, the sub-model information sent by the centralized MIF to the distributed MIF1 and MIF2 may be the same or different.

[0331] S902, the centralized MIF sends model training capability request information to the distributed MIF1 and MIF2.

[0332] Wherein, step S902 may include two sub-steps, namely S902a and S902b, and the above-mentioned model training capability request information may be used to request MIF1 and MIF2 to report their own model training capabilities.

[0333] Optionally, the above-mentioned model training capability request information may include at least one of the following: a single-round training time estimation indication, a computing power margin indication, a memory margin indication, a video memory margin indication, or a video memory bandwidth indication.

[0334] The computing power margin indication can be used to request reporting of computing power margin, and the indication can be for a certain data type, for example, the indication is used to request computing power margin for FP32, or the indication is used to request computing power margin for INT16. The units of the above "computing power" and "computing power margin" can be: FLOPS or TOPS, etc.

[0335] The memory remaining indication may be used to indicate a request to report the memory remaining. The units of the above “memory” and “memory remaining” may be GB and the like.

[0336] The video memory remaining indication can be used to request reporting of the video memory remaining. The units of the above-mentioned "video memory" and "video memory remaining" can be: GB and so on.

[0337] The memory bandwidth indication may be used to request reporting of the memory bandwidth. The unit of the “memory bandwidth” may be terabyte per second (TB / s), etc.

[0338] Optionally, the content included in the model training capability request information sent by the centralized MIF to the distributed MIF1 and MIF2 may be the same or different.

[0339] For example, the model training capability request information sent by the centralized MIF to the distributed MIF1 and MIF2 includes: memory margin, and the values ​​of the memory margin are the same.

[0340] For another example, the model training capability request information sent by the centralized MIF to the distributed MIF1 and MIF2 both includes: memory margin, and the values ​​of the memory margin are different.

[0341] S903, distributed MIF1 and MIF2 send a model training capability report to the centralized MIF.

[0342] Wherein, step S903 may include two sub-steps, namely, S903a and S903b.

[0343] Among them, the above-mentioned model training capability report can be used to indicate the model training capability information of each distributed MIF, that is, the model training capability report can be the model training capability information in method 500.

[0344] The model training information of each distributed MIF may include the same or different contents.

[0345] Optionally, the above-mentioned model training capability may include at least one of the following: an estimated value of a single-round training time, computing power margin, memory margin, video memory margin or video memory bandwidth.

[0346] Among them, the computing power training time estimate can be used to indicate the time required for each MIF to perform a single round of training. The computing power training time estimate can avoid the situation where the training time of some distributed MIFs is too long, resulting in a single round of training time being too long, and then causing the overall model training time to time out.

[0347] The computing power margin can be used to indicate the current computing power margin of the distributed MIF or the computing power that the distributed MIF can currently provide. The unit can be: FLOPS or TOPS, etc. The computing power margin can be for a given data type, for example, the computing power margin for FP32, or the computing power margin for INT16, etc.

[0348] The memory margin may be used to indicate the memory size currently available for the distributed MIF, and the unit may be GB or the like.

[0349] The video memory margin may be used to indicate the current video memory margin of the distributed MIF, or the video memory size that the distributed MIF can currently provide, and the unit may be GB or the like.

[0350] The memory bandwidth may be used to indicate the memory bandwidth currently available for the distributed MIF, and the unit may be TB / s or the like.

[0351] S904, the centralized MIF sends training configuration information to the distributed MIF1 and MIF2.

[0352] Herein, step S904 may include two sub-steps, namely, S904a and S904b.

[0353] Specifically, the configuration information may be used to indicate the training requirements for model training of the distributed MIF, that is, the training configuration information may be the model training requirement information in method 500. The configuration information sent by the centralized MIF to the distributed MIF1 and the distributed MIF2 may include the same or different contents.

[0354] Optionally, the training configuration information may include at least one of the following: the maximum single-round training time, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or a model performance threshold.

[0355] Among them, the maximum value of a single round of training time can be used to indicate the maximum time taken by the distributed MIF to complete a single round of training; the amount of data used for model training can be used to indicate the amount of data used by the distributed MIF for model training; the maximum value of the time taken for model training can be used to indicate the maximum amount of data used by the distributed MIF for model training; the minimum value of the amount of data used for model training can be used to indicate the minimum value of the amount of data used by the distributed MIF for model training; and the model performance threshold can be used to indicate the threshold setting of the performance of the model trained by the distributed MIF.

[0356] Optionally, the model performance threshold may include at least one of the following: maximum and / or minimum value of accuracy, maximum and / or minimum value of precision, maximum and / or minimum value of recall, maximum and / or minimum value of F1 score, maximum and / or minimum value of recall, maximum and / or minimum value of mean absolute error, maximum and / or minimum value of mean square error, maximum and / or minimum value of root mean square error, maximum and / or minimum value of mean absolute percentage error, maximum and / or minimum value of cross entropy.

[0357] S905, distributed MIF1 and MIF2 perform sub-model training respectively.

[0358] Wherein, step S905 may include two sub-steps, namely, S905a and S905b.

[0359] Optionally, when the distributed MIF1 and MIF2 perform model training, they may comply with the indicators determined in step S904.

[0360] S906, the distributed MIF1 and MIF2 send local training parameters to the centralized MIF, and the local training parameters may correspond to the first model training parameters in method 500.

[0361] Herein, step S906 may include two sub-steps, namely, S906a and S906b.

[0362] Optionally, the local training parameters may also be referred to as local training results, which may include gradient values. The content and values ​​of the local training parameters sent by MIF1 and MIF2 to the centralized MIF may be the same or different.

[0363] S907, the centralized MIF sends global training parameters to the distributed MIF1 and MIF2.

[0364] Wherein, step S907 may include two sub-steps, namely, S907a and S907b.

[0365] Specifically, the centralized MIF updates the global training parameters according to the local training parameters sent by MIF1 and MIF2, and sends the global training parameters to MIF1 and MIF2.

[0366] S908, looping through steps S901 to S908 until the model converges.

[0367] It should be understood that in method 900, distributed MIFs are MIF1 and MIF2 as examples, and method 900 is also applicable to the case where the distributed MIF is multiple MIFs. When the distributed MIF is multiple MIFs, the situation of interaction between the centralized MIF and multiple distributed MIFs is basically the same as the description of each step in method 900.

[0368] In an embodiment of the present application, the centralized MIF can coordinate the model training time of each round when the distributed MIF performs model training, thereby ensuring that the model training of the distributed MIF converges quickly to meet the user's model training needs.

[0369] The embodiment of the present application also provides a device for implementing any of the above methods, and the device includes a unit corresponding to executing each step in implementing any of the above methods.

[0370] Fig.10 1 is a schematic diagram of a model training device 1000 provided in an embodiment of the present application, and the device 1000 may include a transceiver unit 1010, a storage unit 1020, and a processing unit 1030. The transceiver unit 1010 is used to receive or send instructions and / or data, and the transceiver unit 1010 may also be called a communication interface or a communication unit; the storage unit 1020 is used to implement the corresponding storage function and store the corresponding instructions and / or data; the processing unit 1030 is used to perform data processing, so that the device 1000 implements the aforementioned model training method.

[0371] In a possible implementation, the apparatus 1000 may include only the transceiver unit 1010 and the processing unit 1030 , but not the storage unit 1020 .

[0372] Optionally, the processing unit 1030 may be located at Figure 2 In the model management function module.

[0373] As a design, device 1000 can execute the actions performed by the first functional entity in the above method embodiment.

[0374] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the transceiver unit 1010 is used to: receive model training requirement information sent by the second functional entity, the model training requirement information includes at least one of the following information: the number of layers of model training, the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; the processing unit 1030 is used to perform model training according to the model training requirement information.

[0375] In a possible implementation, the transceiver unit 1010 is further used to send first information to the second functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0376] In a possible implementation, the first information is used to indicate trained model information, and the first information includes: a first model.

[0377] In a possible implementation, the transceiver unit 1010 is specifically used to receive a first data set sent by a second functional entity; the processing unit 1030 is used to train the second model according to the first data set and model training requirement information to obtain the first model.

[0378] In one possible implementation, the first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0379] In a possible implementation, the transceiver unit 1010 is further used to send model training requirement information confirmation information to the second functional entity, where the model training requirement confirmation information is used to instruct the first functional entity to confirm the model training requirement information.

[0380] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the processing unit 1030 is used to: determine model training requirement information, the model training requirement information includes at least one of the following information: model identification, model structure information, model training layer number information, maximum time used for model training, amount of data used for model training, maximum amount of data used for model training, minimum amount of data used for model training or performance threshold value for model training; the transceiver unit 1010 is used to send the model training requirement information to the second functional entity.

[0381] In a possible implementation, the transceiver unit 1010 is further used to receive first information sent by the second functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0382] In a possible implementation, the first information is used to indicate trained model information, and the first information includes a first model.

[0383] In one possible implementation, the transceiver unit 1010 is also used to receive a first data set sent by a second functional entity; the processing unit 1030 is also used to perform model training based on the first data set to obtain a second model, where the first model is obtained by training the second model based on model training requirement information; the transceiver unit 1010 is also used to send the second model to the second functional entity.

[0384] In one possible implementation, the first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0385] In one possible implementation, the transceiver unit 1010 is also used to receive model training capability information sent by the second functional entity, where the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin or video memory bandwidth; the processing unit 1030 is specifically used to determine the model training requirement information based on the model training capability information.

[0386] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the transceiver unit 1010 is used to: send model training requirement information to multiple second functional entities, the model training requirement information includes at least one of the following: the maximum time of a single round of model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or a performance threshold value for model training; receive multiple first model training parameters sent by multiple second functional entities, the multiple first model training parameters are obtained based on multiple first model trainings; the processing unit 1030 is used to determine global model training parameters based on the multiple first model training parameters; the transceiver unit 1010 is also used to send global model training parameters to multiple second functional entities.

[0387] In one possible implementation, the transceiver unit 1010 is also used to receive multiple model training capability information sent by multiple second functional entities, and the multiple model training capability information is used to indicate at least one of the following contents corresponding to the multiple second functional entities: single-round training time estimate, computing power margin, memory margin, video memory margin or video memory bandwidth; the processing unit 1030 is also used to send model training requirement information to the multiple second functional entities based on the multiple model training capability information.

[0388] In a possible implementation, the transceiver unit 1010 is further configured to send multiple first models to multiple second functional entities.

[0389] As a design, the device 1000 can execute the actions performed by the second functional entity in the above method embodiment.

[0390] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the processing unit 1030 is used to determine model training requirement information, and the model training requirement information includes at least one of the following information: the number of layers of model training, the maximum time used for model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; the transceiver unit 1010 is used to send the model training requirement information to the first functional entity.

[0391] In a possible implementation, the transceiver unit 1010 is further used to receive first information sent by the first functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0392] In one possible implementation, the first information is used to indicate the trained model information, and the first information includes a first model; the processing unit 1030 is specifically used to: obtain a first data set; send the first data set to a first functional entity; train the first model according to the first data set to obtain a third model.

[0393] In one possible implementation, the first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0394] In a possible implementation, the transceiver unit 1010 is further used to receive model training requirement information confirmation information sent by the first functional entity, where the model training requirement confirmation information is used to instruct the first functional entity to confirm the model training requirement information.

[0395] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the transceiver unit 1010 is used to receive model training requirement information sent by the first functional entity, and the model training requirement information includes at least one of the following information: model identification, model structure information, model training layer number information, maximum time used for model training, amount of data used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value for model training; the processing unit 1030 is used to perform model training according to the model training requirement information.

[0396] In a possible implementation, the transceiver unit 1010 is further used to send first information to the first functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

[0397] In a possible implementation, the first information is used to indicate trained model information, and the first information includes: a first model.

[0398] In one possible implementation, the processing unit 1030 is also used to obtain a first data set; the transceiver unit 1010 is also used to: send the first data set to the first functional entity; receive a second model sent by the first functional entity, where the second model is trained based on the first data set; the processing unit 1030 is also used to train the second model according to the first data set and model training requirement information to obtain the first model.

[0399] In one possible implementation, the first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

[0400] In one possible implementation, the transceiver unit 1010 is also used to send model training capability information to the first functional entity, and the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin or video memory bandwidth.

[0401] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the transceiver unit 1010 is used to receive model training requirement information sent by the first functional entity, the model training requirement information includes at least one of the following: the maximum time of a single round of model training, the amount of data used for model training, the maximum amount of data used for model training, the minimum amount of data used for model training, or the performance threshold value of model training; the processing unit 1030 is used to train the first model according to the model training requirement information to obtain first model training parameters; the transceiver unit 1010 is also used to send the first model training parameters to the first functional entity.

[0402] In one possible implementation, the transceiver unit 1010 is also used to send model training capability information to the first functional entity, and the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: single-round training time estimate, computing power margin, memory margin, video memory margin or video memory bandwidth.

[0403] In a possible implementation, the transceiver unit 1010 is further configured to receive a first model sent by the first functional entity.

[0404] Fig.11 It is a schematic diagram of another model training device 1100 provided in an embodiment of the present application.

[0405] The device 1100 includes: a memory 1110, a processor 1120, and a communication interface 1130. The memory 1110, the processor 1120, and the communication interface 1130 are connected through an internal connection path, the memory 1110 is used to store instructions, and the processor 1120 is used to execute the instructions stored in the memory 1110 to control the communication interface 1130 to obtain information, or to enable the device 1100 to implement the aforementioned model training method. Optionally, the memory 1110 can be coupled to the processor 1120 through an interface, or can be integrated with the processor 1120.

[0406] It should be noted that the communication interface 1130 uses a transceiver device such as, but not limited to, a transceiver. The communication interface 1130 may also include an input / output interface.

[0407] The processor 1120 stores one or more computer programs, which include instructions. When the instructions are executed by the processor 1120, the device 1100 executes the model training method in each of the above embodiments.

[0408] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1120 or an instruction in the form of software. The method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in the processor for execution. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1110, and the processor 1120 reads the information in the memory 1110 and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.

[0409] In a possible implementation, the apparatus 1100 may include only the processor 1120 and the communication interface 1130 , but not the memory 1110 .

[0410] Optionally, Fig.11 The communication interface 1130 in the embodiment can be implemented Fig.10 The transceiver unit 1010 in Fig.11 The processor 1120 in the embodiment can implement Fig.10 The processing unit 1030 in.

[0411] The embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program code, and when the computer program code is executed on a computer, the computer executes the above Figures 3 to 9 Any of the methods in .

[0412] The present application also provides a computer program product, which includes a computer program. When the computer program is executed, the computer executes the above Figures 3 to 9 Any of the methods in .

[0413] The present application also provides a chip, including: a circuit, the circuit is used to execute the above Figures 3 to 9 Any of the methods in .

[0414] The embodiment of the present application also provides a system, comprising: a first functional entity and a second functional entity, wherein the first functional entity is used to execute Figures 3 to 9 The actions / steps performed by the first functional entity or the centralized MIF; the second functional entity is used to perform Figures 3 to 9 The actions / steps performed by the second functional entity, distributed MIF or access network equipment.

[0415] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0416] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0417] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0418] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0419] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0420] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0421] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A model training method, characterized in that: The method is applied to a first functional entity, and the method comprises: Receive model training requirement information sent by the second functional entity, where the model training requirement information includes at least one of the following information: number of layers of model training, maximum time used for model training, amount of data used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value of model training; Model training is performed according to the model training requirement information.

2. The method according to claim 1, characterized in that The method further comprises: Sending first information to the second functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

3. The method according to claim 2, characterized in that The first information is used to indicate trained model information, and the first information includes: a first model.

4. The method according to claim 3, characterized in that The performing model training according to the model training requirement information includes: receiving a first data set sent by the second functional entity; A second model is trained according to the first data set and the model training requirement information to obtain the first model.

5. The method according to claim 2, characterized in that The first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

6. The method according to any one of claims 1 to 5, characterized in that: Before performing model training according to the model training requirement information, the method further includes: Model training requirement information confirmation information is sent to the second functional entity, where the model training requirement confirmation information is used to instruct the first functional entity to confirm the model training requirement information.

7. A model training method, characterized in that: The method is applied to a second functional entity, and the method comprises: Determine model training requirement information, where the model training requirement information includes at least one of the following information: number of layers of model training, maximum time used for model training, amount of data used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value of model training; The model training requirement information is sent to the first functional entity.

8. The method according to claim 7, characterized in that The method further comprises: Receive first information sent by the first functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

9. The method according to claim 8, characterized in that The first information is used to indicate trained model information, and the first information includes a first model; Before receiving the first information sent by the first functional entity, the method further includes: Obtaining a first data set; Sending the first data set to the first functional entity; The method further comprises: The first model is trained according to the first data set to obtain a third model.

10. The method according to claim 8, characterized in that The first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

11. The method according to any one of claims 7 to 10, characterized in that The method further comprises: Receive model training requirement information confirmation information sent by the first functional entity, where the model training requirement confirmation information is used to instruct the first functional entity to confirm the model training requirement information.

12. A model training method, characterized in that: The method is applied to a first functional entity, and the method comprises: Determine model training requirement information, where the model training requirement information includes at least one of the following information: model identification, model structure information, model training layer number information, maximum time used for model training, amount of data used for model training, maximum amount of data used for model training, minimum amount of data used for model training, or performance threshold value for model training; The model training requirement information is sent to the second functional entity.

13. The method according to claim 12, characterized in that The method further comprises: Receive first information sent by the second functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

14. The method according to claim 13, characterized in that The first information is used to indicate trained model information, and the first information includes a first model.

15. The method according to claim 14, characterized in that Before receiving the first information sent by the second functional entity, the method further includes: receiving a first data set sent by the second functional entity; Performing model training according to the first data set to obtain a second model, wherein the first model is obtained by training the second model according to the model training requirement information; The second model is sent to the second functional entity.

16. The method according to claim 13, characterized in that The first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

17. The method according to any one of claims 12 to 16, characterized in that Before determining the model training requirement information, the method further includes: Receive model training capability information sent by the second functional entity, where the model training capability information is used to indicate at least one of the following contents corresponding to the second functional entity: computing power margin, memory margin, video memory margin, or video memory bandwidth; The determining of model training requirement information includes: The model training requirement information is determined according to the model training capability information.

18. A model training method, characterized in that: The method is applied to a second functional entity, and the method comprises: Receive model training requirement information sent by the first functional entity, where the model training requirement information includes at least one of the following information: a model identifier, structural information of the model, information on the number of layers of model training, a maximum time used for model training, an amount of data used for model training, a maximum amount of data used for model training, a minimum amount of data used for model training, or a performance threshold value for model training; Model training is performed according to the model training requirement information.

19. The method according to claim 18, characterized in that The method further comprises: Sending first information to the first functional entity, where the first information is used to indicate trained model information, or the first information is used to indicate model training failure.

20. The method of claim 19, wherein: The first information is used to indicate trained model information, and the first information includes: a first model.

21. The method of claim 20, wherein: The performing model training according to the model training requirement information includes: Obtaining a first data set; Sending the first data set to the first functional entity; receiving a second model sent by the first functional entity, where the second model is trained based on the first data set; A second model is trained according to the first data set and the model training requirement information to obtain the first model.

22. The method of claim 19, wherein: The first information is used to indicate that the model training has failed. The first information includes a failure reason value, and the failure reason value is used to indicate at least one of the following: insufficient computing power, insufficient training samples, or the performance of the trained model does not meet the requirements.

23. The method according to any one of claims 18 to 22, characterized in that Before receiving the model training requirement information, the method includes: Model training capability information is sent to the first functional entity, where the model training capability information is used to indicate at least one of the following items corresponding to the second functional entity: computing power margin, memory margin, video memory margin, or video memory bandwidth.

24. A model training method, characterized in that: The method is applied to a first functional entity, and the method comprises: Sending model training requirement information to multiple second functional entities, where the model training requirement information includes at least one of the following: a maximum time for a single round of model training, an amount of data used for model training, a maximum amount of data used for model training, a minimum amount of data used for model training, or a performance threshold value for model training; Receiving a plurality of first model training parameters sent by the plurality of second functional entities, where the plurality of first model training parameters are obtained based on the plurality of first model trainings; Determining a global model training parameter according to the multiple first model training parameters; The global model training parameters are sent to the multiple second functional entities.

25. The method of claim 24, wherein: Before sending the model training requirement information to the multiple second functional entities, the method further includes: Receive multiple model training capability information sent by the multiple second functional entities, where the multiple model training capability information is used to indicate at least one of the following contents corresponding to the multiple second functional entities: single-round training time estimate, computing power margin, memory margin, video memory margin, or video memory bandwidth; The sending of model training requirement information to the plurality of second functional entities includes: Model training requirement information is sent to the multiple second functional entities based on the multiple model training capability information.

26. The method according to claim 24 or 25, characterized in that Before sending the model training requirement information to the plurality of second functional entities, the method includes: The plurality of first models are sent to the plurality of second functional entities.

27. A model training method, characterized in that: The method is applied to a second functional entity, and the method comprises: Receive model training requirement information sent by the first functional entity, where the model training requirement information includes at least one of the following: a maximum time for a single round of model training, an amount of data used for model training, a maximum amount of data used for model training, a minimum amount of data used for model training, or a performance threshold value for model training; Train the first model according to the model training requirement information to obtain first model training parameters; Send the first model training parameters to the first functional entity.

28. The method of claim 27, wherein: The method further comprises: Model training capability information is sent to the first functional entity, where the model training capability information is used to indicate at least one of the following items corresponding to the second functional entity: estimated single-round training time, computing power margin, memory margin, video memory margin, or video memory bandwidth.

29. The method according to claim 27 or 28, characterized in that Before receiving the model training requirement information sent by the first functional entity, the method further includes: Receive the first model sent by the first functional entity.

30. A model training device, characterized in that: The method comprises modules or units for executing the method according to any one of claims 1 to 29.

31. A model training device, characterized in that: include: A processor and a memory, wherein the processor is coupled to the memory and is configured to read and execute instructions in the memory to perform the method according to any one of claims 1 to 29.

32. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program code, and when the computer program code is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 29.

33. A chip, characterized in that: include: A circuit for executing the method according to any one of claims 1 to 29.

34. A system, characterized in that: include: A first functional entity and a second functional entity, wherein the first functional entity is used to execute the method as described in any one of claims 1 to 6, 12 to 17, and 24 to 26; and the second functional entity is used to execute the method as described in any one of claims 7 to 11, 18 to 23, and 27 to 29.

Citation Information

Cited By

  • Model training method and device

    EP4797643A1

  • Model training method and device

    WO2025103214A1