Model training method and device

By transmitting the trained model to the second device in machine learning model training, instead of transmitting training data, the problem of large communication resource overhead in large model training is solved, and efficient update of the second model is achieved.

CN120186040APending Publication Date: 2025-06-20HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311762388.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the process of machine learning model training, directly applying the current model training method leads to a large overhead of communication resource, especially in large-scale data support required in large-scale model training.

Method used

Instead of training data, the trained model is sent to the second device through the first device, which generates an update data set or performs model fusion according to the received model to update the second model.

Benefits of technology

The communication overhead on training data collection during model training is reduced, communication efficiency is improved, and dynamic updates to the second model are realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186040A_ABST
    Figure CN120186040A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and a model training device.The method comprises the steps that first equipment sends first information to second equipment, the first information is used for indicating a first model, the first model is obtained through training of the first equipment, and the first model is used for generating a training data set used for updating a second model, or the first information is used for indicating the second model; the first model is used for being fused with the second model to update the second model, and the second model is obtained by training the second equipment; and the first equipment receives second information, wherein the second information is used for indicating the updated second model. According to the method, the first device sends the first model instead of the training data to the second device, so that communication overhead caused by collection of the training data about the second model between the first device and the second device can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and more particularly, to a model training method and apparatus. Background Art

[0002] Artificial intelligence (AI) is a technology that mimics human cognitive, learning, and reasoning abilities. As a key technology of AI, the machine learning (ML) model has received extensive attention in the research of communication technologies due to its efficient data mining and processing capabilities. Among them, in machine learning training (MLT), usually, the MLT network service (Management Services, MnS) user sends training data to the MLT MnS provider for selection during training. However, as a new AI paradigm, large models require massive data support during the MLT process. If the current model training method is directly applied, it will result in a large communication resource overhead. Summary of the Invention

[0003] This application provides a model training method and apparatus, which can reduce the communication overhead for training data collection during model training.

[0004] In a first aspect, a model training method is provided. This method can be executed by a first device. Here, the first device can refer to the access network device itself, or a processor, module, chip, or chip system in the first device that implements this method. This application does not make any limitations in this regard. The method includes:

[0005] The first device sends first information to a second device. The first information is used to indicate a first model, which is trained by the first device and is used to generate a training data set for updating a second model, or the first model is used to fuse with the second model to update the second model. The second model is trained by the second device; the first device receives second information from the second device. The second information is used to indicate the updated second model.

[0006] By way of example and not limitation, the first device can be an MLT user (such as a communication foundation large model), and the second device can be an MLT provider. The first model can be a data generation model, or the first model can be a model running locally on the first device. For example, when the first device is a base station, the first model can be an on-site model. The second model can be an AI / ML model that implements one or more communication functions in a physical network, or is called a communication foundation large model.

[0007] It should be understood that embodiments of the present application do not limit the specific manner in which the first device obtains the second model.

[0008] By way of example and not limitation, the second model may be pre-configured in the first device. Specifically, the second model may be configured in the first device when the first device leaves the factory, or the second model may be manually configured on the first device before the first device runs.

[0009] By way of example and not limitation, the second model may be sent by the second device to the first device. Specifically, before the first device sends the first information to the second device, the second device sends configuration information to the first device, and the configuration information is used to indicate the second model.

[0010] In a possible implementation, the second device may also send at least one of the performance information of the second model, the evaluation index information of the second model, or the function information of the second model to the first device, so that the first device can determine whether the second model meets the corresponding performance requirements during operation. Among them, the performance information of the second model may include information such as the number of parameters, structure, and number of layers of the second model; the evaluation index information of the second model may include the performance performance of the second model in its available downstream tasks, such as accuracy rate, recall rate, mean square error, etc., and the function information of the second model may include information about the communication function that can be achieved through the model.

[0011] It should be understood that embodiments of the present application do not limit the specific manner in which the first information indicates the first model.

[0012] By way of example and not limitation, the first model is included in the first information, and the second device can obtain the first model after receiving the first information.

[0013] By way of example and not limitation, the storage address of the first model may be included in the first model. After receiving the first information, the second device can obtain the first model according to the storage address of the first model.

[0014] Optionally, when the first device sends the first information to the second device, the first device may also send at least one of the performance information of the first model, the evaluation index information of the first model, or the function information of the first model to the second device.

[0015] Based on the above solution, the first device sends the first model to the second device, enabling the second device to generate a training dataset for updating the second model based on the first model and update the second model based on the training dataset, or enabling the second device to fuse the first model with the second model to update the first model. That is, by sending the first model instead of the training data, the first device can save the communication overhead caused by collecting the training data of the second model between the first device and the second device.

[0016] In combination with the first aspect, in some implementation manners of the first aspect, the method includes: the first device sends third information to the second device, and the third information is used to indicate the usage mode of the first model. The usage mode of the first model includes generating the training dataset based on the first model, or performing model fusion based on the first model.

[0017] By way of example and not limitation, the third information can indicate the usage method of the first model through bit 0 and bit 1. For example, bit 0 is used to indicate that the first model is used to generate a training dataset for updating the second model, and bit 1 is used to indicate that the first model is used to fuse with the second model.

[0018] Based on the above solution, the first device can indicate the usage method of the first model to the second device through the third information, saving the recognition or trial-and-error time of the second device for the usage method of the first model, thereby improving the usage efficiency of the second device for the first model.

[0019] In combination with the first aspect, in some implementation manners of the first aspect, when the first model is used to generate a training dataset for updating the second model, the method further includes: the first device performs model training based on all or part of the local data to generate the first model. The all or part of the local data includes a first dataset, and the first dataset includes data related to a first function, and the first function is one or more functions that the first device expects the second model to implement.

[0020] By way of example and not limitation, the first function may include functions such as beam management and channel measurement required by the first device, or function management such as intention management expected by the first device, or mobility prediction expected by the first device.

[0021] Based on the above solution, the first device can perform model training according to all or part of the local data to generate the above first model, where the all or part of the local data includes data related to the first function, thereby reducing the overhead of the first device for model training and improving the accuracy of model training.

[0022] In combination with the first aspect, in some implementations of the first aspect, when the first model is used to generate a training dataset for updating the second model, the method further includes: the first device fine-tunes the parameters of the second model based on the second dataset according to the first function, where the second dataset includes data related to the first function, and the second dataset is included in the first dataset; determining that the performance of the fine-tuned second model related to at least one function in the first function does not meet the corresponding performance metric requirements.

[0023] It should be understood that the performance of the second model related to at least one function in the first function not meeting the corresponding performance metric requirements can be understood as that among the various performances of the second model related to the first function, at least one performance does not meet the corresponding performance metric requirements.

[0024] It should be understood that the second dataset being included in the first dataset can be understood as that the time period for collecting the second dataset is included in the time period for collecting the first dataset.

[0025] As an example but not a limitation, if the first device is a base station, then the first device can fine-tune the parameters of the second model based on the second dataset according to functions such as required beam management and channel measurement.

[0026] As an example but not a limitation, if the first device is an OAM, then the first device can fine-tune the parameters of the second model based on the second dataset according to functions such as intent management.

[0027] As an example but not a limitation, if the first device is a NWDAF, then the first device can fine-tune the parameters of the second model based on the second dataset according to functions such as mobility prediction.

[0028] It should be understood that the performance metric can be a Key Performance Indicator (KPI) requirement, and the KPI requirement is determined by the first device during the process of fine-tuning the second model.

[0029] It should be understood that when the performance of the fine-tuned second model meets the performance metric requirements corresponding to the first function, the first device can continuously fine-tune the parameters of the second model according to all or local data and the first function, thereby improving the performance of the model deployed locally by the first device.

[0030] Based on the above solution, the first device can fine-tune the second model using all or part of the local data according to the first function, and when the performance of the fine-tuned second model does not meet the performance metric requirements corresponding to the first function, perform model training on the local data to generate the first model, thereby ensuring the availability of the model deployed locally by the first device.

[0031] In combination with the first aspect, in some implementations of the first aspect, when the first model is used to be fused with the second model to update the second model, the method further includes: the first device fine-tunes the parameters of the second model based on all or part of the local data according to the first function, so as to generate the first model, where the first function is one or more functions that the first device expects to achieve through the second model.

[0032] Based on the above solution, the first device can fine-tune the second model using all or part of the local data according to the first function to generate the first model, and send the first model to the second device, so that the second device can fuse the first model and the second model to update the second model, thereby saving the communication overhead caused by the collection of training data of the second model between the first device and the second device and realizing the dynamic update of the second model.

[0033] In combination with the first aspect, in some implementations of the first aspect, the first device is a mobile intelligent network element, and the second device is a device supporting cloud services.

[0034] It should be understood that in this implementation, the training data of the first network element is provided by other devices (such as base stations), and the first device itself is only used for data training.

[0035] In combination with the first aspect, in some implementations of the first aspect, the first device obtains a third data set, where the third data set includes data related to the second function collected by the base station, and the second function includes one or more functions that the base station expects to achieve through the second model; the first device fine-tunes the parameters of the second model based on the third data set according to the second function.

[0036] Based on the above solution, the first device can be a device without computing power (such as a base station) to fine-tune the parameters of the second model according to the second function, thereby improving the performance of the model deployed on the device without computing power.

[0037] In combination with the first aspect, in some implementations of the first aspect, when the first model is used to generate a training data set for updating the second model, before the first device sends the first information to the second device, the method further includes: the first device receives a fourth piece of information, where the fourth piece of information is used to indicate that the performance of the fine-tuned second model is abnormal; the first device determines a fourth data set according to the fourth piece of information, where the fourth data set is stored locally by the first device, or the fourth data set is obtained by the first device through the base station; the first device performs model training based on the fourth data set to generate the first model.

[0038] It should be understood that the fourth data set is a collection of data related to the second function described above. Moreover, the fourth data set includes the third data set, that is, the time period for collecting the fourth data set includes the time period for collecting the third data set.

[0039] It should be understood that the fourth data set for the first device to perform model training may be obtained by the first device requesting the base station when it determines that it is necessary to generate the first model, or the fourth data set for the first device to perform model training may be the data reported by the base station and stored locally by the first device. The embodiments of the present application do not make any limitations in this regard.

[0040] Combined with the first aspect, in some implementation manners of the first aspect, when the first model is used to generate a training data set for updating the second model, the method further includes: the first device sends fifth information to the base station according to the fourth information, and the fifth information is used to request data collected by the base station; the first device receives the fifth information, and the fifth information is used to indicate the fourth data set.

[0041] It should be understood that the fourth information is used to request all the data collected by the base station, or the fourth information is used to request the data related to the second function collected by the base station.

[0042] It should be understood that all the data collected by the foregoing base station may be all the data collected after the base station is put into operation, or all the data collected by the foregoing base station may be all the data collected during the time interval between receiving two request messages, or all the data collected by the foregoing base station may be all the data collected during a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not make any limitations in this regard.

[0043] It should be understood that the data related to the second function collected by the foregoing base station may be all the data related to the second function collected after the base station is put into operation, or the data related to the second function collected by the foregoing base station may be the data related to the second function collected during the time interval between receiving two request messages, or the data related to the second function collected by the foregoing base station may be the data related to the second function collected during a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not make any limitations in this regard.

[0044] By way of example and not limitation, the fourth information may include at least one of identifier #1 and identifier #2, wherein the identifier #1 is used to indicate a specific time period, so that each base station can report the data collected during the specific time period based on the identifier #1, and the identifier #2 is used to indicate a specific function (such as the second function), so that each base station can report the data related to the specific function based on the identifier #2.

[0045] It should be understood that the first device may discard some or all of the data provided by the base station after model training, or the first device may save some or all of the data provided by the base station after model training. The embodiments of the present application do not limit this.

[0046] Based on the above solution, when the first device reports that the model locally deployed by the base station fails, the first device can perform model training based on the data collected by the base station to generate a first model, and send the first model to the second device, so that the second device can dynamically update the second model based on the data collected by the base station, thereby saving the communication overhead brought by the collection of training data of the second model between the first device and the second device while ensuring the model performance of the base station operation.

[0047] Combined with the first aspect, in some implementation manners of the first aspect, before the first device sends the first information to the second device, the method further includes: the first device performs lightweight processing on the first model.

[0048] Based on the above solution, the first device can perform lightweight processing on the first model after generating the first model, thereby further reducing the communication overhead brought by the collection of training data of the second model between the first device and the second device.

[0049] In a second aspect, a model training method is provided. This method can be executed by a second device. Here, the second device may refer to the access network device itself, or may refer to a processor, module, chip, or chip system in the second device that implements this method. The present application does not limit this. The method includes:

[0050] The second device receives first information from the first device. The first information is used to indicate a first model. The first model is trained by the first device. The first model is used to generate a training data set for updating the second model, or the first model is used to fuse with the second model to update the second model. The second model is trained by the second device; the second device updates the second model according to the first model; the second device sends second information to the first device. The second information is used to indicate the updated second model.

[0051] By way of example and not limitation, the first device may be a user of a network management service (such as a communication foundation large model), and the second device may be a provider of the network management service. The first model may be a data generation model, or the first model may be a model running locally on the first device. For example, when the first device is a base station, the first model may be an on-site model. The second model may be an AI / ML model that implements one or more communication functions in a physical network, or is referred to as a communication foundation large model. It should be understood that for the specific manner in which the first device obtains the second model and the specific manner in which the first information indicates the first model, reference may be made to the relevant content in the first aspect, which will not be elaborated here.

[0052] Based on the above solution, the first device sends the first model to the second device, enabling the second device to generate a training data set for updating the second model according to the first model and update the second model based on the training data set, or enabling the second device to fuse the first model with the second model to update the first model. That is, by sending the first model instead of training data, the first device can save the communication overhead caused by collecting the training data of the second model between the first device and the second device.

[0053] Combined with the second aspect, in some implementation manners of the second aspect, the method further includes: the second device receives third information from the first device, and the third information is used to indicate the usage mode of the first model. The usage mode of the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.

[0054] It should be understood that the second device can determine the usage method of the first model according to the third information, and then update the second model according to the first model. For the specific description of the type of the first model, reference may be made to the relevant content in the fifth aspect, which will not be elaborated here.

[0055] Based on the above solution, the first device can indicate the usage method of the first model to the second device through the third information, saving the recognition or trial-and-error time of the second device for the usage method of the first model, thereby improving the usage efficiency of the second device for the first model.

[0056] Combined with the second aspect, in some implementation manners of the second aspect, when the first model is used to generate a training data set for updating the second model, the second device updates the second model according to the first model, including: the second device generates a training data set according to the first model; the second device trains the second model according to the training data set to update the second model.

[0057] In combination with the second aspect, in some implementations of the second aspect, when the first model is used to be fused with the second model to update the second model, the second device updates the second model according to the first model, including: the second device fuses the first model and the second model to update the second model.

[0058] In a third aspect, a model training method is provided. This method can be executed by a first device, where the first device can refer to the access network device itself, or a processor, module, chip, or chip system in the first device that implements this method. This application does not make any limitations in this regard. The method includes:

[0059] The first device sends first information to the second device. The first information is used to indicate a first model, which is trained by the first device. The first model is used to generate a training data set for updating the second model, and the second model is trained by the second device; the first device receives second information, and the second information is used to indicate the updated second model.

[0060] As an example but not a limitation, the first device can be a user of a network management service (such as a communication foundation large model), and the second device can be a provider of the network management service. The first model can be a data generation model, and the second model can be an AI / ML model that implements one or more communication functions in a physical network, or is called a communication foundation large model.

[0061] It should be understood that for the specific manner in which the first device obtains the second model and the specific manner in which the first information indicates the first model, reference can be made to the relevant content of the first aspect, which will not be elaborated here.

[0062] Based on the above solution, the first device sends the first model to the second device, so that the second device can generate a training data set for updating the second model according to the first model, and update the second model based on the training data set. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the first model instead of the training data.

[0063] In combination with the third aspect, in some implementations of the third aspect, before the first device sends the first information to the second device, the method further includes: the first device performs model training based on all or part of the local data to generate the first model. The all or part of the local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the first device expects the second model to implement.

[0064] In combination with the third aspect, in some implementation manners of the third aspect, before the first device performs model training based on all or part of local data, the method further includes: the first device fine-tunes the parameters of the second model based on a second data set according to the first function, where the second data set includes data related to the first function, and the second data set is included in the first data set; determining that the performance of the fine-tuned second model related to at least one function in the first function does not meet the corresponding performance index requirements.

[0065] It should be understood that the performance of the second model related to at least one function in the first function not meeting the corresponding performance index requirements can be understood as that at least one of the performances related to the first function in the second model does not meet the corresponding performance index requirements.

[0066] In combination with the third aspect, in some implementation manners of the third aspect, the first device is a mobile intelligent network element, and the second device is a device supporting cloud services.

[0067] In combination with the third aspect, in some implementation manners of the third aspect, the method further includes: the first device obtains a third data set, where the third data set includes data related to a second function collected by a base station, and the second function includes one or more functions that the base station expects to implement through the second model; the first device fine-tunes the parameters of the second model based on the third data set according to the second function.

[0068] In combination with the third aspect, in some implementation manners of the third aspect, before the first device sends the first information to the second device, the method further includes: the first device receives a fourth information, where the fourth information is used to indicate that the performance of the fine-tuned second model is abnormal; the first device determines a fourth data set according to the fourth information, where the fourth data set is locally stored by the first device, or the fourth data set is obtained by the first device through the base station, and the fourth data set includes the third data set; the first device performs model training based on the fourth data set to generate the first model.

[0069] In combination with the third aspect, in some implementation manners of the third aspect, the method further includes: the first device sends a fifth information to the base station according to the fourth information, where the fifth information is used to request data collected by the base station; the first device receives a sixth information, where the sixth information is used to indicate the fourth data set.

[0070] In combination with the third aspect, in some implementation manners of the third aspect, before the first device sends the first information to the second device, the method further includes: the first device performs lightweight processing on the first model.

[0071] Fourthly, a model training method is provided. This method can be executed by a second device. Here, the second device can refer to the access network device itself, or the processor, module, chip, or chip system in the second device that implements this method. This application does not make any limitations in this regard. The method includes:

[0072] The second device receives first information, which is used to indicate a first model. The first model is trained by the first device and is used to generate a training data set for updating the second model. The second model is trained by the second device; the second device updates the second model according to the first model; the second device sends second information to the first device, and the second information is used to indicate the updated second model.

[0073] By way of example and not limitation, the first device can be a user of a network management service (such as a communication foundation large model), and the second device can be a provider of the network management service. The first model can be a data generation model, and the second model can be an AI / ML model that implements one or more communication functions in a physical network, or is called a communication foundation large model.

[0074] Based on the above solution, the first device sends the first model to the second device, so that the second device can generate a training data set for updating the second model according to the first model, and update the second model based on the training data set. That is, the first device can save the communication overhead brought by the collection of training data for the second model between the first device and the second device by sending the first model instead of the training data.

[0075] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the second device updates the second model according to the first model, including: the second device generates a training data set according to the first model; the second device trains the second model according to the training data set to update the second model.

[0076] Fifthly, a model training method is provided. This method can be executed by a first device. Here, the first device can refer to the access network device itself, or the processor, module, chip, or chip system in the first device that implements this method. This application does not make any limitations in this regard. The method includes:

[0077] The first device sends first information to the second device, and the first information is used to indicate a first model. The first model is trained by the first device and is used to fuse with a second model to update the second model. The second model is trained by the second device; the first device receives second information, and the second information is used to indicate the updated second model.

[0078] By way of example and not limitation, the first device may be a user of a network management service (such as a communication foundation large model), the second device may be a provider of the network management service, the first model may be a model running locally on the first device. For example, when the first device is a base station, the first model may be an on-site model, and the second model may be an AI / ML model that implements one or more communication functions in a physical network, or is referred to as a communication foundation large model.

[0079] It should be understood that for the specific manner in which the first device obtains the second model and the specific manner in which the first information indicates the first model, reference may be made to the relevant content of the first aspect, which will not be elaborated here.

[0080] Based on the above solution, the first device sends the first model to the second device, enabling the second device to fuse the first model with the second model to update the first model. That is, by sending the second model instead of training data, the first device can save the communication overhead caused by collecting the training data of the second model between the first device and the second device.

[0081] Combined with the fifth aspect, in some implementation manners of the fifth aspect, the method further includes: the first device fine-tunes the parameters of the second model based on all or part of the local data according to the first function to generate the first model, where the first function is one or more functions that the first device expects to implement through the second model.

[0082] Combined with the fifth aspect, in some implementation manners of the fifth aspect, the all or local data at least includes a second data set, and the second data set is a set of data related to the first function.

[0083] In a sixth aspect, a model training method is provided. This method may be executed by a second device, where the second device may refer to the access network device itself or a processor, module, chip, or chip system in the second device that implements this method. This application does not make a limitation in this regard. The method includes:

[0084] The second device receives first information from the first device. The first information is used to indicate a first model, which is trained by the first device and is used to fuse with a second model to update the second model, and the second model is trained by the first device; the second device updates the second model according to the first model; the second device sends second information to the first device, and the second information is used to indicate the updated second model.

[0085] By way of example and not limitation, the first device may be a user of a network management service (such as a communication foundation large model), the second device may be a provider of the network management service, the first model may be a model running locally on the first device. For example, when the first device is a base station, the first model may be an on-site model, and the second model may be an AI / ML model that implements one or more communication functions in a physical network, or is referred to as a communication foundation large model.

[0086] Based on the above solution, the first device sends the first model to the second device, enabling the second device to fuse the first model with the second model to update the first model. That is, by sending the second model instead of training data, the first device can save the communication overhead caused by collecting the training data of the second model between the first device and the second device.

[0087] Combined with the sixth aspect, in some implementation manners of the sixth aspect, the second device updates the second model according to the first model, including: the second device fuses the first model and the second model to update the second model.

[0088] In a seventh aspect, a model training device is provided. The device includes: a transceiver unit, configured to receive first information for indicating a first model, where the first model is trained by the first device and is used to fuse with a second model to update the second model, and the second model is trained by the first device; a processing unit, configured to update the second model according to the first model; and the transceiver unit is further configured to send second information to the first device, where the second information is used to indicate the updated second model.

[0089] Combined with the seventh aspect, in some implementation manners of the seventh aspect, the transceiver unit is further configured to send third information to the second device, where the third information is used to indicate the usage mode of the first model, and the usage mode of the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.

[0090] Combined with the seventh aspect, in some implementation manners of the seventh aspect, the model training device further includes a processing unit, configured to perform model training based on all or part of local data to generate the first model, where the all or part of local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the model training device expects to implement through the second model.

[0091] In combination with the seventh aspect, in certain implementations of the seventh aspect, the processing unit is further configured to: based on the first function, fine-tune the parameters of the second model according to a second data set, where the second data set includes data related to the first function, and the second data set is included in the first data set; determine that the performance of the fine-tuned second model related to at least one function in the first function does not meet the corresponding performance metric requirements.

[0092] In combination with the seventh aspect, in certain implementations of the seventh aspect, the processing unit is further configured to: based on the first function, fine-tune the parameters of the second model according to all or part of the local data to generate the first model, where the first function is one or more functions that the first device expects to implement through the second model.

[0093] In combination with the seventh aspect, in certain implementations of the seventh aspect, the model training device is a mobile intelligent network element, and the second device is a device supporting cloud services.

[0094] In combination with the seventh aspect, in certain implementations of the seventh aspect, the transceiver unit is further configured to receive fourth information, where the fourth information is used to indicate that the performance of the fine-tuned second model is abnormal; the processing unit is further configured to determine a fourth data set according to the fourth information, where the fourth data set is locally stored by the first device, or the fourth data set is obtained by the first device through the base station, and the fourth data set includes the third data set; the processing unit is further configured to perform model training based on the fourth data set to generate the first model.

[0095] In combination with the seventh aspect, in certain implementations of the seventh aspect, the processing unit is further configured to send fifth information to the base station according to the fourth information, where the fifth information is used to request data collected by the base station; the transceiver unit is further configured to receive fifth information, where the fifth information is used to indicate the fourth data set.

[0096] In combination with the seventh aspect, in certain implementations of the seventh aspect, before the first device sends the first information to the second device, the processing unit is further configured to perform lightweight processing on the first model.

[0097] In an eighth aspect, there is provided a model training device, including: a transceiver unit, configured to receive first information, where the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to generate a training data set for updating the second model, or the first model is used to fuse with the second model to update the second model, and the second model is trained by the second device; a processing unit, configured to update the second model according to the first model; the transceiver unit is further configured to send second information to the first device, where the second information is used to indicate the updated second model.

[0098] In combination with the eighth aspect, in some implementations of the eighth aspect, the transceiver unit is further configured to receive third information from the first device, where the third information is used to indicate the usage mode of the first model, and the usage mode of the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.

[0099] In combination with the eighth aspect, in some implementations of the eighth aspect, the model training device further includes a processing unit, and the processing unit is configured to: generate a training data set according to the first model; train the second model according to the training data set to update the second model.

[0100] In combination with the eighth aspect, in some implementations of the eighth aspect, the processing unit is further configured to fuse the first model and the second model to update the second model.

[0101] A ninth aspect provides a model training device, including: a transceiver unit, configured to send first information to a second device, where the first information is used to indicate a first model, the first model is trained by the first device, and the first model is used to generate a training data set for updating a second model, and the second model is trained by the second device; the transceiver unit is further configured to receive second information, where the second information is used to indicate the updated second model.

[0102] It should be understood that the ninth aspect is the implementation on the device side corresponding to the third aspect. The descriptions of the supplements, explanations, and beneficial effects of the third aspect also apply to the ninth aspect and will not be elaborated here.

[0103] A tenth aspect provides a model training device, including: a transceiver unit, configured to receive first information, where the first information is used to indicate a first model, the first model is trained by the first device, and the first model is used to generate a training data set for updating the second model, and the second model is trained by the second device; a processing unit, configured to update the second model according to the first model; the transceiver unit is further configured to send second information to the first device, where the second information is used to indicate the updated second model.

[0104] It should be understood that the tenth aspect is the implementation on the device side corresponding to the fourth aspect. The descriptions of the supplements, explanations, and beneficial effects of the fourth aspect also apply to the tenth aspect and will not be elaborated here.

[0105] In an eleventh aspect, a model training apparatus is provided. The apparatus includes a transceiver unit configured to send first information to a second device. The first information is used to indicate a first model that is trained by the first device and is used to fuse with a second model to update the second model, where the second model is trained by the second device. The transceiver unit is further configured to receive second information that is used to indicate the updated second model.

[0106] It should be understood that the eleventh aspect is the implementation on the device side corresponding to the fifth aspect. The supplementary explanations and beneficial effects regarding the fifth aspect also apply to the eleventh aspect and will not be elaborated herein.

[0107] In a twelfth aspect, a model training apparatus is provided. The apparatus includes a transceiver unit configured to receive first information that is used to indicate a first model that is trained by the first device and is used to fuse with a second model to update the second model, where the second model is trained by the second device. A processing unit is configured to update the second model according to the first model. The transceiver unit is further configured to send second information to the first device, where the second information is used to indicate the updated second model.

[0108] It should be understood that the twelfth aspect is the implementation on the device side corresponding to the sixth aspect. The supplementary explanations and beneficial effects regarding the sixth aspect also apply to the twelfth aspect and will not be elaborated herein.

[0109] In a thirteenth aspect, the present application provides a model training apparatus. The model training apparatus includes a processor configured to implement the methods described in the first aspect to the sixth aspect, or any implementation of the first aspect to the sixth aspect. The processor is coupled to a memory, and the memory is used to store instructions and data. When the processor executes the instructions stored in the memory, it can implement the methods described in the first aspect to the sixth aspect, or any implementation of the first aspect to the sixth aspect.

[0110] Optionally, the communication device may further include a memory. Optionally, the memory may be coupled to the processor. Optionally, the communication device may further include a communication interface configured to communicate with other devices. Exemplarily, the communication interface may be a transceiver, a hardware circuit, a bus, a module, a pin, or other types of communication interfaces.

[0111] In one example, the communication device may be the first device, or a device, module, or chip disposed in the first device, or a device that can be used in cooperation with the first device.

[0112] In another example, the communication device may be a second device, or a device, module, or chip disposed in the second device, or a device that can be used in combination with the second device.

[0113] In a fourteenth aspect, the present application provides a system, including: a first device configured to execute the method described in the first aspect, the third aspect, or the fifth aspect, or any implementation manner of the first aspect, the third aspect, or the fifth aspect; a second device configured to execute the method described in the second aspect, the fourth aspect, or the sixth aspect, or any implementation manner of the second aspect, the fourth aspect, or the sixth aspect.

[0114] In a fifteenth aspect, the present application further provides a computer program, which, when running on a computer, causes the computer to execute the method described in any of the first aspect to the sixth aspect, or any implementation manner of the first aspect to the sixth aspect.

[0115] In a sixteenth aspect, the present application further provides a computer program product, including instructions, which, when running on a computer, cause the computer to execute the method described in any of the first aspect to the sixth aspect, or any implementation manner of the first aspect to the sixth aspect.

[0116] In a seventeenth aspect, the present application further provides a computer-readable storage medium, in which a computer program or instructions are stored, which, when running on a computer, cause the computer to execute the method described in any of the first aspect to the sixth aspect, or any implementation manner of the first aspect to the sixth aspect.

[0117] In an eighteenth aspect, the present application further provides a chip, which is configured to read a computer program stored in a memory and execute the method described in any of the first aspect to the sixth aspect, or any implementation manner of the first aspect to the sixth aspect; or, the chip includes means for executing the method described in any of the first aspect to the sixth aspect, or any implementation manner of the first aspect to the sixth aspect.

[0118] In a nineteenth aspect, the present application further provides a chip system, which includes a processor configured to support a device in implementing the method described in any of the first aspect to the sixth aspect, or any implementation manner of the first aspect to the sixth aspect. In a possible design, the chip system further includes a memory configured to store necessary programs and data for the device. The chip system may be composed of chips or may include chips and other discrete devices.

[0119] For the description of the beneficial effects of any of the seventh aspect to the nineteenth aspect, reference may be made to the description of the beneficial effects of the first aspect to the sixth aspect, which will not be elaborated here. Brief Description of the Drawings

[0120] Figure 1 is a schematic diagram of a network architecture;

[0121] Figure 2 is a schematic diagram of the current MLT process;

[0122] Figure 3 is a schematic diagram of a model training method 300 provided by an embodiment of the present application;

[0123] Figure 4 is a schematic diagram of a model training process 400 provided by an embodiment of the present application;

[0124] Figure 5 is a schematic diagram of a model training process 500 provided by an embodiment of the present application;

[0125] Figure 6 is a schematic diagram of a model training process 600 provided by an embodiment of the present application;

[0126] Figure 7 is a schematic diagram of a model training device 1000 provided by an embodiment of the present application;

[0127] Figure 8 is a schematic diagram of another model training device 1100 provided by an embodiment of the present application. Detailed Description of the Embodiments

[0128] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.

[0129] For ease of understanding, first, a communication system to which the embodiments of the present application can be applied will be described.

[0130] Embodiments of the present application can be applied to various communication systems. For example: Long Term Evolution (LTE) systems, LTE Frequency Division Duplex (FDD) systems, LTE Time Division Duplex (TDD), Public Land Mobile Network (PLMN), 5th generation (5G) systems, 6th generation (6G) systems, or future communication systems, etc. The 5G systems in the present application include non-standalone (NSA) 5G mobile communication systems or standalone (SA) 5G mobile communication systems. Embodiments of the present application can also be applied to non-terrestrial network (NTN) communication systems such as satellite communication systems. Embodiments of the present application can also be applied to device-to-device (D2D) communication systems, sidelink (SL) communication systems, machine-to-machine (M2M) communication systems, machine type communication (MTC) systems, Internet of Things (IoT) communication systems, vehicle-to-everything (V2X) communication systems, uncrewed aerial vehicle (UAV) communication systems, or other communication systems.

[0131] As an example, Figure 1 A schematic diagram of a network architecture is shown.

[0132] As Figure 1As shown, this network architecture takes the 5th generation system (5GS) as an example. This network architecture can include three parts, namely the user equipment (UE) part, the data network (DN) part, and the operator network part. Among them, the operator network can include one or more of the following network elements: (radio) access network ((R)AN) equipment, user plane function (UPF) network element, access and mobility management function (AMF) network element, session management function (SMF) network element, network data analytics function (NWDAF) network element, policy control function (PCF) network element, application function (AF) network element, Mobile Intelligence Function (MIF) network element, and network management (Operations, Administration And Management, OAM) network element. In the above operator network, the part other than the RAN part can be called the core network part.

[0133] In this application, the user equipment, (radio) access network equipment, UPF network element, AMF network element, SMF network element, NWDAF network element, PCF network element, AF network element, MIF network element, and OAM network element are respectively abbreviated as UE, (R)AN, UPF, AMF, SMF, NWDAF, PCF, AF, MIF, and OAM.

[0134] The following Figure 1 briefly describes each network element involved.

[0135] 1. UE

[0136] The UE in this application can also be called a terminal, user, access terminal, user unit, user station, mobile station, mobile platform, remote station, remote terminal, mobile device, user terminal, terminal device, wireless communication device, user agent, or user device, etc. For the sake of convenient description, it will be uniformly called a terminal hereinafter.

[0137] A terminal is a device that can access the network. The terminal and the (R)AN can communicate with each other using a certain air interface technology (such as NR or LTE technology). Terminals can also communicate with each other using a certain air interface technology (such as NR or LTE technology). The terminal can be a mobile phone, a pad, a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a terminal in satellite communication, a terminal in an integrated access and backhaul (IAB) system, a terminal in a WiFi communication system, a terminal in industrial control, a terminal in self-driving, a terminal in remote medical, a terminal in smart grid, a terminal in transportation safety, a terminal in smart city, a terminal in smart home, etc.

[0138] Embodiments of this application do not limit the specific technologies and specific device forms adopted by the UE.

[0139] 2. (R)AN

[0140] The (R)AN in this application can be a device for communicating with the terminal or a device for connecting the terminal to the wireless network.

[0141] (R)AN can be a node in a radio access network. (R)AN can be a base station, an evolved NodeB (eNodeB), a transmission reception point (TRP), a home base station (e.g., home evolved NodeB, or home Node B, HNB), a Wi-Fi access point (AP), a mobile switching center, a next generation NodeB (gNB) in a 5G mobile communication system, an access network device in an open radio access network (O-RAN or open RAN), a next generation NodeB in a 6th generation (6G) mobile communication system, or a base station in a future mobile communication system, etc. The network device can also be a module or unit that completes some functions of the base station. For example, it can be a central unit (CU), a distributed unit (DU), a remote radio unit (RRU), or a baseband unit (BBU), etc. (R)AN can also be a device that undertakes the function of the base station in a D2D communication system, a V2X communication system, an M2M communication system, and an IoT communication system, etc. (R)AN can also be a network device in NTN, that is, (R)AN can be deployed on a high-altitude platform or a satellite. (R)AN can be a macro base station, a micro base station or an indoor station, and can also be a relay node or a donor node, etc.

[0142] Embodiments of this application do not limit the specific technologies, device forms, and names adopted by (R)AN.

[0143] 3. UPF

[0144] The main functions of UPF are to route and forward data packets, serve as a mobility anchor, an uplink classifier to support routing service flows to a data network, and a branching point to support multi-homed PDU sessions, etc.

[0145] 4. DN

[0146] DN is mainly used for the operator network that provides data services for terminals. For example, the Internet, a third-party service network, or an IP multimedia service (IMS) network, etc.

[0147] 5. AMF

[0148] The main functions of the AMF include managing user registration, reachability detection, selection of SMF nodes, mobility state transition management, etc.

[0149] 6. SMF

[0150] The main functions of the SMF are to control the establishment, modification, and deletion of sessions, selection of user plane nodes, etc.

[0151] 7. NWDAF

[0152] It has functions such as data collection, model training, data analysis, model inference, etc. It can be used to collect relevant data from network network elements, third-party service servers, terminal devices, or network management systems, perform data analysis or model training based on the relevant data, and provide data analysis results to network network elements, third-party service servers, terminal devices, or network management systems, or provide the trained model to other data analysis function network elements.

[0153] 8. PCF

[0154] The PCF is mainly responsible for policy control decision-making, providing policy rules for control plane functions, and traffic-based charging control functions, etc.

[0155] 9. AF

[0156] The AF mainly supports interacting with the 3GPP core network to provide services, such as influencing data routing decisions, policy control functions, or providing third-party services to the network. The AF can be an AF deployed by the operator's network itself or a third-party AF. The DCAF is a special AF, mainly responsible for collecting data from UE applications and opening it to network elements such as NWDAF in the network.

[0157] 10. OAM

[0158] It mainly completes the daily analysis, prediction, planning, and configuration of the network and services, as well as the testing and fault management of the network and its services, etc. The OAM can interact with the RAN to obtain UE location information measured by the RAN or reported by the UE measurement on the RAN side.

[0159] 11. MIF

[0160] Responsible for the artificial intelligence (AI) or machine learning (ML) functions of the base station, including data management functions, computing power management functions, and model management functions. Among them, the data management functions can include: data collection, data storage, and data analysis; the computing power management functions can include: computing power perception, computing power scheduling, and computing power and transmission coordination; the model management functions can include: model training, model inference, and model life cycle management. The form of the MIF can be: a base station, a network element independent of the base station, a network function independent of the base station, a sub-module inside the base station, or a sub-function inside the base station. For example, in Figure 1 the MIF is a network element / network function independent of the base station.

[0161] In Figure 1 In the network architecture shown, the network elements can communicate through interfaces. The interfaces between the network elements can be point-to-point interfaces or service-based interfaces, which are not limited in this application.

[0162] It should be understood that the network architecture shown above is only an exemplary illustration, and the network architecture applicable to the embodiments of this application is not limited thereto. Any network architecture that can implement the functions of the above-mentioned network elements is applicable to the embodiments of this application.

[0163] It should also be understood that Figure 1 the functions or network elements such as UE, (R)AN, UPF, AMF, SMF, NWDAF, PCF, AF, MIF, OAM shown in

[0164] It should also be understood that the above names are only defined for the convenience of distinguishing different functions and should not constitute any limitation to this application. This application does not exclude the possibility of using other names in 6G networks and future other networks. For example, in 6G networks, some or all of the above-mentioned network elements may continue to use the terms in 5G, or other names may be used.

[0165] For the convenience of understanding the embodiments of this application, several basic concepts involved in the embodiments of this application are briefly described.

[0166] 1. Machine Learning (ML)

[0167] As a key technology of artificial intelligence (AI), machine learning (ML) models have gained wide attention in communication technology research due to their efficient data mining and processing capabilities. The training process of ML models in communication networks is one of the important research directions of the 3rd Generation Partnership Project (3GPP) R18 standard. By efficiently training ML models, their performance in actual communication applications can be improved, and it is also conducive to the subsequent deployment, management and large-scale application of ML models. Among them, SA5TS28.105 clarifies the overall process framework and data requirements for ML model training. Among them, machine learning training (MLTraining, MLT) network service (Management Services, MnS) providers use current and historical relevant data to monitor networks or services related to ML models, prepare data, trigger and execute training.

[0168] 2. Large Model

[0169] A large model is an artificial neural network model with at least 100 million parameters, which needs to be optimized and trained by computers and massive existing data. It has strong reasoning and generalization capabilities. As a new ML paradigm, large models have achieved performance breakthroughs in multiple application fields including natural language processing (NLP). Large models can achieve unprecedented reasoning and generalization performance by utilizing huge parameter scale, massive data and computing resources, and model structures that can capture global data relationships. Only one pre-trained model is needed to adapt to a wide range of application tasks and achieve optimal performance through fine-tuning or small / zero-sample learning. Given that 3GPP R19 standard research has begun to focus on natural language intent, large models may serve as a potential technical solution for natural language intent translation. In addition, relying on its powerful reasoning ability, large models have great application potential in scenarios such as physical layer wireless resource allocation, network management plane operation optimization, and core network intelligent control in communication networks.

[0170] 3. Fine-tuning

[0171] Refers to the process of training a pre-trained artificial neural network model for a specific task or tasks using the specific information of the task. For example, in order to enable an artificial neural network model to have the ability of machine translation, it is necessary to use different language comparison tables to train the pre-trained model for the next step.

[0172] 4. Data Generation Model

[0173] An artificial neural network model with the ability to fit data distribution. After training, for a specific input, it can obtain more output results that conform to this distribution. For example, in the application of the generative adversarial network model in text generation tasks, by using a large amount of natural language to train the model, given any white noise input, it can output a set of natural language output results similar to the training corpus.

[0174] Figure 2 The following is a schematic diagram of the current MLT process. Considering the insufficient computing power deployment on the user side of the current MLT MnS, and the convenience of computing power deployment on the provider side of the MLT MnS (such as OAM), the model training is usually completed by the MLT MnS provider. As Figure 2 shown, the MLT MnS user sends a training request to the MLT MnS provider to initiate the ML entity (model) training process, and at the same time, the MLT MnS user provides historical or existing available training data for the MLT MnS provider to select. The MLT MnS provider responds to the training request of the MLT MnS user, selects the training data for model training, and sends the training result to the MLT MnS user after completion for the latter to select and use.

[0175] As can be seen from the above, during the ML model training process, the MLT MnS user needs to send training data to the MLT MnS provider for the MLT MnS provider to select during training. However, during the training of large models, a huge amount of data is required. If the current ML model training method is directly applied, the data collection process will cause a large communication resource overhead between network elements (especially between the MLT MnS user and the MLT MnS provider). For example, in the reliable wireless transmission considering high time-sensitivity scenarios, when sampling user data with a period of 10 ms for 100,000 users in a wireless network, assuming the required number of user features is 100 and each feature requires 4 bits to represent, then the amount of user data required to be uploaded per minute in this network is about 4 GB, resulting in a large communication resource overhead.

[0176] In view of this, the present application proposes a model training method and device, which can avoid the MLT MnS provider directly loading the huge amount of training data of the MLT MnS user during the large model training process, thereby saving the communication overhead brought by the data collection process.

[0177] To facilitate the understanding of the embodiments of the present application, the following points are explained before introducing the embodiments of the present application.

[0178] In this application, "for indicating" or "indicating" may include direct indication and indirect indication, or in other words, "for indicating" or "indicating" may explicitly and / or implicitly indicate. For example, when describing that a certain piece of information is for indicating information I, it may include that this information directly indicates I or indirectly indicates I, and it does not necessarily mean that I is carried in this information. For another example, implicit indication may be based on the position and / or resources for transmission; explicit indication may be based on one or more parameters, and / or one or more indexes, and / or one or more bit patterns it represents.

[0179] The definitions listed for many features in this application are only for explaining the functions of these features by way of example, and the detailed content can be referred to the prior art.

[0180] In the embodiments shown below, the first, second, third, fourth, and various numerical numbers are only for the convenience of description for distinction, and are not used to limit the scope of the embodiments of this application. For example, to distinguish different fields, different information, etc.

[0181] "Pre-defined" can be implemented by pre-saving corresponding codes, tables, or other ways that can be used to indicate relevant information in the device. This application does not limit its specific implementation method. Among them, "saving" may mean saving in one or more memories. The type of memory can be any form of storage medium, and this application does not limit this.

[0182] The "protocol" involved in the embodiments of this application may refer to the standard protocols in the communication field. For example, it may include the long term evolution (LTE) protocol, the new radio (NR) protocol, and the relevant protocols applied to future communication systems. This application does not limit this.

[0183] This application will present various aspects, embodiments, or features around a system including multiple devices, components, modules, etc. It should be understood and clear that each system may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the drawings. In addition, combinations of these solutions can also be used.

[0184] In the embodiments of this application, words such as "exemplary", "for example", "exemplarily", "as (another) example", etc. are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" in this application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "exemplary" is intended to present concepts in a specific way.

[0185] The terms "comprising", "including", "having" and their variants mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0186] "At least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one (item) of a, b, and c can represent: a, or, b, or, c, or, a and b, or, a and c, or, b and c, or, a, b, and c. Where a, b, and c can be single or multiple respectively.

[0187] In the embodiments of the present application, the related descriptions of the network element A sending messages, information or data to the network element B, and the network element B receiving messages, information or data from the network element A are intended to illustrate which network element the message, information or data is to be sent to, and do not limit whether they are directly sent or indirectly sent via other network elements.

[0188] In the embodiments of the present application, descriptions such as "when...", "in the case of...", "if", and "when" all refer to the device making corresponding processing under a certain objective situation, not limited to time, and do not require the device to have a judgment action when implemented, nor does it mean there are other limitations.

[0189] It should be understood that in the embodiments of the present application, the base station and the access network device can be the same concept, and the two can be used interchangeably.

[0190] Figure 3 It is a schematic diagram of a model training method 300 provided by the embodiments of the present application. As shown in the figure, the method 300 includes the following steps:

[0191] S310, the first device sends the first information to the second device. Correspondingly, the second device receives the first information. Wherein, the first information is used to indicate the first model, and the first model is used to generate a training data set for updating the second model.

[0192] In the embodiments of the present application, the first device can be a user of a network management service (such as a communication foundation large model), the second device can be a network management service provider, the first model can be a data generation model, and the second model can be an AI / ML model that realizes one or more communication functions in a physical network, or is called a communication foundation large model.

[0193] It should be understood that the embodiments of the present application do not limit the specific manner in which the first device obtains the second model.

[0194] As an example rather than a limitation, the second model may be pre-configured in the first device. Specifically, the second model may be configured in the first device when the first device leaves the factory, or the second model may be manually configured on the first device before the first device runs.

[0195] As an example rather than a limitation, the second model may be sent from the second device to the first device. Specifically, before the above step S310, the second device sends configuration information to the first device, and the configuration information is used to indicate the second model.

[0196] Optionally, the second device may further send at least one of the performance information of the second model, the evaluation index information of the second model, or the function information of the second model to the first device, so that the first device can determine whether the second model meets the corresponding performance requirements during operation. Among them, the performance information of the second model may include information such as the number of parameters, structure, and number of layers of the second model; the evaluation index information of the second model may include the performance performance of the second model in its available downstream tasks, such as accuracy rate, recall rate, mean square error, etc., and the function information of the second model may include information on communication functions that can be achieved through the model.

[0197] It should be understood that the embodiments of the present application do not limit the specific manner in which the first information indicates the first model.

[0198] As an example rather than a limitation, if the first model is included in the first information, the second device can obtain the first model after receiving the first information.

[0199] As an example rather than a limitation, the first model may include the storage address of the first model. After receiving the first information, the second device can obtain the first model according to the storage address of the first model.

[0200] Optionally, when the first device sends the first information to the second device, the first device may further send at least one of the performance information of the first model, the evaluation index information of the first model, or the function information of the first model to the second device.

[0201] In some possible implementation manners, the first device may send third information to the second device, and the third information is used to indicate the usage method of the first model. The usage method of the first model includes generating a training data set based on the first model, or performing model fusion based on the first model.

[0202] It should be understood that the embodiments of the present application do not limit the specific manner in which the third information indicates the usage method of the first model.

[0203] By way of example and not limitation, the third information may indicate the usage method of the first model through bit 0 and bit 1. For example, bit 0 is used to indicate generating a training data set based on the first model, and bit 1 is used to indicate model fusion based on the first model.

[0204] It should be understood that the first device may send the third information while sending the above-mentioned first information, for example, carrying the first information and the third information in the same message, or the first device may send the third information after sending the above-mentioned first information. The embodiments of the present application do not limit this.

[0205] It is easy to understand that, distinguished by the usage method of the first model, before the above-mentioned step S310, the above method 300 further includes the following steps:

[0206] Method 1, generating a training data set based on the first model:

[0207] S305, the first device performs model training based on all or part of the local data to generate the first model.

[0208] Among them, all or part of the data of the first device at least includes a first data set, and the first data set is a set of data related to the first function, and the first function includes one or more functions that the first device expects to implement through the second model.

[0209] By way of example and not limitation, the first function may include functions such as beam management and channel measurement required by the first device, or function management such as intention management expected by the first device, or mobility prediction expected by the first device.

[0210] It should be understood that the embodiments of the present application do not limit the triggering conditions for the first device to execute step S305.

[0211] In a possible implementation manner, the first device periodically performs model training according to the local data. Among them, the embodiments of the present application do not limit the period for the first device to perform model training. Exemplarily, the period may be 500 s.

[0212] In another possible implementation manner, before the above-mentioned step S305, the above method 300 further includes the following steps (not shown in the figure):

[0213] S301, the first device fine-tunes the parameters of the second model according to the first function based on the second data set. Among them, the second data set is included in the first data set.

[0214] It should be understood that the second data set includes data related to the above-mentioned first function, and the second data set is included in the first data set. It can be understood that the time period for collecting the second data set is included in the time period for collecting the first data set.

[0215] As an example but not a limitation, if the first device is a base station, the first device can fine-tune the parameters of the second model based on the second data set according to functions such as required beam management and channel measurement.

[0216] As an example but not a limitation, if the first device is an OAM, the first device can fine-tune the parameters of the second model based on the second data set according to functions such as intention management.

[0217] As an example but not a limitation, if the first device is a NWDAF, the first device can fine-tune the parameters of the second model based on the second data set according to functions such as mobility prediction.

[0218] S302, the first device determines whether the performance of the fine-tuned second model meets the requirements of the corresponding key performance indicator (KPI) for the above-mentioned first function.

[0219] It should be understood that the above KPI requirements are determined by the first device during the process of fine-tuning the second model.

[0220] In a possible implementation manner, if at least one performance of the first model fine-tuned through the above step S302 does not meet the corresponding KPI requirements, the first device executes the above step S305.

[0221] In another possible implementation manner, if the performance of the first model fine-tuned through the above step S302 meets the corresponding KPI requirements, the first device sequentially repeats the above steps S301 and S302. In this way, the first device can fine-tune the parameters of the first model according to the first function and local data, and improve the performance of the model running locally on the first device.

[0222] Method 2, model fusion based on the first model:

[0223] S307, the first device fine-tunes the parameters of the second model based on the first function and all or part of the local data to generate the first model.

[0224] In a possible implementation manner, the second model for which the first device performs parameter fine-tuning in step S307 can be the model obtained by the first device performing lightweight processing on the second model according to the above first function.

[0225] It should be understood that all or part of the local data used by the first device includes data related to the first function.

[0226] It should be understood that all or part of the local data may be collected and stored locally by the first device itself, or may be stored locally after being obtained from other devices. The embodiments of the present application do not limit this.

[0227] Based on the above method 1 or method 2, the first device may generate a first model and send the first model through step S310.

[0228] S320. The second device updates the first model according to the second model.

[0229] In a possible implementation manner, when the third information indicates that the first model is used to generate a training data set, the second device generates a training data set for updating the first model according to the second model, and performs model training on the second model according to the training data set, so as to update the second model.

[0230] In another possible implementation manner, when the third information indicates that the first model is used to be fused with the second model, the second device fuses the second model and the first model to update the second model. The specific manner of fusing the second model and the first model will be described in detail below and will not be elaborated here.

[0231] S330. The second device sends second information to the first device. Correspondingly, the first device receives the second information. The second information is used to indicate the updated second model.

[0232] It should be understood that the embodiments of the present application do not limit the specific manner in which the second information indicates the updated second model.

[0233] As an example rather than a limitation, if the updated second model is included in the second information, the first device may obtain the updated second model after receiving the first information.

[0234] As an example rather than a limitation, the storage address of the updated second model may be included in the second information. After receiving the second information, the first device may obtain the first model according to the storage address of the updated second model.

[0235] Optionally, the first device may also send at least one of the performance information of the second model, the evaluation index information of the second model, or the function information of the second model to the second device.

[0236] It is easily understandable that the first device in the above steps S310 to S330 can collect data and perform model training based on the collected data. In some possible embodiments, the first device is a MIF, that is, the training data of the first network element is provided by other devices (such as a base station), and the first device itself is only used for data training.

[0237] In these embodiments, the above method 300 may further include the following steps (not shown in the figure):

[0238] S301’: The first device fine-tunes the parameters of the second model based on the third data set according to the second function.

[0239] It should be understood that the fourth data set is a collection of data related to the second function, and the second function includes one or more functions that the base station expects to achieve through the second model.

[0240] It should be understood that the above fourth data set is requested and obtained by the first device from the base station, that is, before executing the above step S301’, the first device needs to request the data collected by the base station from the base station.

[0241] S302’: The first device sends the fine-tuned second model to the base station. Correspondingly, the base station receives the fine-tuned second model.

[0242] S303’: The first device receives the fourth information from the base station, and the fourth information is used to indicate that the performance of the fine-tuned second model is abnormal.

[0243] In these embodiments, the above step S305 can be replaced by:

[0244] S305’: The first device performs model training based on all or part of the data collected by the base station to generate the first model.

[0245] Among them, all or part of the data collected by the base station at least includes the fourth data set, and the fourth data set is a collection of data related to the above second function. And, the fourth data set includes the third data set, that is, the time period for collecting the fourth data set includes the time period for collecting the third data set.

[0246] It should be understood that the data for the first device to perform model training can be requested and obtained by the first device from the base station when the first device determines that it is necessary to generate the first model, or the data for the first device to perform model training can be the data reported by the base station stored locally by the first device. The embodiments of the present application do not make any limitations in this regard.

[0247] In some possible embodiments, before the first device executes S305’, the first device needs to request all or part of the data collected by the base station from the base station.

[0248] The embodiments of the present application do not limit the specific manner in which the first device requests all or part of the data collected by the base station from the base station.

[0249] As an example rather than a limitation, the first device periodically sends fourth information to the base station, where the fourth information is used to request all the data collected by the base station, or the fourth information is used to request the data related to the second function collected by the base station.

[0250] It should be understood that all the data collected by the aforementioned base station may be all the data collected after the base station is put into operation, or all the data collected by the aforementioned base station may be all the data collected by the base station during the interval between receiving two request messages, or all the data collected by the aforementioned base station may be all the data collected by the base station within a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not limit this.

[0251] It should be understood that the data related to the second function collected by the aforementioned base station may be all the data related to the second function collected after the base station is put into operation, or the data related to the second function collected by the aforementioned base station may be the data related to the second function collected by the base station during the interval between receiving two request messages, or the data related to the second function collected by the aforementioned base station may be the data related to the second function collected by the base station within a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not limit this.

[0252] It should be understood that the first device may discard some or all of the data provided by the base station after model training, or the first device may save some or all of the data provided by the base station after model training. The embodiments of the present application do not limit this.

[0253] It should be understood that the embodiments of the present application do not limit the triggering condition for the first device to execute step S305'.

[0254] In a possible implementation manner, the triggering condition for step S305' is receiving the aforementioned third information.

[0255] Further, after the first device receives the third information, the first device may perform model training based on the data collected by the base station stored locally to generate a first model, or the first device may send fourth information to the base station to request the data collected by the base station and perform model training based on the data to generate a first model.

[0256] In another possible implementation manner, the first device periodically performs model training based on local data. Among them, the embodiments of the present application do not limit the period for the first device to perform model training. Exemplarily, the period may be 500s.

[0257] Based on the above solution, the first device sends a first model to the second device, enabling the second device to generate a training data set for updating the second model according to the first model, and updating the first model based on the training data set. That is, by sending the second model instead of the training data, the first device can save the communication overhead caused by collecting the training data of the second model between the first device and the second device.

[0258] The following is through Figure 4 The scenario where the first device shown below can collect data and train a model is used to illustrate the specific process of the above method 300. In this scenario, the first device can be any one of a base station, OAM, and NWDAF. The following takes the first device as a base station and the second device as a device supporting cloud services (hereinafter referred to as Cloud) as an example for illustration.

[0259] Figure 4 It is a schematic diagram of a model training process 400 provided by an embodiment of the present application. As shown in the figure, this process 400 includes the following steps:

[0260] S401, Cloud sends information #1 to one or more base stations. Correspondingly, the one or more base stations receive the information #1.

[0261] Among them, the information #1 is used to indicate model #1 (an example of the second model). Exemplarily, the model #1 can be a wireless communication basic large model. In the embodiment of the present application, the above step S401 can be understood as Cloud making the wireless communication basic large model available at the base station side.

[0262] It should be understood that the embodiment of the present application does not limit the specific manner in which the information #1 indicates the model #1. For the specific manner, reference can be made to the relevant content of step S310, which will not be elaborated here.

[0263] Optionally, when sending the information #1, Cloud can also send at least one of the model performance information of model #1, the evaluation index information of model #1, and the model function information of model #1 to the one or more base stations. For the specific description of the above information, reference can be made to the relevant content of step S310, which will not be elaborated here.

[0264] S402, each base station fine-tunes the parameters of model #2 using all or part of the locally collected data according to the required functions.

[0265] In a possible implementation manner, the model #2 can be the model determined by the base station through the information #1.

[0266] In another possible implementation, the model #2 can be a model obtained by the base station to lightweight the model #1 according to the required functions (i.e., the first function). Among them, the required functions of the base station can be understood as one or more functions that the base station expects to achieve through the model #1.

[0267] It should be understood that all or part of the above locally collected data at least includes the data set #1, and the data set #1 is a set of data collected by the base station related to its own required functions.

[0268] By way of example and not limitation, the required functions of the model can include functions such as beam management and channel measurement required by the base station.

[0269] In the embodiment of the present application, the model #2 can be an on-site model.

[0270] It should be understood that in the process of fine-tuning the parameters of the model #2 by the base station, the KPI requirements corresponding to the required functions of the base station can be determined.

[0271] S403, each base station determines whether the performance of the fine-tuned model #2 meets the KPI requirements corresponding to the required functions of the above base station.

[0272] In one possible implementation, if the performance of the model #2 meets the KPI requirements of the required functions of the above base station, the base station repeats the above steps S402 and S403. In this way, the base station can fine-tune the parameters of the model #2 according to its own required functions and local data, and improve the performance of the model locally deployed by the base station.

[0273] In another possible implementation, if the performance of the model #2 does not meet at least one of the KPI requirements of the required functions of the base station, the base station executes the following step S404.

[0274] S404, the base station uses all or part of the local data for model training to generate the model #3.

[0275] Among them, all or part of the local data of the base station at least includes the data set #2, and the data set #2 is a set of data collected by the base station related to its own required functions.

[0276] In the embodiment of the present application, the model #3 can be a data generation model.

[0277] S405, the base station sends the information #2 (an example of the first information) and the information #3 (an example of the third information) to the Cloud. Correspondingly, the Cloud receives the information #2 and the information #3.

[0278] Among them, the information #2 is used to indicate the model #3, and the information #3 is used to indicate that the model #3 is used to generate the training data set. In the embodiments of the present application, the above step S405 can be understood as the base station making the model #3 available on the Cloud side.

[0279] In some possible implementation manners, the above information #2 and information #3 may be included in a wireless large model training request message, and the wireless large model training request message is used to request an update of the above model #1.

[0280] Optionally, when sending the information #2, the base station may send at least one of the model performance information of the model #3, the evaluation index information of the model #3, and the model function information of the model #3 to the Cloud.

[0281] Optionally, before the base station sends the information #2, the model #3 generated in step S404 may be lightweight processed, and the lightweight processed model #3 is indicated through the information #2.

[0282] S406, the Cloud updates the model #1 according to the model #3 obtained from one or more base stations.

[0283] Specifically, the Cloud generates a training data set according to the model #3 obtained from one or more base stations, and performs model training on the model #1 according to the training data set to update the model #1.

[0284] S407, the Cloud sends information #4 (an example of the second information) to one or more base stations, and correspondingly, the one or more base stations receive the information #4.

[0285] Among them, the information #4 is used to indicate the updated model #1. In the embodiments of the present application, the above step S407 can be understood as the Cloud making the updated model #1 available on the base station side.

[0286] Optionally, when sending the information #4, the Cloud may also send at least one of the model performance information of the updated model #1, the evaluation index information of the updated model #1, and the model function information of the updated model #1 to the one or more base stations.

[0287] Based on the above solution, the base station / OAM / NWDAF can generate and send the model #3 to the Cloud based on local data, so that the Cloud can generate a training data set for updating the model #1 according to the model #3, and update the model #1 based on the training data set. That is, the base station / OAM / NWDAF can save the communication overhead related to the training data collection of the model #1 between the base station / OAM / NWDAF and the Cloud by sending the model #3 for generating the training data set for updating the model #1 instead of the training data.

[0288] The following uses Figure 5 The scenario where the first device shown below is used to provide computing power for other devices (such as model training) to illustrate the specific process of the above method 300. The following takes the first device as the MIF and the second device as the device supporting cloud services (hereinafter referred to as Cloud) as an example for illustration.

[0289] Figure 5 It is a schematic diagram of a model training process 500 provided by an embodiment of the present application. As shown in the figure, this process 500 includes the following steps:

[0290] S501, Cloud sends information #1' to one or more MIFs. Correspondingly, the one or more MIFs receive information #1'.

[0291] Among them, the information #1' is used to indicate model #1' (another example of the second model). Exemplarily, the model #1' can be a wireless communication basic large model. In the embodiment of the present application, the above step S401 can be understood as Cloud making the wireless communication basic large model available at the MIF side.

[0292] It should be understood that the embodiment of the present application does not limit the specific manner in which the information #1' indicates the model #1'. For the specific manner, reference can be made to the relevant content of step S310, which will not be elaborated here.

[0293] Optionally, when sending the information #1', Cloud can also send at least one of the model performance information of the model #1', the evaluation index information of the model #1', and the model function information of the model #1' to the one or more base stations.

[0294] It should be understood that each of the above MIFs can be used to provide computing power for the corresponding one or more base stations, that is, each MIF can be used to perform model training for the corresponding one or more base stations. The present application does not limit this.

[0295] As an example rather than a limitation, Cloud sends the above information #1' to MIF #1 and MIF #2. Among them, MIF #1 can provide computing power for base station #1 and base station #2, and MIF #2 can provide computing power for base station #3.

[0296] It should be understood that before the above step S501, the model #2' is configured on the base stations corresponding to the one or more MIFs. The model #2' can be the above model #1, or the model #2' can be the model #1' after parameter fine-tuning based on the functions required by the base stations. The present application does not limit the specific manner of configuring the model #2' on the base stations corresponding to the one or more MIFs. For the specific manner, reference can be made to the relevant content of step S310, which will not be elaborated here.

[0297] S502. Each MIF sends information #5 to one or more corresponding base stations. Correspondingly, the one or more base stations receive information #5. The information #5 is used to request data collected by the one or more base stations.

[0298] In a possible implementation, the information #5 includes at least one of identification #1 and identification #2. The identification #1 is used to indicate a specific time period, so that each base station can report the data collected during the specific time period based on the identification #1. The identification #2 is used to indicate a specific function, so that each base station can report the data related to the specific function based on the identification #2.

[0299] S503. Each base station sends all or part of the local data to the corresponding MIF.

[0300] Specifically, each base station samples the local data according to the information #5 and sends the sampled training data to the corresponding MIF. All or part of the local data sent by the base station to the corresponding MIF at least includes dataset #3, and the dataset #3 is data related to the functions required by the base station. The functions required by the base station can be understood as one or more functions that the base station expects to implement through model #1.

[0301] It should be understood that when the information #5 includes the above identification #1, the base station will sample the local data according to the identification #1 and send all or part of the data collected during the specific time period indicated by the identification #1 to the corresponding MIF.

[0302] It should be understood that when the information #5 includes the above identification #2, the base station will sample the local data according to the identification #2 and send all or part of the data related to the specific function indicated by the identification #2 to the corresponding MIF.

[0303] S504. The MIF fine-tunes the parameters of model #1' using all or part of the local data provided by the base station according to the functions required by the base station.

[0304] Specifically, the specific process of the MIF fine-tuning the model #1' can refer to the relevant content of method 300, which will not be elaborated here.

[0305] S505. Each MIF sends information #6 to one or more corresponding base stations, and the information #6 is used to indicate the fine-tuned model #1'.

[0306] Correspondingly, after receiving the fine-tuned model #1', the one or more base stations update the local deployment.

[0307] S506. Each base station determines whether the fine-tuned model #1' meets the KPI requirements corresponding to the required functions.

[0308] In some possible implementation manners, if the performance of the model #1' meets the KPI requirements of the required functions of the base station, the above steps S502 to S506 are repeated. In this way, the MIF can fine-tune the parameters of the model #1' according to the required functions of the base station and the data collected by the base station, and improve the performance of the model locally deployed at the base station.

[0309] In another possible implementation manner, if the performance of the model #1' does not meet at least one KPI requirement of the required functions of the base station, the base station executes the following step S507.

[0310] S507. The base station sends information #7 (an example of the fourth information) to the corresponding MIF. Correspondingly, the corresponding MIF receives the information #7. Among them, the information #6 is used to indicate that the model deployed at the base station is abnormal.

[0311] S508. The MIF uses all or part of the data provided by the corresponding one or more base stations for model training to generate the model #3'.

[0312] Among them, all or part of the data provided by the one or more base stations at least includes the data set #4, and the data set #4 is a set of data related to the required functions of the base station reported by the one or more base stations. And, the data set #4 includes the above data set #3, that is, the time period for collecting the data set #4 includes the time period for collecting the data set #3.

[0313] It should be understood that the embodiments of the present application do not limit the specific manner in which the MIF obtains the data set #4.

[0314] As an example rather than a limitation, after receiving the data reported by the base station through the foregoing step S503, the MIF stores the data locally, so that when executing step S508, the model training is performed based on the locally stored data.

[0315] As an example rather than a limitation, before executing the above step S507, the MIF executes the following steps (not shown in the figure):

[0316] S5081. The MIF sends information #8 (an example of the fifth information) to the corresponding one or more base stations. Correspondingly, the one or more base stations receive the information #8. Among them, the information #8 is used to request the base station to report the collected data.

[0317] Optionally, the information #8 includes at least one of the above identifier #1 or identifier #2, so that the base station can send the corresponding data based on the identifier #1 or identifier #2.

[0318] S5082, each base station reports all or part of its local data to the MIF.

[0319] Optionally, the MIF may also periodically send Information #8 to one or more corresponding base stations to request the base stations to report the collected data.

[0320] Based on the above method, the MIF can obtain the data collected by the base stations and perform model training based on this data to generate Model #3'. Exemplarily, this Model #3' is a data generation model.

[0321] S509, the MIF sends Information #2' (another example of the second information) and Information #3' (another example of the third information) to the Cloud. Correspondingly, the Cloud receives this Information #2' and Information #3'.

[0322] Among them, this Information #2' is used to indicate Model #3', and this Information #3' is used to indicate that the first model is used for fusion with the second model. In the embodiments of the present application, the above step S509 can be understood as making Model #3' available on the Cloud side by the MIF.

[0323] In some possible implementation manners, the above Information #2' and Information #3' may be included in a wireless large model training request message, and this wireless large model training request message is used to request an update of the above Model #1'.

[0324] Optionally, when sending Information #2', the MIF may send at least one of the model performance information of Model #3', the evaluation index information of Model #3', and the model function information of Model #3' to the Cloud.

[0325] Optionally, before the base station sends Information #2', it may perform lightweight processing on the Model #3' generated in step S408 and indicate the lightweight processed Model #3' through Information #2'.

[0326] S510, the Cloud updates Model #1' according to the Model #3' obtained from one or more MIFs.

[0327] Specifically, the Cloud generates a training data set according to the Model #3' obtained from one or more MIFs, and performs model training on Model #1' according to this training data set to update Model #1'.

[0328] S511, the Cloud sends Information #4' to one or more MIFs. Correspondingly, the one or more base stations receive Information #4'.

[0329] Among them, the information #4' is used to indicate the updated model #1'. In the embodiments of the present application, the above step S407 can be understood as Cloud making the updated model #1' available at the MIF side.

[0330] Optionally, when sending the information #4', Cloud can also send at least one of the model performance information of the updated model #1', the evaluation index information of the updated model #1', and the model function information of the updated model #1' to the one or more MIFs.

[0331] Further, the one or more MIFs can send the updated model #1 to the corresponding one or more base stations.

[0332] Optionally, before each MIF sends the updated model #1' to the corresponding one or more base stations, the MIF can also perform lightweight processing on the updated model #1' according to the functions required by the base station, and / or fine-tune the parameters of the updated model #1' using the data collected by the base station according to the functions required by the base station. Among them, the MIF can obtain the data collected by the base station by repeating the above step S502.

[0333] Based on the above solution, the MIF can send the model #3' to Cloud based on the data reported by the base station, so that Cloud can generate a training data set for updating the model #1' according to the model #3', and update the model #1' based on the training data set. That is, the MIF can save the communication overhead related to the collection of the training data of the model #1' between the MIF and Cloud by sending the model #3' for generating the training data set for updating the model #1' instead of the training data.

[0334] The following will be combined with Figure 6 The model training method 600 provided by the embodiments of the present application will be described. In the method 600, the first device can update the second model based on the model locally deployed by the first device by sending the locally deployed model to the second device. In the method 600, the first device can be any one of a base station, OAM, NWDAF, and MIF. Hereinafter, the first device is taken as a base station and the second device is taken as Cloud for illustration.

[0335] Figure 6 It is a schematic diagram of a model training process 600 provided by the embodiments of the present application. As shown in the figure, the process 600 includes the following steps:

[0336] S601, Cloud sends the information #8 to one or more base stations. Correspondingly, the one or more base stations receive the information #8.

[0337] Among them, the information #8 is used to indicate the model #4 (an example of the second model). Exemplarily, the model #4 may be a wireless communication basic large model. In the embodiments of the present application, the above step S601 can be understood as Cloud making the wireless communication basic large model available at the base station side.

[0338] It should be understood that the embodiments of the present application do not limit the specific manner in which the information #8 indicates the model #4. For the specific manner, reference can be made to the relevant content of step S310, which will not be elaborated here.

[0339] Optionally, when sending the information #8, Cloud may also send at least one of the model performance information of the model #4, the evaluation index information of the model #4, and the model function information of the model #4 to the one or more base stations. For the specific description of the above information, reference can be made to the relevant content of step S310, which will not be elaborated here.

[0340] S602, each base station fine-tunes the parameters of the model #5 using all or part of the locally collected data according to the required functions to generate the model #6 (an example of the first model).

[0341] In one possible implementation, the model #5 may be the model determined by the base station through the information #1.

[0342] In another possible implementation, the model #5 may be the model obtained by the base station after lightweight processing of the model #4 according to the required functions (i.e., the first function). Among them, the required functions of the base station can be understood as one or more functions that the base station expects to implement through the model #4.

[0343] It should be understood that the above all or part of the local data at least includes the data set #5, and the data set #5 is a set of data collected by the first device related to its own required functions.

[0344] In the embodiments of the present application, the model #5 may be an on-site model.

[0345] S603, the base station sends the information #9 (an example of the first information) and the information #10 (an example of the third information) to Cloud. Correspondingly, Cloud receives the information #9 and the information #10.

[0346] Among them, the information #9 is used to indicate the model #6, and the information #10 is used to indicate that the model #6 is used for fusion with the model #4. In the embodiments of the present application, the above step S405 can be understood as the base station making the model #6 available at the Cloud side.

[0347] In some possible implementations, the above information #9 and information #10 may be included in a wireless large model training request message, and the wireless large model training request message is used to request an update of the above model #4.

[0348] Optionally, when sending Information #9, the base station may send at least one of the model performance information of Model #6, the evaluation index information of Model #6, and the model function information of Model #6 to the Cloud.

[0349] Optionally, before the base station sends Information #9, the Model #6 generated in step S404 may be lightweight processed, and the lightweight processed Model #6 may be indicated through Information #9.

[0350] S604. The Cloud updates Model #4 according to Model #6 obtained from one or more base stations.

[0351] Specifically, the Cloud fuses the Model #6 obtained from one or more base stations with the Model #4 to update the second model.

[0352] It should be understood that the embodiments of the present application do not limit the fusion method of Model #6 and Model #4.

[0353] As an example rather than a limitation, the Model #6 and Model #4 may perform knowledge distillation, that is, the model parameters of Model #6 and Model #4 are directly used to update the model parameters of Model #4 through calculation methods such as summation and averaging.

[0354] S605. The Cloud sends Information #11 (an example of the second information) to one or more base stations. Correspondingly, the one or more base stations receive Information #11.

[0355] Among them, the Information #11 is used to indicate the updated Model #4. In the embodiments of the present application, the above step S605 may be understood as the Cloud making the updated Model #4 available at the base station side.

[0356] Optionally, when sending Information #11, the Cloud may also send at least one of the model performance information of the updated Model #4, the evaluation index information of the updated Model #4, and the model function information of the updated Model #4 to the one or more base stations.

[0357] Based on the above solution, the base station / OAM / NWDAF / MIF may fine-tune the parameters of Model #4 using all or part of the local data according to its own required functions and generate Model #6, so that the Cloud may fuse the Model #6 with Model #4 to update Model #4, that is, the base station / OAM / NWDAF / MIF may save the communication overhead related to the training data collection of Model #4 between the base station / OAM / NWDAF / MIF and the Cloud by sending Model #6 for fusing with Model #4 instead of the training data to the Cloud.

[0358] The apparatus embodiments corresponding to the method embodiments of the present application are introduced below. Only a brief introduction to the apparatus is provided below. For the specific implementation steps and details of the solution, reference may be made to the foregoing method embodiments.

[0359] To implement each function in the method provided by the present application, both the first device and the second device may include a hardware structure and / or software module, and implement the above functions in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module. Whether a certain function among the above functions is executed in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module depends on the specific application and design constraints of the technical solution.

[0360] Figure 7 FIG. 7 is a schematic diagram of a model training apparatus 1000 provided by an embodiment of the present application. The apparatus 1000 may include a transceiver unit 1010, a storage unit 1020, and a processing unit 1030. The transceiver unit 1010 is configured to receive or send instructions and / or data. The transceiver unit 1010 may also be referred to as a communication interface or a communication unit. The storage unit 1020 is configured to implement a corresponding storage function and store corresponding instructions and / or data. The processing unit 1030 is configured to perform data processing so that the apparatus 1000 implements the foregoing model training method.

[0361] In a possible implementation manner, the apparatus 1000 may only include the transceiver unit 1010 and the processing unit 1030, and does not include the storage unit 1020.

[0362] As a design, the apparatus 1000 may perform the actions performed by the first device in the foregoing method embodiments.

[0363] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010, configured to send first information to a second device, where the first information is used to indicate a first model, the first model is trained by the first device, and the first model is used to generate a training data set for updating a second model, and the second model is trained by the second device; the transceiver unit is further configured to receive second information, where the second information is used to indicate the updated second model.

[0364] In a possible implementation manner, the apparatus 1000 further includes: a processing unit 1020, configured to perform model training based on all or part of local data to generate the first model, where the all or part of local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the model training apparatus expects the second model to implement.

[0365] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010, configured to send first information to a second device, where the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to be fused with a second model to update the second model, and the second model is trained by the second device; the transceiver unit is further configured to receive second information, where the second information is used to indicate the updated second model.

[0366] In a possible implementation, the apparatus 1000 further includes: a processing unit 1030, configured to fine-tune parameters of the second model based on all or part of local data according to a first function, to generate the first model, where the first function is one or more functions that the first device expects to implement through the second model.

[0367] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010, configured to send first information to a second device, where the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to generate a training data set for updating the second model, or the first model is used to be fused with a second model to update the second model, and the second model is trained by the second device; the transceiver unit is further configured to receive second information, where the second information is used to indicate the updated second model.

[0368] In a possible implementation, the apparatus 1000 further includes: a processing unit 1030, configured to perform model training based on all or part of local data to generate the first model, where the all or part of local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the model training apparatus expects to implement through the second model.

[0369] In a possible implementation, the processing unit 1030 is further configured to fine-tune parameters of the second model based on all or part of local data according to a first function, to generate the first model, where the first function is one or more functions that the first device expects to implement through the second model.

[0370] As a design, the apparatus 1000 can perform the actions performed by the second device in the above method embodiments.

[0371] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010 and a processing unit 1030. The transceiver unit 1010 is configured to receive first information for indicating a first model, where the first model is trained by the first device, and the first model is used to generate a training data set for updating a second model, where the second model is trained by the first device; the processing unit 1030 is configured to update the second model according to the first model; the transceiver unit 1010 is further configured to send second information to the first device, where the second information is used to indicate the updated second model.

[0372] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010 and a processing unit 1030. The transceiver unit 1010 is configured to receive first information for indicating a first model, where the first model is trained by the first device, and the first model is used to be fused with a second model to update the second model, where the second model is trained by the first device; the processing unit 1030 is configured to update the second model according to the first model; the transceiver unit 1010 is further configured to send second information to the first device, where the second information is used to indicate the updated second model.

[0373] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010 and a processing unit 1030. The transceiver unit 1010 is configured to receive first information for indicating a first model, where the first model is trained by the first device, and the first model is used to generate a training data set for updating a second model, or the first model is used to be fused with a second model to update the second model, where the second model is trained by the first device; the processing unit 1030 is configured to update the second model according to the first model; the transceiver unit 1010 is further configured to send second information to the first device, where the second information is used to indicate the updated second model.

[0374] Figure 8 It is a schematic diagram of another model training apparatus 1100 provided by an embodiment of the present application.

[0375] The apparatus 1100 includes: a memory 1110, a processor 1120, and a communication interface 1130. Among them, the memory 1110, the processor 1120, and the communication interface 1130 are connected through an internal connection path. The memory 1110 is configured to store instructions, and the processor 1120 is configured to execute the instructions stored in the memory 1110 to control the communication interface 1130 to obtain information, or to enable the apparatus 1100 to implement the foregoing model training method. Optionally, the memory 1110 can be coupled to the processor 1120 through an interface, or can be integrated with the processor 1120.

[0376] It should be noted that the above communication interface 1130 uses a transceiver device such as, but not limited to, a transceiver. The above communication interface 1130 may further include an input / output interface.

[0377] The processor 1120 stores one or more computer programs, and the one or more computer programs include instructions. When the instructions are run by the processor 1120, the device 1100 is caused to execute the model training method in the above embodiments.

[0378] In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 1120 or the instructions in the form of software. The method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1110, and the processor 1120 reads the information in the memory 1110 and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0379] In a possible implementation manner, the device 1100 may only include the processor 1120 and the communication interface 1130, and does not include the memory 1110.

[0380] Optionally, Figure 8 the communication interface 1130 in Figure 7 may implement the transceiver unit 1010 in Figure 8 the processor 1120 in Figure 7 may implement the processing unit 1030 in

[0381] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores program codes. When the computer program codes are run on a computer, the computer is caused to execute any of the above Figures 3 to 7 methods.

[0382] The embodiments of the present application further provide a computer program product. The computer product includes a computer program. When the computer program is run, the computer is caused to execute any of the above Figures 3 to 7 methods.

[0383] The embodiments of the present application further provide a chip, including: a circuit, and the circuit is used to execute any of the above Figures 3 to 7 methods.

[0384] The embodiments of the present application further provide a system, including: a first device and a second device, where the first device is configured to execute Figures 3 to 7 the actions / steps performed by the first device in Figures 3 to 7 ; and the second device is configured to execute the actions / steps performed by the second device in

[0385] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0386] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the system, device, and unit described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0387] In several embodiments provided by the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0388] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0389] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0390] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0391] As described above, the above are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that, Including: A first device sends first information to a second device, where the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to generate a training data set for updating the second model, or the first model is used to fuse with the second model to update the second model, and the second model is trained by the second device; The first device receives second information from the second device, where the second information is used to indicate the updated second model.

2. The method according to claim 1, characterized in that, The method further includes: The first device sends third information to the second device, where the third information is used to indicate the usage mode of the first model, and the usage mode of the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.

3. The method according to claim 1 or 2, characterized in that, When the first model is used to generate a training data set for updating the second model, before the first device sends the first information to the second device, the method further includes: The first device performs model training based on all or part of local data to generate the first model, where all or part of the local data includes a first data set, and the first data set is a set of data related to a first function, and the first function is one or more functions that the first device expects the second model to implement.

4. The method according to claim 3, characterized in that, Before the first device performs model training based on all or part of local data, the method further includes: The first device fine-tunes the parameters of the second model based on the first function and a second data set, where the second data set includes data related to the first function, and the second data set is included in the first data set; The first device determines that the performance of the fine-tuned second model related to at least one function in the first function does not meet the corresponding performance index requirements.

5. The method according to claim 1 or 2, characterized in that, When the first model is used to merge with the second model to update the second model, before the first device sends the first information to the second device, the method further includes: The first device fine-tunes the parameters of the second model based on the first function and all or part of local data to generate the first model, where the first function is one or more functions that the first device expects the second model to implement.

6. The method according to claim 1 or 2, characterized in that, The first device is a mobile intelligent network element, and the second device is a device supporting cloud services.

7. The method according to claim 6, characterized in that, The method further includes: The first device obtains a third data set, where the third data set includes data related to a second function collected by a base station, and the second function includes one or more functions that the base station expects the second model to implement; The first device fine-tunes the parameters of the second model based on the second function and the third data set.

8. The method according to claim 7, characterized in that, Before the first device sends the first information to the second device, the method further includes: The first device receives fourth information, where the fourth information is used to indicate that the performance of the fine-tuned second model is abnormal; The first device determines a fourth data set according to the fourth information, where the fourth data set is locally stored by the first device, or the fourth data set is obtained by the first device through the base station; The first device performs model training based on the fourth data set to generate the first model.

9. The method according to claim 8, characterized in that, When the fourth data set is obtained by the first device through the base station, the first device determines the fourth data set according to the fourth information, including: The first device sends fifth information to the base station according to the fourth information, where the fifth information is used to request data collected by the base station; The first device receives sixth information, where the sixth information is used to indicate the fourth data set.

10. The method according to any one of claims 1 to 9, characterized in that, Before the first device sends the first information to the second device, the method further includes: The first device performs lightweight processing on the first model.

11. A model training method, characterized in that, Including: The second device receives the first information from the first device, where the first information is used to indicate the first model, the first model is trained by the first device, the first model is used to generate a training data set for updating the second model, or the first model is used to fuse with the second model to update the second model, and the second model is trained by the second device; The second device updates the second model according to the first model; The second device sends second information to the first device, where the second information is used to indicate the updated second model.

12. The method according to claim 11, characterized in that The method further includes: The second device receives third information from the first device, where the third information is used to indicate the usage mode of the first model, and the usage mode of the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.

13. The method according to claim 11 or 12, characterized in that When the first model is used to generate a training data set for updating the second model, the second device updates the second model according to the first model, including: The second device generates a training data set according to the first model; The second device trains the second model according to the training data set to update the second model.

14. The method according to claim 11 or 12, characterized in that When the first model is used to fuse with the second model to update the second model, the second device updates the second model according to the first model, including: The second device fuses the first model and the second model to update the second model.

15. A model training device, characterized in that Including a module or unit for executing the method according to any one of claims 1 to 10, or including a module or unit for executing the method according to any one of claims 11 to 14.

16. A model training device, characterized in that Including a processor, where the processor is configured to, by executing a computer program or instruction, or by a logic circuit, enable the model training device to execute the method according to any one of claims 1 to 10, or enable the model training device to execute the method according to any one of claims 11 to 14.

17. The device according to claim 16, characterized in that The communication device further includes a memory for storing the computer program or instruction.

18. The device according to claim 16 or 17, characterized in that The model training device further includes a communication interface, and the communication interface is used for inputting and / or outputting signals.

19. A computer-readable storage medium, characterized in that A computer program or instruction is stored on the computer-readable storage medium, and when the computer program or the instruction runs on a computer, the method described in any one of claims 1 to 10 is caused to be executed, or the method described in any one of claims 11 to 14 is caused to be executed.

20. A computer program product, characterized in that including instructions, and when the instructions run on a computer, the method described in any one of claims 1 to 10 is caused to be executed, or the method described in any one of claims 11 to 14 is caused to be executed.

21. A system, characterized in that including: a first device and a second device, the first device is configured to execute the method described in any one of claims 1 to 10; the second device is configured to execute the method described in any one of claims 11 to 14.

Citation Information

Cited By

  • Model training method and apparatus

    EP4815390A1

  • Model training method and apparatus

    WO2025130539A1