Model training method and apparatus
By sending the trained first model to the second device in machine learning training, generating or fusing the training data set to update the second model, the problem of large communication resource overhead in large model training is solved, and more efficient communication resource utilization is achieved.
Patent Information
- Application Number
- PCT/CN2024/135087
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-26
AI Technical Summary
Large models require massive data support during machine learning training, and directly applying existing model training methods leads to large communication resource overhead.
The trained first model is sent to the second device through the first device, enabling the second device to generate a training data set for updating the second model, or fusing the first model with the second model to update the second model, thereby reducing the communication overhead of the training data.
It effectively reduces the communication overhead for training data collection during model training and improves the utilization efficiency of communication resources.
Smart Images

Figure CN2024135087_26062025_PF_FP_ABST
Abstract
Description
Model training method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 19, 2023, with application number 202311762388.4 and invention name “Model Training Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communications, and more specifically, to a model training method and device. Background Art
[0003] Artificial intelligence (AI) is a technology that mimics human cognitive, learning, and reasoning abilities. As a key AI technology, machine learning (ML) models have garnered widespread attention in communications technology research due to their efficient data mining and processing capabilities. In machine learning training (MLT), MLT network service (MnS) users typically send training data to MLT MnS providers for selection during training. However, as a new AI paradigm, large models require massive amounts of data support in the MLT process. Directly applying current model training methods will result in significant communication resource overhead. Summary of the Invention
[0004] The present application provides a model training method and device that can reduce the communication overhead associated with collecting training data during the model training process.
[0005] In a first aspect, a model training method is provided. The method can be performed by a first device. The first device here can refer to the access network device itself or a processor, module, chip, or chip system that implements the method in the first device, and this application does not limit this. The method includes:
[0006] The first device sends first information to the second device, where the first information is used to indicate a first model, where the first model is trained by the first device and is used to generate a training data set for updating a second model, or where the first model is fused with the second model to update the second model, where the second model is trained by the second device; the first device receives second information from the second device, where the second information is used to indicate the updated second model.
[0007] By way of example and not limitation, the first device may be a user of a network management service (e.g., a communications infrastructure model), and the second device may be a network management service provider. The first model may be a data generation model, or a model running locally on the first device, such as an on-site model when the first device is a base station. The second model may be an AI / ML model that implements one or more communication functions in a physical network, or a communications infrastructure model.
[0008] It should be understood that the embodiment of the present application does not limit the specific manner in which the first device obtains the second model.
[0009] As an example and not a limitation, the second model may be pre-configured in the first device. Specifically, the second model may be configured in the first device when the first device leaves the factory, or the second model may be manually configured in the first device before the first device is put into operation.
[0010] As an example and not a limitation, the second model may be sent by the second device to the first device. Specifically, before the first device sends the first information to the second device, the second device sends configuration information to the first device, where the configuration information is used to indicate the second model.
[0011] In one possible implementation, the second device may also send at least one of the performance information of the second model, the evaluation index information of the second model, or the functional information of the second model to the first device, so that the first device can determine whether the second model meets the corresponding performance requirements during operation. The performance information of the second model may include information such as the parameter quantity, structure, and number of layers of the second model; the evaluation index information of the second model may include the performance of the second model in its applicable downstream tasks, such as accuracy, recall rate, mean square error, etc.; and the functional information of the second model may include information about communication functions that can be implemented by the model.
[0012] It should be understood that the embodiment of the present application does not limit the specific manner in which the first information indicates the first model.
[0013] As an example but not a limitation, the first information includes the first model, and the second device can obtain the first model after receiving the first information.
[0014] As an example but not limitation, the first model may include a storage address of the first model. Then, after receiving the first information, the second device may obtain the first model according to the storage address of the first model.
[0015] Optionally, when the first device sends the first information to the second device, the first device may also send at least one item of performance information of the first model, evaluation index information of the first model, or function information of the first model to the second device.
[0016] Based on the above scheme, the first device sends the first model to the second device, so that the second device can generate a training data set for updating the second model based on the first model, and update the second model based on the training data set, or, the second device can update the first model by fusing the first model with the second model. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the first model instead of training data.
[0017] In combination with the first aspect, in certain implementations of the first aspect, the method includes: the first device sends third information to the second device, the third information is used to indicate how the first model is used, and the way of using the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.
[0018] As an example and not a limitation, the third information can indicate the usage method of the first model through bit 0 and bit 1. For example, bit 0 is used to indicate that the first model is used to generate a training data set for updating the second model, and bit 1 is used to indicate that the first model is used to be fused with the second model.
[0019] Based on the above solution, the first device can indicate to the second device how to use the first model through the third information, saving the second device the time of identifying or trial-and-erroring how to use the first model, thereby improving the efficiency of the second device in using the first model.
[0020] In combination with the first aspect, in certain implementations of the first aspect, when the first model is used to generate a training data set for updating the second model, the method also includes: the first device performs model training based on all or part of the local data to generate the first model, and the all or part of the local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the first device expects to implement through the second model.
[0021] As an example and not a limitation, the first function may include functions such as beam management and channel measurement required by the first device, or functions such as intent management expected by the first device, or functions such as mobility prediction expected by the first device.
[0022] Based on the above scheme, the first device can perform model training based on all or part of the local data to generate the above-mentioned first model, wherein all or part of the local data includes data related to the first function, thereby reducing the overhead of model training for the first device and improving the accuracy of model training.
[0023] In combination with the first aspect, in certain implementations of the first aspect, when the first model is used to generate a training data set for updating the second model, the method also includes: the first device fine-tunes the parameters of the second model based on the second data set according to the first function, the second data set including data related to the first function, and the second data set is included in the first data set; and determines that the performance of the fine-tuned second model related to at least one of the first functions does not meet the corresponding performance indicator requirements.
[0024] It should be understood that the fact that the performance of the second model related to at least one of the first functions does not meet the corresponding performance index requirements can be understood as that the various performances of the second model related to the first function include at least one performance that does not meet the corresponding performance index requirements.
[0025] It should be understood that the second data set being included in the first data set can be understood as the time period for collecting the second data set being included in the time period for collecting the first data set.
[0026] As an example but not limitation, the first device is a base station, and the first device can fine-tune the parameters of the second model based on the second data set according to the required beam management, channel measurement and other functions.
[0027] As an example but not limitation, the first device is an OAM, and the first device can fine-tune the parameters of the second model based on the second data set according to functions such as intent management.
[0028] As an example but not limitation, the first device is an NWDAF, and the first device can fine-tune the parameters of the second model based on the second data set according to functions such as mobility prediction.
[0029] It should be understood that the performance indicator may be a key performance indicator (KPI) requirement, and the KPI requirement is determined by the first device during the process of fine-tuning the second model.
[0030] It should be understood that when the performance of the second model after the above-mentioned fine-tuning meets the performance indicator requirements corresponding to the first function, the first device can continue to fine-tune the parameters of the second model based on all or local data and the first function, thereby improving the performance of the model locally deployed on the first device.
[0031] Based on the above scheme, the first device can use all or part of the local data to fine-tune the second model according to the first function, and when the performance of the fine-tuned second model does not meet the performance indicator requirements corresponding to the first function, the local data is trained to generate the first model, thereby ensuring the availability of the model locally deployed on the first device.
[0032] In combination with the first aspect, in certain implementations of the first aspect, when the first model is used to merge with the second model to update the second model, the method also includes: the first device fine-tunes the parameters of the second model based on all or part of the local data according to the first function to generate the first model, and the first function is one or more functions that the first device expects to achieve through the second model.
[0033] Based on the above solution, the first device can use all or part of the local data to fine-tune the second model according to the first function to generate the first model, and send the first model to the second device, so that the second device can fuse the first model and the second model to update the second model, thereby saving the communication overhead brought by the collection of training data for the second model between the first device and the second device, and realizing dynamic update of the second model.
[0034] In combination with the first aspect, in some implementations of the first aspect, the first device is a mobile intelligent network element, and the second device is a device that supports cloud services.
[0035] It should be understood that in this implementation, the training data of the first network element is provided by other devices (such as a base station), and the first device itself is only used for data training.
[0036] In combination with the first aspect, in certain implementations of the first aspect, the first device obtains a third data set, which includes data related to the second function collected by the base station, and the second function includes one or more functions that the base station expects to implement through the second model; the first device fine-tunes the parameters of the second model based on the third data set according to the second function.
[0037] Based on the above solution, the first device can fine-tune the parameters of the second model based on the second function for a device that does not have computing power (such as a base station), thereby improving the performance of the model deployed on the device that does not have computing power.
[0038] In combination with the first aspect, in certain implementations of the first aspect, when the first model is used to generate a training data set for updating the second model, before the first device sends the first information to the second device, the method also includes: the first device receives fourth information, and the fourth information is used to indicate that the performance of the second model after fine-tuning is abnormal; the first device determines a fourth data set based on the fourth information, and the fourth data set is locally stored by the first device, or the fourth data set is obtained by the first device through the base station; the first device performs model training based on the fourth data set to generate the first model.
[0039] It should be understood that the fourth data set is a collection of data related to the second function, and the fourth data set includes the third data set, that is, the time period for collecting the fourth data set includes the time period for collecting the third data set.
[0040] It should be understood that the fourth data set for the first device to perform model training may be requested by the first device from the base station when the first device determines that the first model needs to be generated, or the fourth data set for the first device to perform model training may be data reported by the base station and stored locally by the first device. This embodiment of the present application is not limited to this.
[0041] In combination with the first aspect, in certain implementations of the first aspect, when the first model is used to generate a training data set for updating the second model, the method also includes: the first device sends fifth information to the base station based on the fourth information, and the fifth information is used to request data collected by the base station; the first device receives the fifth information, and the fifth information is used to indicate the fourth data set.
[0042] It should be understood that the fourth information is used to request the base station to collect all data, or the fourth information is used to request the base station to collect data related to the second function.
[0043] It should be understood that all the data collected by the aforementioned base station may be all the data collected after the base station is put into operation, or all the data collected by the aforementioned base station may be all the data collected by the base station during the interval between receiving two request messages, or all the data collected by the aforementioned base station may be all the data collected within a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not limit this.
[0044] It should be understood that the data related to the second function collected by the aforementioned base station may be all the data related to the second function collected after the base station is put into operation, or, the data related to the second function collected by the aforementioned base station may be the data related to the second function collected by the base station during the interval between receiving two request messages, or, the data related to the second function collected by the aforementioned base station may be the data related to the second function collected within a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not limit this.
[0045] As an example and not a limitation, the fourth information may include at least one of an identifier #1 and an identifier #2, wherein the identifier #1 is used to indicate a specific time period, so that each base station can report data collected within the specific time period based on the identifier #1, and the identifier #2 is used to indicate a specific function (for example, the second function), so that each base station can report data related to the specific function based on the identifier #2.
[0046] It should be understood that the first device may discard part or all of the data provided by the base station after model training, or the first device may save part or all of the data provided by the base station after model training. This embodiment of the present application is not limited to this.
[0047] Based on the above solution, when the base station reports that its locally deployed model has failed, the first device can perform model training based on the data collected by the base station to generate a first model, and send the first model to the second device, so that the second device can dynamically update the second model based on the data collected by the base station, thereby ensuring the model performance of the base station operation while saving the communication overhead caused by the collection of training data for the second model between the first device and the second device.
[0048] In combination with the first aspect, in some implementations of the first aspect, before the first device sends the first information to the second device, the method further includes: the first device performing lightweight processing on the first model.
[0049] Based on the above solution, the first device can perform lightweight processing on the first model after generating the first model, thereby further reducing the communication overhead caused by the collection of training data for the second model between the first device and the second device.
[0050] In a second aspect, a model training method is provided. The method can be performed by a second device. The second device here can refer to the access network device itself or a processor, module, chip, or chip system that implements the method in the second device, which is not limited in this application. The method includes:
[0051] The second device receives first information from the first device, where the first information is used to indicate a first model, where the first model is trained by the first device, and the first model is used to generate a training data set for updating the second model, or the first model is used to be fused with the second model to update the second model, where the second model is trained by the second device; the second device updates the second model based on the first model; and the second device sends second information to the first device, where the second information is used to indicate the updated second model.
[0052] As an example and not a limitation, the first device may be a user of a network management service (e.g., a communication infrastructure model), and the second device may be a network management service provider. The first model may be a data generation model, or the first model may be a model running locally on the first device, for example, when the first device is a base station, the first model may be an on-site model. The second model may be an AI / ML model that implements one or more communication functions in a physical network, or may be referred to as a communication infrastructure model. It should be understood that regarding the specific manner in which the first device obtains the second model, and the specific manner in which the first information indicates the first model, reference may be made to the relevant content of the first aspect, which will not be elaborated here.
[0053] Based on the above scheme, the first device sends the first model to the second device, so that the second device can generate a training data set for updating the second model based on the first model, and update the second model based on the training data set, or, the second device can update the first model by fusing the first model with the second model. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the first model instead of training data.
[0054] In combination with the second aspect, in certain implementations of the second aspect, the method also includes: the second device receives third information from the first device, the third information is used to indicate how the first model is used, and the way of using the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.
[0055] It should be understood that the second device can determine the usage method of the first model based on the third information, and then update the second model based on the first model. Specific descriptions of the type of the first model can refer to the relevant content of the fifth aspect and are not repeated here.
[0056] Based on the above solution, the first device can indicate to the second device how to use the first model through the third information, saving the second device the time of identifying or trial-and-erroring how to use the first model, thereby improving the efficiency of the second device in using the first model.
[0057] In combination with the second aspect, in certain implementations of the second aspect, when the first model is used to generate a training data set for updating the second model, the second device updates the second model based on the first model, including: the second device generates a training data set based on the first model; the second device trains the second model based on the training data set to update the second model.
[0058] In combination with the second aspect, in certain implementations of the second aspect, when the first model is used to be merged with the second model to update the second model, the second device updates the second model based on the first model, including: the second device merges the first model and the second model to update the second model.
[0059] In a third aspect, a model training method is provided. The method can be performed by a first device. The first device here can refer to the access network device itself or a processor, module, chip, or chip system that implements the method in the first device, which is not limited in this application. The method includes:
[0060] The first device sends first information to the second device, where the first information is used to indicate a first model, where the first model is trained by the first device, and where the first model is used to generate a training data set for updating a second model, where the second model is trained by the second device; the first device receives second information, where the second information is used to indicate the updated second model.
[0061] By way of example and not limitation, the first device may be a user of a network management service (e.g., a communications infrastructure model), and the second device may be a network management service provider. The first model may be a data generation model, and the second model may be an AI / ML model that implements one or more communication functions in a physical network, or a communications infrastructure model.
[0062] It should be understood that regarding the specific manner in which the first device obtains the second model, and the specific manner in which the first information indicates the first model, reference may be made to the relevant content of the first aspect and will not be repeated here.
[0063] Based on the above scheme, the first device sends the first model to the second device, so that the second device can generate a training data set for updating the second model based on the first model, and update the second model based on the training data set. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the first model instead of training data.
[0064] In combination with the third aspect, in certain implementations of the third aspect, before the first device sends the first information to the second device, the method also includes: the first device performs model training based on all or part of the local data to generate the first model, and the all or part of the local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the first device expects to implement through the second model.
[0065] In combination with the third aspect, in certain implementations of the third aspect, before the first device performs model training based on all or part of the local data, the method also includes: the first device fine-tunes the parameters of the second model based on the second data set according to the first function, the second data set includes data related to the first function, and the second data set is included in the first data set; determines that the performance of the fine-tuned second model related to at least one function of the first function does not meet the corresponding performance indicator requirements.
[0066] It should be understood that the performance of the second model related to at least one of the first functions does not meet the corresponding performance index requirements, which can be understood as the various performances of the second model related to the first function include at least one performance that does not meet the corresponding performance index requirements.
[0067] In combination with the third aspect, in certain implementations of the third aspect, the first device is a mobile intelligent network element, and the second device is a device that supports cloud services.
[0068] In combination with the third aspect, in certain implementations of the third aspect, the method also includes: the first device obtains a third data set, the third data set includes data related to the second function collected by the base station, and the second function includes one or more functions that the base station expects to implement through the second model; the first device fine-tunes the parameters of the second model based on the third data set according to the second function.
[0069] In combination with the third aspect, in certain implementations of the third aspect, before the first device sends the first information to the second device, the method also includes: the first device receives fourth information, and the fourth information is used to indicate that the performance of the second model after fine-tuning is abnormal; the first device determines a fourth data set based on the fourth information, and the fourth data set is locally stored by the first device, or the fourth data set is obtained by the first device through the base station, wherein the fourth data set includes the third data set; the first device performs model training based on the fourth data set to generate the first model.
[0070] In combination with the third aspect, in certain implementations of the third aspect, the method also includes: the first device sends fifth information to the base station based on the fourth information, and the fifth information is used to request the base station to collect data; the first device receives sixth information, and the sixth information is used to indicate the fourth data set.
[0071] In combination with the third aspect, in certain implementations of the third aspect, before the first device sends the first information to the second device, the method further includes: the first device performing lightweight processing on the first model.
[0072] In a fourth aspect, a model training method is provided, which can be performed by a second device. The second device here can refer to the access network device itself, or a processor, module, chip, or chip system that implements the method in the second device, which is not limited in this application. The method includes:
[0073] The second device receives first information, where the first information is used to indicate a first model, where the first model is trained by the first device, and where the first model is used to generate a training data set for updating the second model, where the second model is trained by the second device; the second device updates the second model based on the first model; and the second device sends second information to the first device, where the second information is used to indicate the updated second model.
[0074] By way of example and not limitation, the first device may be a user of a network management service (e.g., a communications infrastructure model), and the second device may be a network management service provider. The first model may be a data generation model, and the second model may be an AI / ML model that implements one or more communication functions in a physical network, or a communications infrastructure model.
[0075] Based on the above scheme, the first device sends the first model to the second device, so that the second device can generate a training data set for updating the second model based on the first model, and update the second model based on the training data set. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the first model instead of training data.
[0076] In combination with the fourth aspect, in certain implementations of the fourth aspect, the second device updates the second model based on the first model, including: the second device generates a training data set based on the first model; and the second device trains the second model based on the training data set to update the second model.
[0077] In a fifth aspect, a model training method is provided. The method can be performed by a first device. The first device here can refer to the access network device itself, or to a processor, module, chip, or chip system in the first device that implements the method. This application does not limit this. The method includes:
[0078] The first device sends first information to the second device, where the first information is used to indicate a first model, where the first model is trained by the first device, and the first model is used to be fused with a second model to update the second model, where the second model is trained by the second device; the first device receives second information, where the second information is used to indicate the updated second model.
[0079] As an example and not a limitation, the first device may be a user of a network management service (e.g., a communication infrastructure big model), the second device may be a network management service provider, the first model may be a model running locally on the first device, for example, when the first device is a base station, the first model may be an on-site model, and the second model may be an AI / ML model that implements one or more communication functions in a physical network, or is called a communication infrastructure big model.
[0080] It should be understood that regarding the specific manner in which the first device obtains the second model, and the specific manner in which the first information indicates the first model, reference may be made to the relevant content of the first aspect and will not be repeated here.
[0081] Based on the above solution, the first device sends the first model to the second device, so that the second device can fuse the first model with the second model to update the first model. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the second model instead of training data.
[0082] In combination with the fifth aspect, in certain implementations of the fifth aspect, the method also includes: the first device fine-tunes the parameters of the second model based on all or part of the local data according to the first function to generate the first model, and the first function is one or more functions that the first device expects to implement through the second model.
[0083] In combination with the fifth aspect, in certain implementations of the fifth aspect, the entire or local data includes at least a second data set, which is a collection of data related to the first function.
[0084] In a sixth aspect, a model training method is provided, which can be performed by a second device. The second device here can refer to the access network device itself, or a processor, module, chip, or chip system that implements the method in the second device, which is not limited in this application. The method includes:
[0085] The second device receives first information from the first device, where the first information is used to indicate a first model, where the first model is trained by the first device, and the first model is used to be fused with a second model to update the second model, where the second model is trained by the first device; the second device updates the second model based on the first model; and the second device sends second information to the first device, where the second information is used to indicate the updated second model.
[0086] As an example and not a limitation, the first device may be a user of a network management service (e.g., a communication infrastructure big model), the second device may be a network management service provider, the first model may be a model running locally on the first device, for example, when the first device is a base station, the first model may be an on-site model, and the second model may be an AI / ML model that implements one or more communication functions in a physical network, or is called a communication infrastructure big model.
[0087] Based on the above solution, the first device sends the first model to the second device, so that the second device can fuse the first model with the second model to update the first model. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the second model instead of training data.
[0088] In combination with the sixth aspect, in certain implementations of the sixth aspect, the second device updates the second model based on the first model, including: the second device fuses the first model and the second model to update the second model.
[0089] In the seventh aspect, a model training device is provided, which includes: a transceiver unit for receiving first information, where the first information is used to indicate a first model, where the first model is trained by the first device, and where the first model is used to be integrated with a second model to update the second model, where the second model is trained by the first device; a processing unit for updating the second model based on the first model; and the transceiver unit is also used to send second information to the first device, where the second information is used to indicate the updated second model.
[0090] In combination with the seventh aspect, in certain implementations of the seventh aspect, the transceiver unit is also used to send third information to the second device, where the third information is used to indicate how the first model is used, and the way of using the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.
[0091] In combination with the seventh aspect, in certain implementations of the seventh aspect, the model training device also includes a processing unit, which is used to perform model training based on all or part of the local data to generate the first model, and the all or part of the local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the model training device expects to achieve through the second model.
[0092] In combination with the seventh aspect, in certain implementations of the seventh aspect, the processing unit is further used to: fine-tune the parameters of the second model based on a second data set according to the first function, the second data set including data related to the first function, and the second data set included in the first data set; and determine that the performance of the fine-tuned second model related to at least one function of the first function does not meet the corresponding performance indicator requirements.
[0093] In combination with the seventh aspect, in certain implementations of the seventh aspect, the processing unit is also used to: according to the first function, fine-tune the parameters of the second model based on all or part of the local data to generate the first model, where the first function is one or more functions that the first device expects to implement through the second model.
[0094] In combination with the seventh aspect, in certain implementations of the seventh aspect, the model training device is a mobile intelligent network element, and the second device is a device that supports cloud services.
[0095] In combination with the seventh aspect, in certain implementations of the seventh aspect, the transceiver unit is further used to receive fourth information, where the fourth information is used to indicate that the performance of the second model after fine-tuning is abnormal; the processing unit is further used to determine a fourth data set based on the fourth information, where the fourth data set is locally stored by the first device, or the fourth data set is obtained by the first device through the base station, wherein the fourth data set includes the third data set; the processing unit is also used to perform model training based on the fourth data set to generate the first model.
[0096] In combination with the seventh aspect, in certain implementations of the seventh aspect, the processing unit is also used to send fifth information to the base station based on the fourth information, and the fifth information is used to request the base station to collect data; the transceiver unit is also used to receive fifth information, and the fifth information is used to indicate the fourth data set.
[0097] In combination with the seventh aspect, in some implementations of the seventh aspect, before the first device sends the first information to the second device, the processing unit is further used to perform lightweight processing on the first model.
[0098] In the eighth aspect, a model training device is provided, which includes: a transceiver unit for receiving first information, the first information being used to indicate a first model, the first model being trained by the first device, the first model being used to generate a training data set for updating the second model, or the first model being used to fuse with the second model to update the second model, the second model being trained by the second device; a processing unit being used to update the second model based on the first model; the transceiver unit is also used to send second information to the first device, the second information being used to indicate the updated second model.
[0099] In combination with the eighth aspect, in certain implementations of the eighth aspect, the transceiver unit is further used to receive third information from the first device, where the third information is used to indicate how the first model is used, and the way of using the first model includes generating the training data set based on the first model, or performing model fusion based on the first model.
[0100] In combination with the eighth aspect, in certain implementations of the eighth aspect, the model training device also includes a processing unit, which is used to: generate a training data set based on the first model; and train the second model based on the training data set to update the second model.
[0101] In combination with the eighth aspect, in some implementations of the eighth aspect, the processing unit is further used to fuse the first model and the second model to update the second model.
[0102] In the ninth aspect, a model training device is provided, which includes: a transceiver unit, which is used to send first information to a second device, the first information is used to indicate a first model, the first model is trained by the first device, and the first model is used to generate a training data set for updating a second model, the second model is trained by the second device; the transceiver unit is also used to receive second information, the second information is used to indicate the updated second model.
[0103] It should be understood that the ninth aspect is an implementation method on the device side corresponding to the third aspect. The supplement, explanation and beneficial effects of the third aspect are also applicable to the ninth aspect and will not be repeated here.
[0104] In the tenth aspect, a model training device is provided, which includes: a transceiver unit for receiving first information, the first information being used to indicate a first model, the first model being trained by the first device, the first model being used to generate a training data set for updating the second model, the second model being trained by the second device; a processing unit for updating the second model based on the first model; the transceiver unit is also used to send second information to the first device, the second information being used to indicate the updated second model.
[0105] It should be understood that the tenth aspect is an implementation method on the device side corresponding to the fourth aspect. The supplement, explanation and beneficial effects of the fourth aspect are also applicable to the tenth aspect and will not be repeated here.
[0106] In the eleventh aspect, a model training device is provided, which includes: a transceiver unit, which is used to send first information to a second device, the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to be integrated with the second model to update the second model, the second model is trained by the second device; the transceiver unit is also used to receive second information, the second information is used to indicate the updated second model.
[0107] It should be understood that the eleventh aspect is an implementation method on the device side corresponding to the fifth aspect. The supplement, explanation and beneficial effects of the fifth aspect are also applicable to the eleventh aspect and will not be repeated here.
[0108] In the twelfth aspect, a model training device is provided, which includes: a transceiver unit for receiving first information, the first information being used to indicate a first model, the first model being trained by the first device, the first model being used to be integrated with a second model to update the second model, the second model being trained by the second device; a processing unit for updating the second model based on the first model; the transceiver unit is also used to send second information to the first device, the second information being used to indicate the updated second model.
[0109] It should be understood that the twelfth aspect is an implementation method on the device side corresponding to the sixth aspect. The supplement, explanation and beneficial effects of the sixth aspect are also applicable to the twelfth aspect and will not be repeated here.
[0110] In a thirteenth aspect, the present application provides a model training device, comprising a processor configured to implement the method described in any one of the first to sixth aspects, or any one of the implementations of the first to sixth aspects. The processor is coupled to a memory configured to store instructions and data, and when the processor executes the instructions stored in the memory, the method described in any one of the first to sixth aspects, or any one of the implementations of the first to sixth aspects, can be implemented.
[0111] Optionally, the communication device may further include a memory. Optionally, the memory may be coupled to the processor. Optionally, the communication device may further include a communication interface, which is used for the device to communicate with other devices. Exemplarily, the communication interface may be a transceiver, hardware circuit, bus, module, pin, or other type of communication interface.
[0112] In an example, the communication device may be a first device, or may be a device, module, chip, etc. provided in the first device, or may be a device that can be used in conjunction with the first device.
[0113] In another example, the communication device may be the second device, or may be a device, module, chip, etc. provided in the second device, or may be a device that can be used in conjunction with the second device.
[0114] In the fourteenth aspect, the present application provides a system comprising: a first device for executing the method described in the first aspect, the third aspect, or the fifth aspect, or any implementation of the first aspect, the third aspect, or the fifth aspect; a second device for executing the method described in the second aspect, the fourth aspect, or the sixth aspect, or any implementation of the second aspect, the fourth aspect, or the sixth aspect.
[0115] In the fifteenth aspect, the present application also provides a computer program, which, when executed on a computer, enables the computer to execute the method described in any one of the implementations of the first to sixth aspects above, or the first to sixth aspects.
[0116] In the sixteenth aspect, the present application also provides a computer program product, comprising instructions, which, when executed on a computer, enable the computer to execute the method described in any one of the implementations of the first to sixth aspects above, or the first to sixth aspects.
[0117] In the seventeenth aspect, the present application also provides a computer-readable storage medium, which stores a computer program or instructions. When the computer program or instructions are run on a computer, the computer executes the method described in any one of the implementations of the first to sixth aspects above, or the first to sixth aspects.
[0118] In the eighteenth aspect, the present application also provides a chip, which is used to read the computer program stored in the memory and execute the method described in any implementation of the above-mentioned first to sixth aspects, or the first to sixth aspects; or, the chip includes a method for executing the above-mentioned first to sixth aspects, or the method described in any implementation of the first to sixth aspects.
[0119] In a nineteenth aspect, the present application further provides a chip system, which includes a processor for supporting a device to implement the method described in any of the above-mentioned aspects 1 to 6, or any of the implementations of aspects 1 to 6. In one possible design, the chip system also includes a memory for storing programs and data necessary for the device. The chip system can be composed of a chip, or it can include a chip and other discrete devices.
[0120] The description of the beneficial effects of any of the seventh to nineteenth aspects can refer to the description of the beneficial effects of the first to sixth aspects, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0121] FIG1 is a schematic diagram of a network architecture;
[0122] Figure 2 is a schematic diagram of the current MLT process;
[0123] FIG3 is a schematic diagram of a model training method 300 provided in an embodiment of the present application;
[0124] FIG4 is a schematic diagram of a model training process 400 provided in an embodiment of the present application;
[0125] FIG5 is a schematic diagram of a model training process 500 provided in an embodiment of the present application;
[0126] FIG6 is a schematic diagram of a model training process 600 provided in an embodiment of the present application;
[0127] FIG7 is a schematic diagram of a model training device 1000 provided in an embodiment of the present application;
[0128] FIG8 is a schematic diagram of another model training device 1100 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0129] The technical solution in this application will be described below with reference to the accompanying drawings.
[0130] To facilitate understanding, a communication system to which the embodiments of the present application can be applied is first described.
[0131] The embodiments of the present application can be applied to various communication systems. For example: long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD) system, public land mobile network (PLMN), fifth generation (5G) system, sixth generation (6G) system or future communication system. The 5G system in the present application includes a non-standalone (NSA) 5G mobile communication system or a standalone (SA) 5G mobile communication system. The embodiments of the present application can also be applied to non-terrestrial network (NTN) communication systems such as satellite communication systems. The embodiments of the present application can also be applied to device-to-device (D2D) communication systems, sidelink (SL) communication systems, machine-to-machine (M2M) communication systems, machine type communication (MTC) systems, Internet of Things (IoT) communication systems, vehicle-to-everything (V2X) communication systems, uncrewed aerial vehicle (UAV) communication systems, or other communication systems.
[0132] As an example, FIG1 shows a schematic diagram of a network architecture.
[0133] As shown in Figure 1, the network architecture takes the 5G system (5GS) as an example. The network architecture may include three parts, namely the user equipment (UE) part, the data network (DN) part and the operator network part. Among them, the operator network may include one or more of the following network elements: (radio) access network (RAN) equipment, user plane function (UPF) network element, access and mobility management function (AMF) network element, session management function (SMF) network element, network data analysis function (NWDAF) network element, policy control function (PCF) network element, application function (AF) network element, mobile intelligence function (MIF) network element and network management (Operations, Administration And Management, OAM) network element. In the above-mentioned operator network, the part other than the RAN part can be called the core network part.
[0134] In this application, user equipment, (wireless) access network equipment, UPF network element, AMF network element, SMF network element, NWDAF network element, PCF network element, AF network element, MIF network element, and OAM network element are respectively referred to as UE, (R)AN, UPF, AMF, SMF, NWDAF, PCF, AF, MIF, and OAM.
[0135] The following briefly describes the network elements involved in FIG1 .
[0136] 1.UE
[0137] The UE in this application may also be referred to as a terminal, user, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal device, wireless communication device, user agent or user device, etc. For the sake of convenience of description, it is collectively referred to as a terminal below.
[0138] A terminal is a device that can access a network. Terminals and (R)ANs can communicate with each other using an air interface technology (such as NR or LTE). Terminals can also communicate with each other using an air interface technology (such as NR or LTE). A terminal can be a mobile phone, tablet, computer with wireless transceiver capabilities, virtual reality (VR) terminal, augmented reality (AR) terminal, satellite communication terminal, integrated access and backhaul (IAB) system terminal, WiFi communication system terminal, industrial control terminal, self-driving terminal, remote medical terminal, smart grid terminal, transportation safety terminal, smart city terminal, smart home terminal, etc.
[0139] The embodiments of the present application do not limit the specific technology and specific device form adopted by the UE.
[0140] 2. (R)AN
[0141] The (R)AN in this application may be a device used to communicate with a terminal, or may be a device that connects a terminal to a wireless network.
[0142] The (R)AN can be a node in a radio access network. The (R)AN can be a base station, an evolved NodeB (eNodeB), a transmission reception point (TRP), a home base station (e.g., home evolved NodeB, or home NodeB, HNB), a Wi-Fi access point (AP), a mobile switching center, a next-generation NodeB (gNB) in a 5G mobile communication system, access network equipment in an open radio access network (O-RAN or open RAN), a next-generation base station in a sixth-generation (6G) mobile communication system, or a base station in a future mobile communication system. A network device can also be a module or unit that performs some of the functions of a base station, such as a central unit (CU), a distributed unit (DU), a remote radio unit (RRU), or a baseband unit (BBU). The (R)AN can also function as a base station in D2D, V2X, M2M, and IoT communication systems. It can also be a network device within an NTN, meaning it can be deployed on a high-altitude platform or satellite. It can be a macro base station, a micro base station, an indoor station, a relay node, or a donor node.
[0143] The embodiments of this application do not limit the specific technology, device form and name adopted by (R)AN.
[0144] 3. UPF
[0145] The main functions of UPF are packet routing and forwarding, mobility anchor, uplink classifier to support routing service flows to data networks, branch point to support multi-homed PDU sessions, etc.
[0146] 4. DN
[0147] DN is mainly used for operator networks that provide data services to terminals, such as the Internet, third-party service networks, or IP Multimedia Service (IMS) networks.
[0148] 5. AMF
[0149] The main functions of AMF include managing user registration, reachability detection, SMF node selection, and mobile state transition management.
[0150] 6. SMF
[0151] The main functions of SMF are to control the establishment, modification and deletion of sessions, the selection of user plane nodes, etc.
[0152] 7. NWDAF
[0153] It has functions such as data collection, model training, data analysis, and model reasoning. It can be used to collect relevant data from network elements, third-party service servers, terminal devices, or network management systems, perform data analysis or model training based on relevant data, and provide data analysis results to network elements, third-party service servers, terminal devices, or network management systems, or provide trained models to other data analysis function elements.
[0154] 8. PCF
[0155] PCF is mainly responsible for policy control decisions, policy rules for providing control plane functions, and flow-based charging control functions.
[0156] 9. AF
[0157] The AF primarily interacts with the 3GPP core network to deliver services, such as influencing data routing decisions, implementing policy control functions, or providing third-party services to the network. The AF can be deployed within the operator's network or a third-party AF. The DCAF is a specialized AF that collects data from UE applications and makes it available to network elements such as the NWDAF.
[0158] 10. OAM
[0159] It mainly performs daily network and service analysis, forecasting, planning, and configuration, as well as network and service testing and fault management. OAM can interact with the RAN to obtain UE location information measured by the RAN or reported by the UE.
[0160] 11. MIF
[0161] Responsible for the artificial intelligence (AI) or machine learning (ML) functions of the base station, including data management functions, computing power management functions, and model management functions. Among them, the data management function may include: data collection, data storage, and data analysis; the computing power management function may include: computing power perception, computing power scheduling, and computing power and transmission coordination; the model management function may include: model training, model reasoning, and model lifecycle management. The MIF can be in the form of: a base station, a network element independent of the base station, a network function independent of the base station, a submodule within the base station, or a subfunction within the base station. For example, in Figure 1, the MIF is a network element / network function independent of the base station.
[0162] In the network architecture shown in Figure 1, each network element can communicate with each other through an interface. The interface between each network element can be a point-to-point interface or a service-oriented interface, which is not limited in this application.
[0163] It should be understood that the network architecture shown above is only an exemplary illustration, and the network architecture applicable to the embodiments of the present application is not limited to this. Any network architecture that can realize the functions of the above-mentioned network elements is applicable to the embodiments of the present application.
[0164] It should also be understood that the functions or network elements such as UE, (R)AN, UPF, AMF, SMF, NWDAF, PCF, AF, MIF, OAM shown in Figure 1 can be understood as network elements for implementing different functions, for example, they can be combined into network slices as needed. These network elements can be independent devices, or they can be integrated into the same device to implement different functions, or they can be network elements in hardware devices, or they can be software functions running on dedicated hardware, or they can be virtualized functions instantiated on a platform (for example, a cloud platform). This application does not limit the specific form of the above network elements.
[0165] It should also be understood that the above naming is defined only to facilitate the distinction between different functions and should not constitute any limitation to this application. This application does not exclude the possibility of adopting other naming in 6G networks and other future networks. For example, in a 6G network, some or all of the above network elements may continue to use the terminology used in 5G, or may adopt other names.
[0166] To facilitate understanding of the embodiments of the present application, several basic concepts involved in the embodiments of the present application are briefly explained.
[0167] 1. Machine Learning (ML)
[0168] As a key technology in artificial intelligence (AI), machine learning (ML) models have garnered widespread attention in communications technology research due to their efficient data mining and processing capabilities. The training process for ML models in communication networks is a key research direction in the 3rd Generation Partnership Project (3GPP) Release 18 standard. Efficiently training ML models can improve their performance in practical communication applications and facilitate subsequent deployment, management, and large-scale application of ML models. SA5 TS28.105 defines the overall process framework and data requirements for ML model training. Providers of machine learning training (MLT) network services (MnS) utilize current and historical data to monitor networks or services related to ML models, prepare data, and trigger and execute training.
[0169] 2. Large Model
[0170] A large model is an artificial neural network model containing at least hundreds of millions of parameters. It requires computers and massive amounts of existing data for optimization and training. It possesses powerful reasoning and generalization capabilities. As a new ML paradigm, large models are enabling performance breakthroughs in multiple application areas, including natural language processing (NLP). Leveraging their massive parameter size, vast amounts of data and computing resources, and a model structure that captures global data relationships, large models achieve unprecedented reasoning and generalization capabilities. A single pre-trained model can be adapted to a wide range of application tasks, achieving optimal performance through fine-tuning or small- or zero-shot learning. Given that 3GPP Release 19 standardization research is focusing on natural language intent, large models may serve as a potential technical solution for natural language intent translation. Furthermore, with their powerful reasoning capabilities, large models have significant application potential in communication network scenarios such as physical layer radio resource allocation, network management plane operation optimization, and core network intelligent control.
[0171] 3. Fine-tuning
[0172] This refers to the process of training a pre-trained artificial neural network model for a specific task or tasks using information specific to that task. For example, to enable a neural network model to perform machine translation, it is necessary to use a table of language mappings to further train the pre-trained model.
[0173] 4. Data Generation Model
[0174] An artificial neural network model capable of fitting a data distribution. After training, given a given input, it can produce more outputs that conform to that distribution. For example, generative adversarial networks are used in text generation tasks. By training the model using a large amount of natural language, it is able to output a set of natural language outputs similar to the training corpus given any white noise input.
[0175] Figure 2 is a schematic diagram of the current MLT process. Considering the current insufficient computing power deployment on the MLT MnS user side and the ease of computing power deployment on the MLT MnS provider side (e.g., OAM), model training is typically completed by the MLT MnS provider. As shown in Figure 2, the MLT MnS user sends a training request to the MLT MnS provider to initiate the ML entity (model) training process. The MLT MnS user also provides historical or currently available training data for the MLT MnS provider to select. The MLT MnS provider responds to the MLT MnS user's training request, selects training data for model training, and, upon completion, sends the training results to the MLT MnS user for selection and use.
[0176] As can be seen from the above, the ML model training process requires MLTMnS users to send training data to MLTMnS providers for selection by MLTMnS providers during training. However, large-scale model training requires massive amounts of data. If the current ML model training method is directly applied, the data collection process will result in large communication resource overhead between network elements (especially between MLTMnS users and MLTMnS providers). For example, in a reliable wireless transmission scenario with high time sensitivity, data sampling is performed on users in a 100,000-user wireless network with a period of 10ms. Assuming that the number of required user features is 100 and each feature requires 4 bits to represent, the amount of user data required to be uploaded per minute in this network is approximately 4GB, resulting in large communication resource overhead.
[0177] In view of this, the present application proposes a model training method and device, which can avoid the MLTMnS provider directly loading the massive training data of the MLTMnS user during the large model training process, thereby saving the communication overhead caused by the data collection process.
[0178] To facilitate understanding of the embodiments of the present application, the following points are explained before introducing the embodiments of the present application.
[0179] In this application, "used to indicate" or "indicate" can include direct indication and indirect indication, or "used to indicate" or "indicate" can indicate explicitly and / or implicitly. For example, when describing that a certain information is used to indicate information I, it can include that the information directly indicates I or indirectly indicates I, but it does not mean that the information necessarily carries I. For another example, implicit indication can be based on the location and / or resources used for transmission; explicit indication can be based on one or more parameters, and / or one or more indexes, and / or one or more bit patterns represented by it.
[0180] The definitions of many characteristics listed in this application are only used to explain the functions of the characteristics by way of example. For details, please refer to the prior art.
[0181] In the embodiments shown below, the first, second, third, fourth, and various numbers are only used for the convenience of description and are not intended to limit the scope of the embodiments of the present application. For example, they are used to distinguish different fields, different information, etc.
[0182] "Pre-definition" can be achieved by pre-storing corresponding codes, tables, or other methods that can be used to indicate relevant information in the device. This application does not limit the specific implementation method. Here, "storage" can mean storing in one or more memories. The type of memory can be any form of storage medium, which is not limited by this application.
[0183] The “protocol” involved in the embodiments of the present application may refer to a standard protocol in the field of communications, for example, it may include a long term evolution (LTE) protocol, a new radio (NR) protocol, and related protocols used in future communication systems, which are not limited in this application.
[0184] This application will present various aspects, embodiments, or features around systems including multiple devices, components, modules, etc. It should be understood and appreciated that each system may include additional devices, components, modules, etc., and / or may not include all of the devices, components, modules, etc. discussed in conjunction with the figures. Furthermore, combinations of these aspects may also be used.
[0185] In the embodiments of this application, words such as "exemplary," "for example," "illustratively," and "as another example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as an "exemplary" in this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner.
[0186] The terms "include", "comprising", "having" and variations thereof mean "including but not limited to", unless specifically emphasized otherwise.
[0187] "At least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c. Where a, b and c can be single or multiple, respectively.
[0188] In the embodiments of the present application, the descriptions of network element A sending a message, information or data to network element B, and network element B receiving a message, information or data from network element A are intended to illustrate to which network element the message, information or data is to be sent, but do not limit whether they are sent directly or indirectly via other network elements.
[0189] In the embodiments of the present application, descriptions such as "when...", "in the case of...", "if" and "if" all mean that the device will perform corresponding processing under certain objective circumstances. It does not limit the time, nor does it require the device to perform judgment actions when implemented, nor does it mean that there are other limitations.
[0190] It should be understood that in the embodiments of the present application, the base station and the access network device may be the same concept and may be used interchangeably.
[0191] FIG3 is a schematic diagram of a model training method 300 provided in an embodiment of the present application. As shown in the figure, the method 300 includes the following steps:
[0192] S310: A first device sends first information to a second device, and the second device receives the first information. The first information is used to indicate a first model, and the first model is used to generate a training data set for updating a second model.
[0193] In an embodiment of the present application, the first device may be a user of a network management service (e.g., a communication infrastructure big model), the second device may be a network management service provider, the first model may be a data generation model, and the second model may be an AI / ML model that implements one or more communication functions in a physical network, or may be called a communication infrastructure big model.
[0194] It should be understood that the embodiment of the present application does not limit the specific manner in which the first device obtains the second model.
[0195] As an example and not a limitation, the second model may be pre-configured in the first device. Specifically, the second model may be configured in the first device when the first device leaves the factory, or the second model may be manually configured in the first device before the first device is put into operation.
[0196] As an example and not a limitation, the second model may be sent by the second device to the first device. Specifically, before the above step S310, the second device sends configuration information to the first device, where the configuration information is used to indicate the second model.
[0197] Optionally, the second device may also send at least one of the performance information of the second model, the evaluation index information of the second model, or the functional information of the second model to the first device, so that the first device can determine whether the second model meets the corresponding performance requirements during operation. The performance information of the second model may include information such as the parameter quantity, structure, and number of layers of the second model; the evaluation index information of the second model may include the performance of the second model in its applicable downstream tasks, such as accuracy, recall rate, mean square error, etc.; the functional information of the second model may include information about communication functions that can be implemented by the model.
[0198] It should be understood that the embodiment of the present application does not limit the specific manner in which the first information indicates the first model.
[0199] As an example but not a limitation, the first information includes the first model, and the second device can obtain the first model after receiving the first information.
[0200] As an example but not limitation, the first model may include a storage address of the first model. Then, after receiving the first information, the second device may obtain the first model according to the storage address of the first model.
[0201] Optionally, when the first device sends the first information to the second device, the first device may also send at least one item of performance information of the first model, evaluation index information of the first model, or function information of the first model to the second device.
[0202] In some possible implementations, the first device may send third information to the second device, where the third information is used to indicate a method for using the first model. The method for using the first model includes generating a training data set based on the first model, or performing model fusion based on the first model.
[0203] It should be understood that the embodiment of the present application does not limit the specific manner in which the third information indicates the method of using the first model.
[0204] As an example and not a limitation, the third information can indicate the method of using the first model through bit 0 and bit 1. For example, bit 0 is used to indicate that a training data set is generated based on the first model, and bit 1 is used to indicate model fusion based on the first model.
[0205] It should be understood that the first device can send the third information at the same time as sending the above-mentioned first information, for example, carrying the first information and the third information in the same message, or the first device can send the third information after sending the above-mentioned first information. The embodiment of the present application is not limited to this.
[0206] It is easy to understand that, based on the usage of the first model, before the step S310, the method 300 further includes the following steps:
[0207] Method 1: Generate a training dataset based on the first model:
[0208] S305: The first device performs model training based on all or part of the local data to generate the first model.
[0209] Among them, all or part of the data of the first device includes at least a first data set, which is a collection of data related to a first function, and the first function includes one or more functions that the first device expects to implement through the second model.
[0210] As an example and not a limitation, the first function may include functions such as beam management and channel measurement required by the first device, or functions such as intent management expected by the first device, or functions such as mobility prediction expected by the first device.
[0211] It should be understood that the embodiment of the present application does not limit the triggering conditions for the first device to execute step S305.
[0212] In one possible implementation, the first device periodically performs model training based on local data. The embodiment of the present application does not limit the period of model training performed by the first device. Exemplarily, the period may be 500 seconds.
[0213] In another possible implementation, before step S305, the method 300 further includes the following steps (not shown):
[0214] S301: The first device fine-tunes parameters of the second model based on a second data set according to a first function, wherein the second data set is included in the first data set.
[0215] It should be understood that the second data set includes data related to the above-mentioned first function, and the second data set being included in the first data set can be understood as the time period for collecting the second data set being included in the time period for collecting the first data set.
[0216] As an example but not limitation, the first device is a base station, and the first device can fine-tune the parameters of the second model based on the second data set according to the required beam management, channel measurement and other functions.
[0217] As an example but not limitation, the first device is an OAM, and the first device can fine-tune the parameters of the second model based on the second data set according to functions such as intent management.
[0218] As an example but not limitation, the first device is an NWDAF, and the first device can fine-tune the parameters of the second model based on the second data set according to functions such as mobility prediction.
[0219] S302: The first device determines whether the performance of the fine-tuned second model meets the key performance indicator (KPI) requirement corresponding to the first function.
[0220] It should be understood that the above KPI requirements are determined by the first device during the process of fine-tuning the second model.
[0221] In a possible implementation, if at least one performance of the first model fine-tuned in step S302 does not meet the corresponding KPI requirement, the first device executes step S305.
[0222] In another possible implementation, if the performance of the first model after fine-tuning through the above step S302 meets the corresponding KPI requirements, the first device repeats the above steps S301 and S302 in sequence. In this way, the first device can fine-tune the parameters of the first model according to the first function and local data, thereby improving the performance of the model running locally on the first device.
[0223] Method 2: Perform model fusion based on the first model:
[0224] S307: The first device fine-tunes the parameters of the second model based on all or part of the local data according to the first function to generate the first model.
[0225] In a possible implementation, the second model whose parameters are fine-tuned by the first device in step S307 may be a model obtained by the first device performing lightweight processing on the second model according to the first function.
[0226] It should be understood that all or part of the local data used by the first device includes data related to the first function.
[0227] It should be understood that all or part of the local data may be collected and stored locally by the first device itself, or may be obtained from other devices and then stored locally. This embodiment of the present application does not limit this.
[0228] Based on the above-mentioned method 1 or method 2, the first device can generate a first model and send the first model through step S310.
[0229] S320: The second device updates the first model according to the second model.
[0230] In one possible implementation, when the third information indicates that the first model is used to generate a training data set, the second device generates a training data set for updating the first model based on the second model, and performs model training on the second model based on the training data set, thereby updating the second model.
[0231] In another possible implementation, when the third information indicates that the first model is to be fused with the second model, the second device fuses the second model with the first model to update the second model. The specific manner in which the second model and the first model are fused will be described in detail below and is not detailed here.
[0232] S330: The second device sends second information to the first device, and the first device receives the second information accordingly, wherein the second information is used to indicate the updated second model.
[0233] It should be understood that the embodiment of the present application does not limit the specific manner in which the second information indicates the updated second model.
[0234] As an example but not a limitation, the second information includes the updated second model, and the first device can obtain the updated second model after receiving the first information.
[0235] As an example but not limitation, the second information may include the updated storage address of the second model. After receiving the second information, the first device may obtain the first model according to the updated storage address of the second model.
[0236] Optionally, the first device may also send at least one item of performance information of the second model, evaluation index information of the second model, or function information of the second model to the second device.
[0237] It is easy to understand that the first device in steps S310 to S330 above can collect data and perform model training based on the collected data. In some possible embodiments, the first device is a MIF, that is, the training data of the first network element is provided by other devices (such as a base station), and the first device itself is only used for data training.
[0238] In these embodiments, the method 300 may further include the following steps (not shown):
[0239] S301′: The first device fine-tunes the parameters of the second model based on the third data set according to the second function.
[0240] It should be understood that the fourth data set is a collection of data related to the second function, and the second function includes one or more functions that the base station expects to implement through the second model.
[0241] It should be understood that the fourth data set is obtained by the first device through a request from the base station, that is, before executing the above step S301 ′, the first device needs to request the base station for the data collected by the base station.
[0242] S302': the first device sends the fine-tuned second model to the base station, and correspondingly, the base station receives the fine-tuned second model.
[0243] S303': The first device receives fourth information from the base station, where the fourth information is used to indicate that the performance of the fine-tuned second model is abnormal.
[0244] In these embodiments, the above step S305 can be replaced by:
[0245] S305': The first device performs model training based on all or part of the data collected by the base station to generate the first model.
[0246] The entire or partial data collected by the base station includes at least a fourth data set, which is a collection of data related to the second function. Furthermore, the fourth data set includes the third data set, i.e., the time period for collecting the fourth data set includes the time period for collecting the third data set.
[0247] It should be understood that the data for the first device to perform model training may be the data requested by the first device from the base station when the first device determines that the first model needs to be generated, or the data for the first device to perform model training may be the data reported by the base station and stored locally by the first device. This embodiment of the present application is not limited to this.
[0248] In some possible embodiments, before the first device executes S305 ′, the first device needs to request the base station for all or part of the data collected by the base station.
[0249] The embodiment of the present application does not limit the specific manner in which the first device requests the base station to request all or part of the data collected by the base station.
[0250] As an example but not limitation, the first device periodically sends fourth information to the base station, where the fourth information is used to request all data collected by the base station, or the fourth information is used to request data collected by the base station that is related to the second function.
[0251] It should be understood that all the data collected by the aforementioned base station may be all the data collected after the base station is put into operation, or all the data collected by the aforementioned base station may be all the data collected by the base station during the interval between receiving two request messages, or all the data collected by the aforementioned base station may be all the data collected within a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not limit this.
[0252] It should be understood that the data related to the second function collected by the aforementioned base station may be all the data related to the second function collected after the base station is put into operation, or, the data related to the second function collected by the aforementioned base station may be the data related to the second function collected by the base station during the interval between receiving two request messages, or, the data related to the second function collected by the aforementioned base station may be the data related to the second function collected within a specific time period agreed upon by the first device and the base station. The embodiments of the present application do not limit this.
[0253] It should be understood that the first device may discard part or all of the data provided by the base station after model training, or the first device may save part or all of the data provided by the base station after model training. This embodiment of the present application is not limited to this.
[0254] It should be understood that the embodiment of the present application does not limit the triggering conditions for the first device to execute step S305'.
[0255] In a possible implementation, the triggering condition of step S305 ′ is receiving the third information.
[0256] Furthermore, after the first device receives the third information, the first device may perform model training based on the locally stored data collected by the base station to generate a first model, or the first device may send a fourth information to the base station to request the data collected by the base station, and perform model training based on the data to generate a first model.
[0257] In another possible implementation, the first device periodically performs model training based on local data. The embodiment of the present application does not limit the period of model training performed by the first device. Exemplarily, the period may be 500 seconds.
[0258] Based on the above scheme, the first device sends the first model to the second device, so that the second device can generate a training data set for updating the second model based on the first model, and update the first model based on the training data set. That is, the first device can save the communication overhead caused by the collection of training data for the second model between the first device and the second device by sending the second model instead of training data.
[0259] The following describes the specific process of the above method 300 through the scenario in which the first device shown in Figure 4 is capable of data collection and model training. In this scenario, the first device can be any one of a base station, an OAM, and an NWDAF. The following description takes the first device as a base station and the second device as a device supporting cloud services (hereinafter referred to as Cloud) as an example.
[0260] FIG4 is a schematic diagram of a model training process 400 provided in an embodiment of the present application. As shown in the figure, the process 400 includes the following steps:
[0261] S401, Cloud sends information #1 to one or more base stations, and correspondingly, the one or more base stations receive the information #1.
[0262] The information #1 is used to indicate model #1 (an example of the second model). For example, the model #1 may be a wireless communication infrastructure large model. In the embodiment of the present application, the above step S401 may be understood as the Cloud making the wireless communication infrastructure large model available at the base station.
[0263] It should be understood that the embodiment of the present application does not limit the specific manner in which the information #1 indicates the model #1. The specific manner can be referred to the relevant content of step S310 and will not be repeated here.
[0264] Optionally, when sending information #1, Cloud may also send at least one of the model performance information of model #1, the evaluation index information of model #1, and the model function information of model #1 to the one or more base stations. For specific descriptions of the above information, please refer to the relevant content of step S310 and will not be repeated here.
[0265] S402 , each base station uses all or part of the locally collected data to fine-tune the parameters of model #2 according to the required functions.
[0266] In a possible implementation, the model #2 may be a model determined by the base station through information #1.
[0267] In another possible implementation, model #2 may be a model obtained by performing lightweight processing on model #1 according to the required function (i.e., the first function) of the base station. The required function of the base station may be understood as one or more functions that the base station expects to implement through model #1.
[0268] It should be understood that all or part of the above-mentioned locally collected data includes at least data set #1, which is a collection of data collected by the base station and related to its own required functions.
[0269] As an example and not a limitation, the functions required by the model may include beam management, channel measurement, and other functions required by the base station.
[0270] In an embodiment of the present application, the model #2 may be an on-site model.
[0271] It should be understood that in the process of the base station fine-tuning the parameters of the model #2, the KPI requirements corresponding to the required functions of the base station can be determined.
[0272] S403 , each base station determines whether the performance of the fine-tuned model #2 meets the KPI requirements corresponding to the functions required by the base station.
[0273] In one possible implementation, if the performance of model #2 meets the KPI requirements of the functions required by the above-mentioned base station, the base station repeats the above-mentioned steps S402 and S403. In this way, the base station can fine-tune the parameters of model #2 according to its own required functions and local data, thereby improving the performance of the model locally deployed by the base station.
[0274] In another possible implementation, if the performance of model #2 fails to meet at least one KPI requirement of the function required by the base station, the base station executes the following step S404.
[0275] S404: The base station uses all or part of the local data to perform model training to generate model #3.
[0276] The entire or partial local data of the base station includes at least data set #2, and the data set #2 is a collection of data collected by the base station and related to the functions required by the base station.
[0277] In an embodiment of the present application, the model #3 may be a data generation model.
[0278] In step S405 , the base station sends information #2 (an example of the first information) and information #3 (an example of the third information) to Cloud. In response, Cloud receives information #2 and information #3.
[0279] The information #2 is used to indicate the model #3, and the information #3 is used to indicate that the model #3 is used to generate the training data set. In the embodiment of the present application, the above step S405 can be understood as the base station making the model #3 available on the Cloud side.
[0280] In some possible implementations, the information #2 and information #3 may be included in a wireless large model training request message, which is used to request an update of the model #1.
[0281] Optionally, the base station may send at least one of the model performance information of model #3, the evaluation index information of model #3, and the model function information of model #3 to the Cloud when sending information #2.
[0282] Optionally, before sending information #2, the base station may perform lightweight processing on the model #3 generated in step S404, and indicate the lightweighted model #3 through information #2.
[0283] S406, Cloud updates model #1 based on model #3 obtained from one or more base stations.
[0284] Specifically, Cloud generates a training data set based on model #3 obtained from one or more base stations, and performs model training on model #1 based on the training data set to update model #1.
[0285] S407, Cloud sends information #4 (an example of the second information) to one or more base stations, and correspondingly, the one or more base stations receive information #4.
[0286] The information #4 is used to indicate the updated model #1. In the embodiment of the present application, the above step S407 can be understood as the Cloud making the updated model #1 available on the base station side.
[0287] Optionally, when sending information #4, Cloud may also send at least one of the updated model performance information of model #1, the updated evaluation index information of model #1, and the updated model function information of model #1 to the one or more base stations.
[0288] Based on the above solution, the base station / OAM / NWDAF can generate and send Model #3 based on local data to the Cloud. The Cloud can then generate a training dataset for updating Model #1 based on Model #3 and update Model #1 based on the training dataset. That is, the base station / OAM / NWDAF can save the communication overhead related to collecting training data for Model #1 between the base station / OAM / NWDAF and the Cloud by sending Model #3, which is used to generate the training dataset for updating Model #1, instead of the training data, to the Cloud.
[0289] The specific process of the above method 300 is described below through the scenario in which the first device shown in Figure 5 is used to provide computing power for other devices (for example, for model training). The following description is based on the example that the first device is MIF and the second device is a device that supports cloud services (hereinafter referred to as Cloud).
[0290] FIG5 is a schematic diagram of a model training process 500 provided in an embodiment of the present application. As shown in the figure, the process 500 includes the following steps:
[0291] S501, Cloud sends information #1' to one or more MIFs, and correspondingly, the one or more MIFs receive information #1'.
[0292] The information #1' is used to indicate the model #1' (another example of the second model). For example, the model #1' can be a wireless communication basic large model. In the embodiment of the present application, the above step S401 can be understood as the Cloud making the wireless communication basic large model available on the MIF side.
[0293] It should be understood that the embodiment of the present application does not limit the specific manner in which the information #1' indicates the model #1'. The specific manner can be referred to the relevant content of step S310 and will not be described in detail here.
[0294] Optionally, when sending information #1', Cloud may also send at least one of model performance information of model #1', evaluation index information of model #1', and model function information of model #1' to the one or more base stations.
[0295] It should be understood that each of the above-mentioned MIFs can be used to provide computing power for the corresponding one or more base stations, that is, each MIF can be used to perform model training for the corresponding one or more base stations, and this application does not limit this.
[0296] As an example and not a limitation, Cloud sends the above information #1' to MIF#1 and MIF#2, where MIF#1 can provide computing power for base station #1 and base station #2, and MIF#2 can provide computing power for base station #3.
[0297] It should be understood that before the above-mentioned step S501, model #2' is configured on the base station corresponding to the one or more MIFs. The model #2' may be the above-mentioned model #1, or the model #2' may be the model #1' after parameter fine-tuning based on the required functions of the base station. The present application does not limit the specific method of configuring the model #2' for the base station corresponding to the one or more MIFs. The specific method can be referred to the relevant content of step S310 and will not be repeated here.
[0298] S502: Each MIF sends information #5 to one or more corresponding base stations, and the one or more base stations receive information #5 accordingly. The information #5 is used to request data collected by the one or more base stations.
[0299] In one possible implementation, the information #5 includes at least one of an identifier #1 and an identifier #2, wherein the identifier #1 is used to indicate a specific time period, so that each base station can report data collected within the specific time period based on the identifier #1, and the identifier #2 is used to indicate a specific function, so that each base station can report data related to the specific function based on the identifier #2.
[0300] S503: Each base station sends all or part of the local data to the corresponding MIF.
[0301] Specifically, each base station samples local data based on information #5 and sends the sampled training data to the corresponding MIF. The total or local data sent by the base station to the corresponding MIF includes at least dataset #3, which is data related to the functions required by the base station. The functions required by the base station can be understood as one or more functions that the base station expects to implement using model #1.
[0302] It should be understood that when the information #5 includes the above-mentioned identifier #1, the base station will sample local data according to the identifier #1 and send all or part of the data collected within the specific time period indicated by the identifier #1 to the corresponding MIF.
[0303] It should be understood that when the information #5 includes the above-mentioned identifier #2, the base station will sample local data according to the identifier #2 and send all or part of the data related to the specific function indicated by the identifier #2 to the corresponding MIF.
[0304] S504: The MIF fine-tunes the parameters of the model #1' using all or part of the local data provided by the base station according to the required functions of the base station.
[0305] Specifically, the specific process of MIF fine-tuning the model #1' can be referred to the relevant content of method 300, which will not be described in detail here.
[0306] S505 , each MIF sends information # 6 to one or more corresponding base stations, where the information # 6 is used to indicate the fine-tuned model # 1 ′.
[0307] Correspondingly, after receiving the fine-tuned model #1', the one or more base stations update the local deployment.
[0308] S506 , each base station determines whether the fine-tuned model # 1 ′ meets the KPI requirements corresponding to the required functions.
[0309] In some possible implementations, if the performance of model #1' meets the KPI requirements of the functions required by the above-mentioned base station, then the above-mentioned steps S502 to S506 are repeated. In this way, the MIF can fine-tune the parameters of model #1' according to the functions required by the base station and the data collected by the base station, thereby improving the performance of the model locally deployed by the base station.
[0310] In another possible implementation, if the performance of the model #1' fails to meet at least one KPI requirement of the function required by the base station, the base station executes the following step S507.
[0311] S507, the base station sends information #7 (an example of the fourth information) to the corresponding MIF, and the corresponding MIF receives the information #7. The information #6 is used to indicate that the model deployed by the base station is abnormal.
[0312] S508, MIF uses all or part of the data provided by the corresponding one or more base stations to perform model training and generate model #3'.
[0313] The data provided by the one or more base stations, in whole or in part, includes at least data set #4, which is a collection of data related to the functions required by the base stations and reported by the one or more base stations. Furthermore, data set #4 includes data set #3, meaning that the time period during which data set #4 was collected includes the time period during which data set #3 was collected.
[0314] It should be understood that the embodiment of the present application does not limit the specific method in which the MIF obtains the data set #4.
[0315] As an example and not a limitation, the MIF stores the data reported by the base station locally after receiving it in the aforementioned step S503, so that model training is performed based on the locally stored data when executing step S508.
[0316] As an example and not a limitation, the MIF performs the following steps (not shown) before performing the above step S507:
[0317] S5081: MIF sends information #8 (an example of the fifth information) to one or more corresponding base stations, and the one or more base stations receive the information #8. The information #8 is used to request the base station to report the collected data.
[0318] Optionally, the information #8 includes at least one of the above-mentioned identifier #1 or identifier #2, so that the base station can send corresponding data based on the identifier #1 or identifier #2.
[0319] S5082: Each base station reports all or part of the local data to the MIF.
[0320] Optionally, the MIF may also periodically send information #8 to one or more corresponding base stations to request the base stations to report collected data.
[0321] Based on the above method, MIF can obtain data collected by the base station and perform model training based on the data to generate model #3'. Exemplarily, the model #3' is a data generation model.
[0322] S509 , MIF sends information # 2 ′ (another example of the second information) and information # 3 ′ (another example of the third information) to Cloud. Correspondingly, Cloud receives information # 2 ′ and information # 3 ′.
[0323] The information #2' is used to indicate the model #3', and the information #3' is used to indicate that the first model is used to be fused with the second model. In the embodiment of the present application, the above step S509 can be understood as the MIF making the model #3' available on the Cloud side.
[0324] In some possible implementations, the information #2' and the information #3' may be included in a wireless large model training request message, where the wireless large model training request message is used to request an update of the model #1'.
[0325] Optionally, the MIF may send at least one of the model performance information of model #3', the evaluation index information of model #3', and the model function information of model #3' to the Cloud when sending information #2'.
[0326] Optionally, before sending the information #2', the base station may perform lightweight processing on the model #3' generated in step S408, and indicate the lightweighted model #3' through the information #2'.
[0327] S510 , Cloud updates model # 1 ′ according to model # 3 ′ obtained from one or more MIFs.
[0328] Specifically, Cloud generates a training data set based on model #3' obtained from one or more MIFs, and performs model training on model #1' based on the training data set to update model #1'.
[0329] S511, Cloud sends information #4' to one or more MIFs, and correspondingly, the one or more base stations receive information #4'.
[0330] The information #4' is used to indicate the updated model #1'. In the embodiment of the present application, the above step S407 can be understood as the Cloud making the updated model #1' available on the MIF side.
[0331] Optionally, when sending information #4', Cloud may also send at least one of the updated model performance information of model #1', the updated evaluation index information of model #1', and the updated model function information of model #1' to the one or more MIFs.
[0332] Furthermore, the one or more MIFs may send the updated model #1 to the corresponding one or more base stations.
[0333] Optionally, before each MIF sends the updated model #1' to the corresponding one or more base stations, the MIF may also perform lightweight processing on the updated model #1' based on the functions required by the base station, and / or fine-tune the parameters of the updated model #1' using data collected by the base station based on the functions required by the base station. The MIF may obtain the data collected by the base station by repeating step S502 above.
[0334] Based on the above solution, MIF can send model #3' to Cloud based on the data reported by the base station, so that Cloud can generate a training data set for updating model #1' based on model #3', and update model #1' based on the training data set. That is, MIF can save the communication overhead related to the collection of training data for model #1' between MIF and Cloud by sending model #3' instead of training data to Cloud to generate the training data set for updating model #1'.
[0335] The following describes a model training method 600 provided in an embodiment of the present application in conjunction with Figure 6. In method 600, a first device can send a locally deployed model to a second device, so that the second device can update the second model based on the locally deployed model of the first device. In method 600, the first device can be any one of a base station, an OAM, an NWDAF, and a MIF. The following description takes the first device as a base station and the second device as a Cloud as an example.
[0336] FIG6 is a schematic diagram of a model training process 600 provided in an embodiment of the present application. As shown in the figure, the process 600 includes the following steps:
[0337] S601, Cloud sends information #8 to one or more base stations, and correspondingly, the one or more base stations receive the information #8.
[0338] The information #8 is used to indicate model #4 (an example of the second model). For example, the model #4 may be a wireless communication infrastructure large model. In the embodiment of the present application, the above step S601 can be understood as the Cloud making the wireless communication infrastructure large model available on the base station side.
[0339] It should be understood that the embodiment of the present application does not limit the specific manner in which the information #8 indicates the model #4. The specific manner can be referred to the relevant content of step S310 and will not be repeated here.
[0340] Optionally, when sending information #8, Cloud may also send at least one of the model performance information of model #4, the evaluation index information of model #4, and the model function information of model #4 to the one or more base stations. For specific descriptions of the above information, please refer to the relevant content of step S310 and will not be repeated here.
[0341] S602: Each base station uses all or part of the locally collected data to fine-tune the parameters of model #5 according to the required functions to generate model #6 (an example of the first model).
[0342] In a possible implementation, the model #5 may be a model determined by the base station through information #1.
[0343] In another possible implementation, model #5 may be a model obtained by performing lightweight processing on model #4 according to the required function (i.e., the first function) of the base station. The required function of the base station may be understood as one or more functions that the base station expects to implement through model #4.
[0344] It should be understood that all or part of the above-mentioned local data includes at least data set #5, which is a collection of data collected by the first device and related to its own required functions.
[0345] In an embodiment of the present application, the model #5 may be an on-site model.
[0346] In step S603 , the base station sends information #9 (an example of the first information) and information #10 (an example of the third information) to Cloud. In response, Cloud receives information #9 and information #10.
[0347] The information #9 is used to indicate the model #6, and the information #10 is used to indicate that the model #6 is used to be merged with the model #4. In the embodiment of the present application, the above step S405 can be understood as the base station making the model #6 available on the Cloud side.
[0348] In some possible implementations, the information #9 and the information #10 may be included in a wireless large model training request message, which is used to request an update of the model #4.
[0349] Optionally, the base station may send at least one of the model performance information of model #6, the evaluation index information of model #6, and the model function information of model #6 to the Cloud when sending information #9.
[0350] Optionally, before sending information #9, the base station may perform lightweight processing on the model #6 generated in step S404, and indicate the lightweight model #6 through information #9.
[0351] S604, Cloud updates model #4 based on model #6 obtained from one or more base stations.
[0352] Specifically, Cloud fuses model #6 obtained from one or more base stations with model #4 to update the second model.
[0353] It should be understood that the embodiment of the present application does not limit the fusion method of model #6 and model #4.
[0354] As an example and not a limitation, the model #6 and the model #4 can perform knowledge distillation, that is, the model parameters of the model #6 and the model #4 are directly updated by summing, averaging, and other calculation methods.
[0355] S605 , Cloud sends information #11 (an example of second information) to one or more base stations, and correspondingly, the one or more base stations receive information #11.
[0356] The information #11 is used to indicate the updated model #4. In the embodiment of the present application, the above step S605 can be understood as the Cloud making the updated model #4 available on the base station side.
[0357] Optionally, when sending information #11, Cloud may also send at least one of the updated model performance information of model #4, the updated evaluation index information of model #4, and the updated model function information of model #4 to the one or more base stations.
[0358] Based on the above solution, the base station / OAM / NWDAF / MIF can use all or part of the local data to fine-tune the parameters of Model #4 based on its required functions and generate Model #6. The Cloud can then fuse Model #6 with Model #4 to update Model #4. That is, the base station / OAM / NWDAF / MIF can send Model #6 to the Cloud for fusion with Model #4 instead of training data, saving the communication overhead associated with collecting training data for Model #4 between the base station / OAM / NWDAF / MIF and the Cloud.
[0359] The following is an introduction to the device embodiment corresponding to the method embodiment of the present application. The following is only a brief introduction to the device, and the specific implementation steps and details of the solution can be referred to the method embodiment above.
[0360] To implement the various functions of the method provided herein, the first device and the second device may each include hardware structures and / or software modules, and implement the aforementioned functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular one of the aforementioned functions is implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0361] FIG7 is a schematic diagram of a model training device 1000 provided in an embodiment of the present application. The device 1000 may include a transceiver unit 1010, a storage unit 1020, and a processing unit 1030. The transceiver unit 1010 is used to receive or send instructions and / or data, and the transceiver unit 1010 may also be referred to as a communication interface or communication unit; the storage unit 1020 is used to implement corresponding storage functions and store corresponding instructions and / or data; and the processing unit 1030 is used to perform data processing, so that the device 1000 implements the aforementioned model training method.
[0362] In a possible implementation, the apparatus 1000 may include only the transceiver unit 1010 and the processing unit 1030 , but not the storage unit 1020 .
[0363] As a design, the apparatus 1000 may execute the actions executed by the first device in the above method embodiment.
[0364] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010, for sending first information to a second device, the first information being used to indicate a first model, the first model being trained by the first device, the first model being used to generate a training data set for updating a second model, the second model being trained by the second device; the transceiver unit is also used to receive second information, the second information being used to indicate the updated second model.
[0365] In one possible implementation, the device 1000 also includes: a processing unit 1020, which is used to perform model training based on all or part of the local data to generate the first model, wherein all or part of the local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the model training device expects to implement through the second model.
[0366] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010, for sending first information to a second device, the first information being used to indicate a first model, the first model being trained by the first device, the first model being used to fuse with a second model to update the second model, the second model being trained by the second device; the transceiver unit is also used to receive second information, the second information being used to indicate the updated second model.
[0367] In one possible implementation, the device 1000 also includes: a processing unit 1030, which is used to fine-tune the parameters of the second model based on all or part of the local data according to the first function to generate the first model, where the first function is one or more functions that the first device expects to implement through the second model.
[0368] In one embodiment, the apparatus 1000 includes: a transceiver unit 1010, for sending first information to a second device, the first information being used to indicate a first model, the first model being trained by the first device, the first model being used to generate a training data set for updating a second model, or the first model being used to fuse with a second model to update the second model, the second model being trained by the second device; the transceiver unit is also used to receive second information, the second information being used to indicate the updated second model.
[0369] In one possible implementation, the device 1000 also includes: a processing unit 1030, which is used to perform model training based on all or part of the local data to generate the first model, wherein all or part of the local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the model training device expects to implement through the second model.
[0370] In one possible implementation, the processing unit 1030 is also used to fine-tune the parameters of the second model based on all or part of the local data according to the first function to generate the first model, where the first function is one or more functions that the first device expects to implement through the second model.
[0371] As a design, the apparatus 1000 may execute the actions executed by the second device in the above method embodiment.
[0372] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the transceiver unit 1010 is used to receive first information, the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to generate a training data set for updating the second model, the second model is trained by the first device; the processing unit 1030 is used to update the second model based on the first model; the transceiver unit 1010 is also used to send second information to the first device, the second information is used to indicate the updated second model.
[0373] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the transceiver unit 1010 is used to receive first information, the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to be integrated with a second model to update the second model, the second model is trained by the first device; the processing unit 1030 is used to update the second model based on the first model; the transceiver unit 1010 is also used to send second information to the first device, the second information is used to indicate the updated second model.
[0374] In one embodiment, the device 1000 includes: a transceiver unit 1010 and a processing unit 1030, the transceiver unit 1010 is used to receive first information, the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to generate a training data set for updating the second model, or the first model is used to merge with the second model to update the second model, the second model is trained by the first device; the processing unit 1030 is used to update the second model based on the first model; the transceiver unit 1010 is also used to send second information to the first device, the second information is used to indicate the updated second model.
[0375] FIG8 is a schematic diagram of another model training device 1100 provided in an embodiment of the present application.
[0376] The device 1100 includes a memory 1110, a processor 1120, and a communication interface 1130. The memory 1110, processor 1120, and communication interface 1130 are connected via an internal connection path. The memory 1110 is used to store instructions, and the processor 1120 is used to execute the instructions stored in the memory 1110 to control the communication interface 1130 to obtain information or enable the device 1100 to implement the aforementioned model training method. Optionally, the memory 1110 can be coupled to the processor 1120 via an interface or integrated with the processor 1120.
[0377] It should be noted that the communication interface 1130 may be a transceiver device such as, but not limited to, a transceiver. The communication interface 1130 may also include an input / output interface.
[0378] The processor 1120 stores one or more computer programs, which include instructions. When the instructions are executed by the processor 1120, the device 1100 executes the model training method in each of the above embodiments.
[0379] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 1120 or by instructions in the form of software. The method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1110, and the processor 1120 reads the information in the memory 1110 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0380] In one possible implementation, the apparatus 1100 may include only the processor 1120 and the communication interface 1130 , but not the memory 1110 .
[0381] Optionally, the communication interface 1130 in FIG. 8 may implement the transceiver unit 1010 in FIG. 7 , and the processor 1120 in FIG. 8 may implement the processing unit 1030 in FIG. 7 .
[0382] An embodiment of the present application further provides a computer-readable storage medium storing a program code. When the computer program code is executed on a computer, the computer executes any one of the methods in FIG. 3 to FIG. 7 .
[0383] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed, the computer executes any one of the methods in Figures 3 to 7 above.
[0384] An embodiment of the present application further provides a chip, comprising: a circuit, wherein the circuit is used to execute any one of the methods in FIG. 3 to FIG. 7 above.
[0385] An embodiment of the present application also provides a system, including: a first device and a second device, the first device is used to execute the actions / steps executed by the first device in Figures 3 to 7; the second device is used to execute the actions / steps executed by the second device in Figures 3 to 7.
[0386] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0387] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0388] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0389] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0390] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0391] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0392] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A model training method, characterized in that: include: The first device sends first information to the second device, where the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to generate a training data set for updating a second model, or the first model is used to merge with the second model to update the second model, and the second model is trained by the second device; The first device receives second information from the second device, where the second information is used to indicate the updated second model.
2. The method according to claim 1, characterized in that: The method further comprises: The first device sends third information to the second device, where the third information is used to indicate how the first model is used, where the way the first model is used includes generating the training data set based on the first model, or performing model fusion based on the first model.
3. The method according to claim 1 or 2, characterized in that: When the first model is used to generate a training data set for updating a second model, before the first device sends the first information to the second device, the method further includes: The first device performs model training based on all or part of the local data to generate the first model, wherein all or part of the local data includes a first data set, and the first data set is a collection of data related to a first function, and the first function is one or more functions that the first device expects to implement through the second model.
4. The method according to claim 3, characterized in that Before the first device performs model training based on all or part of the local data, the method further includes: The first device fine-tunes the parameters of the second model based on a second data set according to the first function, the second data set includes data related to the first function, and the second data set is included in the first data set; The first device determines that the performance of the fine-tuned second model related to at least one of the first functions does not meet the corresponding performance indicator requirements.
5. The method according to claim 1 or 2, characterized in that: When the first model is used to be merged with the second model to update the second model, before the first device sends the first information to the second device, the method further includes: The first device fine-tunes the parameters of the second model based on all or part of the local data according to a first function to generate the first model, where the first function is one or more functions that the first device expects to achieve through the second model.
6. The method according to claim 1 or 2, characterized in that: The first device is a mobile intelligent network element, and the second device is a device supporting cloud services.
7. The method according to claim 6, characterized in that The method further comprises: The first device acquires a third data set, where the third data set includes data collected by the base station and related to a second function, where the second function includes one or more functions that the base station expects to implement through the second model; The first device fine-tunes the parameters of the second model based on the third data set according to the second function.
8. The method according to claim 7, characterized in that Before the first device sends the first information to the second device, the method further includes: The first device receives fourth information, where the fourth information is used to indicate that performance of the fine-tuned second model is abnormal; The first device determines a fourth data set according to the fourth information, where the fourth data set is locally stored by the first device, or the fourth data set is acquired by the first device through the base station; The first device performs model training based on the fourth data set to generate the first model.
9. The method according to claim 8, characterized in that When the fourth data set is acquired by the first device through the base station, the first device determines the fourth data set according to the fourth information, including: The first device sends fifth information to the base station according to the fourth information, where the fifth information is used to request the base station to collect data; The first device receives sixth information, where the sixth information is used to indicate the fourth data set.
10. The method according to any one of claims 1 to 9, characterized in that Before the first device sends the first information to the second device, the method further includes: The first device performs lightweight processing on the first model.
11. A model training method, characterized in that: include: The second device receives first information from the first device, where the first information is used to indicate a first model, the first model is trained by the first device, the first model is used to generate a training data set for updating a second model, or the first model is used to merge with the second model to update the second model, and the second model is trained by the second device; The second device updates the second model according to the first model; The second device sends second information to the first device, where the second information is used to indicate the updated second model.
12. The method according to claim 11, characterized in that The method further comprises: The second device receives third information from the first device, where the third information is used to indicate how the first model is used, where the way the first model is used includes generating the training data set based on the first model, or performing model fusion based on the first model.
13. The method according to claim 11 or 12, characterized in that: When the first model is used to generate a training data set for updating a second model, the second device updates the second model according to the first model, including: The second device generates a training data set according to the first model; The second device trains the second model according to the training data set to update the second model.
14. The method according to claim 11 or 12, characterized in that: When the first model is used to be merged with the second model to update the second model, the second device updates the second model according to the first model, including: The second device fuses the first model and the second model to update the second model.
15. A model training method, characterized in that: include: The first device sends first information to the second device, where the first information is used to indicate a first model, the first model is obtained by training the first device, the first model is used to generate a training data set for updating a second model, and the second model is obtained by training the second device; The first device receives second information from the second device, where the second information is used to indicate the updated second model.
16. The method according to claim 15, characterized in that Before the first device sends the first information to the second device, the method further includes: The first device performs model training based on all or part of the local data to generate the first model, wherein all or part of the local data includes a first data set, and the first data set includes data related to a first function, and the first function is one or more functions that the first device expects to implement through the second model.
17. The method according to claim 16, characterized in that Before the first device performs model training based on all or part of the local data, the method further includes: The first device fine-tunes the parameters of the second model based on a second data set according to the first function, the second data set includes data related to the first function, and the second data set is included in the first data set; It is determined that the performance of the fine-tuned second model related to at least one of the first functions does not meet the corresponding performance indicator requirement.
18. The method according to any one of claims 15 to 17, characterized in that The first device is a mobile intelligent network element, and the second device is a device supporting cloud services.
19. The method according to claim 18, characterized in that The method further comprises: The first device acquires a third data set, where the third data set includes data collected by the base station and related to a second function, where the second function includes one or more functions that the base station expects to implement through the second model; The first device fine-tunes the parameters of the second model based on the third data set according to the second function.
20. The method according to claim 19, characterized in that Before the first device sends the first information to the second device, the method further includes: The first device receives fourth information, where the fourth information is used to indicate that performance of the fine-tuned second model is abnormal; The first device determines a fourth data set according to the fourth information, where the fourth data set is locally stored by the first device, or the fourth data set is acquired by the first device through the base station, wherein the fourth data set includes the third data set; The first device performs model training based on the fourth data set to generate the first model.
21. The method according to claim 20, characterized in that The method further comprises: The first device sends fifth information to the base station according to the fourth information, where the fifth information is used to request the base station to collect data; The first device receives sixth information, where the sixth information is used to indicate the fourth data set.
22. The method according to any one of claims 15 to 21, characterized in that Before the first device sends the first information to the second device, the method further includes: The first device performs lightweight processing on the first model.
23. A model training method, characterized in that: include: The second device receives first information, where the first information is used to indicate a first model, the first model is obtained by training the first device, the first model is used to generate a training data set for updating the second model, and the second model is obtained by training the second device; The second device updates the second model according to the first model; The second device sends second information to the first device, where the second information is used to indicate the updated second model.
24. The method according to claim 23, characterized in that The second device updates the second model according to the first model, including: The second device generates a training data set according to the first model; The second device trains the second model according to the training data set to update the second model.
25. A model training method, characterized in that: include: The first device sends first information to the second device, where the first information is used to indicate a first model, the first model is obtained by training the first device, the first model is used to be merged with a second model to update the second model, and the second model is obtained by training the second device; The first device receives second information, where the second information is used to indicate the updated second model.
26. The method according to claim 25, characterized in that The method further comprises: The first device fine-tunes the parameters of the second model based on all or part of the local data according to a first function to generate the first model, where the first function is one or more functions that the first device expects to achieve through the second model.
27. The method according to claim 26, characterized in that The whole or local data includes at least a second data set, which is a collection of data related to the first function.
28. A model training method, characterized in that: include: The second device receives first information from the first device, where the first information is used to indicate a first model, where the first model is trained by the first device, and the first model is used to be merged with a second model to update the second model, where the second model is trained by the first device; The second device updates the second model according to the first model; The second device sends second information to the first device, where the second information is used to indicate the updated second model.
29. The method according to claim 28, characterized in that The second device updates the second model according to the first model, including: the second device fuses the first model and the second model to update the second model.
30. A model training device, characterized in that: Comprising a module or unit for executing the method of any one of claims 1 to 10, or comprising a module or unit for executing the method of any one of claims 11 to 14, or comprising a module or unit for executing the method of any one of claims 15 to 22, or comprising a module or unit for executing the method of claim 23 or 24, or comprising a module or unit for executing the method of any one of claims 25 to 27, or comprising a module or unit for executing the method of claim 28 or 29.
31. A model training device, characterized in that: comprising a processor configured to, by executing a computer program or instructions, or, by means of a logic circuit, The model training device executes the method according to any one of claims 1 to 10, or The model training device executes the method according to any one of claims 11 to 14, or The model training device executes the method according to any one of claims 15 to 22, or, The model training device executes the method of claim 23 or 24, or The model training device executes the method according to any one of claims 25 to 27, or, The model training device executes the method described in claim 28 or 29.
32. The device according to claim 31, characterized in that The communication device further comprises a memory for storing the computer program or instructions.
33. The device according to claim 31 or 32, characterized in that The model training device also includes a communication interface, which is used to input and / or output signals.
34. A computer-readable storage medium, characterized in that: The computer readable storage medium stores a computer program or instruction. When the computer program or instruction is executed on a computer, so that the method of any one of claims 1 to 10 is performed, or, so that the method of any one of claims 11 to 14 is performed, or, so that the method of any one of claims 15 to 22 is performed, or, so that the method of claim 23 or 24 is performed, or, so that the method of any one of claims 25 to 27 is performed, or, The method of claim 28 or 29 is performed.
35. A computer program product, characterized in that Contains instructions that, when executed on a computer, so that the method of any one of claims 1 to 10 is performed, or, so that the method of any one of claims 11 to 14 is performed, or, so that the method of any one of claims 15 to 22 is performed, or, so that the method of claim 23 or 24 is performed, or, so that the method of any one of claims 25 to 27 is performed, or, The method of claim 28 or 29 is performed.
36. A system, characterized in that: include: A first device and a second device, wherein the first device is used to execute the method as described in any one of claims 1 to 10, 15 to 22, and 25 to 27; and the second device is used to execute the method as described in any one of claims 11 to 14, 23 or 24, 28 or 29.
Citation Information
Patent Citations
Model training method and device
CN120186040A
Model training method, device and system
CN115718868A
Communication method and device
CN115802370A
Wireless communication method and device
CN116249166A
Model training method and device and communication equipment
CN116432013A