Model training method and apparatus

By sending public data sets and sharing model weights between the central node and the worker, the local model that constrains workers has close results on the same public data set, solving the problems of model overfitting and user privacy leaks in the prior art, and achieving efficient model training and support from multiple types of workers.

WO2025108005A1PCT designated stage expired Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/127799
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-10-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing model training methods based on federated learning have the risk of model overfitting and user privacy leaks in communication systems, and cannot effectively support different types of workers.

Method used

By sending public datasets and sharing model weights between central nodes and workers, the local model that constrains workers to have close results on the same public dataset, avoiding model overfitting and supporting different types of workers.

Benefits of technology

It effectively improves the training effect of the model, avoids model overfitting, and supports different types of workers while protecting user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127799_30052025_PF_FP_ABST
    Figure CN2024127799_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and provides a model training method and apparatus. The method comprises: sending to at least one worker a first public data set, or the first public data set and a first shared model weight, the at least one worker comprising a first worker, the first shared model weight being used for determining a weight of a local model of the first worker, the first public data set being used for training the local model of the first worker, and / or the first public data set being used for determining a first data set, and the first data set being an output data set obtained after the first public data set is processed by the local model of the first worker. The method in embodiments of the present application is beneficial for improving the training effect of models.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method and device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 21, 2023, with application number 202311563129.9 and application name “Model Training Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and specifically to a method and device for model training. Background Art

[0003] With the rapid development of artificial intelligence (AI) technology, the application scenarios of AI models are becoming more and more extensive. For example, AI models can be introduced into certain communication systems to achieve network intelligence.

[0004] In communication systems, the private data of each device often involves user privacy. To protect user privacy and data security, federated learning (FL) can be used to train AI models. However, existing FL-based model training methods still have some problems.

[0005] Summary of the Invention

[0006] This application provides a method and device for model training, which helps to improve the training effect of the model.

[0007] In a first aspect, a model training method is provided, which is applied to a central node, including: sending a first public data set to at least one worker, or the first public data set and a first shared model weight, wherein the at least one worker includes a first worker, the first shared model weight is used to determine the weight of the first worker's local model, the first public data set is used to train the first worker's local model, and / or the first public data set is used to determine a first data set, which is an output data set obtained after the first worker's local model processes the first public data set.

[0008] In an embodiment of the present application, the central node sends a first public data set, or a first public data set and a first shared model weight to at least one worker, so as to facilitate the first worker to train a local model and / or determine the first data set based on the first public data set, and helps to constrain the local model of at least one worker to have similar working methods by constraining the local model of at least one worker to have similar results on the same public data set (such as the first public data set), thereby helping to avoid overfitting of the local model and improve the training effect of the model.

[0009] In some possible implementations, after sending the first public dataset to at least one worker, or the first public dataset and the first shared model weight, the method further includes: sending a second public dataset to the at least one worker, or the second public dataset and the second shared model weight, wherein the second shared model weight is used to determine the weight of the local model of the first worker, the second public dataset is used to train the local model of the first worker, and / or the second public dataset is used to determine a second dataset, which is an output dataset obtained after the local model of the first worker processes the second public dataset.

[0010] In some possible implementations, the second public data set is determined based on the first output data set fed back by the at least one worker, the second shared model weight is determined based on the first weight update value fed back by the at least one worker and / or the first output data set, the first output data set is the output data set obtained by the first worker using a local model to process the first public data set, and the first weight update value is the update value of the first shared model weight in the process of the first worker training the local model.

[0011] In some possible implementations, after sending the first public dataset to at least one worker, or the first public dataset and the first shared model weight, the method further includes: receiving a first output dataset from the first worker, where the first output dataset is an output dataset obtained by the first worker using a local model to process the first public dataset.

[0012] In some possible implementations, before receiving the first output data set from the first worker, the method further includes: sending first information to the first worker, where the first information is used to indicate a first time-frequency resource position for feeding back the first output data set to the central node.

[0013] In some possible implementations, when the first output data set includes explanation labels, the method further includes: sending indication information to the first worker for indicating the type of explanation labels in the first output data set, wherein the explanation labels in the first output data set are used to explain the processing of the output data in the first output data set by the worker's local model.

[0014] In some possible implementations, the first output data set includes output data, labels corresponding to the output data, and / or explanation labels corresponding to the output data, where the explanation labels are used to explain the processing of the output data by the worker's local model.

[0015] In some possible implementations, after sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: receiving a first weight update value from the first worker, the first weight update value being an update value of the first shared model weight during the process of the first worker training the local model.

[0016] In some possible implementations, before receiving the first weight update value from the first worker, the method further includes: sending second information to the first worker, where the second information is used to indicate a second time-frequency resource location for feeding back the first weight update value to the central node.

[0017] In some possible implementations, the method further includes: sending third information to the first worker, where the third information is used to indicate that the first worker needs to feed back a first output data set and / or a first weight update value.

[0018] In some possible implementations, before sending the first public dataset, or the first public dataset and the first shared model weight to at least one worker, the method further includes: receiving fourth information from the first worker, wherein the fourth information is used to indicate that the first worker needs to use the first public dataset and / or the first shared model weight.

[0019] In some possible implementations, before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: sending fifth information to the first worker, wherein the fifth information is used to instruct the central node to send a third time-frequency resource location of the first public data set.

[0020] In some possible implementations, before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: sending sixth information to the first worker, wherein the sixth information is used to indicate the fourth time-frequency resource location of the central node to send the first shared model weight.

[0021] In some possible implementations, the method further includes: sending instruction information to the first worker, which is used to instruct the first worker to use specific public data in the first public data set to determine the first data set.

[0022] In some possible implementations, when the first public dataset includes explanation tags, the method further includes: sending indication information to the first worker for indicating the type of explanation tags in the first public dataset, wherein the explanation tags in the first public dataset are used to explain the processing of the public data by the worker's local model.

[0023] In some possible implementations, the first public data set includes public data, labels corresponding to the public data, and / or explanation labels corresponding to the public data, where the explanation labels are used to explain the processing of the public data by the worker's local model.

[0024] In some possible implementations, the method further includes: sending the first shared model weight to a second worker, where the first shared model weight is further used to determine a weight of a local model of the second worker.

[0025] In some possible implementations, the method further includes: sending seventh information to the second worker, where the seventh information is used to instruct the central node to send a second time-frequency resource location of the first shared model weight.

[0026] In some possible implementations, after sending the first shared model weight to the second worker, the method further includes: sending a second shared model weight to the second worker, the second shared model weight being used to determine the weight of the local model of the second worker, the second shared model weight being determined based on the first weight update value and / or the first output data set fed back by the at least one worker, and the second weight update value fed back by at least one second worker, the first weight update value being an update value of the first shared model weight in the process of the first worker training the local model, the first output data set being an output data set obtained by the first worker using the local model to process the first public data set, and the second weight update value being an update value of the first shared model weight in the process of the second worker training the local model.

[0027] In some possible implementations, after sending the first shared model weight to the second worker, the method further includes: receiving a second weight update value from the second worker, where the second weight update value is an update value of the first shared model weight during the process of the second worker training the local model.

[0028] In some possible implementations, the method further includes: sending eighth information to the second worker, where the eighth information is used to indicate a fifth time-frequency resource position for feeding back a second weight update value to the central node.

[0029] In some possible implementations, the method further includes: sending ninth information to the second worker, where the ninth information is used to indicate that the second worker needs to feed back the second weight update value.

[0030] In some possible implementations, before sending the first shared model weight to the second worker, the method further includes: receiving tenth information from the second worker, where the tenth information is used to indicate that the second worker needs to use the first shared model weight.

[0031] In a second aspect, a model training method is provided, which is applied to a first worker and is characterized in that it includes: receiving a first public data set from a central node, or the first public data set and a first shared model weight, wherein the first shared model weight is used to determine the weight of the local model of the first worker, and the first public data set is used to train the local model of the first worker, and / or the first public data set is used to determine a first data set, which is an output data set obtained after the local model of the first worker processes the first public data set.

[0032] In an embodiment of the present application, the first worker receives a first public data set from a central node, or the first public data set and a first shared model weight, so that the first worker can train a local model and / or determine the first data set based on the first public data set, which helps to constrain the local model of at least one worker to have similar working methods by constraining the local model of at least one worker to have similar results on the same public data set (such as the first public data set), thereby helping to avoid overfitting of the local model and improve the training effect of the model.

[0033] In some possible implementations, after receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: receiving a second public data set from the central node, or the second public data set and the second shared model weight, wherein the second shared model weight is used to determine the weight of the local model of the first worker, the second public data set is used to train the local model of the first worker, and / or the second public data set is used to determine a second data set, which is an output data set obtained after the local model of the first worker processes the second public data set.

[0034] In some possible implementations, the second public data set is determined based on the first output data set fed back by at least one worker, the at least one worker includes the first worker, the second shared model weight is determined based on the first weight update value fed back by the at least one worker and / or the first output data set, the first output data set is the output data set obtained by the first worker using a local model to process the first public data set, and the first weight update value is the update value of the first shared model weight in the process of the first worker training the local model.

[0035] In some possible implementations, the second shared model weight is further determined according to a second weight update value fed back by at least one second worker.

[0036] In some possible implementations, after receiving the first public dataset from the central node, or the first public dataset and the first shared model weight, the method further includes: sending a first output dataset to the central node, where the first output dataset is an output dataset obtained by the first worker using a local model to process the first public dataset.

[0037] In some possible implementations, before sending the first output data set to the central node, the method further includes: receiving first information from the central node, where the first information is used to indicate a first time-frequency resource position for feeding back the first output data set to the central node.

[0038] In some possible implementations, when the first output data set includes explanation labels, the method further includes: receiving indication information from the central node for indicating the type of explanation labels in the first output data set, and the explanation labels in the first output data set are used to explain the processing of the output data in the first output data set by the worker's local model.

[0039] In some possible implementations, the first output data set includes output data, labels corresponding to the output data, and / or explanation labels corresponding to the output data, where the explanation labels are used to explain the processing of the output data by the worker's local model.

[0040] In some possible implementations, after receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: sending a first weight update value to the central node, where the first weight update value is an update value of the first shared model weight during the process of the first worker training the local model.

[0041] In some possible implementations, before sending the first weight update value to the central node, the method further includes: receiving second information from the central node, where the second information is used to indicate a second time-frequency resource location for feeding back the first weight update value to the central node.

[0042] In some possible implementations, the method further includes: receiving third information from the central node, where the third information is used to indicate that the first worker needs to feed back a first output data set and / or a first weight update value.

[0043] In some possible implementations, before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: sending fourth information to the central node, wherein the fourth information is used to indicate that the first worker needs to use the first public data set and / or the first shared model weight.

[0044] In some possible implementations, before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method also includes: receiving fifth information from the central node, and the fifth information is used to indicate the third time-frequency resource position of the central node to send the first public data set.

[0045] In some possible implementations, before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method also includes: receiving sixth information from the central node, the sixth information being used to indicate the fourth time-frequency resource location of the central node to send the first shared model weight.

[0046] In some possible implementations, when the first public data set includes explanation tags, the method further includes: receiving indication information from the central node for indicating the type of explanation tags in the first public data set, and the explanation tags in the first public data set are used to explain the processing of the public data by the local model of the explanation worker.

[0047] In some possible implementations, the method further includes: receiving instruction information from the central node for instructing the first worker to determine the first data set using specific public data in the first public data set.

[0048] In some possible implementations, the first public data set includes public data, labels corresponding to the public data, and / or explanation labels corresponding to the public data, where the explanation labels are used to explain the processing of the public data by the worker's local model.

[0049] In a third aspect, a model training device is provided, which is applied to a central node and includes: a sending unit, which is used to send a first public data set to at least one worker, or the first public data set and a first shared model weight, wherein the at least one worker includes a first worker, the first shared model weight is used to determine the weight of the local model of the first worker, the first public data set is used to train the local model of the first worker, and / or the first public data set is used to determine a first data set, which is an output data set obtained after the local model of the first worker processes the first public data set.

[0050] In an embodiment of the present application, the central node sends a first public data set, or a first public data set and a first shared model weight to at least one worker, so as to facilitate the first worker to train a local model and / or determine the first data set based on the first public data set, and helps to constrain the local model of at least one worker to have similar working methods by constraining the local model of at least one worker to have similar results on the same public data set (such as the first public data set), thereby helping to avoid overfitting of the local model and improve the training effect of the model.

[0051] In a fourth aspect, a model training device is provided, which is applied to a first worker, including: a receiving unit, used to receive a first public data set from a central node, or the first public data set and a first shared model weight, wherein the first shared model weight is used to determine the weight of the local model of the first worker, and the first public data set is used to train the local model of the first worker, and / or the first public data set is used to determine a first data set, which is an output data set obtained after the local model of the first worker processes the first public data set.

[0052] In an embodiment of the present application, the first worker receives a first public data set from a central node, or the first public data set and a first shared model weight, so that the first worker can train a local model and / or determine the first data set based on the first public data set, which helps to constrain the local model of at least one worker to have similar working methods by constraining the local model of at least one worker to have similar results on the same public data set (such as the first public data set), thereby helping to avoid overfitting of the local model and improve the training effect of the model.

[0053] In the fifth aspect, a device for model training is provided, comprising: a processor and a memory, wherein the processor is coupled to the memory, and the memory is used to store a computer program (also referred to as code or instructions). When the computer program is executed by the processor, the device executes the method in the first aspect or any possible implementation of the first aspect.

[0054] In some possible implementations, the apparatus further includes a memory coupled to the processor.

[0055] In some possible implementations, there are one or more processors and / or one or more memories.

[0056] In some possible implementations, the memory may be integrated with the processor, or the memory may be provided separately from the processor.

[0057] In the sixth aspect, a device for model training is provided, comprising: a processor and a memory, wherein the processor is coupled to the memory, and the memory is used to store a computer program (also referred to as code, or instructions). When the computer program is executed by the processor, the device executes the method in the second aspect or any possible implementation of the second aspect.

[0058] In some possible implementations, the apparatus further includes a memory coupled to the processor.

[0059] In some possible implementations, there are one or more processors and / or one or more memories.

[0060] In some possible implementations, the memory may be integrated with the processor, or the memory may be provided separately from the processor.

[0061] In the seventh aspect, a computer-readable storage medium is provided, which stores a computer program (also referred to as code, or instructions). When the computer program runs on a computer, the computer executes the method in any one of the above aspects or any possible implementation of any one of the aspects.

[0062] In an eighth aspect, a computer program product is provided, comprising: a computer program (also referred to as code, or instructions), which, when executed on a computer, enables the computer to execute a method in any one of the above aspects or any one of the possible implementations of any one of the aspects.

[0063] In the ninth aspect, a chip is provided, comprising: a processor and a memory, wherein the memory is used to store a computer program (also referred to as code, or instruction), and the processor is used to call and run the computer program stored in the memory, so that a device or equipment equipped with the chip executes the method in any one of the above aspects or any possible implementation of any one of the aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] FIG1 is a schematic block diagram of a wireless communication system applicable to the present application.

[0065] FIG2 is a schematic block diagram of another wireless communication system applicable to the present application.

[0066] FIG3 is a schematic structural diagram of a neural network in this application.

[0067] FIG4 is a schematic flowchart of a model training method provided in one embodiment of the present application.

[0068] FIG5 is a schematic diagram of various types of workers supported by an embodiment of the present application.

[0069] FIG6 is a schematic diagram of data feedback by various types of workers in one embodiment of the present application.

[0070] FIG7 is a schematic flowchart of a model training method provided in another embodiment of the present application.

[0071] FIG8 is a schematic flowchart of a model training method provided in yet another embodiment of the present application.

[0072] FIG9 is a schematic flowchart of a model training method provided in yet another embodiment of the present application.

[0073] FIG10 is a schematic structural diagram of a model training device provided in one embodiment of the present application.

[0074] FIG11 is a schematic structural diagram of a model training device provided in another embodiment of the present application.

[0075] FIG12 is a schematic structural diagram of a device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0076] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0077] In the description of this application, unless otherwise specified, " / " indicates that the objects associated with each other are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, in the description of this application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. In addition, to facilitate the clear description of the technical solutions of the embodiments of this application, in the embodiments of this application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily limit differences. It should be understood that in this application, similar expressions such as "under the circumstances of...", "if...", "when...", and "if..." can be used interchangeably.

[0078] The technical solutions of the embodiments of the present application can be applied to various communication systems, such as: fifth generation (5G) system or new radio (NR), long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD) system, etc. The technical solutions provided by the present application can also be applied to future communication systems, such as the sixth generation mobile communication system, satellite communication system, etc.

[0079] The communication system used in the embodiments of the present application may include a first entity and a second entity. The first entity and the second entity may communicate, for example, the first entity may send configuration information to the second entity and send data to the second entity, or receive data sent by the second entity; the second entity may receive the configuration information sent by the first entity and receive data sent by the first entity based on the configuration information, or send data to the first entity based on the configuration information.

[0080] The first entity may be the network device 110 in FIG. 1 , and the second entity may be the terminal device 120 or the terminal device 130 in FIG. 1 ; or, the first entity may be the terminal device 210 in FIG. 2 , and the second entity may be the terminal device 220 or the terminal device 230 in FIG. 2 ; or, the first entity and the second entity may be other network entities in the wireless communication system, which is not limited in the embodiments of the present application.

[0081] The wireless communication system 100 in FIG. 1 and the wireless communication system 200 in FIG. 2 are introduced below.

[0082] FIG1 illustrates a wireless communication system 100 used in an embodiment of the present application. The wireless communication system 100 may include a network device 110, a terminal device 120, and a terminal device 130. The network device 110 may provide communication coverage (or network coverage) for a specific geographic area. The terminal devices 120 and 130 may access a network (e.g., a wireless network) through the network device 110. The network device 110 may communicate with the terminal devices 120 and 130 via a Uu interface, and the terminal devices 120 and 130 may communicate via a PC5 interface.

[0083] Figure 1 exemplarily shows a network device and two terminal devices. Optionally, the wireless communication system 100 may include multiple network devices, and each network device may include a different number of terminal devices within its coverage area, which is not limited in the embodiments of the present application. Optionally, the wireless communication system 100 may also include other network elements or network entities such as a network controller and a mobility management entity, which is not limited in the embodiments of the present application.

[0084] FIG2 is a wireless communication system 200 used in an embodiment of the present application. The wireless communication system 200 may include a terminal device 210, a terminal device 220, and a terminal device 230. The terminal device 210 may provide a sidelink (SL) signal to the terminal device 220 and the terminal device 230. The terminal device 210 may communicate with the terminal device 220 and the terminal device 230 on the sidelink via the PC5 interface. The terminal device 220 and the terminal device 230 may also communicate on the sidelink via the PC5 interface. Optionally, one or more of the terminal device 210, the terminal device 220, and the terminal device 230 may be within the network coverage mentioned by the network device.

[0085] Figure 2 exemplarily shows three terminal devices. Optionally, the wireless communication system 200 may include multiple terminal devices that provide sidelink signals, and each terminal device may provide sidelink signals for another number of terminal devices, which is not limited in the embodiments of the present application. Optionally, the wireless communication system 200 may also include other network elements or network entities such as network devices, network controllers, and mobility management entities, which is not limited in the embodiments of the present application.

[0086] The terminal device in the embodiment of the present application may refer to user equipment (UE), station, access terminal, user unit, user station, mobile station, mobile station (MS), remote station, remote terminal, mobile terminal (MT), user terminal, terminal, wireless communication device, user agent or user device, etc., and this is not limited in the embodiment of the present application. The terminal device in the embodiment of the present application may also be a mobile phone, a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a large screen, a vehicle-mounted device, a wearable device, a terminal device in a 5G network or a terminal device in a future evolved public land mobile communication network (PLMN), etc., and this is not limited in the embodiment of the present application. The terminal device in the embodiments of the present application may also be a tablet computer (Pad), a laptop computer, a PDA, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc., and this is not limited in the embodiments of the present application.

[0087] In some embodiments, the terminal device can be configured to act as a base station. Optionally, the terminal device can act as a dispatching entity, providing sidelink signals between terminal devices in vehicle-to-everything (V2X) or device-to-device (D2D) communications. For example, a cell phone and a car can communicate using sidelink signals, or a cell phone and a smart home device can communicate using sidelink signals without relaying the communication signals through a base station.

[0088] The network device in the embodiment of the present application may refer to a radio access network (RAN) node (or device) that connects a terminal device to a wireless network, and may also be referred to as a base station. For example, the network device may be a NodeB, an evolved NodeB (eNodeB), a next generation NodeB (gNB) in a 5G mobile communication system, a transmission reception point (TRP), an access point (AP), a base station in a future mobile communication system or an access point (AP) in a WiFi system, a wireless controller in a cloud radio access network (CRAN) scenario, a relay station, an access point, a vehicle-mounted device, a wearable device, a network device in other communication systems that will evolve in the future, and the like.

[0089] In some embodiments, multiple RAN nodes may collaborate to assist terminal devices in achieving wireless access, and different RAN nodes may respectively implement part of the functions of a base station. For example, a RAN node (i.e., a network device in this application) may be a centralized unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU), etc. The CU and DU may be set separately, or may be included in the same network element, such as a baseband unit (BBU). The RU may be included in a radio frequency device or radio frequency unit, such as a remote radio unit (RRU), an active antenna unit (AAU), or a remote radio head (RRH). In different systems, CU (or CU-CP and CU-UP), DU, or RU may also have different names, but those skilled in the art may understand their meanings. For example, in an open radio access network (ORAN) system, CU may also be referred to as an open CU (O-CU), DU may also be referred to as an open DU (O-DU), CU-CP may also be referred to as O-CU-CP, CU-UP may also be referred to as O-CU-UP, and RU may also be referred to as O-RU. Any unit of the CU (or CU-CP, CU-UP), DU and RU in this application may be implemented by a software module, a hardware module, or a combination of a software module and a hardware module. It should be understood that this application does not limit the specific technology and specific device form adopted by the network device.

[0090] In some embodiments, the network device may be fixed or mobile, which is not limited in the embodiments of the present application. For example, a helicopter or drone may be configured as a mobile network device, and one or more cells may be moved based on the location of the mobile network device. In other examples, a helicopter or drone may be configured to communicate with another network device.

[0091] In some embodiments, network devices may be deployed on land or in the air, which is not limited in the embodiments of the present application. For example, network devices may be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; on water; or in the air on aircraft, balloons, and satellites.

[0092] In an embodiment of the present application, a terminal device or a network device may include a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system may be any one or more computer operating systems that implement business processing through processes, such as a Linux operating system, a Unix operating system, an Android operating system, an iOS operating system, or a Windows operating system. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Furthermore, in the embodiment of the present application, the specific structure of the execution subject of the method provided in the embodiment of the present application is not particularly limited, as long as it is possible to communicate according to the method provided in the embodiment of the present application by running a program that records the code of the method provided in the embodiment of the present application.

[0093] In addition, various aspects or features of the present application can be implemented as methods, devices or products using standard programming and / or engineering techniques. The term "product" as used in this application covers computer programs that can be accessed from any computer-readable device, carrier or medium. For example, computer-readable media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks or magnetic tapes, etc.), optical disks (e.g., compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks or key drives, etc.). In addition, the various storage media described herein may represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing and / or carrying instructions and / or data.

[0094] Before describing the embodiments of the present application, the relevant terms involved in the present application are first introduced.

[0095] 1. Artificial intelligence (AI): AI enables machines to learn, accumulate experience, and solve problems that humans can solve through experience, such as natural language understanding, image recognition, and chess.

[0096] 2. Machine Learning (ML): Machine learning is an implementation of artificial intelligence. It empowers machines to perform tasks that cannot be accomplished through direct programming. In practical terms, machine learning involves training models using data and then using these models to make predictions.

[0097] 3. Neural Network (NN): A neural network is a specific embodiment of machine learning. A neural network is a mathematical model that processes information by imitating the behavioral characteristics of animal neural networks. As shown in Figure 3, a neural network can be composed of three types of computational layers: an input layer, a hidden layer, and an output layer. Each layer can have one or more logical judgment units, which are called neurons. Common neural network structures include feedforward neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). These network structures are all based on neurons. Each neuron can perform a weighted sum operation on its input values ​​and apply this weighted summation to a nonlinear function to produce an output. Therefore, the weights of the weighted summation operation and the nonlinear function in the neural network can be called the parameters of the neural network. The connections between neurons in a neural network can be called the neural network weights, and the weights of all neurons in a neural network constitute the parameters of the neural network (also called the neural network weights).

[0098] 4. Deep neural network (DNN): A neural network with multiple hidden layers.

[0099] 5. Deep learning (DL): Machine learning using deep neural networks.

[0100] With the rapid development of AI technology, AI models (such as neural network models) are finding increasingly widespread application in a wide range of scenarios, including smart healthcare, finance, smart cities, and intelligent networks. For example, AI models can be introduced into certain communication systems to achieve network intelligence. Typically, AI models require training before use, so effectively training AI models for specific application scenarios has become an important research topic.

[0101] AI model training involves three key factors: training data, models, and computing platforms (such as neural network computing platforms). With the widespread deployment of graphics processing units (GPUs), more and more devices are equipped with computing platforms. Therefore, obtaining training data has become critical for training AI models on these devices. However, the private data contained in a single device is often limited. Using only limited private data for model training can easily lead to model overfitting. Furthermore, private data often involves user privacy. While obtaining private data from other devices to expand the training set can avoid model overfitting, it carries the risk of user privacy leakage.

[0102] Federated learning (FL) is a distributed model training method that enables multiple devices to collaborate on AI model training while protecting user privacy and data security. Therefore, training AI models based on FL can address the aforementioned issues.

[0103] For example, AI models can be trained using the following federated learning systems:

[0104] (1) Basic Federated Learning System

[0105] A basic federated learning system can include one central service node (also referred to as the central node) and N workers, where N is a positive integer greater than 1. The N workers can implement similar functions, and each of the N workers can contain certain private data that can be used for local model training. At the same time, the central service node can interact with each of the N workers.

[0106] The basic idea of ​​basic federated learning can be: each worker deploys a copy of the same public model (denoted as w), and each worker can use its own private data to update the deployed copy of the public model and obtain the updated value of the public model. For example, worker i can update the public model and obtain the updated value w of the public model. new,i =w+Δ i , i is a positive integer greater than or equal to 1; further, N workers can each update the public model Δ i Feedback to the central service node, which performs weighted average of these N update values ​​to obtain the average update value Furthermore, the central service node can send the average update value Δ to N workers, and these N workers can update their copies to w+Δ respectively.

[0107] After the replica is updated, the N workers may iteratively execute the above model update or local model training operation locally repeatedly until the model converges, or until other conditions are met.

[0108] It can be seen that in the basic federated learning system, different computing nodes (i.e., workers) fully share the model, and it is impossible to adapt the model to the local data (i.e., private data) in the computing nodes. When the private data in different computing nodes differ greatly, fully sharing the model parameters may cause the model's performance to be impaired. At the same time, the basic federated learning system requires that all computing nodes have the same model processing capabilities.

[0109] (2) Federated Learning System with Partial Model Sharing

[0110] A federated learning system with partial model sharing can include one central service node and N workers, where N is a positive integer greater than 1. The N workers can implement similar functions, and each of the N workers contains certain private data that can be used for local model training. At the same time, the central service node can interact with each of the N workers.

[0111] The basic idea of ​​federated learning with partial model sharing is as follows: each worker deploys a copy of a public model (denoted as w), and each worker can combine the public model copy with the private model in the worker to obtain a new model, and use its own private data to update the new model to obtain the updated value of the new model. For example, worker i can combine the public model copy with the private model q in the worker. i A new model is obtained by splicing, and the new model is updated with its own private data to obtain the updated value w of the new model new,i =w+Δ i ,q new,i =q i +Δ q,i ; Further, N workers can update the public model's update value Δ i Feedback to the central service node, which performs weighted average of these N update values ​​to obtain the average update value Furthermore, the central service node can send the average update value Δ to N workers, and these N workers can update their copies to w+Δ respectively.

[0112] After the replica is updated, the N workers may iteratively execute the above model update or local model training operation locally repeatedly until the model converges, or until other conditions are met.

[0113] In a federated learning system with partial model sharing, the adaptability to differences in computing nodes can be improved by reducing the weight of the globally shared part (i.e., the public model). However, when the weight of the private model is too high, it is easy to cause the model to overfit on the local data set. Therefore, it is necessary to fine-tune the proportion of public and private models. At the same time, a federated learning system with partial model sharing requires that all computing nodes have the ability to process the public model part.

[0114] In summary, existing federated learning-based model training methods still have some problems. For example, the private data of workers may vary greatly, resulting in poor performance of the trained model; they cannot support different types of workers; and when there are differences between a worker and other workers, the ratio of public and private models in the worker needs to be fine-tuned, which is difficult to adjust and increases training costs.

[0115] In order to solve one or more of the above technical problems, the present application proposes a model training method and device, and a communication method and device. The model training method in the embodiment of the present application is described in detail with reference to FIG4 .

[0116] FIG4 is a schematic flow chart of a model training method provided by an embodiment of the present application. The method 400 shown in FIG4 may include step S410, which is as follows:

[0117] S410: The central node sends a first public data set, or the first public data set and a first shared model weight to a first worker.

[0118] The central node may refer to the first entity in the aforementioned embodiment, and the first worker may refer to the second entity in the aforementioned embodiment. For example, the central node may be the network device 110 in Figure 1, and the first worker may be the terminal device 120 or the terminal device 130 in Figure 1; alternatively, the central node may be the terminal device 210 in Figure 2, and the first worker may be the terminal device 220 or the terminal device 230 in Figure 2; alternatively, the central node and the first worker may also be other network entities in the wireless communication system, which is not limited in the embodiments of the present application.

[0119] In some embodiments, the first worker may be a type 3 worker or a type 4 worker as shown in FIG. 5 .

[0120] Optionally, the first worker may be one of the at least one worker. Optionally, the central node may send the first public dataset, or the first public dataset and the first shared model weight to the at least one worker.

[0121] In some embodiments, the first shared model weight (also referred to as the first shared model) can be used to determine the weight of the first worker's local model. The weights here can be understood as parameters of the neural network model. The first worker's local model can be understood as the model on the first worker that needs to be trained.

[0122] For example, the first shared model weight can be used as the initial value of the weight of the first worker's local model; or, the first shared model weight can be spliced ​​with the weight of the first worker's private model to obtain a local model, that is, the first shared model weight and the weight of the first worker's private model are spliced ​​together and used as the initial value of the weight of the first worker's local model.

[0123] Alternatively, the weight of the first worker's private model may be used as the initial value of the weight of the first worker's local model.

[0124] In some embodiments, the first public dataset can be used to train the local model of the first worker, and / or the first public dataset can be used to determine the first dataset, wherein the first dataset can be an output dataset obtained after the local model of the first worker processes the first public dataset.

[0125] For example, the first public dataset can be used as a training dataset to train the local model of the first worker; alternatively, the first public dataset can be input into the local model of the first worker (such as the local model of the first worker trained based on the first public dataset) to obtain the first dataset; alternatively, the first public dataset can be first used as a training dataset to train the local model of the first worker, and then the first public dataset can be input into the trained local model of the first worker to obtain the first dataset.

[0126] Among them, the first public data set may include public data, labels corresponding to the public data, and / or explanation labels corresponding to the public data. The explanation labels in the first public data set can be used to explain the processing of the public data by the worker's local model, or can also be used to explain the reason why the worker's local model obtains a specific output result after processing the public data. For example, the explanation labels in the first public data set can be used to explain: the processing of the public data in the first public data set by the first worker in the subsequent process of training the local model using the first public data set, or the reason why the first worker's local model obtains specific output data in the first output data set after processing the public data in the first public data set.

[0127] Optionally, the explanation labels in the first public dataset can be obtained by methods such as local interpretable model-agnostic explanations (LIME) and Shapley additive explanations (SHAPLEY).

[0128] In some embodiments, when the first public dataset includes explanation tags, the central node may indicate to the first worker the type of explanation tags in the first public dataset. For example, the method 400 may further include step S401, which is as follows:

[0129] At step S401, the central node sends an indication message indicating the type of an explanation tag in a first public dataset to a first worker. The explanation tag in the first public dataset can be used to explain the processing performed by the worker's local model on the public data, or can also be used to explain why the worker's local model obtained a specific output result after processing the public data. For example, explanation tags (in the first public dataset) obtained by different methods can be of different types.

[0130] For example, the explanation tag in the first public dataset can be used to explain the method by which the first worker's local model processes the first public dataset, or the explanation tag can also be used to explain what processing the first worker's local model performs on the first public dataset, or the explanation tag can also be used to explain why the first worker's local model obtains specific output data in the first output dataset after processing the first public dataset, or the explanation tag can also be used to explain other content (or information) in the process of the first worker's local model processing the first public dataset. The type of the explanation tag in the first public dataset can indicate the content (or information) indicated by the explanation tag or the method used to obtain the explanation tag.

[0131] In an embodiment of the present application, the central node sends a first public data set, or a first public data set and a first shared model weight to at least one worker, so as to facilitate the first worker to train a local model and / or determine the first data set based on the first public data set, and helps to constrain the local model of at least one worker to have similar working methods by constraining the local model of at least one worker to have similar results on the same public data set (such as the first public data set), thereby helping to avoid overfitting of the local model and improve the training effect of the model.

[0132] In some embodiments, before S410, the first worker may report the data type it needs to the central node. For example, before S410, the method 400 may further include step S402, which is as follows:

[0133] S402: The first worker sends fourth information to the central node. The fourth information may be used to indicate that the first worker needs to use the first public dataset and / or the first shared model weight.

[0134] In some embodiments, before S410, the central node may indicate to the first worker the time-frequency resource location to which the first public data set is to be sent. In embodiments of the present application, the time-frequency resource location may include the location of the time domain resource and / or the location of the frequency domain resource. For example, before S410, method 400 may further include step S404, specifically as follows:

[0135] S404: The central node sends fifth information to the first worker. The fifth information may be used to instruct the central node to send the third time-frequency resource location of the first public data set.

[0136] In some embodiments, before S410, the central node may indicate to the first worker the time-frequency resource location at which to send the first shared model weight. For example, before S410, the method 400 may further include step S406, which is as follows:

[0137] S406: The central node sends sixth information to the first worker. The sixth information may be used to instruct the central node to send a fourth time-frequency resource location of the first shared model weight.

[0138] In some embodiments, the central node may instruct the first worker to determine the first data set using specific public data in the first public data set. For example, the method 400 may further include step S408, which is as follows:

[0139] S408: The central node sends instruction information to the first worker for instructing the first worker to use specific public data in the first public data set to determine the first data set.

[0140] In some embodiments, after S410, the first worker may report to the central node the output data set obtained by processing the first public data set using the local model (such as the trained local model). For example, after S410, the method 400 may further include step S412, which is as follows:

[0141] S412: The first worker sends a first output dataset to the central node. The first output dataset may be an output dataset obtained by the first worker processing the first public dataset using a local model (eg, the first worker's local model trained based on the first public dataset).

[0142] The first output data set may include output data and / or explanation labels corresponding to the output data. The explanation labels in the first output data set may be used to explain the processing performed by the worker's local model to obtain the output data, or may also be used to explain the reason why the worker's local model obtains the output data. For example, the explanation labels in the first output data set may be used to explain: the processing performed by the first worker's local model on the public data in the first public data set to obtain the first output data set, or the reason why the first worker's local model processes the public data in the first public data set to obtain specific output data in the first output data set.

[0143] Optionally, the explanation labels in the first output dataset can be obtained by methods such as LIME and SHAPLEY.

[0144] In some embodiments, before S412, the central node may indicate to the first worker the time-frequency resource location for sending the first output data set. For example, before S412, the method 400 may further include step S411, which is as follows:

[0145] S411: The central node sends first information to a first worker. The first information may be used to indicate a first time-frequency resource position of a first output data set to be fed back to the central node.

[0146] In some embodiments, when the first output data set includes an explanation label, the central node may indicate to the first worker the type of the explanation label in the first output data set. For example, the method 400 may further include step S414, which is as follows:

[0147] At step S414, the central node sends, to the first worker, information indicating the type of the explanation label in the first output dataset. The explanation label in the first output dataset can be used to explain the processing performed by the worker's local model to obtain the output data in the first output dataset, or can also be used to explain why the worker's local model obtained specific output data in the first output dataset. For example, explanation labels (in the first output dataset) obtained by different methods can be of different types.

[0148] For example, the explanation tag in the first output dataset can be used to explain the method by which the first worker's local model obtains the first output dataset, or the explanation tag can also be used to explain what processing the first worker's local model performs to obtain the first output dataset, or the explanation tag can also be used to explain why the first worker's local model obtains specific output data in the first output dataset, or the explanation tag can also be used to explain other content (or information) in the process of the first worker's local model obtaining the first output dataset. The type of the explanation tag in the first output dataset can indicate the content (or information) indicated by the explanation tag or the method used to obtain the explanation tag.

[0149] In some embodiments, after S410, the first worker may report to the central node the updated value of the first shared model weight during the training of the local model. For example, after S410, the method 400 may further include step S416, which is as follows:

[0150] S416: The first worker sends a first weight update value to the central node. The first weight update value may be an update value of the first shared model weight during the first worker's training of the local model.

[0151] For example, the first worker can use the first shared model weight as the initial value of the local model weight, train the local model, and update the local model weight based on the loss value obtained after training. At this time, the updated value of the local model weight can be used as the updated value of the first shared model weight, that is, the first weight update value.

[0152] For another example, the first worker can use the first shared model weight as the initial value of the global shared model weight, splice the global shared model with the private model on the first worker to obtain a model, use the spliced ​​model as the initial value of the local model, train the local model, and update the local model weight (including the global shared model weight and the private model weight) based on the loss value obtained after training (the loss value includes the loss value of the global shared model and the loss value of the private model). At this time, the updated value of the global shared model weight can be used as the updated value of the first shared model weight, that is, the first weight update value.

[0153] In some embodiments, before S416, the central node may indicate to the first worker the time-frequency resource location for sending the first weight update value. For example, before S416, the method 400 may further include step S415, which is as follows:

[0154] S415: The central node sends second information to the first worker. The second information may be used to indicate a second time-frequency resource location for feeding back the first weight update value to the central node.

[0155] In some embodiments, the central node may indicate that the first worker needs to feedback a first output data set and / or a first weight update value.

[0156] For example, the method 400 may further include step S418, which is specifically as follows:

[0157] S418: The central node sends third information to the first worker. The third information may be used to indicate that the first worker needs to feed back the first output data set and / or the first weight update value.

[0158] In some embodiments, after S410, method 400 may further include step S420, which is as follows:

[0159] S420: The central node sends the second public data set, or the second public data set and the second shared model weight, to the first worker.

[0160] Optionally, the central node may send a second public dataset, or a second public dataset and a second shared model weight to at least one worker.

[0161] In some embodiments, the second shared model weight may be used to determine the weight of the first worker's local model.

[0162] For example, the second shared model weight can be used as the initial value of the weight of the first worker's local model; or, the second shared model weight can be spliced ​​with the weight of the first worker's private model to obtain a local model, that is, the second shared model weight and the weight of the first worker's private model are spliced ​​together and used as the initial value of the weight of the first worker's local model.

[0163] In some embodiments, the second public dataset can be used to train the first worker's local model, and / or the second public dataset can be used to determine the second dataset, wherein the second dataset can be an output dataset obtained by processing the second public dataset by the first worker's local model.

[0164] For example, the second public dataset can be used as a training dataset to train the local model of the first worker; alternatively, the second public dataset can be input into the local model of the first worker (such as the local model of the first worker trained based on the second public dataset) to obtain the second dataset; alternatively, the second public dataset can be first used as a training dataset to train the local model of the first worker, and then the second public dataset can be input into the trained local model of the first worker to obtain the second dataset.

[0165] In some embodiments, the second public dataset may be determined based on an output dataset (such as the first output dataset) fed back by the first worker. Alternatively, the second public dataset may be determined based on the first output dataset fed back by at least one worker (including the first worker).

[0166] Optionally, after receiving the first output data sets fed back by N workers, the central node may perform a weighted average on the N first output data sets to obtain a second public data set, where N is a positive integer. For example, the first output data set fed back by the nth worker (among the N workers) may contain M output data, where the mth output data x m,n The output data obtained by the nth worker after processing the mth data in the first public data set; m,1 ,…,x m,N}After weighted averaging, the output corresponding to the mth data in the second public data set can be obtained, where n and m are positive integers.

[0167] It should be noted that this embodiment is only an example and not a limitation. In this application, the second public data set can also be determined by other methods, which is not limited in the embodiments of this application.

[0168] In some embodiments, the second shared model weight may be determined based on an updated value of the shared model weight fed back by the first worker during the training of the local model and / or an output dataset obtained after the first worker's local model processes the first public dataset. Alternatively, the second shared model weight may be determined based on an updated value of the first weight fed back by at least one worker (including the first worker) and / or the first output dataset.

[0169] For example, after receiving the first weight update values ​​fed back by N workers, the central node may perform a weighted average on the N first weight update values ​​to obtain a second shared model weight, where N is a positive integer; or, after receiving the first weight update values ​​and the first output data set fed back by N workers, the central node may perform a weighted average on the N first output data sets to obtain a second public data set, and update the N first weight update values ​​based on the second public data set to obtain a second shared model weight. Updating the N first weight update values ​​based on the second public data set may include: updating the N first weight update values ​​based on the second public data set to obtain N updated weight values, and then performing a weighted average on the N updated weight values ​​to obtain a second shared model weight; or performing a weighted average on the N first weight update values ​​to obtain a weighted average weight value, and then updating the weighted average weight value based on the second public data set to obtain a second shared model weight.

[0170] It should be noted that this embodiment is only an example and not a limitation. The second sharing model weight can also be determined by other methods in this application, and this is not limited in the embodiment of this application.

[0171] It should be noted that, in an embodiment of the present application, the time-frequency resource location at which the central node sends the public data set and / or shared model weights to the first worker may be statically configured, semi-statically configured, or dynamically configured. For example, before S420, the central node may instruct the first worker to send the time-frequency resource location at which the second public data set is sent, as well as the time-frequency resource location at which the second shared model weights are sent; alternatively, the central node may directly send the second public data set to the first worker via the third time-frequency resource location, and send the second shared model weights to the first worker via the fourth time-frequency resource location.

[0172] In some embodiments, when the second worker reports to the central node the updated value of the first shared model weight during the process of training the local model, the second shared model weight may be determined based on the updated value of the shared model weight during the process of training the local model fed back by the first worker and / or the output data set obtained after the first worker's local model processes the first public data set, and the updated value of the shared model weight during the process of training the local model fed back by the second worker. Optionally, the second shared model weight may be determined based on the first weight update value and / or the first output data set fed back by at least one first worker, and the second weight update value fed back by at least one second worker.

[0173] For example, after receiving the first weight update values ​​fed back by N workers and the second weight update values ​​fed back by M second workers, the central node can perform a weighted average on the N first weight update values ​​and the M second weight update values ​​to obtain a second shared model weight, where N and M are positive integers; or, after receiving the first weight update values ​​and the first output data set fed back by N workers, and the second weight update values ​​fed back by M second workers, the central node can perform a weighted average on the N first output data sets to obtain a second common data set, and update the N first weight update values ​​and the M second weight update values ​​based on the second common data set to obtain a second shared model weight. Among them, updating the N first weight update values ​​and the M second weight update values ​​based on the second public data set may include: based on the second public data set, updating the N first weight update values ​​to obtain N updated weight values, updating the M first weight update values ​​to obtain M updated weight values, and then performing weighted averaging on the N updated weight values ​​and the M updated weight values ​​to obtain the second shared model weight; or, performing weighted averaging on the N first weight update values ​​and the M second weight update values ​​to obtain the weighted average weight value, and then updating the weighted average weight value based on the second public data set to obtain the second shared model weight.

[0174] It should be noted that this embodiment is only an example and not a limitation. The second sharing model weight can also be determined by other methods in this application, and this is not limited in the embodiment of this application.

[0175] In some embodiments, the central node may also send the shared model weights to the second workers (e.g., at least one second worker). For example, the method 400 may further include step S430, which is as follows:

[0176] S430: The central node sends the first shared model weight to the second worker.

[0177] The second worker (eg, at least one second worker) may refer to the second entity in the aforementioned embodiment. In some embodiments, the second worker may be a type 1 worker or a type 2 worker as shown in FIG. 5 .

[0178] In some embodiments, the first shared model weight can be used to determine the weight of the second worker's local model. In other words, the first shared model weight can be used to determine the weight of the first worker's local model as well as the weight of the second worker's local model.

[0179] For example, the first shared model weight can be used as the initial value of the weight of the second worker's local model; or, the first shared model weight can be spliced ​​with the weight of the second worker's private model to obtain a local model, that is, the first shared model weight and the weight of the second worker's private model are spliced ​​together and used as the initial value of the weight of the second worker's local model.

[0180] In some embodiments, before S430, the central node may indicate to the second worker the time-frequency resource location at which to send the first shared model weight. For example, before S430, the method 400 may further include step S422, which is as follows:

[0181] S422: The central node sends seventh information to the second worker.

[0182] The seventh information may be used to instruct the central node to send the second time-frequency resource location of the first shared model weight.

[0183] In some embodiments, the second time-frequency resource location and the fourth time-frequency resource location may be the same. That is, the central node may send the first shared model weight to the first worker and the second worker via the same time-frequency resource location.

[0184] In some embodiments, before S430, the second worker may report the data type it needs to the central node. For example, before S430, the method 400 may further include step S424, which is as follows:

[0185] S424: The second worker sends the tenth message to the central node.

[0186] The tenth information may be used to indicate that the second worker needs to use the first shared model weight.

[0187] In some embodiments, after S430, the second worker may report to the central node the updated value of the first shared model weight during the training of the local model. For example, after S430, the method 400 may further include step S432, which is as follows:

[0188] S432: The second worker sends a second weight update value to the central node.

[0189] The second weight update value may be an update value of the first shared model weight during the process of the second worker training the local model.

[0190] For example, the second worker can use the first shared model weight as the initial value of the local model weight, train the local model using a private dataset, and update the local model weight based on the loss value obtained after training. At this time, the updated value of the local model weight can be used as the updated value of the first shared model weight, that is, the second weight update value.

[0191] For another example, the second worker can use the weight of the first shared model as the initial value of the weight of the global shared model, splice the global shared model with the private model on the second worker to obtain a model, use the spliced ​​model as the initial value of the local model, train the local model, and update the weight of the local model (including the weight of the global shared model and the weight of the private model) based on the loss value obtained after training (the loss value includes the loss value of the global shared model and the loss value of the private model). At this time, the updated value of the weight of the global shared model can be used as the updated value of the weight of the first shared model, that is, the second weight update value.

[0192] In some embodiments, before S432, the central node may indicate to the second worker the time-frequency resource location for sending the second weight update value. For example, before S432, the method 400 may further include step S431, which is as follows:

[0193] S431: The central node sends the eighth information to the second worker.

[0194] The eighth information may be used to indicate the fifth time-frequency resource position for feeding back the second weight update value to the central node.

[0195] In some embodiments, the central node may indicate that the second worker needs to feedback the second weight update value. For example, the method 400 may further include step S434, which is as follows:

[0196] S434: The central node sends ninth information to the second worker.

[0197] The ninth information may be used to indicate that the second worker needs to feed back a second weight update value.

[0198] In some embodiments, after step S430, the method 400 may further include step S440, which is as follows:

[0199] S440: The central node sends the second shared model weight to the second worker.

[0200] The second shared model weight can be used to determine the weight of the second worker's local model. In other words, the second shared model weight can be used to determine the weight of the first worker's local model, and can also be used to determine the weight of the second worker's local model.

[0201] In some embodiments, the second shared model weight may be determined based on an updated value of the shared model weight fed back by the first worker during the training of the local model and / or an output dataset obtained after the first worker's local model processes the first public dataset, and an updated value of the shared model weight fed back by the second worker during the training of the local model. Optionally, the second shared model weight may be determined based on a first weight update value and / or a first output dataset fed back by at least one worker, and a second weight update value fed back by at least one second worker.

[0202] It should be noted that, in the embodiment of the present application, the time-frequency resource location at which the central node sends the shared model weights to the second worker may be statically configured, semi-statically configured, or dynamically configured. For example, before S440, the central node may indicate to the second worker the time-frequency resource location at which it sends the second shared model weights; alternatively, the central node may directly send the second shared model weights to the second worker via the fourth time-frequency resource location.

[0203] In the above-mentioned method 400, although the second worker does not directly use the public dataset in the process of training the local model, the shared model weight used by the second worker when determining the weight of the local model can be determined by the output dataset fed back by the first worker. In this way, although the second worker does not directly use the public dataset, the same public dataset can be applied to both the first worker and the second worker at the same time through the shared model weight (determined based on the output dataset fed back by the first worker). Therefore, the solution in the embodiment of the present application can support different types of workers to join federated learning.

[0204] 5 and 6 , an exemplary description is given below of how the embodiment of the present application supports multiple types of workers.

[0205] Figure 5 is a schematic diagram of multiple types of workers supported by an embodiment of the present application. The federated learning system shown in Figure 5 can include four types of workers, as follows:

[0206] Type 1: Update the local model based on the global shared model

[0207] Type 1 workers can receive the global shared model w (such as the first shared model weight or the second shared model weight in the aforementioned embodiment) sent by the central node, and use the global shared model w as the initial value of the local model to calculate the loss function of the local model and the reverse gradient Δ of the loss function on the private dataset (i.e., the local dataset of the type 1 worker). i (It can also be understood as the weight update value of the global shared model, such as the second weight update value in the aforementioned embodiment), and the local model is updated to w new,i =w+Δi .

[0208] Furthermore, as shown in Figure 6, the type 1 worker can update the value Δ i Feedback to the central node.

[0209] Type 2: Update the local model based on the global shared model and private model

[0210] Type 2 workers can receive the global shared model w (such as the first shared model weight or the second shared model weight in the aforementioned embodiment) sent by the central node, and combine the global shared model w with a private model q on the worker. i The concatenated model is used as the initial value of the local model, and the loss function of the local model and the reverse gradient Δ of the loss function are calculated on the private dataset. i (It can also be understood as the weight update value of the global shared model, such as the second weight update value in the aforementioned embodiment) and Δ q,i , and update the local model. The updated local model includes w new,i =w+Δ i ,q new,i =q i +Δ q,i .

[0211] Furthermore, as shown in Figure 6, the type 2 worker can update the value Δ i Feedback to the central node.

[0212] Type 3: Updating local models based on public and private datasets

[0213] Type 3 workers can receive public datasets sent by the central node and use private models q i As a local model, the loss function of the local model and the reverse gradient Δ of the loss function are calculated on the private dataset and the public dataset. q,i , update the local model q new,i =q i +Δ q,i .

[0214] Type 3 workers can also feed public datasets into the updated local model q new,i Obtaining an output data set (such as the first output data set in the aforementioned embodiment);

[0215] Furthermore, as shown in Figure 6, type 3 workers can feed back the output dataset to the central node.

[0216] Type 4: Update the local model based on the global shared model and the local model, public dataset and private dataset

[0217] Type 4 workers can receive the global shared model w and public dataset sent by the central node, and use the global shared model w and a private model q on the worker i The concatenated model is used as the initial value of the local model. The loss function of the local model and the reverse gradient Δ of the loss function are calculated on the local dataset and the public dataset. i (It can also be understood as the weight update value of the global shared model, such as the second weight update value in the aforementioned embodiment) and Δ q,i , and update the local model. The updated local model includes w new,i =w+Δ i ,q new,i =q i +Δ q,i .

[0218] Type 4 can also input public datasets into the updated local model (such as by w new,i and q new,i splicing to obtain) an output data set (such as the first output data set in the aforementioned embodiment).

[0219] Furthermore, as shown in Figure 6, the type 4 worker can update the value Δ i And the output data set is fed back to the central node.

[0220] After receiving data uploaded by multiple workers, the central node can perform two types of fusion, as follows:

[0221] Data Fusion:

[0222] The central node can fuse the output datasets fed back by multiple workers (such as type 3 workers and / or type 4 workers). For example, the central node can perform a weighted average of the output datasets fed back by multiple workers to obtain an output result, which can be used as the public dataset that the central node will send to multiple workers next time.

[0223] Since different types of workers in federated learning may not share model structures, in an embodiment of the present application, the effect of constraining different workers' models to have similar working methods can be achieved by constraining them to have similar results on the same public dataset. For example, the central node can request different workers to calculate output datasets for a public dataset and upload the calculated output datasets to the central node. The output datasets calculated by different workers can then be fused to obtain an output result, which can be regarded as an output result agreed upon by multiple workers. This output result will then be used for the next update of each worker.

[0224] Model fusion:

[0225] The central node can fuse the updated values ​​of the global shared model weights fed back by multiple workers (such as type 1 workers, type 2 workers and / or type 4 workers).

[0226] For example, the central node may perform weighted averaging on the updated values ​​of the global shared model weights fed back by multiple workers to obtain a new global shared model weight; or the central node may fuse the output data sets fed back by multiple workers to obtain a new common data set, and update the updated values ​​of the global shared model weights fed back by multiple workers based on the new common data set to obtain a new global shared model weight. Updating the updated values ​​of the global shared model weights fed back by multiple workers based on the new common data set may include: updating the updated values ​​of the global shared model weights fed back by multiple workers based on the new common data set to obtain multiple updated weight values, and then performing weighted averaging on the multiple updated weight values ​​to obtain a new global shared model weight; or performing weighted averaging on the updated values ​​of the global shared model weights fed back by multiple workers to obtain a weighted average weight value, and then updating the weighted average weight value based on the new common data set to obtain a new global shared model weight.

[0227] The following is an exemplary description of the model training process for the above-mentioned type 1 workers and type 2 workers with reference to FIG7 .

[0228] FIG7 is a schematic flow chart of a model training method provided by another embodiment of the present application. The method 700 shown in FIG7 may include steps S710 to S760, which are as follows:

[0229] S710, worker A reports the required data type to the central node.

[0230] Worker A may include type 1 workers and / or type 2 workers.

[0231] Worker A can report to the central node that it needs to use the shared model weights.

[0232] S720, the central node sends data configuration to worker A.

[0233] The data configuration may include indication information of the time-frequency resource location for instructing the central node to send data (shared model weights) to worker A.

[0234] For example, the central node can determine to send the shared model weights to worker A based on the required data type reported by worker A; further, the central node can send data configuration to worker A to indicate at which time-frequency resource locations worker A receives the required shared model weights.

[0235] S730: The central node sends feedback configuration to worker A.

[0236] The feedback configuration may include indication information for indicating the type of data that the central node expects worker A to feedback. For example, the feedback configuration may include indication information for instructing worker A to feedback a weight update value.

[0237] The feedback configuration may further include indication information for indicating the time-frequency resource location of the feedback data of worker A. For example, the feedback configuration may include indication information for indicating the time-frequency resource location of the feedback weight update value of worker A.

[0238] S740: The central node sends the shared model weights to worker A.

[0239] The central node can send the shared model weights to worker A through the time-frequency resource location indicated by the data configuration.

[0240] S750: Worker A updates the local model based on the shared model weights.

[0241] Worker A can use the shared model weight as the initial value of the local model and update the local model to obtain the updated weight value.

[0242] Alternatively, worker A can concatenate the shared model weights with the private model, use the concatenated model as the initial value of the local model, and update the local model to obtain the updated weight value.

[0243] S760: Worker A feeds back the updated weight value to the central node.

[0244] Worker A can send the shared model weights to the central node by feeding back the time-frequency resource location indicated by the configuration.

[0245] The following is an exemplary description of the model training process of the above-mentioned type 3 worker in conjunction with Figure 8.

[0246] FIG8 is a schematic flow chart of a model training method provided by another embodiment of the present application. The method 800 shown in FIG8 may include steps S810 to S860, which are as follows:

[0247] S810, worker B reports the required data type to the central node.

[0248] Worker B may include type 3 workers.

[0249] Worker B can report to the central node that it needs to use the public dataset.

[0250] S820, the central node sends data configuration to worker B.

[0251] The data configuration may include indication information for instructing the central node to send the time-frequency resource location of the data (public data set) to the worker B.

[0252] For example, the central node can determine to send a public data set to worker B based on the required data type reported by worker B; further, the central node can send a data configuration to worker B to instruct worker B at which time-frequency resource locations to receive the required public data set.

[0253] The data configuration may further include indication information for indicating the type of explanation labels in the public dataset.

[0254] The data configuration may further include instruction information for instructing worker B to use specific data in the public dataset to determine the output dataset.

[0255] S830: The central node sends feedback configuration to worker B.

[0256] The feedback configuration may include indication information for indicating (as expected by the central node) the type of data to be fed back by worker B. For example, the feedback configuration may include indication information for instructing worker B to feed back an output data set.

[0257] The feedback configuration may further include indication information for indicating the type of explanation labels in the output dataset.

[0258] The feedback configuration may further include indication information for indicating the time-frequency resource location of the feedback data of worker B. For example, the feedback configuration may include indication information for indicating the time-frequency resource location of the feedback output data set of worker B.

[0259] S840: The central node sends the public data set to worker B.

[0260] The central node can send the public dataset to worker B through the time-frequency resource location indicated by the data configuration.

[0261] S850, worker B updates the local model based on the public dataset.

[0262] Worker B can update the local model based on the public dataset and the private dataset, and input the public dataset into the updated local model to obtain the output dataset.

[0263] S860: Worker B feeds back the output data set to the central node.

[0264] Worker B can send the output data set to the central node by feeding back the time-frequency resource location indicated by the configuration.

[0265] The following is an exemplary description of the model training process of the above-mentioned type 4 worker in conjunction with FIG9 .

[0266] FIG9 is a schematic flow chart of a model training method provided by another embodiment of the present application. The method 900 shown in FIG9 may include steps S910 to S960, which are as follows:

[0267] S910, worker C reports the required data type to the central node.

[0268] Worker C may include type 4 workers.

[0269] Worker C can report to the central node that it needs to use shared model weights and public datasets.

[0270] S920: The central node sends data configuration to worker C.

[0271] The data configuration may include indication information for instructing the central node to send the time-frequency resource location of the shared model weight to the worker C.

[0272] The data configuration may further include indication information for instructing the central node to send the time-frequency resource location of the public data set to the worker C.

[0273] For example, the central node can determine to send shared model weights and public data sets to worker C based on the required data type reported by worker C; further, the central node can send data configuration to worker C to instruct worker C at which time-frequency resource locations to receive the required shared model weights and public data sets.

[0274] The data configuration may further include indication information for indicating the type of explanation labels in the public dataset.

[0275] The data configuration may further include instruction information for instructing the worker C to use specific data in the public dataset to determine the output dataset.

[0276] S930: The central node sends feedback configuration to worker C.

[0277] The feedback configuration may further include indication information for indicating (as expected by the central node) the type of data to be fed back by the worker C. For example, the feedback configuration may include indication information for instructing the worker C to feed back a weight update value and an output data set.

[0278] The feedback configuration may further include indication information for indicating the type of explanation labels in the output dataset.

[0279] The feedback configuration may also include indication information for indicating the time-frequency resource location of the feedback data by worker C. For example, the feedback configuration may include indication information for indicating the time-frequency resource location of the feedback weight update value by worker C; the feedback configuration may also include indication information for indicating the time-frequency resource location of the feedback output data set by worker C.

[0280] S940: The central node sends the shared model weights and public dataset to worker C.

[0281] The central node can send the shared model weights and public dataset to worker C through the time-frequency resource location indicated by the data configuration.

[0282] S950, worker C updates the local model based on the shared model weights and public dataset.

[0283] Worker C can use the shared model weight as the initial value of the local model, update the local model based on the public dataset and private dataset to obtain the updated weight value, and input the public dataset into the updated local model to obtain the output dataset.

[0284] Alternatively, worker C can concatenate the shared model weights with the private model, use the concatenated model as the initial value of the local model, update the local model based on the public dataset, obtain the updated weight value, and input the public dataset into the updated local model to obtain the output dataset.

[0285] S960, worker C feeds back the weight update value and output data set to the central node.

[0286] Worker C can send weight update values ​​and output data sets to the central node by feeding back the time-frequency resource locations indicated by the configuration.

[0287] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0288] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 9 . The device embodiment of the present application is described in detail below in conjunction with Figures 10 to 12 . It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.

[0289] FIG10 is a schematic structural diagram of a device for model training provided in an embodiment of the present application. As shown in FIG10 , the device 1000 includes a sending unit 1010, which is specifically as follows:

[0290] A sending unit 1010 is configured to send a first public dataset, or the first public dataset and a first shared model weight, to at least one worker, wherein the at least one worker includes a first worker, the first shared model weight is used to determine the weight of the first worker's local model, the first public dataset is used to train the first worker's local model, and / or the first public dataset is used to determine a first dataset, which is an output dataset obtained after the first worker's local model processes the first public dataset.

[0291] Optionally, after sending the first public dataset, or the first public dataset and the first shared model weight to at least one worker, the sending unit 1010 is further used to: send a second public dataset, or the second public dataset and the second shared model weight to the at least one worker, wherein the second shared model weight is used to determine the weight of the local model of the first worker, the second public dataset is used to train the local model of the first worker, and / or the second public dataset is used to determine a second dataset, which is an output dataset obtained after the local model of the first worker processes the second public dataset.

[0292] Optionally, the second public data set is determined based on the first output data set fed back by the at least one worker, and the second shared model weight is determined based on the first weight update value fed back by the at least one worker and / or the first output data set, the first output data set is the output data set obtained by the first worker using the local model to process the first public data set, and the first weight update value is the update value of the first shared model weight in the process of the first worker training the local model.

[0293] Optionally, after sending the first public dataset to at least one worker, or the first public dataset and the first shared model weight, the device 1000 also includes a receiving unit 1020, which is used to: receive a first output dataset from the first worker, where the first output dataset is an output dataset obtained by the first worker using a local model to process the first public dataset.

[0294] Optionally, before receiving the first output data set from the first worker, the sending unit 1010 is further used to: send first information to the first worker, where the first information is used to indicate a first time-frequency resource position of the first output data set to be fed back to the central node.

[0295] Optionally, in the case where the first output data set includes an explanation label, the sending unit 1010 is further used to: send indication information to the first worker for indicating the type of the explanation label in the first output data set, and the explanation label in the first output data set is used to explain the processing of the output data in the first output data set performed by the worker's local model.

[0296] Optionally, the first output data set includes output data, labels corresponding to the output data, and / or explanation labels corresponding to the output data, and the explanation labels are used to explain the processing of the output data by the worker's local model.

[0297] Optionally, after sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the device 1000 also includes a receiving unit 1020, which is used to: receive a first weight update value from the first worker, where the first weight update value is an update value of the first shared model weight during the process of the first worker training the local model.

[0298] Optionally, before receiving the first weight update value from the first worker, the sending unit 1010 is further used to: send second information to the first worker, where the second information is used to indicate a second time-frequency resource location for feeding back the first weight update value to the central node.

[0299] Optionally, the sending unit 1010 is further configured to send third information to the first worker, where the third information is configured to indicate that the first worker needs to feed back a first output data set and / or a first weight update value.

[0300] Optionally, before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the device 1000 also includes a receiving unit 1020, which is used to: receive fourth information from the first worker, and the fourth information is used to indicate that the first worker needs to use the first public data set and / or the first shared model weight.

[0301] Optionally, before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the sending unit 1010 is further used to: send fifth information to the first worker, and the fifth information is used to instruct the central node to send the third time-frequency resource location of the first public data set.

[0302] Optionally, before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the sending unit 1010 is further used to: send sixth information to the first worker, and the sixth information is used to indicate the fourth time-frequency resource location of the central node to send the first shared model weight.

[0303] Optionally, the sending unit 1010 is further configured to send, to the first worker, instruction information for instructing the first worker to use specific public data in the first public data set to determine the first data set.

[0304] Optionally, in the case where the first public data set includes explanation tags, the sending unit 1010 is further used to: send indication information to the first worker for indicating the type of explanation tags in the first public data set, and the explanation tags in the first public data set are used to explain the processing of the public data performed by the worker's local model.

[0305] Optionally, the first public data set includes public data, labels corresponding to the public data, and / or explanation labels corresponding to the public data, and the explanation labels are used to explain the processing of the public data by the worker's local model.

[0306] Optionally, the sending unit 1010 is further used to: send the first shared model weight to the second worker, and the first shared model weight is further used to determine the weight of the local model of the second worker.

[0307] Optionally, the sending unit 1010 is further used to: send seventh information to the second worker, where the seventh information is used to instruct the central node to send the second time-frequency resource position of the first shared model weight.

[0308] Optionally, after sending the first shared model weight to the second worker, the sending unit 1010 is further used to: send a second shared model weight to the second worker, the second shared model weight is used to determine the weight of the local model of the second worker, the second shared model weight is determined based on the first weight update value and / or the first output data set fed back by the at least one worker, and the second weight update value fed back by at least one second worker, the first weight update value is the update value of the first shared model weight in the process of the first worker training the local model, the first output data set is the output data set obtained by the first worker using the local model to process the first public data set, and the second weight update value is the update value of the first shared model weight in the process of the second worker training the local model.

[0309] After sending the first shared model weight to the second worker, the device 1000 also includes a receiving unit 1020, which is used to: receive a second weight update value from the second worker, where the second weight update value is an update value of the first shared model weight during the process of the second worker training the local model.

[0310] Optionally, the sending unit 1010 is further used to: send eighth information to the second worker, where the eighth information is used to indicate a fifth time-frequency resource position for feeding back a second weight update value to the central node.

[0311] Optionally, the sending unit 1010 is further configured to send ninth information to the second worker, where the ninth information is configured to indicate that the second worker needs to feed back the second weight update value.

[0312] Optionally, before sending the first shared model weight to the second worker, the device 1000 further includes a receiving unit 1020, which is used to: receive tenth information from the second worker, where the tenth information is used to indicate that the second worker needs to use the first shared model weight.

[0313] FIG11 is a schematic structural diagram of a device for model training provided in an embodiment of the present application. As shown in FIG11 , the device 1100 includes a receiving unit 1100, which is specifically as follows:

[0314] The receiving unit 1110 is used to receive a first public data set from a central node, or the first public data set and a first shared model weight, wherein the first shared model weight is used to determine the weight of the local model of the first worker, and the first public data set is used to train the local model of the first worker, and / or the first public data set is used to determine a first data set, which is an output data set obtained after the local model of the first worker processes the first public data set.

[0315] Optionally, after receiving the first public data set from the central node, or the first public data set and the first shared model weight, the receiving unit 1110 is further used to: receive a second public data set from the central node, or the second public data set and the second shared model weight, wherein the second shared model weight is used to determine the weight of the local model of the first worker, the second public data set is used to train the local model of the first worker, and / or the second public data set is used to determine a second data set, which is an output data set obtained after the local model of the first worker processes the second public data set.

[0316] Optionally, the second public data set is determined based on the first output data set fed back by at least one worker, the at least one worker includes the first worker, the second shared model weight is determined based on the first weight update value fed back by the at least one worker and / or the first output data set, the first output data set is the output data set obtained by the first worker using a local model to process the first public data set, and the first weight update value is the update value of the first shared model weight in the process of the first worker training the local model.

[0317] Optionally, the second shared model weight is further determined according to a second weight update value fed back by at least one second worker.

[0318] Optionally, after receiving the first public data set from the central node, or the first public data set and the first shared model weight, the device 1100 also includes a sending unit 1120, which is used to: send a first output data set to the central node, where the first output data set is an output data set obtained by the first worker using a local model to process the first public data set.

[0319] Optionally, before sending the first output data set to the central node, the receiving unit 1110 is further used to: receive first information from the central node, where the first information is used to indicate a first time-frequency resource position of the first output data set to be fed back to the central node.

[0320] Optionally, in the case where the first output data set includes an explanation label, the receiving unit 1110 is further used to: receive indication information from the central node for indicating the type of the explanation label in the first output data set, and the explanation label in the first output data set is used to explain the processing of the output data in the first output data set performed by the worker's local model.

[0321] Optionally, the first output data set includes output data, labels corresponding to the output data, and / or explanation labels corresponding to the output data, and the explanation labels are used to explain the processing of the output data by the worker's local model.

[0322] Optionally, after receiving the first public data set from the central node, or the first public data set and the first shared model weight, the device 1100 also includes a sending unit 1120, which is used to: send a first weight update value to the central node, where the first weight update value is an update value of the first shared model weight during the process of the first worker training the local model.

[0323] Optionally, before sending the first weight update value to the central node, the receiving unit 1110 is further used to: receive second information from the central node, where the second information is used to indicate a second time-frequency resource location for feeding back the first weight update value to the central node.

[0324] Optionally, the receiving unit 1110 is further configured to: receive third information from the central node, where the third information is configured to indicate that the first worker needs to feed back a first output data set and / or a first weight update value.

[0325] Optionally, before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the device 1100 also includes a sending unit 1120, which is used to: send fourth information to the central node, and the fourth information is used to indicate that the first worker needs to use the first public data set and / or the first shared model weight.

[0326] Optionally, before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the receiving unit 1110 is also used to: receive fifth information from the central node, and the fifth information is used to indicate the third time-frequency resource position of the central node to send the first public data set.

[0327] Optionally, before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the receiving unit 1110 is also used to: receive sixth information from the central node, and the sixth information is used to indicate the fourth time-frequency resource location of the central node to send the first shared model weight.

[0328] Optionally, in the case where the first public data set includes explanation tags, the receiving unit 1110 is further used to: receive indication information from the central node for indicating the type of explanation tags in the first public data set, and the explanation tags in the first public data set are used to explain the processing of the public data by the local model of the explanation worker.

[0329] Optionally, the receiving unit 1110 is further configured to: receive instruction information from the central node for instructing the first worker to use specific public data in the first public data set to determine the first data set.

[0330] Optionally, the first public data set includes public data, labels corresponding to the public data, and / or explanation labels corresponding to the public data, and the explanation labels are used to explain the processing of the public data by the worker's local model.

[0331] FIG12 is a schematic diagram of the structure of an apparatus provided in one embodiment of the present application. The dotted lines in FIG12 indicate that the unit or module is optional. Apparatus 1200 may be used to implement the method described in the above method embodiment. Apparatus 1200 may be a chip or model training apparatus.

[0332] The device 1200 may include one or more processors 1210. The processor 1210 may support the device 1200 to implement the method described in the above method embodiment. The processor 1210 may be a general-purpose processor or a special-purpose processor. For example, the processor may be a central processing unit (CPU). Alternatively, the processor may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0333] The apparatus 1200 may further include one or more memories 1220. The memories 1220 store programs that can be executed by the processor 1210, causing the processor 1210 to perform the methods described in the above method embodiments. The memories 1220 may be independent of the processor 1210 or integrated into the processor 1210.

[0334] The apparatus 1200 may further include a transceiver 1230. The processor 1210 may communicate with other devices or chips via the transceiver 1230. For example, the processor 1210 may transmit and receive data with other devices or chips via the transceiver 1230.

[0335] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0336] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0337] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed by a computer, the computer implements the steps in the above-mentioned various method embodiments.

[0338] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device (such as a server or a terminal device), the electronic device implements the steps in the above-mentioned various method embodiments.

[0339] An embodiment of the present application provides a chip, which includes a processor and a memory, wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs stored in the memory, so that an electronic device (such as a server or terminal device) equipped with the chip executes the steps in the above-mentioned method embodiments.

[0340] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may at least include: any entity or device that can carry the computer program code to the device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable storage medium cannot be an electric carrier signal or a telecommunication signal.

[0341] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0342] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0343] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0344] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0345] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A model training method, applied to a central node, characterized in that: include: A first public data set is sent to at least one worker, or the first public data set and a first shared model weight, wherein the at least one worker includes a first worker, the first shared model weight is used to determine the weight of a local model of the first worker, the first public data set is used to train the local model of the first worker, and / or the first public data set is used to determine a first data set, and the first data set is an output data set obtained after the local model of the first worker processes the first public data set.

2. The method according to claim 1, characterized in that After sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: Sending a second public data set to the at least one worker, or the second public data set and a second shared model weight, wherein the second shared model weight is used to determine the weight of the local model of the first worker, and the second public data set is used to train the local model of the first worker, and / or the second public data set is used to determine a second data set, and the second data set is an output data set obtained after the local model of the first worker processes the second public data set.

3. The method according to claim 2, characterized in that The second public data set is determined based on the first output data set fed back by the at least one worker, the second shared model weight is determined based on the first weight update value fed back by the at least one worker and / or the first output data set, the first output data set is the output data set obtained by the first worker using the local model to process the first public data set, and the first weight update value is the update value of the first shared model weight in the process of the first worker training the local model.

4. The method according to any one of claims 1 to 3, characterized in that After sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: A first output dataset is received from the first worker, where the first output dataset is an output dataset obtained by the first worker using a local model to process the first public dataset.

5. The method according to claim 4, characterized in that Prior to receiving the first output data set from the first worker, the method further includes: Sending first information to the first worker, where the first information is used to indicate feeding back a first time-frequency resource position of the first output data set to the central node.

6. The method according to claim 4 or 5, characterized in that: In the case where the first output data set includes explanation labels, the method further comprises: Indication information for indicating the type of explanation labels in the first output data set is sent to the first worker, where the explanation labels in the first output data set are used to explain processing of output data in the first output data set by the worker's local model.

7. The method according to any one of claims 4 to 6, characterized in that The first output data set includes output data, labels corresponding to the output data, and / or explanation labels corresponding to the output data, where the explanation labels are used to explain the processing of the output data by the worker's local model.

8. The method according to any one of claims 1 to 7, characterized in that After sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: A first weight update value is received from the first worker, where the first weight update value is an update value of the first shared model weight during a process in which the first worker trains a local model.

9. The method according to claim 8, characterized in that Before receiving the first weight update value from the first worker, the method further includes: Sending second information to the first worker, where the second information is used to indicate a second time-frequency resource location for feeding back the first weight update value to the central node.

10. The method according to any one of claims 4 to 9, characterized in that The method further comprises: Sending third information to the first worker, where the third information is used to indicate that the first worker needs to feed back a first output data set and / or a first weight update value.

11. The method according to any one of claims 1 to 10, characterized in that Before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: Fourth information is received from the first worker, where the fourth information is used to indicate that the first worker needs to use the first public data set and / or the first shared model weight.

12. The method according to any one of claims 1 to 11, characterized in that Before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: Send fifth information to the first worker, where the fifth information is used to instruct the central node to send a third time-frequency resource position of the first public data set.

13. The method according to any one of claims 1 to 12, characterized in that Before sending the first public data set to at least one worker, or the first public data set and the first shared model weight, the method further includes: Send sixth information to the first worker, where the sixth information is used to instruct the central node to send a fourth time-frequency resource location of the first shared model weight.

14. The method according to any one of claims 1 to 13, characterized in that The method further comprises: Instruction information for instructing the first worker to determine the first data set using specific public data in the first public data set is sent to the first worker.

15. The method according to any one of claims 1 to 14, characterized in that In the case where the first public dataset includes explanation labels, the method further includes: Indication information for indicating the type of explanation labels in the first public data set is sent to the first worker, where the explanation labels in the first public data set are used to explain the processing of the public data by the local model of the worker.

16. The method according to any one of claims 1 to 15, characterized in that The first public data set includes public data, labels corresponding to the public data, and / or explanation labels corresponding to the public data, where the explanation labels are used to explain the processing of the public data by the worker's local model.

17. The method according to any one of claims 1 to 16, characterized in that The method further comprises: The first shared model weight is sent to a second worker, where the first shared model weight is also used to determine a weight of a local model of the second worker.

18. The method according to claim 17, characterized in that The method further comprises: Sending seventh information to the second worker, where the seventh information is used to instruct the central node to send a second time-frequency resource location of the first shared model weight.

19. The method according to claim 17 or 18, characterized in that After sending the first shared model weight to the second worker, the method further includes: A second shared model weight is sent to the second worker, where the second shared model weight is used to determine the weight of the local model of the second worker, and the second shared model weight is determined based on the first weight update value and / or the first output data set fed back by at least one of the workers, and at least one second weight update value fed back by the second worker, where the first weight update value is an update value of the first shared model weight during the process of the first worker training the local model, the first output data set is an output data set obtained by the first worker using the local model to process the first public data set, and the second weight update value is an update value of the first shared model weight during the process of the second worker training the local model.

20. The method according to any one of claims 17 to 19, characterized in that After sending the first shared model weight to the second worker, the method further includes: Receive a second weight update value from the second worker, where the second weight update value is an update value of the first shared model weight during the process of the second worker training the local model.

21. The method according to claim 20, characterized in that The method further comprises: Sending eighth information to the second worker, where the eighth information is used to indicate a fifth time-frequency resource position for feeding back a second weight update value to the central node.

22. The method according to claim 20 or 21, characterized in that The method further comprises: Ninth information is sent to the second worker, where the ninth information is used to indicate that the second worker needs to feed back the second weight update value.

23. The method according to any one of claims 17 to 22, characterized in that Before sending the first shared model weight to the second worker, the method further includes: A tenth message is received from the second worker, where the tenth message is used to indicate that the second worker needs to use the first shared model weight.

24. A model training method, applied to a first worker, characterized in that: include: Receive a first public data set from a central node, or the first public data set and a first shared model weight, wherein the first shared model weight is used to determine the weight of the first worker's local model, the first public data set is used to train the first worker's local model, and / or the first public data set is used to determine a first data set, and the first data set is an output data set obtained after the first worker's local model processes the first public data set.

25. The method according to claim 24, characterized in that After receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: Receive a second public data set from a central node, or, the second public data set and a second shared model weight, wherein the second shared model weight is used to determine the weight of the local model of the first worker, and the second public data set is used to train the local model of the first worker, and / or, the second public data set is used to determine a second data set, and the second data set is an output data set obtained after the local model of the first worker processes the second public data set.

26. The method according to claim 25, characterized in that The second public data set is determined based on the first output data set fed back by at least one worker, the at least one worker includes the first worker, the second shared model weight is determined based on the first weight update value fed back by the at least one worker and / or the first output data set, the first output data set is the output data set obtained by the first worker using a local model to process the first public data set, and the first weight update value is the update value of the first shared model weight in the process of the first worker training the local model.

27. The method according to claim 26, characterized in that The second shared model weight is also determined according to a second weight update value fed back by at least one second worker.

28. The method according to any one of claims 24 to 27, characterized in that After receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: A first output data set is sent to the central node, where the first output data set is an output data set obtained by the first worker using a local model to process the first public data set.

29. The method according to claim 28, characterized in that Before sending the first output data set to the central node, the method further includes: First information is received from the central node, where the first information is used to indicate a first time-frequency resource position of the first output data set to be fed back to the central node.

30. The method according to claim 28 or 29, characterized in that In the case where the first output data set includes explanation labels, the method further comprises: Receive indication information from the central node for indicating the type of explanation labels in the first output data set, where the explanation labels in the first output data set are used to explain the processing of output data in the first output data set by a local model of the worker.

31. The method according to any one of claims 28 to 30, characterized in that The first output data set includes output data, labels corresponding to the output data, and / or explanation labels corresponding to the output data, where the explanation labels are used to explain the processing of the output data by the worker's local model.

32. The method according to any one of claims 24 to 31, characterized in that After receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: A first weight update value is sent to the central node, where the first weight update value is an update value of the first shared model weight during the process of the first worker training the local model.

33. The method according to claim 32, characterized in that Before sending the first weight update value to the central node, the method further includes: Second information is received from the central node, where the second information is used to indicate a second time-frequency resource position for feeding back the first weight update value to the central node.

34. The method according to any one of claims 28 to 33, characterized in that The method further comprises: Receive third information from the central node, where the third information is used to indicate that the first worker needs to feed back a first output data set and / or a first weight update value.

35. The method according to any one of claims 24 to 34, characterized in that Before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: Sending fourth information to the central node, where the fourth information is used to indicate that the first worker needs to use the first public data set and / or the first shared model weight.

36. The method according to any one of claims 24 to 35, characterized in that Before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: Fifth information is received from the central node, where the fifth information is used to instruct the central node to send a third time-frequency resource position of the first public data set.

37. The method according to any one of claims 24 to 36, characterized in that Before receiving the first public data set from the central node, or the first public data set and the first shared model weight, the method further includes: Receive sixth information from the central node, where the sixth information is used to instruct the central node to send a fourth time-frequency resource position of the first shared model weight.

38. The method according to any one of claims 24 to 37, characterized in that In the case where the first public dataset includes explanation labels, the method further includes: Indication information for indicating the type of explanation labels in the first public data set is received from the central node, where the explanation labels in the first public data set are used to explain the processing of the public data by the local model of the worker.

39. The method according to any one of claims 24 to 38, characterized in that The method further comprises: Instruction information is received from the central node for instructing the first worker to determine the first data set using specific public data in the first public data set.

40. The method according to any one of claims 24 to 39, characterized in that The first public data set includes public data, labels corresponding to the public data, and / or explanation labels corresponding to the public data, where the explanation labels are used to explain the processing of the public data by the worker's local model.

41. A model training device, applied to a central node, characterized in that: include: A sending unit, used to send a first public data set to at least one worker, or the first public data set and a first shared model weight, wherein the at least one worker includes a first worker, the first shared model weight is used to determine the weight of a local model of the first worker, the first public data set is used to train the local model of the first worker, and / or the first public data set is used to determine a first data set, and the first data set is an output data set obtained after the local model of the first worker processes the first public data set.

42. A model training device, applied to a first worker, characterized in that: include: A receiving unit, configured to receive a first public data set from a central node, or the first public data set and a first shared model weight, wherein the first shared model weight is used to determine the weight of the local model of the first worker, and the first public data set is used to train the local model of the first worker, and / or the first public data set is used to determine a first data set, and the first data set is an output data set obtained after the local model of the first worker processes the first public data set.

43. A model training device, characterized in that: include: A processor and a memory, wherein the processor is coupled to the memory, and the memory is used to store a computer program, and when the computer program is executed by the processor, the device performs the method according to any one of claims 1 to 23.

44. A model training device, characterized in that: include: A processor and a memory, wherein the processor is coupled to the memory, and the memory is used to store a computer program, and when the computer program is executed by the processor, the device performs the method according to any one of claims 24 to 40.

45. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 40.

46. ​​A computer program product, characterized in that include: A computer program, when the computer program is run on a computer, causes the computer to perform the method according to any one of claims 1 to 40.

47. A chip, characterized in that: include: A processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, so that a device or equipment equipped with the chip executes the method as described in any one of claims 1 to 40.

Citation Information

Patent Citations

  • Model training method and device

    CN120031070A

  • Federal learning-based model training method and apparatus, and related device

    CN113837397A

  • Federal learning model training method for large-scale industrial chain privacy calculation

    CN114169412A

  • Data model training method and device

    CN114548416A

  • Federal learning-based artificial intelligence model training method, device and system

    CN114692868A