Federated learning method, device, communication device, and readable storage medium
The method addresses the challenge of selecting members in federated learning by using agreement, state, and performance information to enhance training efficiency and reduce dropout.
Patent Information
- Application Number
- JP2025500245
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-08
- Filing Date
- 2023-07-06
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2043-07-06
AI Technical Summary
In federated learning systems, selecting appropriate members for participation in the training process is challenging due to varying availability, performance, and task priorities, leading to inefficiencies and potential dropout.
A method and device for determining member participation in federated learning based on information such as agreement, current state, and model performance, allowing for informed selection and efficient training.
Enables reasonable selection of members, improving training efficiency by reducing dropout and ensuring higher performance through proactive management of member participation.
Smart Images

Figure 2025520969000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This application claims the priority of Chinese Patent Application No. 202210815546.7 filed in China on July 8, 2022, and all of the content of the said application is incorporated herein by reference.
[0002] This application belongs to the field of communication technologies, and specifically relates to a federated learning method, apparatus, communication device, and readable storage medium.
Background Art
[0003] In related communication networks, in order to improve the model effect, the training of the model can be carried out based on federated learning. However, members participating in federated learning may, for various reasons in the process of federated learning, such as not wanting to participate in federated learning because other more important tasks have arrived, or withdrawing from federated learning first because there are too many tasks to process, etc., may not want to participate in federated learning or may no longer be suitable as federated learning members. In such cases, how to reasonably select members participating in federated learning is a problem that needs to be urgently solved at present.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Embodiments of this application provide a federated learning method, apparatus, communication device, and readable storage medium that can solve the problem of how to reasonably select members participating in federated learning.
Means for Solving the Problems
[0005] The first aspect provides a federated learning method, and this method is The first communication device receives first information from the second communication device, where the first information includes at least one of second information for indicating whether the second communication device agrees to participate in the federated learning, state information of the second communication device's current round of federated learning, and model performance information of the current round of federated learning. The first communication device determines, based on the first information, whether the second communication device participates in the next round of federated learning.
[0006] A second aspect provides a federated learning method, and this method includes: The second communication device determines first information, where the first information includes at least one of second information for indicating whether the second communication device agrees to participate in the federated learning, state information of the second communication device's current round of federated learning, and model performance information of the current round of federated learning. The second communication device transmits the first information to the first communication device, where the first information is used for the first communication device to determine whether the second communication device participates in the next round of federated learning.
[0007] A third aspect provides a federated learning device used for a first communication device, and this device includes: A first receiving module for receiving first information from the second communication device, where the first information includes at least one of second information for indicating whether the second communication device agrees to participate in the federated learning, state information of the second communication device's current round of federated learning, and model performance information of the current round of federated learning; A first determining module for determining, based on the first information, whether the second communication device participates in the next round of federated learning.
[0008] A fourth aspect provides a federated learning device used for a second communication device, and this device includes: A second determination module for determining first information, where the first information includes at least one of second information for instructing whether the second communication device agrees to participate in federated learning, state information of the second communication device's current round of federated learning, and model performance information of the current round of federated learning; a second determination module. A second transmission module for transmitting the first information to a first communication device, where the first information is used by the first communication device to determine whether the second communication device participates in the next round of federated learning, including a second transmission module.
[0009] A fifth aspect provides a communication device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, it realizes the steps of the method described in the first aspect or the steps of the method described in the second aspect.
[0010] A sixth aspect provides a communication device, which includes a processor and a communication interface. For example, when this communication device is a first communication device, the communication interface is used to receive first information from a second communication device, and the processor is used to determine whether the second communication device participates in the next round of federated learning based on the first information. Or, when this communication device is a second communication device, the processor is used to determine the first information, and the communication interface is used to transmit the first information to the first communication device. Here, the first information includes at least one of second information for instructing whether the second communication device agrees to participate in federated learning, state information of the second communication device's current round of federated learning, and model performance information of the current round of federated learning.
[0011] The seventh aspect provides a communication system, which includes the first communication device and the second communication device. The first communication device may be used to execute the steps of the federated learning method described in the first aspect, and the second communication device may be used to execute the steps of the federated learning method described in the second aspect.
[0012] The eighth aspect provides a readable storage medium, which stores a program or instructions. When the program or instructions are executed by a processor, the steps of the method described in the first aspect are realized, or the steps of the method described in the second aspect are realized.
[0013] The ninth aspect provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor runs a program or instructions and is used to realize the steps of the method described in the first aspect, or the steps of the method described in the second aspect.
[0014] The tenth aspect provides a computer program / program product, which is stored in a storage medium. When the computer program / program product is executed by at least one processor, the steps of the method described in the first aspect are realized, or the steps of the method described in the second aspect are realized.
Advantages of the Invention
[0015] In an embodiment of the present application, the first information is received from a second communication device, and based on the first information, it can be determined whether the second communication device participates in the next round of federated learning. The first information includes at least one of second information for instructing whether the second communication device agrees to participate in federated learning, status information of the second communication device in this round of federated learning, and model performance information of this round of federated learning. Thereby, by combining the willingness, status information, and / or model performance of the second communication device, etc., and determining whether the second communication device participates in the next round of federated learning, a reasonable selection of members participating in federated learning can be realized.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Modes for Carrying Out the Invention
[0017] The following clearly describes the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, not all of them. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art shall fall within the protection scope of the present application.
[0018] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and are not for describing a specific order or sequence. It should be understood that such terms are interchangeable when appropriate, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, without limiting the number of objects. For example, the first object may be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the related objects before and after are in an "or" relationship.
[0019] It should be noted that the technology described in the embodiments of this application is not limited to the Long Term Evolution (LTE) / LTE-Advanced (LTE-A) system, but can also be applied to other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency Division Multiple Access (SC-FDMA), and other systems. The terms "system" and "network" in the embodiments of this application are always used interchangeably, and the described technology may be used in the systems and radio technologies mentioned above, or in other systems and radio technologies. The following description describes the New Radio (NR) system for illustrative purposes and uses NR terms in most of the following descriptions. However, these technologies may also be applied to applications other than NR system applications, such as the Sixth Generation (6 th Generation, 6G) communication system.
[0020] FIG. 1 shows a block diagram of a wireless communication system to which the embodiments of the present application are applicable. The wireless communication system includes a terminal 11 and a network-side device 12. Here, the terminal 11 may be a mobile phone, a tablet personal computer, a laptop computer (or called a notebook computer), a personal digital assistant (PDA), a palm-top computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, a vehicle user equipment (VUE), a pedestrian user equipment (PUE), a smart home (home devices having a wireless communication function, such as a refrigerator, a television, a washing machine or furniture, etc.), a game console, a personal computer (PC), a deposit and payment machine or a self-service machine, etc. The wearable device may include a smart watch, a smart trist band, smart earphones, smart glasses, smart accessories (such as a smart bracelet, a smart hand chain, a smart ring, a smart necklace, a smart ankle bracelet, a smart anklet, etc.), a smart band, smart clothing, etc. It should be noted that the terminal 11 in the embodiments of the present application is not limited to a specific type. The network-side device 12 may include an access network device or a core network device. The access network device may be called a wireless access network device, a radio access network (RAN), a wireless access network function or a wireless access network unit.The access network device may include a base station, a Wireless Local Area Networks (WLAN) access point, or a WiFi node, etc. The base station may be referred to as Node B, evolved Node B (eNB), access point, Base Transceiver Station (BTS), radio base station, radio transceiver, Basic Service Set (BSS), Extended Service Set (ESS), home B node, home evolved B node, Transmitting Receiving Point (TRP), or other appropriate terms in the art. As long as the same technical effect is achieved, the base station is not limited to specific technical terms. For the sake of explanation, in the embodiments of this application, only the base station in the NR system is taken as an example for introduction, and the specific type of the base station is not limited.The core network device may include, but is not limited to, at least one of a Network Data Analytic Function (NWDAF), a core network node, a core network function, a Mobility Management Entity (MME), an Access and Mobility Management Function (AMF), a Session Management Function (SMF), a User Plane Function (UPF), a Policy Control Function (PCF), a Policy and Charging Rules Function unit (PCRF), an Edge Application Server Discovery Function (EASDF), a Unified Data Management (UDM), a Unified Data Repository (UDR), a Home Subscriber Server (HSS), a Centralized network configuration (CNC), a Network Repository Function (NRF), a Network Exposure Function (NEF), a Local NEF (L-NEF), a Binding Support Function (BSF), an Application Function (AF), etc. It should be noted that in the embodiments of this application, only the core network device in the NR system is taken as an example for introduction, and the specific type of the core network device is not limited.
[0021] Optionally, in the embodiments of the present application, the network data analysis function NWDAF may be divided into two network elements, for example, a model training logical function (MTLF) and an analytics logical function (AnLF). Here, the model training logical function MTLF is mainly used to generate a model and perform model training, and may be a central server in federated learning or a member (clients) in federated learning. The analytics logical function AnLF is mainly used to perform inferences to generate prediction information or a model, etc., and may request a model from the MTLF, and this model may be generated by federated learning.
[0022] Optionally, the model in the embodiments of the present application may be an artificial intelligence (AI) model. There are various algorithm implementation methods for AI models, such as neural networks, decision trees, support vector machines, Bayesian classifiers, etc. The present application will be described by taking a neural network as an example, but does not limit the specific type of the AI module.
[0023] For example, a schematic diagram of a neural network may be shown as in Figure 2. Here, X1, X2…Xn, etc. are input values, Y is the output result, and one "〇" represents one neuron, which is a place for performing operations, and the result continues to be transmitted to the next layer. The input layer, hidden layer, and output layer composed of these many neurons are a neural network. The number of hidden layers and the number of neurons in each layer are the "network structure" of the neural network.
[0024] Also for example, a neural network is composed of neurons, and a schematic diagram of a neuron may be shown as in Figure 3. Here, a1, a k …a K(i.e., X1, X2... shown in Figure 2) is the input, w is the weight value (which may also be called the multiplication coefficient), b is the bias (which may also be called the addition coefficient), σ() is the activation function, z is the output value, and the corresponding calculation process is
Number
[0025] In the actual usage process, an AI model is a file that includes elements such as the network structure and parameter information. The trained AI model can be directly reused by its framework platform without the need for repeated construction or learning and can directly perform intelligent functions such as judgment and / or identification.
[0026] Federated learning aims to establish a federated learning model based on a distributed dataset. In the process of model training, the information related to the model can be exchanged (or exchanged in an encrypted form) among each party, but the raw data cannot be exchanged. This exchange does not expose any protected privacy parts of the data on each training node.
[0027] Optionally, the federated learning according to the embodiments of the present application is horizontal federated learning. The essence of horizontal federated learning is the cooperation of samples, which is applicable to scenarios where the business states among participants are the same, but the customers reached are different, that is, there are many overlapping features and few overlapping users. For example, the same service provided to different users by the CN domain and the RAN domain in a communication network (for example, each UE, that is, the samples are different), such as a mobility management (MM) service, a session management (SM) service, or a certain service. By associating the same data features of different samples of participants, horizontal federation can obtain a better model by increasing the number of training samples.
[0028] In the embodiments of the present application, the server (which may also be referred to as a central server or an organizer) in federated learning may be an MTLF divided by a network element device in the network, such as NWDAF. The members (which may also be referred to as participants) involved in federated learning may be network element devices in the network, such as MTLF divided by NWDAF, or may be terminals, etc. When performing federated learning, the server in federated learning may first select the members involved in federated learning. For example, it sends a request to a storage information network element such as NRF to request the acquisition of the capability information of each intelligent network element device such as MTLF, matches whether it can participate in federated learning according to the capability information, and then sends information such as the initialization model of federated learning to each selected member. Each member feeds back intermediate results, such as gradients, to the server after performing local model training. Then, the server aggregates the received intermediate results to update the global model. The steps of member selection - model distribution - local model training - intermediate result feedback - global model aggregation and update are repeated multiple times, and the model training can be stopped when the model converges or other situations occur.
[0029] In the following, with reference to the drawings, the federated learning method, apparatus, communication device, and readable storage medium according to the embodiments of the present application will be described in detail by way of several embodiments and their application scenarios.
[0030] Referring to FIG. 4, FIG. 4 is a flowchart of the federated learning method according to the embodiment of the present application. This method is used in a first communication device, which is specifically a server in federated learning and includes, but is not limited to, intelligent network element devices such as MTLF. As shown in FIG. 4, this method includes the following steps.
[0031] Step 41: The first communication device receives first information from the second communication device.
[0032] Step 42: The first communication device determines whether the second communication device participates in the next round of federated learning based on the first information.
[0033] In this embodiment, the above first information may include at least one of second information for instructing whether the second communication device agrees to participate in federated learning, the status information of the second communication device in this round of federated learning, the model performance information of this round of federated learning, etc., but is not limited thereto. For example, this second information may optionally be willingness information for instructing whether the second communication device has the willingness to participate in federated learning.
[0034] In addition, the above first information may further include the capability information of the second communication device. For example, this capability information is the capability information after the completion of the model training in this round, and includes, but is not limited to, whether it can become a participant (member) in federated learning, the accuracy information related to participating in the training model, etc. For example, after one round of local training is completed, the capability information of a member is that it can become a participant in federated learning, has the ability to perform local training, and the accuracy information related to participating in the training model is X, etc.
[0035] The above-mentioned second communication device is specifically a member (client) device in federated learning, and may include, but is not limited to, intelligent network element devices such as terminals and MTLF.
[0036] In some embodiments, the above-mentioned first information may be spontaneously reported by the second communication device (i.e., the member in federated learning). For example, by feeding back to the server in federated learning together with the results of local training, the consumption of signaling and the number of interactions can be reduced.
[0037] In some embodiments, when a member in federated learning no longer wants to participate in federated learning, for example, when other more important tasks arrive and the member no longer wants to participate in federated learning, or when there are too many tasks that need to handle situations such as exiting federated learning first, information indicating that the member does not want to participate in federated learning, that is, the intention information to exit federated learning, is fed back to the server in federated learning, which can assist in the selection of members in the federated learning process and realize a reasonable selection of members participating in federated learning. On the other hand, if this member has the intention to continue participating in federated learning, it may not be necessary to feed back information for instructing consent to participate in federated learning. At this time, the server tacitly approves that the member has the intention to continue participating in federated learning. It should be noted that a member in federated learning may also directly instruct the server that it wants to participate in federated learning.
[0038] In some other embodiments, when the state of a member in federated learning changes (for example, the load becomes heavy, etc.), the computing power required for local model training becomes insufficient and it becomes unqualified to be selected as a member participating in the next round of federated learning. Therefore, the member in federated learning sends the state information of this round of federated learning to the server in federated learning. By the server determining whether this member will participate in the next round of federated learning, it can assist in the selection of members in the federated learning process, realize a reasonable selection of members participating in federated learning, improve the training efficiency, for example, avoid member dropout with a deteriorating state (for example, when a member with a deteriorating state does not feedback the result within a predetermined time), and select members that can bring higher efficiency.
[0039] In some other embodiments, when the data of a member in federated learning has already been learned multiple times or has already been incorporated into the global model of federated learning, this global model may be overfitted to the environment of this member, and this member may become unqualified to be selected as a member participating in the next round of federated learning. Therefore, by temporarily suspending the training of this member for several rounds at this time, model convergence can be realized more quickly. Therefore, the member in federated learning sends the model performance information in this round of federated learning to the server in federated learning. By the server determining whether this member will participate in the next round of federated learning, it can assist in the selection of members in the federated learning process, realize a reasonable selection of members participating in federated learning, improve the training efficiency, for example, avoid member dropout with a deteriorating state (for example, when a member with a deteriorating state does not feedback the result within a predetermined time), and select members that can bring higher efficiency.
[0040] Optionally, after receiving the above first information, the first communication device may select a third communication device that participates in the next round of federated learning based on this first information. Different from the second communication device, this third communication device is a new member (client) device specifically involved in federated learning, and may include, but is not limited to, intelligent network element devices such as terminals and MTLF. For example, if it is determined based on the received first information that more member devices are not suitable for participating in the next round of federated learning, new members participating in the next round of federated learning can be selected to ensure the smooth progress of federated learning.
[0041] In the embodiments of the present application, the above state information may be used to explain the state information after the local training of the second communication device (i.e., the member in federated learning) in this round of federated learning, and may include, but is not limited to, at least one of the following.
[0042] 1) Load information of the second communication device in this round of federated learning.
[0043] In this embodiment, this load information may be understood as load status information and may represent the load status of a network element, for example, a network function (NF).
[0044] Optionally, this load information may include at least one of average load information and peak load information, etc. The average load information may be understood as the average load value within the scope of this round of federated learning. For example, in one local training, the average load of a certain member is 70% and the peak load is 80%.
[0045] 2) Resource usage information of the second communication device in this round of federated learning.
[0046] In this embodiment, this resource usage information may be understood as resource usage status information.
[0047] Optionally, this resource usage information may include at least one of average resource usage information and peak resource usage information. The average resource usage information may be understood as the average resource usage situation within the scope of the federated learning in this round.
[0048] For example, the resource usage end corresponding to this resource usage information (for example, resource usage) may include, but is not limited to, a central processing unit (CPU), memory, magnetic disk, graphics processing unit (GPU), etc. This resource usage information may include power consumption information, etc.
[0049] For example, in one local training, the average resource usage situation of a certain member is CPU usage rate of 60%, GPU usage rate of 80%, memory usage rate of 70% (for example, occupying 12GB, that is, when represented by a numerical value), and magnetic disk capacity usage rate of 40%, and the peak resource usage situation of this member is CPU usage rate of 80%, GPU usage rate of 100%, memory usage rate of 80% (for example, occupying 14GB, that is, when represented by a numerical value), and magnetic disk capacity usage rate of 50%.
[0050] In the embodiments of the present application, the above model performance information is optionally the model performance information before and / or after the start of local model training, the first model performance information after the completion of local model training, and may include at least one of the second model performance information before the start of local model training.
[0051] Optionally, the model performance information may include at least one of accuracy and mean absolute error (MAE). Note that it may further include, but is not limited to, at least one of precision, recall, F1 score, area under curve (AUC), sum of squares due to error (SSE), sum of variances, mean squared error (MSE), variance, root mean squared error (RMSE), standard deviation, and coefficient of determination (R-Squared).
[0052] In some embodiments, the first model performance information may include accuracy, mean absolute error MAE, and the like. The second model performance information may include accuracy, mean absolute error MAE, and the like.
[0053] As can be understood, the first model performance information is mainly used to explain the performance of the model based on its local data after the local model training in this round of federated learning is completed. It may include a certain statistical parameter and the numerical value corresponding to this parameter, such as the accuracy of the model and a specific value (e.g., 80%), and the mean absolute error MAE and its value (e.g., 0.1). The second model performance information is mainly used to explain the performance of the model based on its local data before the local model training in this round of federated learning starts. That is, it is necessary to perform a statistical calculation of the model performance once after receiving the model. It may include a certain statistical calculation parameter and the numerical value corresponding to this parameter, such as the accuracy of the model and a specific value (e.g., 70%), and the mean absolute error MAE and its value (e.g., 0.15).
[0054] What should be noted is that the accuracy rate is the percentage of the number of correct predictions to the total number of predictions. In the model training stage, the dataset includes input data and labels (label data), and the two are in a corresponding relationship. A set of input data corresponds to one or a set of labels. By comparing the predicted value generated by the model with the label corresponding to this training, it is determined whether this training is correct. The mean absolute error MAE represents the average value of the absolute error between the predicted value and the true value, and the calculation method is as follows.
[0055] [Number] Here, h(x i ) represents the predicted value of the model, y i represents the corresponding true value, and m represents the number of training samples.
[0056] In the embodiments of the present application, whether the second communication device feeds back the first information may be determined by the first communication device. Optionally, the first communication device may also send third information to the second communication device, and the third information is used to identify that the second communication device needs to feed back the first information. On the other hand, when the first communication device does not send the third information, that is, when the second communication device does not receive the third information, the second communication device does not need to feed back the first information.
[0057] Optionally, the third information is information for identifying that the second communication device needs to feed back the second information (for example, this information is an identifier for which the second information needs to be fed back), and information for identifying that the second communication device needs to feed back status information (for example, this information is an identifier for which status information needs to be fed back), and Information for identifying that a second communication device needs to feedback model performance information (for example, this information may include, but is not limited to, an identifier indicating that the second communication device needs to feedback model performance information after local model training is completed, and / or an identifier indicating that the second communication device needs to feedback model performance information before local model training starts).
[0058] It should be noted that the information for identifying that the second communication device needs to feedback status information is mainly used to explain that the second communication device needs to feedback its status information after the local model training of the current round of federated learning is completed. It may be specified that the specific status information can be at least one of the member's load status (for example, NF load), the member's resource usage status (for example, resource usage includes CPU, memory, disk, and / or GPU, etc.).
[0059] The information for identifying that the second communication device needs to feedback model performance information is mainly used to explain that the second communication device needs to feedback the model performance information before and / or after the completion of local model training after the local model training of the current round of federated learning is completed, and this model performance information includes the above-mentioned first model performance information and / or second model performance information.
[0060] Optionally, transmitting the above third information may include at least one of the following.
[0061] The first communication device transmits third information to the second communication device based on a preset policy. Here, this preset policy may refer to when or under what circumstances the first communication device transmits the third information to the second communication device. For example, after every five rounds of training, the third information is transmitted to the second communication device, or after a certain second communication device participates in five rounds of training, the third information is transmitted to this second communication device. This preset policy can not only indicate the necessity of feedback from the second device, but also indicate when or under what circumstances to request feedback. For example, when it is indicated by the preset policy that the second communication device needs to feedback the first information, the first communication device may transmit the third information to the second communication device. However, when it is indicated by the preset policy that the second communication device does not need to feedback the first information, the first communication device does not transmit the third information to the second communication device. This preset policy may be something predefined, agreed upon by a protocol, etc.
[0062] The first communication device transmits third information to the second communication device according to the requirements in the model training process based on federated learning. For example, when the first communication device expects to determine whether the second communication device will participate in the next round of federated learning based on the willingness, status, and / or model performance of the second communication device, etc., the first communication device may transmit the third information to the second communication device. Otherwise, the first communication device does not transmit the third information to the second communication device. That is, the first communication device may autonomously decide whether to transmit the third information to the second communication device.
[0063] Optionally, transmitting the above third information may include transmitting a first request to a second communication device, where the first request is used to request the second communication device to participate in federated learning, and the third information is carried in the first request. In this way, by transmitting the third information by means of the first request for requesting the second communication device to participate in federated learning, the consumption of signaling and the number of interactions can be reduced.
[0064] In an embodiment of the present application, when a first communication device receives a plurality of pieces of model performance information from a plurality of second communication devices, first, the plurality of pieces of model performance information are summarized to obtain third model performance information, and based on this third model performance information, it is determined whether model training has ended, for example, it may be determined whether the model has converged. For example, when the third model performance information includes accuracy and this accuracy is higher than a preset threshold, it is determined that model training has ended, otherwise model training may continue, or when the third model performance information includes the mean absolute error MAE and this MAE is lower than a preset threshold, it is determined that model training has ended, otherwise model training may continue.
[0065] Optionally, the above summarization method includes, but is not limited to, calculating the average value of a plurality of pieces of model performance information, calculating the weighted average value of a plurality of pieces of model performance information, etc. When calculating the weighted average value, the corresponding weight may be determined by the first communication device, for example, a preset weight or a weight calculated by itself may be adopted.
[0066] In some embodiments, the first communication device summarizes a plurality of pieces of first model performance information (i.e., model performance information after local model training is completed), and determines whether model training has ended based on the model performance information after summarization.
[0067] Furthermore, after obtaining the third model performance information, the first communication device may feedback this third model performance information to the model user to facilitate the model user's understanding of the model performance.
[0068] The above embodiments mainly explain the present application from the perspective of the first communication device (i.e., the server in federated learning). The following will explain the present application from the perspective of the second communication device (i.e., the member in federated learning).
[0069] Referring to FIG. 5, FIG. 5 is a flowchart of a federated learning method according to an embodiment of the present application. This method is used for a second communication device, which is specifically a member (client) in federated learning and includes, but is not limited to, intelligent network element devices such as terminals and MTLF. As shown in FIG. 5, this method includes the following steps.
[0070] Step 51: The second communication device determines first information.
[0071] Step 52: The second communication device transmits the first information to the first communication device, and the first information is used for the first communication device to determine whether the second communication device participates in the next round of federated learning.
[0072] In this embodiment, the above first information may include, but is not limited to, at least one of second information for instructing whether the second communication device agrees to participate in federated learning, the status information of the second communication device in this round of federated learning, and the model performance information of this round of federated learning.
[0073] The above first communication device is specifically a server in federated learning and may include, but is not limited to, intelligent network element devices such as MTLF.
[0074] In some embodiments, the above first information may be spontaneously reported by the second communication device (i.e., a member in the federated learning), for example, by feeding back to the server in the federated learning together with the results of local training, so as to reduce the consumption of signaling and the number of interactions.
[0075] In the federated learning method according to the embodiments of the present application, by sending first information including at least one of second information for instructing whether the second communication device agrees to participate in the federated learning, state information of the second communication device in this round of federated learning, first model performance information after the local model training of this round of federated learning is completed, and second model performance information before the local model training of this round of federated learning to the first communication device, the first communication device can combine the willingness, state information, and / or model performance of the second communication device, etc., to determine whether the second communication device will participate in the next round of federated learning, thereby realizing a reasonable selection of members participating in the federated learning, improving the training efficiency, for example, avoiding member dropout (i.e., in the case where there are members who do not feedback results within a predetermined time), and selecting members who can bring higher efficiency.
[0076] In the embodiments of the present application, the above state information may be used to explain the state information of the second communication device (i.e., a member in the federated learning) after the local training of this round of federated learning is completed, and may include at least one of the following, but is not limited thereto.
[0077] 1) Load information of the second communication device in this round of federated learning.
[0078] In this embodiment, this load information may be understood as load status information and may represent the NF load status.
[0079] Optionally, this load information may include at least one of average load information and peak load information, etc. The average load information may be understood as the average load value within the scope of the current round of federated learning. For example, in one local training, the average load of a certain member is 70%, and the peak load is 80%.
[0080] 2) Resource usage information of the second communication device in the current round of federated learning.
[0081] Optionally, this resource usage information may include at least one of average resource usage information and peak resource usage information. The average resource usage information may be understood as the average resource usage status within the scope of the current round of federated learning.
[0082] For example, the resource users corresponding to this resource usage information (e.g., resource usage) may include, but are not limited to, a central processing unit (CPU), memory, magnetic disk, graphics processing unit (GPU), etc. This resource usage information may include power consumption information, etc.
[0083] For example, in one local training, the average resource usage status of a certain member is 60% CPU usage, 80% GPU usage, 70% memory usage (e.g., occupying 12GB, i.e., when represented by a numerical value), 40% magnetic disk capacity usage, and the peak resource usage status of this member is 80% CPU usage, 100% GPU usage, 80% memory usage (e.g., occupying 14GB, i.e., when represented by a numerical value), 50% magnetic disk capacity usage.
[0084] In the embodiments of the present application, the above model performance information is optionally the model performance information before and / or after the start of local model training, the first model performance information after the completion of local model training, It may include at least one of the second model performance information before the start of local model training.
[0085] Optionally, the model performance information may include at least one of accuracy and mean absolute error (MAE), etc. For example, the first model performance information described above may include accuracy and mean absolute error (MAE). The second model performance information described above may include accuracy and mean absolute error (MAE).
[0086] In an embodiment of the present application, whether the second communication device feeds back the first information may be determined by the first communication device. Determining the first information described above may first include receiving third information from the first communication device, where the third information is used to identify that the second communication device needs to feed back the first information, and based on the third information, determining the first information.
[0087] Optionally, the third information may be information for identifying that the second communication device needs to feed back the second information (for example, this information is an identifier for which the second communication device needs to feed back the second information), and information for identifying that the second communication device needs to feed back status information (for example, this information is an identifier for which the second communication device needs to feed back status information), and information for identifying that the second communication device needs to feed back model performance information (for example, this information includes an identifier for which the second communication device needs to feed back model performance information after local model training is completed and / or an identifier for which the second communication device needs to feed back model performance information before local model training starts), but is not limited thereto.
[0088] Optionally, receiving the third information from the first communication device may include receiving a first request from the first communication device, where the first request is used to request the second communication device to participate in the federated learning, and the third information is carried in the first request. In this way, by transmitting the third information by the first request for requesting the second communication device to participate in the federated learning, the consumption of signaling and the number of interactions can be reduced.
[0089] Hereinafter, the federated learning process in the embodiments of the present application will be described with reference to FIG. 6.
[0090] In the embodiments of the present application, the federated learning server is an NWDAF (e.g., MTLF), and the federated learning members (clients) are NWDAFs (e.g., MTLF). As shown in FIG. 6, the specific federated learning process includes the following.
[0091] Step 61: The federated learning consumer (e.g., NWDAF (AnLF)) sends a model request (e.g., Nnwdaf_MLModelProvision_Subscribe) to the federated learning server (e.g., NWDAF (MTLF)), and this model request is used to request to obtain a model for completing its own task. At this time, the server determines whether to trigger the federated learning based on situations such as the local configuration or the request of the federated learning consumer, and performs the initialization and member selection of the federated learning.
[0092] Step 62: If the federated learning is triggered, when the server selects members, it may initialize and formulate the policy for the federated learning. For example, it may stipulate how many rounds of training are performed before collecting the state information once, and / or how many rounds of training are performed before collecting the model performance information, etc.
[0093] Step 63: The server sends an association learning task request (e.g., Nnwdaf_MLModelTraining_Subscribe) to each member (clients) to request participation in the association learning, and performs local training for the association learning based on the global model and the local data of each member. This task request may include a task identifier (e.g., analytic ID), model initialization information (e.g., including training parameters), information indicating the need to feedback status information / information for identifying model performance information (i.e., feedback requirement), etc.
[0094] Here, the analytic ID is mainly used to indicate what task the corresponding model is used for. The model initialization information is used to describe the model and the configuration information in this round of association learning, etc. The described model refers to the model itself. For example, it describes how the model is composed in terms of algorithms, architectures, parameters, and hyperparameters, etc., or describes the model itself, such as model files, address information of model files, etc. The configuration information in this round of association learning refers to information such as the number of rounds for local training and the data type to be used in the local training process of this round of association learning. Regarding the information indicating the need to feedback status information / information for identifying model performance information, reference may be made to the description of the above embodiments, and no further explanation will be given here.
[0095] Step 64: The member sends a data acquisition request (e.g., Ndccf_DataManagement_Subscribe / Nnf_EventExposure_Subscribe) to the area where it is located or the data source to which it belongs to collect data and perform local model training. Depending on the task, the network elements providing the data are also different, such as UPF, OAM, UDM, etc.
[0096] Step 65: The data source returns a response to the corresponding member, and this response includes the requested data. This response is, for example, Ndccf_DataManagement_Notify / Nnf_EventExposure_Notify.
[0097] Step 66: Each member performs local model training using the data obtained based on Steps 64 and 65, generates intermediate results, and feeds back to the server in subsequent steps. The server aggregates and updates the global model and analyzes the model performance using local data.
[0098] For example, the analysis of model performance may be to calculate accuracy or MAE using the model after local training and local data. If the identifier information that needs to feedback the model performance information before local training is carried in the task request in Step 63, the member needs to perform statistical calculations of model performance before performing local training.
[0099] In one implementation, the member NWDAF uses the number of times the model prediction result is correct divided by the total number of predictions as the local training accuracy of the model. That is, the calculation formula is: local training accuracy = number of correct result times ÷ total number of times. Specifically, the member NWDAF may set one verification dataset for evaluating the local training accuracy. This verification dataset includes the input data for the model (input data) and the true label data (label / ground truth). The member NWDAF inputs the input data into the trained model to obtain output data. The member NWDAF further compares whether the output data matches the true label data, and further uses the above calculation formula to obtain the value of the local training accuracy. Explanation: The concept that the prediction result is correct does not necessarily mean that the result completely matches the label data. There was a certain difference between the two before, but when this difference is within the allowable range, the prediction result may be considered correct.
[0100] In one implementation, the member NWDAF calculates the mean value of the sum of squared point errors corresponding to the prediction data and the label data (label value, raw data) to obtain the MAE, as shown in the following calculation formula: local training
Number
Number
[0101] Step 67: Each member, either spontaneously or in response to the request in Step 63, feeds back to the server the intermediate result after the completion of local training and information such as motivation, status, and / or model performance (the above first information). For example, information such as motivation, status, and / or model performance may be fed back by a feedback message corresponding to the request message of the federated learning training process. This feedback message may optionally be a notify message.
[0102] In one implementation, when the member client discovers that its training situation is good, for example, the accuracy reaches a certain threshold (this threshold may be carried in the model initialization information in Step 63, may be carried in the model request, or may be obtained / configured in advance), it may spontaneously feed back the model performance information. Or, the member client may spontaneously feed back its motivation, status, and / or model performance information, etc. in each round. The server can use the intermediate result to update the global model and use information such as motivation, status, and / or model performance to assist in the selection decision of members in the next round of federated learning.
[0103] Step 68: The server aggregates the intermediate results and updates the global model based on the fed-back intermediate results, and determines whether the corresponding member needs to further participate in the next round of federated learning based on information such as the motivation, status, and / or model performance of the feedback. Note that by aggregating the model performance information, the overall / global training situation of the model can be obtained.
[0104] For example, after obtaining the intermediate results fed back from each client, the server can aggregate these intermediate results using algorithms of the server, such as average, weighted average, etc., and further update the global model using these intermediate results. Also, for example, based on information such as the motivation, status, and / or model performance fed back from the client, the server can determine whether this client can participate in the next round of federated learning. For example, if this client indicates in its motivation information that it wants to withdraw from federated learning, the server will not select this client to participate in the next round of federated learning. Also, for example, if the status information fed back from this client shows that the CPU usage rate is 90%, the GPU usage rate is 100%, the memory usage rate is 80% (e.g., 14GB when represented by a numerical value), and the magnetic disk capacity usage rate is 50%, the server may consider that this client is not a good member to participate in federated learning. When performing local training, the GPU is already fully utilized, and the next training may take a long time or the connection may be disconnected. Therefore, this client will not be selected in the member selection for the next round. Also, for example, if the model performance fed back from this client is 98% accuracy, but at the same time, the model performance fed back from other clients is basically between 60% and 80%, the server may consider that the model is already overfitted in the environment of this client and it is necessary to pause the training of this client, and there is a possibility that this client will not be selected in the member selection for the next round.
[0105] What should be pointed out is that aggregating model performance information to obtain the overall / global training situation of the model means that the server collects the model performance information of each client and generates a single global training situation by means of methods such as average or weighted average. For example, if 5 clients participate in federated learning and they feedback their model performance, and assume that the accuracies are 70%, 72%, 75%, 68% and 65% respectively, then the server can obtain the global training situation by calculating the average value of these accuracies, that is, the accuracy of the global training situation is (70% + 72% + 75% + 68% + 65%) / 5 = 70% That is.
[0106] After the re-selection of members is completed, steps 63 to 68 may be repeatedly executed until the model converges.
[0107] Step 69: After the model training of federated learning is completed, the server feeds back the trained model and the overall / global model performance to the consumer (for example, AnLF).
[0108] In the federated learning method according to the embodiments of the present application, the execution body may be a federated learning device. In the embodiments of the present application, taking the federated learning device executing the federated learning method as an example, the federated learning device according to the embodiments of the present application will be described.
[0109] Referring to FIG. 7, FIG. 7 is a schematic structural diagram of a federated learning device according to an embodiment of the present application. This device is used for a first communication device, and this first communication device is specifically a server in federated learning, including but not limited to intelligent network element devices such as MTLF. As shown in FIG. 7, the federated learning device 70 includes A first receiving module 71 for receiving first information from a second communication device, wherein the first information includes at least one of second information for indicating whether the second communication device agrees to participate in federated learning, state information of the second communication device in this round of federated learning, and model performance information of this round of federated learning. The first receiving module 71 And a first decision module 72 for determining whether the second communication device participates in the next round of federated learning based on the first information.
[0110] Optionally, the state information includes At least one of load information and Resource usage information.
[0111] Optionally, the load information includes at least one of average load information and peak load information, and The resource usage information includes at least one of average resource usage information and peak resource usage information.
[0112] Optionally, the model performance information includes At least one of first model performance information after local model training is completed and Second model performance information before local model training starts.
[0113] Optionally, the model performance information includes at least one of accuracy, mean absolute error, precision, and mean squared error.
[0114] Optionally, the federated learning device 70 further includes A first transmitting module for transmitting third information to the second communication device, where the third information is used to identify that the second communication device needs to feedback the first information.
[0115] Optionally, the third information includes Information for identifying that the second communication device needs to feedback the second information (for example, this information is an identifier indicating that the second information needs to be fed back), and information for identifying that the second communication device needs to feedback status information, and includes at least one of information for identifying that the second communication device needs to feedback model performance information.
[0116] Optionally, the first transmission module is specifically used for at least one of transmitting third information to the second communication device based on a preset policy and transmitting third information to the second communication device according to the requirements in the model training process based on federated learning.
[0117] Optionally, the first transmission module is specifically used for transmitting a first request to the second communication device, the first request is used for requesting the second communication device to participate in federated learning, and the third information is carried in the first request.
[0118] Optionally, the federated learning device 70 further includes a processing module for determining whether model training is completed based on third model performance information obtained by summarizing a plurality of the model performance information when the first communication device receives a plurality of model performance information from a plurality of second communication devices.
[0119] Optionally, the federated learning device 70 further includes a feedback module for feeding back the third model performance information to the model user.
[0120] Optionally, the federated learning device 70 It further includes a selection module for selecting a third communication device to participate in the next round of federated learning based on the first information. This third communication device is different from the second communication device and is specifically a new member (client) device involved in federated learning, and may include, but is not limited to, intelligent network element devices such as terminals and MTLF. For example, based on the received first information, if it is determined that more members are no longer suitable to participate in the next round of federated learning, new members participating in the next round of federated learning can be selected to ensure the smooth progress of federated learning.
[0121] The federated learning device 70 according to the embodiment of the present application realizes each process realized by the method embodiment shown in FIG. 4 and can achieve the same technical effect, and will not be described further here to avoid repetition of the description.
[0122] Referring to FIG. 8, FIG. 8 is a schematic structural diagram of a federated learning device according to an embodiment of the present application. This device is used for a second communication device, and this second communication device is specifically a member (client) in federated learning and includes, but is not limited to, intelligent network element devices such as terminals and MTLF. As shown in FIG. 8, the federated learning device 80 includes a second determination module 81 for determining first information, where the first information includes at least one of second information for instructing whether the second communication device agrees to participate in federated learning, status information of the second communication device in this round of federated learning, and model performance information of this round of federated learning; and the second determination module 81 a second transmission module 82 for transmitting the first information to the first communication device, where the first information is used for the first communication device to determine whether the second communication device participates in the next round of federated learning; and the second transmission module 82
[0123] Optionally, the status information includes load information, and It includes at least one of the resource usage information.
[0124] Optionally, the load information includes at least one of the average load information and the peak load information. The resource usage information includes at least one of the average resource usage information and the peak resource usage information.
[0125] Optionally, the model performance information includes at least one of the first model performance information after the completion of the local model training and the second model performance information before the start of the local model training.
[0126] Optionally, the model performance information includes at least one of the accuracy, the mean absolute error, the precision, and the mean squared error.
[0127] Optionally, the federated learning device 80 includes a second receiving module for receiving third information from the first communication device, and the third information is used to identify that the second communication device needs to feedback the first information. Specifically, the second decision module 81 is used to determine the first information based on the third information.
[0128] Optionally, the second receiving module is further used to receive a first request from the first communication device, and the first request is used to request the second communication device to participate in the federated learning, and the third information is carried in the first request.
[0129] The federated learning device 80 according to the embodiments of the present application realizes each process realized by the embodiment of the method shown in FIG. 5 and can achieve the same technical effects. To avoid repeated description, it will not be described further here.
[0130] Optionally, as shown in FIG. 9, the embodiment of the present application further provides a communication device 90, which includes a processor 91 and a memory 92. Programs or instructions that can run on the processor 91 are stored on the memory 92. For example, when the communication device 90 is a first communication device, when the program or instruction is executed by the processor 91, each step of the embodiment of the collaborative learning method shown in FIG. 4 above can be realized, and the same technical effect can be achieved. When the communication device 90 is a second communication device, when the program or instruction is executed by the processor 91, each step of the embodiment of the collaborative learning method shown in FIG. 5 above can be realized, and the same technical effect can be achieved. To avoid repetition of the description, it will not be described further here.
[0131] The embodiment of the present application further provides a communication device, which includes a processor and a communication interface. For example, when the communication device is a first communication device, the communication interface is used to receive first information from a second communication device, and the processor is used to determine whether the second communication device participates in the next round of collaborative learning based on the first information. Or when the communication device is a second communication device, the processor is used to determine the first information, and the communication interface is used to send the first information to the first communication device. The first information includes at least one of second information for instructing whether the second communication device agrees to participate in collaborative learning, state information of the second communication device in this round of collaborative learning, first model performance information after the local model training of this round of collaborative learning is completed, and second model performance information before the local model training of this round of collaborative learning starts. This embodiment corresponds to the embodiment of the above method, and each implementation process and realization method of the embodiment of the above method can be applied to this embodiment, and the same technical effect can be achieved.
[0132] Specifically, FIG. 10 is a schematic diagram of the hardware structure for realizing the terminal of the embodiment of the present application.
[0133] This terminal 1000 includes, but is not limited to, at least some of the components such as a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010.
[0134] As can be understood by those skilled in the art, the terminal 1000 may further include a power source (e.g., a battery) for supplying power to each component. The power source may be logically connected to the processor 1010 by a power management system, thereby enabling functions such as charge and discharge management and power consumption management to be realized by the power management system. The terminal structure shown in FIG. 10 does not constitute a limitation on the terminal. The terminal may include more or fewer components than those shown, or a combination of some components, or a different arrangement of components, which will not be further described herein.
[0135] It should be understood that in the embodiments of the present application, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The graphics processing unit 10041 processes the image data of a still image or a video obtained by an image capture device (e.g., a camera) in a video capture mode or an image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (e.g., volume control buttons, switch buttons, etc.), a trackball, a mouse, and an operation lever, which will not be further described herein.
[0136] In an embodiment of the present application, after receiving downlink data from a network-side device, the radio frequency unit 1001 can transmit it to the processor 1010 for processing. Also, the radio frequency unit 1001 can transmit uplink data to the network-side device. Generally, the radio frequency unit 1001 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc.
[0137] Memory 1009 may be used to store software programs or instructions and various data. Memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Here, the first storage area can store an operating system, application programs or instructions required for at least one function (for example, a voice playback function, an image playback function, etc.). Note that Memory 1009 may include volatile memory or non-volatile memory, or Memory 1009 may include both volatile and non-volatile memory. Here, the non-volatile memory may be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory may be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), or a Direct Rambus RAM (DRRAM). The Memory 1009 in the embodiments of this application includes these and any other suitable types of memory, but is not limited thereto.
[0138] Processor 1010 may include one or more processing units. Optionally, processor 1010 integrates an application processor and a modem processor, where the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, for example, a baseband processor. As can be understood, the above modem processor may not be integrated into processor 1010.
[0139] Optionally, terminal 1000 may also be a member in federated learning, and processor 1010 is used to determine first information. It is used to transmit the first information to a server in federated learning. The first information is used for the server to determine whether terminal 1000 participates in the next round of federated learning. The first information includes at least one of second information for instructing whether terminal 1000 agrees to participate in federated learning, state information of terminal 1000 in this round of federated learning, model performance information of this round of federated learning, etc.
[0140] The terminal 1000 according to the embodiments of the present application realizes each process realized by the method embodiments shown in FIG. 5 and can achieve the same technical effects. To avoid repetition of description, it will not be described further here.
[0141] Specifically, the embodiments of the present application further provide a network-side device. As shown in FIG. 11, this network-side device 110 includes a processor 111, a network interface 112, and a memory 113. Here, the network interface 112 is, for example, a common public radio interface (CPRI).
[0142] Specifically, the network-side device 110 in the embodiments of the present application further includes instructions or programs stored in the memory 113 and executable on the processor 111. The processor 111 calls the instructions or programs in the memory 113, executes the methods executed by the respective modules shown in FIGS. 7 and / or 8, and achieves the same technical effects. To avoid repetition of the description, it will not be further described herein.
[0143] The embodiments of the present application further provide a readable storage medium, in which programs or instructions are stored. When these programs or instructions are executed by a processor, each process of the embodiments of the above collaborative learning method is realized, and the same technical effects can be achieved. To avoid repetition of the description, it will not be further described herein.
[0144] Here, this processor is the processor in the terminal described in the above embodiments. This readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical disks.
[0145] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor runs programs or instructions and is used to realize each process of the embodiments of the above collaborative learning method, and the same technical effects can be achieved. To avoid repetition of the description, it will not be further described herein.
[0146] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-level chip, system chip, chip system, or system-on-chip, etc.
[0147] Embodiments of the present application further provide a computer program / program product, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to implement each process of the embodiments of the above-described collaborative learning method and achieve the same technical effects, and for the sake of avoiding repetition of description, it will not be further described herein.
[0148] Embodiments of the present application further provide a communication system, the communication system including the first communication device and the second communication device as described above, the first communication device may be used to execute the steps of the collaborative learning method shown in FIG. 4, and the second communication device may be used to execute the steps of the collaborative learning method shown in FIG. 5.
[0149] It should be noted that in this specification, the term "including", "comprising" or any other variation thereof is intended to cover non-exclusive "including", whereby a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements specific to such a process, method, article or device. In the case of no further limitation, for an element defined by the phrase "comprising one...", it is not excluded that there are also other same elements in the process, method, article or device including this element. It should be pointed out that the scope of the method and device in the embodiments of the present application is not limited to executing functions in the order shown or discussed, and may include executing functions in a basically simultaneous manner or in a reverse order based on the related functions, for example, it is possible to execute a method described in a procedure different from that described, and various steps can be added, omitted or combined. Also, features described with reference to some examples can be combined in other examples.
[0150] As can be clearly understood by those skilled in the art from the description of the above embodiments, the method of the above embodiments can be implemented in the form of software and the necessary general-purpose hardware platform. Of course, it may also be implemented by hardware, but in many cases, the former is a more preferred embodiment. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the related technology, may be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0151] The above has described the embodiments of the present application in conjunction with the drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely exemplary and not restrictive. Those skilled in the art can make many forms without departing from the spirit of the present application and the scope of the claims, and all belong to the protection scope of the present application.
Claims
1. A federated learning method, comprising: a first communication device receiving first information from a second communication device, the first information including at least one of second information for instructing whether the second communication device agrees to participate in federated learning, status information of the second communication device in this round of federated learning, model performance information of this round of federated learning, and willingness information for instructing the second communication device to withdraw from federated learning; the first communication device determining, based on the first information, whether the second communication device participates in the next round of federated learning.
2. The status information includes at least one of load information and resource usage information. The method according to claim 1.
3. The load information includes at least one of average load information and peak load information, and the resource usage information includes at least one of average resource usage information and peak resource usage information. The method according to claim 2.
4. The model performance information includes at least one of first model performance information after local model training is completed and second model performance information before local model training starts. The method according to claim 1.
5. The model performance information includes at least one of accuracy, mean absolute error, precision, and mean squared error. The method according to claim 1.
6. The method further includes the first communication device transmitting third information to the second communication device, where the third information is used to identify that the second communication device needs to feedback the first information. The method according to claim 1.
7. The third information includes information for identifying that the second communication device needs to feedback the second information, information for identifying that the second communication device needs to feedback status information, and at least one of information for identifying that the second communication device needs to feedback model performance information. The method according to claim 6.
8. The transmitting the third information to the second communication device includes the first communication device transmitting the third information to the second communication device based on a preset policy. The method according to claim 6, comprising at least one of: transmitting third information to the second communication device according to the demand in the model training process based on federated learning by the first communication device.
9. The transmitting the third information to the second communication device as described above The method according to claim 6, comprising: the first communication device transmitting a first request to the second communication device, where the first request is used to request the second communication device to participate in federated learning, and the third information is carried in the first request.
10. When the first communication device receives a plurality of pieces of the model performance information from a plurality of second communication devices, the method further comprises: the first communication device summarizing the plurality of pieces of the model performance information to obtain third model performance information; and the first communication device determining whether model training has ended based on the third model performance information. The method according to claim 1.
11. After obtaining the third model performance information, the method further comprises: the first communication device feeding back the third model performance information to a model user. The method according to claim 10.
12. After receiving the first information as described above, the method further comprises: the first communication device selecting a third communication device to participate in the next round of federated learning based on the first information, where the third communication device is different from the second communication device and is a new member device participating in federated learning. The method according to claim 1.
13. A federated learning method, comprising: a second communication device determining first information, where the first information includes at least one of: second information for indicating whether the second communication device agrees to participate in federated learning, status information of the second communication device in this round of federated learning, model performance information of this round of federated learning, and intention information for instructing the second communication device to withdraw from federated learning; and the second communication device transmitting the first information to a first communication device, where the first information is used for the first communication device to determine whether the second communication device will participate in the next round of federated learning.
14. The status information includes load information, The method according to claim 13, comprising at least one of resource usage information.
15. The model performance information is at least one of first model performance information after local model training is completed and second model performance information before local model training starts, the method according to claim 13.
16. Determining the first information is the second communication device receiving third information from the first communication device, the third information being used to identify that the second communication device needs to feedback the first information, and the second communication device determining the first information based on the third information, the method according to claim 13.
17. Receiving the third information from the first communication device is the second communication device receiving a first request from the first communication device, the first request being used to request the second communication device to participate in federated learning, and the third information being carried in the first request, the method according to claim 16.
18. A federated learning device, comprising a first receiving module for receiving first information from a second communication device, the first information including at least one of second information for indicating whether the second communication device agrees to participate in federated learning, state information of the second communication device in this round of federated learning, model performance information of this round of federated learning, and intention information for instructing the second communication device to withdraw from federated learning; and a first decision module for determining whether the second communication device participates in the next round of federated learning based on the first information.
19. A federated learning device, comprising a second decision module for determining first information, the first information including at least one of second information for indicating whether a second communication device agrees to participate in federated learning, state information of the second communication device in this round of federated learning, model performance information of this round of federated learning, and intention information for instructing the second communication device to withdraw from federated learning. A federated learning device including a second transmission module for transmitting the first information to the first communication device, where the first information is used by the first communication device to determine whether the second communication device participates in the next round of federated learning.
20. A communication device including a processor and a memory, where the memory stores a program or instructions that can run on the processor, and when the program or instructions are executed by the processor, the steps of the federated learning method according to any one of claims 1 to 12 are realized, or the steps of the federated learning method according to any one of claims 13 to 17 are realized.
21. A readable storage medium that stores a program or instructions, and when the program or instructions are executed by a processor, the steps of the federated learning method according to any one of claims 1 to 12 are realized, or the steps of the federated learning method according to any one of claims 13 to 17 are realized.
Citation Information
Patent Citations
Federal learning method and device
CN114079902A