Model training method and device

CN120035833APending Publication Date: 2025-05-23BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101122.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In hierarchical federated learning, due to the heterogeneity of computing resources and data resources of terminal devices, the quality of global model aggregation is reduced, training accuracy is low, and energy consumption and delay are large. Existing technologies cannot effectively solve this problem. .

Method used

By receiving resource information sent by multiple terminal devices, the first algorithm is used to determine the target terminal device participating in model training, and the target terminal device identification is sent to the second network device, so that the second network device sends the model to the target terminal device, thereby in The approximately optimal solution is selected in each round of model training to reduce the heterogeneity of data distribution and training overhead.

Benefits of technology

It improves the model prediction accuracy, accelerates the model convergence speed, reduces energy consumption, and effectively solves the problems of terminal device data distribution heterogeneity and resource heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035833A_ABST
    Figure CN120035833A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training method and device, and the method comprises the steps: receiving resource information sent by a plurality of terminal devices, and determining at least one target terminal device participating in the model training of each round from the plurality of terminal devices through employing a first algorithm according to the resource information, and sending the identifier of the at least one target terminal device to a second network device, the identifier of the at least one target terminal device being used by the second network device to send the model of each round to the at least one target terminal device, so that the network side can find an approximately optimal solution in each round of model training. According to the method, the terminal equipment participating in model training is selected, so that the data distribution heterogeneous degree of the overall selected terminal equipment and the weighted sum of the training overhead are minimized, the model prediction precision can be effectively improved, the model convergence speed can be increased, and the energy consumption can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method and device Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a model training method and device. Background Art

[0002] Future 6G communication technology is an artificial intelligence (AI) technology exemplified by machine learning (ML), and the two are inextricably linked. A major trend from 5G to 6G is the interconnection of everything. All terminal devices can function as intelligent entities (smart entities) for analysis and computing, forming an intelligent network. Mobile phones, laptops, sensors, and other terminal devices generate massive amounts of heterogeneous data locally. Effectively utilizing this data within the network will significantly contribute to intelligent analysis and resource optimization.

[0003] Hierarchical Federated Learning (HFL) technology can differentiate tasks based on their priority and category. Tasks with low computing power requirements are executed locally, common tasks with low latency requirements are executed at the edge, and some tasks requiring concentrated computing resources are transferred to the cloud center for computation. This hierarchical architecture effectively saves energy and reduces latency. However, due to the extremely complex real-world scenarios and the significant differences in computing and data resources between different terminals, the algorithm design of hierarchical federated learning in related technologies does not fully consider these differences, resulting in deviations in the aggregated global model and reducing the quality of global model aggregation.

[0004] Summary of the Invention

[0005] The first embodiment of the present application provides a model training method, which is performed by a first network device and includes:

[0006] Receive resource information sent by multiple terminal devices;

[0007] Determine, based on the resource information, at least one target terminal device that participates in each round of model training from the multiple terminal devices using a first algorithm;

[0008] The identifier of the at least one target terminal device is sent to the second network device, where the identifier of the at least one target terminal device is used by the second network device to send the model of each round to the at least one target terminal device.

[0009] A second aspect of the present application provides a model training method, which is performed by a second network device and includes:

[0010] Receiving an identifier of at least one target terminal device participating in each round of model training sent by a first network device, where the at least one target terminal device is determined by the first network from the multiple terminal devices using a first algorithm based on resource information of the multiple terminal devices;

[0011] The model of each round is sent to the at least one target terminal device.

[0012] A third aspect of the present application provides a model training method, which is executed by a terminal device and includes:

[0013] Sending resource information to the first network device, where the resource information is used by the first network device to determine, based on the resource information and using a first algorithm, at least one target terminal device from a plurality of terminal devices to participate in each round of model training;

[0014] In response to the terminal device being a target terminal device participating in the current round of model training, the current round model sent by the second network device is received.

[0015] A fourth aspect of the present application provides a model training device, which is applied to a first network device and includes:

[0016] A transceiver unit, configured to receive resource information sent by multiple terminal devices;

[0017] a processing unit, configured to determine, based on the resource information and using a first algorithm, at least one target terminal device from the plurality of terminal devices to participate in each round of model training;

[0018] The transceiver unit is further configured to send an identifier of the at least one target terminal device to the second network device, and the identifier of the at least one target terminal device is used by the second network device to send the model of each round to the at least one target terminal device.

[0019] A fifth aspect of the present application provides a model training device, which is applied to a second network device and includes:

[0020] a transceiver unit, configured to receive an identifier of at least one target terminal device participating in each round of model training, sent by a first network device, where the at least one target terminal device is determined by the first network from the plurality of terminal devices using a first algorithm based on resource information of the plurality of terminal devices;

[0021] The transceiver unit is further configured to send the model of each round to the at least one target terminal device.

[0022] A sixth aspect of the present application provides a model training device, which is applied to a terminal device and includes:

[0023] a transceiver unit, configured to send resource information to a first network device, wherein the resource information is used by the first network device to determine, based on the resource information, at least one target terminal device to participate in each round of model training from a plurality of terminal devices using a first algorithm;

[0024] The transceiver unit is further configured to receive the current round model sent by the second network device in response to the terminal device being a target terminal device participating in the current round model training.

[0025] The seventh aspect embodiment of the present application proposes a communication device, which includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program stored in the memory so that the device executes the model training method described in the first aspect embodiment above, or so that the device executes the model training method described in the second aspect embodiment above.

[0026] The eighth embodiment of the present application proposes a communication device, which includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program stored in the memory so that the device executes the model training method described in the third embodiment above.

[0027] The ninth aspect embodiment of the present application proposes a communication device, which includes a processor and an interface circuit, wherein the interface circuit is used to receive code instructions and transmit them to the processor, and the processor is used to run the code instructions to enable the device to execute the model training method described in the first aspect embodiment above, or the processor is used to run the code instructions to enable the device to execute the model training method described in the second aspect embodiment above.

[0028] The tenth embodiment of the present application proposes a communication device, which includes a processor and an interface circuit. The interface circuit is used to receive code instructions and transmit them to the processor. The processor is used to run the code instructions to enable the device to execute the model training method described in the third embodiment above.

[0029] The eleventh embodiment of the present application proposes a computer-readable storage medium for storing instructions. When the instructions are executed, the model training method described in the first embodiment above is implemented, or the model training method described in the second embodiment above is implemented.

[0030] The twelfth embodiment of the present application proposes a computer-readable storage medium for storing instructions. When the instructions are executed, the model training method described in the third embodiment above is implemented.

[0031] The thirteenth aspect of the present application proposes a computer program, which, when running on a computer, enables the computer to execute the measurement allocation method described in the first aspect of the embodiment, or enables the computer to execute the model training method described in the second aspect of the embodiment.

[0032] The fourteenth embodiment of the present application proposes a computer program, which, when run on a computer, enables the computer to execute the model training method described in the third embodiment.

[0033] An embodiment of the present application provides a model training method and device, which receives resource information sent by multiple terminal devices, and uses a first algorithm based on the resource information to determine at least one target terminal device participating in each round of model training from the multiple terminal devices, and sends an identifier of the at least one target terminal device to a second network device. The identifier of the at least one target terminal device is used by the second network device to send each round of model to the at least one target terminal device, so that the network side can find an approximate optimal solution in each round of model training and select terminal devices participating in model training, so that the weighted sum of the data distribution heterogeneity and training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0034] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.

[0036] FIG1 is a schematic diagram of the architecture of a communication system provided in an embodiment of the present application;

[0037] FIG2 is a flow chart of a model training method provided in an embodiment of the present application;

[0038] FIG3 is a flow chart of a model training method provided in an embodiment of the present application;

[0039] FIG4 is a flow chart of a model training method provided in an embodiment of the present application;

[0040] FIG5 is a flow chart of a model training method provided in an embodiment of the present application;

[0041] FIG6 is a flow chart of a model training method provided in an embodiment of the present application;

[0042] FIG7 is a flow chart of a model training method provided in an embodiment of the present application;

[0043] FIG8 is a schematic structural diagram of a model training device provided in an embodiment of the present application;

[0044] FIG9 is a schematic structural diagram of a model training device provided in an embodiment of the present application;

[0045] FIG10 is a schematic structural diagram of a model training device provided in an embodiment of the present application;

[0046] FIG11 is a schematic structural diagram of another model training device provided in an embodiment of the present application;

[0047] FIG12 is a schematic structural diagram of a chip provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0048] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numbers in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible implementations consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0049] The terms used in the embodiments of this application are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of this application. The singular forms "a" and "the" used in the embodiments of this application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0050] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to a determination."

[0051] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be understood as limiting the present application.

[0052] In order to better understand the model training method disclosed in the embodiment of the present application, the communication system to which the embodiment of the present application is applicable is first described below.

[0053] Please refer to Figure 1, which is a schematic diagram of the architecture of a communication system provided in an embodiment of the present application. The communication system may include, but is not limited to, one network device and one terminal device. The number and form of devices shown in Figure 1 are for example purposes only and do not constitute a limitation on the embodiments of the present application. In actual applications, two or more network devices and two or more terminal devices may be included. The communication system shown in Figure 1 includes, for example, a first network device 101, multiple second network devices 102, a third network device 103, and multiple terminal devices 104.

[0054] It should be noted that the technical solutions of the embodiments of the present application can be applied to various communication systems, such as Long Term Evolution (LTE) systems, fifth-generation mobile communication systems, 5G new air interface systems, 6G communication systems, or other future new mobile communication systems.

[0055] The network devices 101-103 in the embodiments of the present application are all entities on the network side for transmitting or receiving signals. For example, the network device 101 can be an evolved NodeB (eNB), a transmission point (TRP), a next generation NodeB (gNB) in an NR system, a base station in other future mobile communication systems, or an access node in a wireless fidelity (WiFi) system. The embodiments of the present application do not limit the specific technology and specific device form adopted by the network device. The network device provided in the embodiment of the present application can be composed of a centralized unit (CU) and a distributed unit (DU), wherein the CU can also be called a control unit (Control Unit). The CU-DU structure can be used to split the protocol layer of the network device, such as the base station, and the functions of some protocol layers are placed in the CU for centralized control, and the functions of the remaining part or all of the protocol layers are distributed in the DU, and the DU is centrally controlled by the CU.

[0056] The terminal device 104 in the embodiment of the present application is an entity on the user side for receiving or transmitting signals, such as a mobile phone. The terminal device can also be referred to as a terminal device (terminal), user equipment (UE), mobile station (MS), mobile terminal device (MT), etc. The terminal device can be a car with communication function, a smart car, a mobile phone (Mobile Phone), an Internet of Things (IoT) terminal, a wearable device, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control (Industrial Control), a wireless terminal device in self-driving (Self-Driving), a wireless terminal device in remote medical surgery (Remote Medical Surgery), a wireless terminal device in smart grid (Smart Grid), a wireless terminal device in transportation safety (Transportation Safety), a wireless terminal device in smart city (Smart City), a wireless terminal device in smart home (Smart Home), etc. The embodiment of the present application does not limit the specific technology and specific device form adopted by the terminal device.

[0057] In some embodiments of the present application, the first network device 101 may be, for example, a system scheduling manager that can schedule each terminal device and is a trusted network device that can communicate with the terminal device. The second network device 102 may be an edge server. In one example, an edge server may be a network device that exists at the logical extreme or "edge" of a network and can provide a channel for terminal devices to enter the network and the function of communicating with other server devices. The edge server can store content as close as possible to the terminal device that issued the request, thereby reducing latency and shortening page loading time. The third network device 103 may be a central server or a cloud server. A central server or a cloud server is a network device that exists in the logical "cloud" of the network and has strong computing power and can perform computing tasks that concentrate computing resources.

[0058] In the application scenario of the embodiment of the present application, it can be considered to include multiple terminal devices (indexed as n, n∈{1,....,N}, N is a positive integer), multiple edge servers (indexed as k, k∈{1,...,K}, K is a positive integer) and a central server (or cloud server). In the embodiment of the present application, it is assumed that each terminal device can only communicate with one edge server, and each terminal device n has a local data set D n , the global data set is D={D1,...,D n ,...,D N}.

[0059] The core technology of future 6G communications is artificial intelligence (AI), exemplified by machine learning (ML). The two are inextricably linked. A major trend from 5G to 6G is the interconnectedness of everything. All devices can function as intelligent entities, performing analysis and computing, forming an intelligent network. Mobile phones, laptops, sensors, and other devices generate massive amounts of heterogeneous data locally. Effectively utilizing this data within the network will significantly contribute to intelligent analysis and resource optimization.

[0060] In traditional cloud-server-centric computing, user data is typically sent to a centralized server for computation and storage. This approach not only raises numerous security and privacy issues, but also incurs significant energy consumption and latency during transmission, contradicting the promise of 6G green communications. It also consumes significant bandwidth, hindering the normal operation of other task queues in the network. With the widespread adoption of mobile edge computing and federated learning technologies, and the increased computing power of terminal devices, more computing tasks are being offloaded to edge servers or local users. This intelligent network architecture offers greater possibilities for orchestration and deployment.

[0061] In federated learning (FL), end devices use their local data to train and update the ML model required by the server. The end devices then send the updated model parameters, rather than the original data, to the server for aggregation. The aggregated model is then distributed to the end devices, repeating this process until the model converges.

[0062] Hierarchical Federated Learning (HFL) technology can differentiate tasks based on priority and category. Tasks with low computing power requirements are executed locally, common tasks with low latency requirements are executed at the edge, and some tasks that require intensive computing resources are transferred to the cloud center for computation. This layered architecture effectively saves energy and reduces latency.

[0063] However, real-world scenarios are extremely complex, particularly those in the Internet of Vehicles (IoV) and smart transportation, where a large number of heterogeneous end devices, such as different car models, sensors, and complex smart devices, exist. These devices possess vastly different computing and data resources, necessitating the full consideration of the performance differences arising from this resource and data heterogeneity in HFL algorithm design. Because the data from end devices is not independently and identically distributed (IID), this inherently leads to low classification accuracy. While various approaches have been proposed to address this issue, such as personalized local loss functions and model distillation, they lack a comprehensive consideration of client data distribution, local model contribution, and communication overhead, which are crucial for determining the quality of global model aggregation. Diverse training data distributions cause local updates on clients to deviate from each other, resulting in a biased aggregated global model. Furthermore, lagging devices with limited computing power or poor channel conditions can significantly slow model convergence. Therefore, how to comprehensively account for the device and resource heterogeneity of each client during FL training to achieve improved training accuracy while minimizing energy consumption and latency overhead remains a pressing issue.

[0064] It can be understood that the communication system described in the embodiment of the present application is for the purpose of more clearly illustrating the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided by the embodiment of the present application is also applicable to similar technical problems.

[0065] The model training method and device provided in this application are described in detail below with reference to the accompanying drawings.

[0066] Please refer to Figure 2, which is a flow chart of a model training method provided in an embodiment of the present application. It should be noted that the model training method in the embodiment of the present application is performed by the first network device. This method can be performed independently or in combination with any other embodiment of the present application. As shown in Figure 2, the method may include the following steps:

[0067] Step 201: Receive resource information sent by multiple terminal devices.

[0068] In an embodiment of the present application, the first network device is capable of receiving resource information sent by multiple terminal devices.

[0069] The resource information includes at least one of the following: an Earth Mover's Distance (EMD) value of the local data distribution of the terminal device; a processor frequency of the terminal device, such as a central processing unit (CPU) frequency; and a battery level of the terminal device.

[0070] Among them, the pile-up distance EMD can be used to measure the distance between two distributions. In an embodiment of the present application, the pile-up distance EMD can be used to measure the degree of heterogeneity of the data label distribution of each terminal device, that is, the difference from the global distribution.

[0071] The resource information of the multiple terminal devices can be used by the first network device to determine at least one target terminal device from the multiple terminal devices, and the at least one target terminal device is a terminal device participating in each round of model training.

[0072] It should be noted that in the embodiments of this application, data heterogeneity of terminal devices is understood as category imbalance, where each terminal device does not follow a common data distribution, that is, the data distribution in the terminal devices is non-independent and identically distributed. The value of the earth distance (EMD) can be used to measure the degree of heterogeneity of the data label distribution of each terminal device.

[0073] In this embodiment of the present application, the first network device may be a system scheduling manager.

[0074] Step 202: Based on the resource information, a first algorithm is used to determine at least one target terminal device that participates in each round of model training from the multiple terminal devices.

[0075] It should be noted that in the embodiment of the present application, the first network device may determine the target terminal device for each round at one time, or it may determine the target terminal device for one or several rounds and then execute subsequent steps. During the execution of the subsequent steps, the first network device continues to determine the target terminal for other rounds. The embodiment of the present application does not limit this.

[0076] In an embodiment of the present application, the first network device can use a first algorithm to determine at least one target terminal device participating in each round of model training from the multiple terminal devices based on resource information of the multiple terminal devices.

[0077] In some embodiments, the first network device determines the target terminal devices participating in each round of model training for a purpose of minimizing a weighted sum of an EMD value of a local data distribution of at least one target terminal device participating in all rounds of model training, a total energy loss of the at least one target terminal device participating in all rounds of model training, and a total time delay of the at least one target terminal device participating in all rounds of model training. It will be appreciated that the purpose of the first network device can be viewed as an optimization problem.

[0078] The first network device can use the first algorithm to solve the above optimization problem and obtain the optimal solution, that is, to obtain at least one target terminal device participating in each round of model training.

[0079] In some embodiments, the first algorithm may be a TD3 (Twin Delayed Deep Deterministic policy gradient) algorithm based on , and the use of this algorithm to solve the above optimization problem will be described in detail in the following embodiments.

[0080] Step 203: Send the identifier of the at least one target terminal device to the second network device. The identifier of the at least one target terminal device is used by the second network device to send the model of each round to the at least one target terminal device.

[0081] In an embodiment of the present application, after determining at least one target terminal device participating in each round of model training, the first network device can send the identifier of the at least one target terminal device to the second network device, and the identifier of the at least one target terminal device can be used by the second network device to send the model of each round of secondary training to the at least one target terminal device.

[0082] In an embodiment of the present application, the model of each round can be used for model training on the target terminal device to obtain model update parameters, and the model update parameters can be sent by the target terminal device to the second network device so that the second network device updates the model for the next round of model training.

[0083] In an embodiment of the present application, the model updated by the second network device can be aggregated into a global model. A global model refers to a model obtained by aggregating models updated by multiple second network devices. After all rounds of model training, the accuracy of the global model converges.

[0084] In an embodiment of the present application, after all rounds of model training, the accuracy of the global model can reach a predefined accuracy.

[0085] In this embodiment of the present application, the second network device may be an edge server.

[0086] In some embodiments, the updated model of the second network device can be sent by the second network device to the central server for aggregation into a global model. After each round of training, the second network device can receive the model update parameters, update the model, and send the updated model to the central server for aggregation into a global model.

[0087] In summary, by receiving resource information sent by multiple terminal devices, based on the resource information, a first algorithm is adopted to determine at least one target terminal device participating in each round of model training from the multiple terminal devices, and an identifier of the at least one target terminal device is sent to the second network device. The identifier of the at least one target terminal device is used by the second network device to send each round of model to the at least one target terminal device, so that the network side can find an approximate optimal solution in each round of model training and select terminal devices participating in model training, so that the weighted sum of the data distribution heterogeneity and training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0088] Please refer to Figure 3, which is a flow chart of a model training method provided in an embodiment of the present application. It should be noted that the model training method in the embodiment of the present application is performed by the first network device. This method can be performed independently or in combination with any other embodiment of the present application. As shown in Figure 3, the method may include the following steps:

[0089] Step 301: Receive resource information sent by multiple terminal devices.

[0090] In an embodiment of the present application, the first network device is capable of receiving resource information sent by multiple terminal devices.

[0091] The resource information includes at least one of the following: an EMD value of the local data distribution of the terminal device; a processor frequency of the terminal device, such as a central processing unit (CPU) frequency; and a battery charge of the terminal device.

[0092] The resource information of the multiple terminal devices can be used by the first network device to determine at least one target terminal device participating in each round of model training from the multiple terminal devices.

[0093] It should be noted that in the embodiments of this application, data heterogeneity of terminal devices is understood as category imbalance, where each terminal device does not follow a common data distribution, that is, the data distribution in the terminal devices is non-independent and identically distributed. The value of the earth distance (EMD) can be used to measure the degree of heterogeneity of the data label distribution of each terminal device.

[0094] It should also be noted that in the application scenario of the embodiment of the present application, it can be considered to include multiple terminal devices (indexed as n, n∈{1,....,N}, N is a positive integer), multiple second network devices (indexed as k, k∈{1,...,K}, K is a positive integer) and a central server (or cloud server). In the embodiment of the present application, it is assumed that each terminal device can only communicate with one second network device, and each terminal device n has a local data set D n , the global data set is D={D1,...,D n ,...,D N}.

[0095] In the embodiment of the present application, the EMD value of the terminal device n can be expressed as: Among them, C represents the number of tag types, represents the proportion of the i-th data type of terminal device n, p y=i Indicates the global proportion of the number of data of type i.

[0096] In this embodiment of the present application, the first network device may be a system scheduling manager.

[0097] Step 302: Based on the resource information, a first algorithm is used to determine at least one target terminal device participating in each round of model training from the multiple terminal devices, so as to minimize the weighted sum of the EMD value of the local data distribution of the at least one target terminal device participating in all rounds of model training, the total energy loss of the at least one target terminal device participating in all rounds of model training, and the total time delay of the at least one target terminal device participating in all rounds of model training.

[0098] In an embodiment of the present application, the first network device can use a first algorithm to determine at least one target terminal device participating in each round of model training from the multiple terminal devices based on the resource information of the multiple terminal devices, with the aim of minimizing the weighted sum of the EMD value of the local data distribution and the energy consumption and delay overhead of the selected terminal device.

[0099] Let a n,t Represents a binary indicator. If the terminal device n is selected as the target terminal device by the first network device in the tth round of model training, then a n,t =1; otherwise, a n,t = 0. Since the EMD value of terminal device n can be expressed as: Then the data label distribution heterogeneity of the terminal devices that participate in all T (T is a positive integer) rounds of model training can be expressed as:

[0100] Then the optimization problem can be modeled as:

[0101] P1:

[0102] Constrained by: T t ≤T max ,

[0103] E n,t ≤E battery ,

[0104]

[0105]

[0106] a n,t ∈{0,1},

[0107] Among them, β1, β2 are weighted coefficients, E g represents the total energy loss of all target terminal devices participating in the training in the tth round of training, T g E represents the total delay of global iterations in the tth round of training. g and T g The expression and derivation of are given later and will not be repeated here.

[0108] Among them, constraint 1 requires that the time taken by the terminal device to complete model training and upload in each round of model training must be less than or equal to the deadline T max Constraint 2 ensures that the energy consumed by terminal device n participating in each training round is less than the remaining battery capacity. Constraint 3 is the CPU cycle frequency range of each terminal device. Constraint 4 is the transmission power range of each terminal device in round t. Constraint 5 is the decision of terminal device selection, which is a binary variable.

[0109] Since P1 is a collaboration problem between multiple terminal devices, it is an NLMIP (nonlinear mixed integer programming) problem. The first algorithm (a reinforcement learning algorithm based on online optimization, namely the TD3 algorithm) is used here to solve the problem. The action space in P1 contains two types, one is the selection of the target terminal device, and the other is power allocation, so it contains both continuous actions and discrete actions. As a policy-based algorithm, TD3 can effectively deal with continuous action decisions, and can solve the over-estimation problem by adopting two target networks. In an embodiment of the present application, the formulated problem can be modeled as a Markov decision process (MDP), and the specific details are as follows.

[0110] An MDP can be represented by a 4-tuple (S, A, P, R), where S is the state space, A is the action space, P is the transition probability, and R is the reward function.

[0111] State space: According to problem P1, the state space of MDP includes the EMD value of each terminal device, the channel gain of the uplink, the channel bandwidth from the terminal device to the second network device, and the remaining power of the terminal device. State space S t The definition is as follows: S t ={(I n,t , G n,t , b n,t , E n,t ), n=1,…,N}.

[0112] Among them, I n,t represents the EMD value of terminal device n in the tth round of training, G n,t represents the uplink channel gain of terminal device n in the tth round of training, b n,t represents the channel bandwidth from terminal device n to the second network device in the tth round of training, E n,t Indicates the remaining power of terminal device n in the tth round of training.

[0113] Action space: The action of MDP is the choice of terminal device a n,t , and the transmission power p of the selected target terminal device n,t Action space A t The definition is as follows: A t ∈A={(a n,t , p n,t ), n=1,…,N}, where a n,t ∈{0,1}, and

[0114] Reward function: Let the reward function R(S t+1 |S t ,A t ) indicates that in state S t Next, take action A t ∈A receives the timely reward. The reward is defined as the sum of the data quality scores of the terminal devices selected by the federated learning in the current round and the weighted sum of the energy consumption and delay overhead, that is, This means that from the current state S t Through action A t Transition to the next state S t+1 The rewards received.

[0115] The selected action can then be evaluated using the evaluation policy μ, where μ is a mapping from states to actions. Our goal is to maximize the expected total reward, i.e., the action value function Q(St ,A t |θμ): Among them, γ∈{0,1} is the discount factor of the future state, and E{·} represents the expectation. According to the Bellman formula, a state-action pair (S t ,A t ) and subsequent state-action pairs (S t′ ,A t′ ) can be expressed as Optimal Action It can be expressed as: in, Provides the current state of the largest federated learning target ratio.

[0116] In the embodiment of the present application, a first algorithm is used to find an approximately optimal solution to problem P1. The first algorithm may adopt an Actor-Critic network structure, and the first network device, as the decision maker of the algorithm, is trained to optimize actions in a continuous action space.

[0117] The first algorithm uses a policy gradient-based method to map the network state to a specific action. Among them, the Actor network (also called policy network) μ(S|θ μ ) deterministically maps the state S to a specific continuous action, and the Critic network is used to approximate the actor-value function. The first algorithm uses an experience replay mechanism in each training step to use a register B to store the previous action taken by the first network device, the current state, the current action, the reward value, and the next state information. Each time a mini-batch of data is sampled from the experience pool to train the Actor-Critic network. After obtaining the result of each action, since the action selected by the client output from the policy network is a continuous value, the continuous value can be converted into a binary value by setting a threshold T as shown below.

[0118]

[0119] For continuous client selection and power allocation, the goal is to find the optimal policy μ to maximize the expected reward function J, where the policy μ can be obtained by taking the expected return with respect to the network parameters θ μ The policy gradient of the Actor network can be calculated by the chain rule:

[0120]

[0121] This is the parameter θ from the starting distribution J relative to the Actor network μThe gradient of the expected return is calculated and averaged over the sampled mini-batch.

[0122] In addition, the optimal action value function can be obtained through the Critic network Q(S t ,A t |θ Q ) to approximate. In order to update Q(S t ,A t |θ Q ), the Critic network adjusts the parameter θ Q To minimize the current target value and Q(S t ,A t |θ Q ), and the mean square error (MSE) is used to measure this error.

[0123]

[0124] However, the update of the critic network will lead to overestimation problem, which will produce suboptimal strategy of the policy network. The first algorithm can be used to solve the overestimation problem by using two approximately independent critic networks {Q ω1 ,Q ω2} to estimate the value function:

[0125]

[0126] The smaller of the two is used to update the value function: y a =min{y a,1 ,y a,2}.

[0127] For the target network, a soft update method is used to update it, introducing a learning rate τ. Every update interval d, the old target network parameters and the new corresponding network parameters are weighted averaged and then assigned to the target network:

[0128] θ Q′ ←τθ Q +(1-τ)θ Q′ θ μ′ ←τθ μ +(1-τ)θ μ′ .

[0129] As an example, in each embodiment of the present application, the overall process of the first algorithm can be as follows:

[0130] Randomly initialize the first critic network Second Critic Network And the weight parameter is and The policy network μ(S|θ μ ).

[0131] Use weights θ μ′ ←θ μ Initialize the target network Q′ and μ′.

[0132] Initialize playback register B.

[0133] In the t=1,...,T rounds of iteration, if t≤T E , using a random strategy to explore T E Round, get the reward function R t And T E The state transition in rounds (S t ,A t ,R t ,S t+1 ) is stored in register B.

[0134] Otherwise, select action A based on the current policy and exploration noise t =μ(S t |θ μ )+ε,ε~N(0,σ).

[0135] The first network device can convert the corresponding power p n,t Assign to the selected target terminal device (execute action A t ).

[0136] The selected target terminal device uses local data to perform model training and can send the model update parameters obtained after training to the second network device.

[0137] The second network device can perform model aggregation to obtain an updated model, and transmit the updated model to the third network device (central server) for global aggregation.

[0138] The first network device is able to observe the reward value R t And the new state S t+1 .

[0139] The experience (S t ,A t ,R t ,S t+1 ) is stored in register B.

[0140] From the experience pool of N rounds stored in register B Sample a random mini-batch of data from .

[0141] Get target action

[0142] Update the Critic network by minimizing the loss L:

[0143] If t mod d = 0, then the action policy is updated using the sampled policy gradient as follows:

[0144]

[0145] Update target network: θ Q′ ←τθQ+(1-τ)θ Q′ θ μ′ ←τθμ+(1-τ)θ μ′ .

[0146] Step 303: Send the identifier of the at least one target terminal device to the second network device. The identifier of the at least one target terminal device is used by the second network device to send the model of each round to the at least one target terminal device.

[0147] In an embodiment of the present application, after determining at least one target terminal device participating in each round of model training, the first network device can send the identifier of the at least one target terminal device to the second network device, and the identifier of the at least one target terminal device can be used by the second network device to send the model of each round of secondary training to the at least one target terminal device.

[0148] In an embodiment of the present application, the model of each round can be used for model training on the target terminal device to obtain model update parameters, and the model update parameters can be sent by the target terminal device to the second network device so that the second network device updates the model for the next round of model training.

[0149] In an embodiment of the present application, the model updated by the second network device can be aggregated into a global model. After all rounds of model training, the accuracy of the global model converges.

[0150] In this embodiment of the present application, the second network device may be an edge server.

[0151] In some embodiments, the updated model of the second network device can be sent by the second network device to the central server for aggregation into a global model. After each round of training, the second network device can receive the model update parameters, update the model, and send the updated model to the central server for aggregation into a global model.

[0152] In summary, by receiving resource information sent by multiple terminal devices, and based on the resource information, adopting the first algorithm, at least one target terminal device participating in each round of model training is determined from the multiple terminal devices, so that the EMD value of the local data distribution of the at least one target terminal device participating in all rounds of model training, the total energy loss of the at least one target terminal device participating in all rounds of model training, and the weighted sum of the total time delay of the at least one target terminal device participating in all rounds of model training are minimized, and the identifier of the at least one target terminal device is sent to the second network device. The identifier of the at least one target terminal device is used by the second network device to send the model of each round to the at least one target terminal device, so that the network side can find an approximate optimal solution in each round of model training, select the terminal devices participating in the model training, so that the weighted sum of the data distribution heterogeneity and the training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0153] In the aforementioned embodiment, E g and T g The expression and derivation are as follows:

[0154] For the target terminal device n selected in the tth round of model training, let Indicates its local training energy consumption in round t. n represents the number of CPU cycles required for the nth terminal device to execute one data sample, which is known a priori. Assuming that all data samples have the same data size, i.e., the number of bits, let D n is the data volume of terminal device n, then the number of CPU cycles required for terminal device n to run a local epoch (one iteration) is c n D n . E n Indicates the number of local training iterations before each upload of model parameter updates by terminal device n. n represents the CPU cycle frequency of terminal device n, β n is the effective capacitance coefficient of the computing chip of terminal device n. Then the local training energy consumption formula of terminal device n in the tth round of global iteration is as follows:

[0155] After local training is complete, we use an Orthogonal Frequency Division Multiple Access (OFDMA) scheme to send the local model to the second network device with a total bandwidth of B, in round t, where the system bandwidth is divided into n subchannels equal to the number of selected terminal devices. OFDMA is a multiple access scheme based on OFDM. In OFDMA, different users are assigned different subcarriers, so multiple users can transmit their data simultaneously.

[0156] In the tth round of global iteration, bandwidth b is allocated to terminal device n. n,t According to Shannon's theorem, the achievable transmission rate of terminal device n is defined as: Where B is the bandwidth, N0 is the power spectral density of Gaussian white noise, and p n,t is the transmission power of terminal device n in round t, G n,t is the channel gain between the terminal device n and the second network device k.

[0157] The data size of the model update parameters sent by the target terminal device is S n , assuming S n The size is constant, then the time taken by the nth terminal device to send the model update parameters to the second network device k in round t is:

[0158] Then, during the t-th round of model update parameter transmission, the energy consumption of terminal device n is:

[0159] The total energy consumption of terminal device n is: The total energy loss of all terminal devices participating in the training in the tth round of training is

[0160] In each round of global iteration, the total delay of terminal device n mainly includes the local model training time of terminal device n and the parameter result upload time from terminal device n to the second network device and from the second network device to the central server. Since the downlink bandwidth is much larger than the uplink bandwidth, the time it takes for the central server to send the model to the second network device and the time it takes for the second network device to send the model to the terminal device can be ignored. and They represent the local model training time and result upload time of client n in round t respectively.

[0161] in, As mentioned above, Since each terminal device performs local computation in parallel and starts sending model update parameters once the computation is completed, the total latency in the tth round of training follows the stragglers effect, that is: Second, during the aggregation phase from network device k to the central server, the model upload latency is: Therefore, the total delay of the tth round of global iteration is

[0162] Please refer to Figure 4, which is a flow chart of a model training method provided in an embodiment of the present application. It should be noted that the model training method in the embodiment of the present application is performed by the second network device. This method can be performed independently or in combination with any other embodiment of the present application. As shown in Figure 4, the method may include the following steps:

[0163] Step 401: Receive the identifier of at least one target terminal device participating in each round of model training sent by the first network device, where the at least one target terminal device is determined by the first network from the multiple terminal devices based on resource information of the multiple terminal devices using a first algorithm.

[0164] In an embodiment of the present application, the second network device is capable of receiving an identifier of at least one target terminal device participating in each round of model training, sent by the first network device. The at least one target terminal device is determined by the first network device from among the multiple terminal devices using a first algorithm based on resource information of the multiple terminal devices.

[0165] The resource information includes at least one of the following: an EMD value of the local data distribution of the terminal device; a processor frequency of the terminal device, such as a central processing unit (CPU) frequency; and a battery charge of the terminal device.

[0166] The resource information of the multiple terminal devices can be used by the first network device to determine at least one target terminal device participating in each round of model training from the multiple terminal devices.

[0167] It should be noted that in the embodiments of this application, data heterogeneity of terminal devices is understood as category imbalance, where each terminal device does not follow a common data distribution, that is, the data distribution in the terminal devices is non-independent and identically distributed. The value of the earth distance (EMD) can be used to measure the degree of heterogeneity of the data label distribution of each terminal device.

[0168] In an embodiment of the present application, the second network device may be an edge server, and the first network device may be a system scheduling manager.

[0169] In some embodiments, the first network device determines the target terminal devices participating in each round of model training for a purpose of minimizing a weighted sum of an EMD value of a local data distribution of at least one target terminal device participating in all rounds of model training, a total energy loss of the at least one target terminal device participating in all rounds of model training, and a total time delay of the at least one target terminal device participating in all rounds of model training. It will be appreciated that the purpose of the first network device can be viewed as an optimization problem.

[0170] The first network device can use the first algorithm to solve the above optimization problem and obtain the optimal solution, that is, to obtain at least one target terminal device participating in each round of model training.

[0171] In some implementations, the first algorithm may be an algorithm based on TD3. Specifically, the first algorithm may be as described in the above embodiment, which will not be described in detail here.

[0172] Step 402: Send the model of each round to the at least one target terminal device.

[0173] In an embodiment of the present application, the second network device can send the model of each round to the at least one target terminal device according to the identifier of the at least one target terminal device.

[0174] In some embodiments, the second network device is capable of downloading the model of each round from the central server and sending the model to the target terminal device selected by the first network device to participate in each round of model training, so that the target terminal device uses local data to train the model.

[0175] In some embodiments, the second network device is further capable of receiving model update parameters sent by the at least one target terminal device after model training, and the model update parameters are used by the second network device to update the model for the next round of model training.

[0176] In an embodiment of the present application, the model updated by the second network device can be aggregated into a global model. After all rounds of model training, the accuracy of the global model converges.

[0177] In some embodiments, the updated model of the second network device can be sent by the second network device to the central server for aggregation into a global model. After each round of training, the second network device can receive the model update parameters, update the model, and send the updated model to the central server for aggregation into a global model.

[0178] In some embodiments, the second network device may also aggregate the model update parameters received from at least one target terminal device to obtain aggregated model update parameters, and send the aggregated model update parameters to the central server for aggregation to perform model update and obtain a new global model.

[0179] In summary, by receiving the identifier of at least one target terminal device participating in each round of model training sent by the first network device, the at least one target terminal device is determined by the first network from the multiple terminal devices based on the resource information of the multiple terminal devices using the first algorithm, and sending the model of each round to the at least one target terminal device, the network side can find an approximate optimal solution in each round of model training, select the terminal devices participating in the model training, and minimize the weighted sum of the data distribution heterogeneity and training overhead of the overall selected terminal devices, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0180] Please refer to Figure 5, which is a flow chart of a model training method provided in an embodiment of the present application. It should be noted that the model training method in the embodiment of the present application is performed by the second network device. This method can be performed independently or in combination with any other embodiment of the present application. As shown in Figure 5, the method may include the following steps:

[0181] Step 501: Receive the identifier of at least one target terminal device participating in each round of model training sent by the first network device, where the at least one target terminal device is determined by the first network from the multiple terminal devices based on resource information of the multiple terminal devices using a first algorithm.

[0182] In an embodiment of the present application, the second network device is capable of receiving an identifier of at least one target terminal device participating in each round of model training, sent by the first network device. The at least one target terminal device is determined by the first network device from among the multiple terminal devices using a first algorithm based on resource information of the multiple terminal devices.

[0183] The resource information includes at least one of the following: an EMD value of the local data distribution of the terminal device; a processor frequency of the terminal device, such as a central processing unit (CPU) frequency; and a battery charge of the terminal device.

[0184] The resource information of the multiple terminal devices can be used by the first network device to determine at least one target terminal device participating in each round of model training from the multiple terminal devices.

[0185] It should be noted that in the embodiments of this application, data heterogeneity of terminal devices is understood as category imbalance, where each terminal device does not follow a common data distribution, that is, the data distribution in the terminal devices is non-independent and identically distributed. The value of the earth distance (EMD) can be used to measure the degree of heterogeneity of the data label distribution of each terminal device.

[0186] In an embodiment of the present application, the second network device may be an edge server, and the first network device may be a system scheduling manager.

[0187] In some embodiments, the first network device determines the target terminal devices participating in each round of model training for a purpose of minimizing a weighted sum of an EMD value of a local data distribution of at least one target terminal device participating in all rounds of model training, a total energy loss of the at least one target terminal device participating in all rounds of model training, and a total time delay of the at least one target terminal device participating in all rounds of model training. It will be appreciated that the purpose of the first network device can be viewed as an optimization problem.

[0188] The first network device can use the first algorithm to solve the above optimization problem and obtain the optimal solution, that is, to obtain at least one target terminal device participating in each round of model training.

[0189] In some implementations, the first algorithm may be an algorithm based on TD3. Specifically, the first algorithm may be as described in the above embodiment, which will not be described in detail here.

[0190] Step 502: Send the model of each round to the at least one target terminal device.

[0191] In an embodiment of the present application, the second network device can send the model of each round to the at least one target terminal device according to the identifier of the at least one target terminal device.

[0192] In some embodiments, the second network device is capable of downloading the model of each round from the central server and sending the model to the target terminal device selected by the first network device to participate in each round of model training, so that the target terminal device uses local data to train the model.

[0193] It can be understood that the model of each round of the central server is a global model obtained by aggregating the updated model sent to the central server by the second network device after the central server completes the previous round of model training.

[0194] Step 503: Receive model update parameters sent by the at least one target terminal device after model training, and the model update parameters are used by the second network device to update the model for the next round of model training.

[0195] In an embodiment of the present application, the second network device is also capable of receiving model update parameters sent by the at least one target terminal device after model training, and the model update parameters are used by the second network device to update the model for the next round of model training.

[0196] In an embodiment of the present application, the model updated by the second network device can be aggregated into a global model. After all rounds of model training, the accuracy of the global model converges.

[0197] In some embodiments, the updated model of the second network device can be sent by the second network device to the central server for aggregation into a global model. After each round of training, the second network device can receive the model update parameters, update the model, and send the updated model to the central server for aggregation into a global model.

[0198] As an example, each target terminal device participating in the current round of model training receives the model sent by the second network device. The target terminal device n can use the mini-batch gradient descent (MBGD) algorithm to train and update the model to obtain the updated model: Where, e∈[1,E]. After the target terminal device performs E rounds of local iterative updates, the target terminal device n will obtain the model update parameters The second network device k is sent to the second network device k. The second network device k can update the model according to the model update parameters received from at least one target terminal device, and can update the model by averaging aggregation: in, is the set of terminal devices that communicate with the second network device k, and is a subset of the set of all terminal devices. E After the round edge aggregation, the central server performs aggregated updates of the global model: The second network device can obtain the updated global model and then send it to the target terminal device participating in the next round of training.

[0199] In summary, by receiving the identifier of at least one target terminal device participating in each round of model training sent by the first network device, the at least one target terminal device is determined by the first network from the multiple terminal devices based on the resource information of the multiple terminal devices using the first algorithm, sending the model of each round to the at least one target terminal device, and receiving the model update parameters sent by the at least one target terminal device after the model training, the model update parameters are used by the second network device to update the model for the next round of model training, so that the network side can find an approximate optimal solution in each round of model training, select the terminal devices participating in the model training, so that the weighted sum of the data distribution heterogeneity and the training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0200] Please refer to Figure 6, which is a flow chart of a model training method provided in an embodiment of the present application. It should be noted that the model training method in the embodiment of the present application is executed by a terminal device. The method can be executed independently or in combination with any other embodiment of the present application. As shown in Figure 6, the method may include the following steps:

[0201] Step 601: Send resource information to a first network device. The resource information is used by the first network device to determine at least one target terminal device participating in each round of model training from multiple terminal devices based on the resource information and using a first algorithm.

[0202] In an embodiment of the present application, the second network device is capable of receiving an identifier of at least one target terminal device participating in each round of model training, sent by the first network device. The at least one target terminal device is determined by the first network device from among the multiple terminal devices using a first algorithm based on resource information of the multiple terminal devices.

[0203] The resource information includes at least one of the following: an EMD value of the local data distribution of the terminal device; a processor frequency of the terminal device, such as a central processing unit (CPU) frequency; and a battery charge of the terminal device.

[0204] The resource information of the multiple terminal devices can be used by the first network device to determine at least one target terminal device participating in each round of model training from the multiple terminal devices.

[0205] It should be noted that in the embodiments of this application, data heterogeneity of terminal devices is understood as category imbalance, where each terminal device does not follow a common data distribution, that is, the data distribution in the terminal devices is non-independent and identically distributed. The value of the earth distance (EMD) can be used to measure the degree of heterogeneity of the data label distribution of each terminal device.

[0206] In an embodiment of the present application, the second network device may be an edge server, and the first network device may be a system scheduling manager.

[0207] In some embodiments, the first network device determines the target terminal devices participating in each round of model training for a purpose of minimizing a weighted sum of an EMD value of a local data distribution of at least one target terminal device participating in all rounds of model training, a total energy loss of the at least one target terminal device participating in all rounds of model training, and a total time delay of the at least one target terminal device participating in all rounds of model training. It will be appreciated that the purpose of the first network device can be viewed as an optimization problem.

[0208] The first network device can use the first algorithm to solve the above optimization problem and obtain the optimal solution, that is, to obtain at least one target terminal device participating in each round of model training.

[0209] In some implementations, the first algorithm may be an algorithm based on TD3. Specifically, the first algorithm may be as described in the above embodiment, which will not be described in detail here.

[0210] Step 602: In response to the terminal device being a target terminal device participating in the current round of model training, receiving the current round of model training sent by the second network device.

[0211] In an embodiment of the present application, the terminal device selected to participate in the current round of model training can receive the current round of model sent by the second network device and use local data to perform model training on the model.

[0212] In some implementations, the terminal device can use a mini-batch gradient descent algorithm to train and update the model.

[0213] In some embodiments, the terminal device can also send model update parameters obtained after model training to the second network device, and the model update parameters are used by the second network device to update the model for the next round of model training.

[0214] In an embodiment of the present application, the model updated by the second network device can be aggregated into a global model. After all rounds of model training, the accuracy of the global model converges.

[0215] In some embodiments, the updated model of the second network device can be sent by the second network device to the central server for aggregation into a global model. After each round of training, the second network device can receive the model update parameters, update the model, and send the updated model to the central server for aggregation into a global model.

[0216] In summary, by sending resource information to the first network device, the resource information is used by the first network device to determine at least one target terminal device participating in each round of model training from multiple terminal devices based on the resource information and using the first algorithm. In response to the terminal device being the target terminal device participating in the current round of model training, the model of the current round sent by the second network device is received, so that the network side can find an approximate optimal solution in each round of model training and select the terminal devices participating in the model training, so that the weighted sum of the data distribution heterogeneity and training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0217] Please refer to Figure 7, which is a flow chart of a model training method provided in an embodiment of the present application. This method can be performed independently or in combination with any other embodiment of the present application. As shown in Figure 7, the method may include the following steps:

[0218] In step 701, the third network device can obtain an initial model that needs to be trained.

[0219] In step 702 , the third network device can send the model to multiple second network devices for model training.

[0220] Step 703: The first network device is capable of receiving resource information sent by multiple terminal devices.

[0221] The resource information includes at least one of the following: an EMD value of the local data distribution of the terminal device; a processor frequency of the terminal device, such as a central processing unit (CPU) frequency; and a battery charge of the terminal device.

[0222] In step 704, the first network device can use the first algorithm to determine at least one target terminal device participating in the current round of model training from the multiple terminal devices based on the resource information of the multiple terminal devices, so as to minimize the weighted sum of the EMD value, total energy loss and total time delay of the local data distribution of at least one target terminal device participating in all rounds of model training.

[0223] The first algorithm may be as described in any of the aforementioned embodiments and will not be described again here.

[0224] Step 705: The first network device sends the determined identifier of the target terminal device to the second network device.

[0225] Step 706: The second network device can send the model to be trained in this round to the at least one target terminal device.

[0226] Step 707: The at least one target terminal device can receive the model and use local data to perform model training on the model to obtain model update parameters.

[0227] Step 708: The at least one target terminal device can send the model update parameters to the second network device.

[0228] Step 709: After receiving the model update parameters uploaded by the target terminal device, the second network device can update the model according to the model update parameters.

[0229] In step 710, the third network device can receive updated models sent by multiple second network devices, and aggregate the multiple updated models to generate a new global model, which can be used for the next round of model training.

[0230] Repeat steps 702 to 710 until the accuracy of the generated global model converges and reaches a predefined accuracy.

[0231] It should be noted that in the embodiment of the present application, the first network device may determine the target terminal device for each round at one time, or it may determine the target terminal device for one or several rounds and then execute subsequent steps. During the execution of the subsequent steps, the first network device continues to determine the target terminal for other rounds. The embodiment of the present application does not limit this.

[0232] Corresponding to the model training methods provided in the above-mentioned embodiments, the present application also provides a model training device. Since the model training device provided in the embodiments of the present application corresponds to the methods provided in the above-mentioned embodiments, the implementation method of the model training method is also applicable to the model training device provided in the following embodiments and will not be described in detail in the following embodiments.

[0233] Please refer to Figure 8, which is a structural diagram of a model training device provided in an embodiment of the present application.

[0234] As shown in FIG8 , the model training device 800 includes a transceiver unit 810 and a processing unit 820 , wherein:

[0235] The transceiver unit 810 is configured to receive resource information sent by multiple terminal devices;

[0236] The processing unit 820 is configured to determine, based on the resource information and using a first algorithm, at least one target terminal device from the plurality of terminal devices to participate in each round of model training;

[0237] The transceiver unit 810 is further configured to send the identifier of the at least one target terminal device to the second network device, where the identifier of the at least one target terminal device is used by the second network device to send the model of each round to the at least one target terminal device.

[0238] Optionally, the resource information includes at least one of the following:

[0239] The EMD value of the local data distribution of the terminal device;

[0240] The central processing unit (CPU) frequency of the terminal device;

[0241] The battery level of the terminal device.

[0242] Optionally, the weighted sum of the EMD value of the local data distribution of at least one target terminal device participating in all rounds of model training, the total energy loss of at least one target terminal device participating in all rounds of model training, and the total time delay of at least one target terminal device participating in all rounds of model training is minimized.

[0243] Optionally, the model of each round is used for model training on the target terminal device to obtain model update parameters, and the model update parameters are used by the second network device to update the model for the next round of model training.

[0244] Optionally, the model updated by the second network device is used to aggregate into a global model; after all rounds of model training, the accuracy of the global model converges.

[0245] The model training device of this embodiment can receive resource information sent by multiple terminal devices, and based on the resource information, use a first algorithm to determine at least one target terminal device participating in each round of model training from the multiple terminal devices, and send an identifier of the at least one target terminal device to the second network device. The identifier of the at least one target terminal device is used by the second network device to send each round of model to the at least one target terminal device, so that the network side can find an approximate optimal solution in each round of model training and select terminal devices participating in model training, so that the weighted sum of the data distribution heterogeneity and training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0246] Please refer to Figure 9, which is a structural diagram of a model training device provided in an embodiment of the present application.

[0247] As shown in FIG9 , the model training device 900 includes a transceiver unit 910 , wherein:

[0248] The transceiver unit 910 is configured to receive an identifier of at least one target terminal device participating in each round of model training, sent by a first network device, where the at least one target terminal device is determined by the first network from among the multiple terminal devices using a first algorithm based on resource information of the multiple terminal devices.

[0249] The transceiver unit 910 is further configured to send the model of each round to the at least one target terminal device.

[0250] Optionally, the resource information includes at least one of the following:

[0251] The EMD value of the local data distribution of the terminal device;

[0252] The central processing unit (CPU) frequency of the terminal device;

[0253] The battery level of the terminal device.

[0254] Optionally, the weighted sum of the EMD value of the local data distribution of at least one target terminal device participating in all rounds of model training, the total energy loss of at least one target terminal device participating in all rounds of model training, and the total time delay of at least one target terminal device participating in all rounds of model training is minimized.

[0255] Optionally, the transceiver unit 910 is further configured to:

[0256] Receive model update parameters sent by the at least one target terminal device after model training, and use the model update parameters for the second network device to update the model for the next round of model training.

[0257] Optionally, the model updated by the second network device is used to aggregate into a global model; after all rounds of model training, the accuracy of the global model converges.

[0258] The model training device of this embodiment can receive the identifier of at least one target terminal device participating in each round of model training sent by the first network device, and the at least one target terminal device is determined by the first network from the multiple terminal devices based on the resource information of the multiple terminal devices using the first algorithm. The model of each round is sent to the at least one target terminal device, so that the network side can find an approximate optimal solution in each round of model training and select the terminal devices participating in the model training, so that the weighted sum of the data distribution heterogeneity and training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0259] Please refer to Figure 10, which is a structural diagram of a model training device provided in an embodiment of the present application.

[0260] As shown in FIG10 , the model training device 1000 includes a transceiver unit 1010 , wherein:

[0261] The transceiver unit 1010 is configured to send resource information to the first network device, where the resource information is used by the first network device to determine, based on the resource information and using a first algorithm, at least one target terminal device from a plurality of terminal devices to participate in each round of model training;

[0262] The transceiver unit 1010 is further configured to receive the current round model sent by the second network device in response to the terminal device being a target terminal device participating in the current round model training.

[0263] Optionally, the resource information includes at least one of the following:

[0264] The EMD value of the local data distribution of the terminal device;

[0265] The central processing unit (CPU) frequency of the terminal device;

[0266] The battery level of the terminal device.

[0267] Optionally, the weighted sum of the EMD value of the local data distribution of at least one target terminal device participating in all rounds of model training, the total energy loss of at least one target terminal device participating in all rounds of model training, and the total time delay of at least one target terminal device participating in all rounds of model training is minimized.

[0268] Optionally, the transceiver unit 1010 is further used to: send model update parameters obtained after model training to the second network device, and the model update parameters are used by the second network device to update the model for the next round of model training.

[0269] Optionally, the model updated by the second network device is used to aggregate into a global model; after all rounds of model training, the accuracy of the global model converges.

[0270] The model training device of this embodiment can send resource information to the first network device, and the resource information is used by the first network device to determine at least one target terminal device participating in each round of model training from multiple terminal devices based on the resource information and using a first algorithm. In response to the terminal device being the target terminal device participating in the current round of model training, the model of the current round sent by the second network device is received, so that the network side can find an approximate optimal solution in each round of model training and select the terminal devices participating in the model training, so that the weighted sum of the data distribution heterogeneity and training overhead of the overall selected terminal devices is minimized, which can effectively improve the model prediction accuracy, while also accelerating the model convergence speed and reducing energy consumption.

[0271] In order to implement the above-mentioned embodiments, the embodiments of the present application also propose a communication device, including: a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the computer program stored in the memory so that the device executes the method shown in the embodiments of Figures 2 to 3, or executes the method shown in the embodiments of Figures 4 to 5.

[0272] In order to implement the above embodiment, the embodiment of the present application also proposes a communication device, including: a processor and a memory, the memory storing a computer program, and the processor executing the computer program stored in the memory so that the device executes the method shown in the embodiment of Figure 6.

[0273] In order to implement the above embodiments, the embodiments of the present application also propose a communication device, including: a processor and an interface circuit, the interface circuit is used to receive code instructions and transmit them to the processor, and the processor is used to run the code instructions to execute the method shown in the embodiments of Figures 2 to 3, or execute the method shown in the embodiments of Figures 4 to 5.

[0274] In order to implement the above embodiment, the embodiment of the present application also proposes a communication device, including: a processor and an interface circuit, the interface circuit is used to receive code instructions and transmit them to the processor, and the processor is used to run the code instructions to execute the method shown in the embodiment of Figure 6.

[0275] Please refer to Figure 11, which is a schematic diagram of the structure of another model training device provided in an embodiment of the present disclosure. Model training device 1100 can be a network device, a terminal device, a chip, a chip system, or a processor that supports the network device to implement the above method, or a chip, a chip system, or a processor that supports the terminal device to implement the above method. This device can be used to implement the method described in the above method embodiment. For details, please refer to the description of the above method embodiment.

[0276] The model training device 1100 may include one or more processors 1101. The processor 1101 may be a general-purpose processor or a dedicated processor. For example, it may be a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control the model training device (e.g., a base station, a baseband chip, a terminal device, a terminal device chip, a DU or CU, etc.), execute computer programs, and process computer program data.

[0277] Optionally, the model training device 1100 may further include one or more memories 1102, on which a computer program 1103 may be stored. The processor 1101 executes the computer program 1103 to enable the model training device 1100 to perform the method described in the above method embodiment. The computer program 1103 may be fixed in the processor 1101, in which case the processor 1101 may be implemented by hardware.

[0278] Optionally, data may also be stored in the memory 1102. The model training device 1100 and the memory 1102 may be provided separately or integrated together.

[0279] Optionally, the model training device 1100 may further include a transceiver 1105 and an antenna 1106. The transceiver 1105 may be referred to as a transceiver unit, a transceiver, or a transceiver circuit, and is configured to implement transceiver functions. The transceiver 1105 may include a receiver and a transmitter. The receiver may be referred to as a receiver or a receiving circuit, and is configured to implement a receiving function; the transmitter may be referred to as a transmitter or a transmitting circuit, and is configured to implement a transmitting function.

[0280] Optionally, the model training device 1100 may further include one or more interface circuits 1107. The interface circuit 1107 is configured to receive code instructions and transmit the instructions to the processor 1101. The processor 1101 executes the code instructions to enable the model training device 1100 to perform the method described in the above method embodiment.

[0281] In one implementation, processor 1101 may include a transceiver for implementing receiving and transmitting functions. For example, the transceiver may be a transceiver circuit, an interface, or an interface circuit. The transceiver circuit, interface, or interface circuit for implementing the receiving and transmitting functions may be separate or integrated. The transceiver circuit, interface, or interface circuit may be used for reading and writing code / data, or may be used for transmitting or delivering signals.

[0282] In one implementation, the model training device 1100 may include a circuit that can implement the functions of sending, receiving, or communicating in the aforementioned method embodiments. The processor and transceiver described in this disclosure can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit RFIC, a mixed signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (nMetal-oxide-semiconductor, NMOS), P-type metal oxide semiconductor (positive channel metal oxide semiconductor, PMOS), bipolar junction transistor (bipolar junction transistor, BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0283] The model training device described in the above embodiments may be a network device or a terminal device, but the scope of the model training device described in this disclosure is not limited thereto, and the structure of the model training device may not be limited by Figures 8-10. The model training device may be an independent device or may be part of a larger device. For example, the model training device may be:

[0284] (1) An independent integrated circuit (IC), or chip, or chip system or subsystem;

[0285] (2) a collection of one or more ICs, optionally including a storage component for storing data and computer programs;

[0286] (3) ASIC, such as modem;

[0287] (4) Modules that can be embedded in other devices;

[0288] (5) Receivers, terminal devices, intelligent terminal devices, cellular phones, wireless devices, handheld devices, mobile units, vehicle-mounted devices, network devices, cloud devices, artificial intelligence devices, etc.;

[0289] (6)Others, etc.

[0290] In the case where the model training device can be a chip or a chip system, please refer to the structural diagram of the chip shown in Figure 12. The chip shown in Figure 12 includes a processor 1201 and an interface 1202. The number of processors 1201 can be one or more, and the number of interfaces 1202 can be multiple.

[0291] For the case where the chip is used to implement the functions of the network device in the embodiments of the present disclosure:

[0292] Interface 1202, used for transmitting code instructions to the processor;

[0293] The processor 1201 is configured to execute code instructions to execute the method shown in FIG. 2 to FIG. 3 , or to execute the method shown in FIG. 4 to FIG. 5 .

[0294] For the case where the chip is used to implement the functions of the terminal device in the embodiments of the present disclosure:

[0295] Interface 1202, used for transmitting code instructions to the processor;

[0296] The processor 1201 is configured to execute code instructions to perform the method shown in FIG6 .

[0297] Optionally, the chip further includes a memory 1203, which is used to store necessary computer programs and data.

[0298] Those skilled in the art will also appreciate that the various illustrative logical blocks and steps listed in the embodiments of the present disclosure may be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement functionality for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present disclosure.

[0299] An embodiment of the present disclosure also provides a communication system, which includes the model training device as a network device and the model training device as a terminal device in the embodiments of Figures 7 to 9 above, or the system includes the model training device as a terminal device and the model training device as a network device in the embodiment of Figure 11 above.

[0300] The present disclosure also provides a readable storage medium having instructions stored thereon, which implement the functions of any of the above method embodiments when executed by a computer.

[0301] The present disclosure also provides a computer program product, which implements the functions of any of the above method embodiments when executed by a computer.

[0302] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs. When the computer program is loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. Available media may be magnetic media (eg, floppy disks, hard disks, tapes), optical media (eg, high-density digital video discs (DVDs)), or semiconductor media (eg, solid state disks (SSDs)).

[0303] Those skilled in the art will understand that the various numerical numbers such as first and second involved in the present disclosure are only for the convenience of description and are not used to limit the scope of the embodiments of the present disclosure, and also indicate the order of precedence.

[0304] The at least one in the present disclosure can also be described as one or more, and the multiple can be two, three, four or more, which is not limited in the present disclosure. In the embodiments of the present disclosure, for a technical feature, the technical features in the technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", and there is no order of precedence or size between the technical features described by "first", "second", "third", "A", "B", "C" and "D".

[0305] The correspondences shown in the tables of the present disclosure can be configured or predefined. The values ​​of the information in each table are merely examples and can be configured to other values, which are not limited by the present disclosure. When configuring the correspondences between information and parameters, it is not necessarily required to configure all the correspondences shown in each table. For example, in the tables of the present disclosure, the correspondences shown in certain rows may not be configured. For another example, appropriate deformation adjustments can be made based on the above tables, such as splitting, merging, etc. The names of the parameters shown in the titles of the above tables may also adopt other names that can be understood by the communication device, and the values ​​or representations of the parameters may also adopt other values ​​or representations that can be understood by the communication device. When implementing the above tables, other data structures may also be used, such as arrays, queues, containers, stacks, linear lists, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables or hash tables, etc.

[0306] The predefined in the present disclosure may be understood as defined, predefined, stored, pre-stored, pre-negotiated, pre-configured, solidified, or pre-burned.

[0307] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0308] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0309] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the embodiments of the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and this document is not limited here.

[0310] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A model training method, characterized in that: The method is performed by a first network device, and the method includes: Receiving resource information sent by multiple terminal devices; According to the resource information, using a first algorithm, determining at least one target terminal device participating in each round of model training from the multiple terminal devices; The identifier of the at least one target terminal device is sent to the second network device, and the identifier of the at least one target terminal device is used by the second network device to send the model of each round to the at least one target terminal device.

2. The method according to claim 1, characterized in that The resource information includes at least one of the following: The soil pile distance EMD value of the local data distribution of the terminal device; The central processing unit CPU frequency of the terminal device; The battery level of the terminal device.

3. The method according to claim 2, characterized in that The EMD value of the local data distribution of at least one target terminal device participating in all rounds of model training, the weighted sum of the total energy loss of at least one target terminal device participating in all rounds of model training, and the total time delay of at least one target terminal device participating in all rounds of model training is minimized.

4. The method according to claim 1, characterized in that: The model of each round is used for the target terminal device to perform model training to obtain model update parameters, and the model update parameters are used by the second network device to update the model for the next round of model training.

5. The method according to claim 4, characterized in that The model updated by the second network device is used to aggregate into a global model; after all rounds of model training, the accuracy of the global model converges.

6. A model training method, characterized in that: The method is performed by a second network device, and the method includes: Receiving an identifier of at least one target terminal device participating in each round of model training sent by the first network device, wherein the at least one target terminal device is determined by the first network from the multiple terminal devices according to resource information of the multiple terminal devices using a first algorithm; The model of each round is sent to the at least one target terminal device.

7. The method according to claim 6, characterized in that The resource information includes at least one of the following: The soil pile distance EMD value of the local data distribution of the terminal device; The central processing unit CPU frequency of the terminal device; The battery level of the terminal device.

8. The method according to claim 7, characterized in that The EMD value of the local data distribution of at least one target terminal device participating in all rounds of model training, the weighted sum of the total energy loss of at least one target terminal device participating in all rounds of model training, and the total time delay of at least one target terminal device participating in all rounds of model training is minimized.

9. The method according to claim 6, characterized in that The method further comprises: Receive model update parameters sent by the at least one target terminal device after model training, and use the model update parameters for the second network device to update the model for the next round of model training.

10. The method according to claim 9, characterized in that The model updated by the second network device is used to aggregate into a global model; after all rounds of model training, the accuracy of the global model converges.

11. A model training method, characterized in that: The method is performed by a terminal device, and the method includes: Sending resource information to the first network device, where the resource information is used by the first network device to determine at least one target terminal device participating in each round of model training from a plurality of terminal devices according to the resource information and using a first algorithm; In response to the terminal device being a target terminal device participating in the current round of model training, the current round of model training sent by the second network device is received.

12. The method according to claim 11, characterized in that The resource information includes at least one of the following: The soil pile distance EMD value of the local data distribution of the terminal device; The central processing unit CPU frequency of the terminal device; The battery level of the terminal device.

13. The method according to claim 12, characterized in that The EMD value of the local data distribution of at least one target terminal device participating in all rounds of model training, the weighted sum of the total energy loss of at least one target terminal device participating in all rounds of model training, and the total time delay of at least one target terminal device participating in all rounds of model training is minimized.

14. The method according to claim 11, characterized in that The method further comprises: The model update parameters obtained after the model training are sent to the second network device, where the model update parameters are used by the second network device to update the model for the next round of model training.

15. The method according to claim 14, characterized in that The model updated by the second network device is used to aggregate into a global model; after all rounds of model training, the accuracy of the global model converges.

16. A model training device, characterized in that: The device is applied to a first network device, and the device includes: A transceiver unit, used for receiving resource information sent by multiple terminal devices; A processing unit, configured to determine at least one target terminal device participating in each round of model training from the plurality of terminal devices according to the resource information and using a first algorithm; The transceiver unit is further used to send the identifier of the at least one target terminal device to the second network device, and the identifier of the at least one target terminal device is used by the second network device to send each round of models to the at least one target terminal device.

17. A model training device, characterized in that: The device is applied to a second network device, and the device includes: a transceiver unit, configured to receive an identifier of at least one target terminal device participating in each round of model training sent by a first network device, wherein the at least one target terminal device is determined by the first network from the plurality of terminal devices using a first algorithm according to resource information of the plurality of terminal devices; The transceiver unit is further used to send the model of each round to the at least one target terminal device.

18. A model training device, characterized in that: The device is applied to a terminal device, and the device includes: A transceiver unit, configured to send resource information to a first network device, wherein the resource information is used by the first network device to determine at least one target terminal device participating in each round of model training from a plurality of terminal devices according to the resource information and using a first algorithm; The transceiver unit is also used to receive the current round model sent by the second network device in response to the terminal device being a target terminal device participating in the current round model training.

19. A communication device, characterized in that: The device includes a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the computer program stored in the memory so that the device performs the method as claimed in any one of claims 1 to 5, or performs the method as claimed in any one of claims 6 to 10, or performs the method as claimed in any one of claims 11 to 15.

20. A communication device, characterized in that: include: processor and interface circuits; The interface circuit is used to receive code instructions and transmit them to the processor; The processor is configured to execute the code instructions to execute the method according to any one of claims 1 to 5, or to execute the method according to any one of claims 6 to 10, or to execute the method according to any one of claims 11 to 15.

21. A computer-readable storage medium storing instructions, which, when executed, implement the method according to any one of claims 1 to 5, or the method according to any one of claims 6 to 10, or the method according to any one of claims 11 to 15.

22. A communication system, characterized in that: The communication system comprises: A first network device, configured to execute the method according to any one of claims 1 to 5; A second network device, configured to perform the method according to any one of claims 6 to 10; A terminal device, configured to execute the method as claimed in any one of claims 11 to 15.