Federal Learning Device Scheduling Optimization Method and Device Based on Double-Layer Reinforcement Learning

Through the equipment scheduling optimization method of double-layer reinforcement learning, the global model performance degradation, energy consumption and unfair participation caused by equipment heterogeneity and dynamics is solved, and efficient, fair and stable scheduling of equipment resources is achieved.

CN120068994BActive Publication Date: 2025-07-22QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510512002.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the environment of device heterogeneity and dynamics, traditional federated learning scheduling algorithms fail to effectively utilize the computing power of high-performance devices, resulting in low-performance devices becoming a system bottleneck and uneven energy consumption control, affecting the performance of the global model and fairness of device participation.

Method used

The equipment scheduling optimization method based on double-layer reinforcement learning is adopted, and through the upper reinforcement learning grouping and the participation rate is allocated, the lower reinforcement learning selects equipment, and combined with dynamic weight adjustment, the equipment scheduling strategy is optimized to improve the performance of the global model, reduce energy consumption and improve the fairness of equipment participation.

Benefits of technology

The efficiency and robustness of resource scheduling are achieved in the heterogeneous equipment environment, the performance of the global model is improved, the energy consumption of equipment is reduced, and the fairness of equipment participation is improved, and the dynamic changes in equipment status is adapted to the dynamic changes of equipment status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068994B_ABST
    Figure CN120068994B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of federated learning, and particularly relates to a method and device for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning. The method includes: obtaining the current state characteristics of the devices; dividing the devices into groups, and allocating participation rates for each group of devices through upper-layer reinforcement learning; selecting the devices participating in federated learning within each group through lower-layer reinforcement learning; constructing a device scheduling objective function, and initializing the global model parameters, and performing federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, the device scheduling objective function is maximized by adjusting the participation rates of each group of devices and the scoring weights of the optimization objectives, so as to determine the optimal device scheduling strategy. The present invention uses a double-layer reinforcement learning strategy to optimize the scheduling of federated learning devices, aiming to improve the performance of the global model, reduce the energy consumption of the devices, and improve the fairness of device participation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of federated learning, and more specifically, relates to a method and device for optimizing device scheduling in federated learning based on double-layer reinforcement learning. Background Art

[0002] Traditional centralized training methods not only bring huge communication overhead but also have serious data privacy leakage risks because they require uploading a large amount of raw data to the server for unified training. To solve such problems, federated learning, as a distributed machine learning method that can protect data privacy, has emerged.

[0003] Federated learning allows devices to complete model training locally and achieve global aggregation by uploading model parameters instead of raw data, which not only effectively reduces the risk of data privacy leakage but also significantly reduces the bandwidth requirements for data transmission. Due to its unique privacy protection and communication efficiency advantages, federated learning has been widely applied in scenarios such as the Internet of Things, mobile terminals, and edge computing in recent years.

[0004] However, in an environment where there are significant differences in the resource states of devices such as computing power, battery power, and network latency, how to design an efficient resource scheduling strategy for device heterogeneity to make full use of the computing power of high-performance devices while avoiding low-performance devices from becoming system bottlenecks, thereby improving the global training efficiency and model performance, has become a key issue restricting the development of federated learning technology.

[0005] Traditional federated learning scheduling algorithms usually adopt a unified training strategy without designing an optimization mechanism for device heterogeneity and dynamics. This leads to the following problems: on the one hand, the computing power and network bandwidth of high-performance devices are often not fully utilized, while low-performance devices may become system bottlenecks, dragging down the training efficiency of the global model. On the other hand, there is a lack of pertinence in energy consumption control, and some devices may be forced to withdraw from training due to excessive power consumption or insufficient power, affecting the convergence stability of the global model.

[0006] For example, Chinese patent document CN115033382A discloses a method for device scheduling in a multi-task federated learning system, including: constructing a system model for multi-task federated learning; establishing an optimization problem with the goal of minimizing the time of the multi-task federated learning process; scheduling devices to participate in the training process of the federated learning task; transforming the device scheduling process into a multi-armed bandit and matching process; and designing a device scheduling algorithm. The present invention is to schedule the most suitable device for each task in federated learning, thereby minimizing the latency of the multi-task federated learning process.

[0007] In addition, the unfairness problem of device scheduling cannot be ignored. In a heterogeneous environment, devices that participate frequently dominate model training. However, due to insufficient participation, the local data of devices with low participation cannot be effectively learned by the global model, resulting in a decline in the generalization performance of the model. More importantly, device states usually change dynamically over time, such as battery power consumption and network condition fluctuations. Existing scheduling strategies lack flexibility and are difficult to adapt to the dynamically changing device environment, further restricting the practical application of federated learning.

[0008] For example, Chinese Patent Document CN119299398A discloses a method for selecting federated learning terminals and resource scheduling based on dynamic priorities, including: obtaining the basic model structure and resource information parameters exchanged between the federated learning central controller and users; calculating the priority value of each user according to the resource information parameters, and sorting the users from largest to smallest according to the priority value; based on the user priority order, selecting a subset of users that meet the training delay requirements; for the subset of users, calculating the minimum bandwidth allocation ratio required to complete the training task, considering the data volume, computing power, and network conditions of the users; using the dichotomy method to allocate bandwidth resources to the subset of users, and the dichotomy method adjusts the bandwidth allocation through iteration; determining whether the total energy consumption of the subset of users exceeds a given energy constraint threshold, and if it exceeds the threshold, deleting the user with the lowest priority from the subset of users.

[0009] Based on this, the present invention designs an optimization method for federated learning device scheduling based on double-layer reinforcement learning, aiming to solve the problems of the decline in the performance of the global model, uneven energy consumption distribution, and insufficient fairness of device participation caused by unreasonable device resource scheduling during the federated learning training process in a heterogeneous device environment. Summary of the Invention

[0010] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides an optimization method for federated learning device scheduling based on double-layer reinforcement learning, which is used to optimize the federated learning process in a heterogeneous device environment and achieve comprehensive optimization of the performance of the global model, energy consumption efficiency, and fairness of device participation.

[0011] The present invention also discloses a device loaded with the optimization method for federated learning device scheduling based on double-layer reinforcement learning.

[0012] The detailed technical solution of the present invention is as follows:

[0013] An optimization method for federated learning device scheduling based on double-layer reinforcement learning, the method includes:

[0014] S1. Obtain the current state characteristics of the device , the current state characteristics of the device include computing power and battery power and network latency ;

[0015] S2. Divide the devices into groups, and allocate participation rates for each group of devices through upper-layer reinforcement learning , including: based on the current state characteristics of each device in each group of devices and the change in global model accuracy and the total energy consumption of the devices, construct the state of the upper-layer reinforcement learning, and allocate participation rates for each group of devices based on this state , and determine whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-layer reward function ;

[0016] S3. Select the devices participating in federated learning within each group through lower-layer reinforcement learning, including: based on the current state characteristics of each device in each group of devices and the change in the local model accuracy of each device and the energy consumption of the device, construct the state of the lower-layer reinforcement learning, calculate the scoring mechanism of each device based on this state to select the devices participating in federated learning within each group, and measure the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function;

[0017] S4. Construct a device scheduling objective function and initialize the global model parameters , and perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rates of each group of devices and the weight coefficients of the optimization objectives maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein, the optimization objectives include improving the global model accuracy, controlling the system energy consumption, and ensuring fairness in device participation.

[0018] Preferably according to the present invention, in S2, based on the current state characteristics of each device in each group of devices and the change in global model accuracy and the total energy consumption of the devices, construct the state of the upper-layer reinforcement learning, and allocate participation rates for each group of devices based on this state , specifically, the upper-layer reinforcement learning maps the comprehensive state of each group of devices to the participation rate through the policy function , that is: (1);

[0019] (1);

[0020] wherein, the comprehensive state of each group of devices includes the computing power of the devices in the group , battery power , network latency , and the current change in global model accuracy and the total energy consumption of the device .

[0021] Preferably according to the present invention, in the step S2, based on the change in the global model accuracy , the total energy consumption of the device and the data quality to construct an upper-layer reward function, that is:

[0022] (2);

[0023] In formula (2): represents the upper-layer reward function; represents the change in the global model accuracy; represents the total energy consumption of the selected device; represents the scoring weight corresponding to the total energy consumption of the device; represents the data quality coefficient of the nth group of devices.

[0024] Preferably according to the present invention, in the step S3, the scoring mechanism is:

[0025] (3);

[0026] In formula (3): represents the scoring result of the device in the training round ; represents the maximum value of the computing power of the devices in the system; represents the fully charged state value of the device; , , , are the scoring weight coefficients of the computing performance, energy consumption status, communication quality and fair participation of the device respectively; is a time decay function, representing a time decay term set considering the historical participation of the device, and:

[0027] (4);

[0028] In formula (4): represents the participation frequency of the device in the training round ; represents the time decay coefficient, used to control the decay speed; represents the time interval between the current training round and the device last participation in training.

[0029] Preferably according to the present invention, in the step S3, the lower-layer reward function is:

[0030] (5);

[0031] In formula (5): represents the lower-layer reward function; represents the change in accuracy of the local model in the training round ; represents the network stability of the device ; represents the device in the training round energy consumption; represents the scoring weight corresponding to the change in local model accuracy; represents the scoring weight corresponding to the device energy consumption.

[0032] According to the preference of the present invention, in step S4, the constructed device scheduling objective function is:

[0033] (6);

[0034] In formula (6): represents the device scheduling objective function; represents the training round increment of the global model accuracy; represents in the training round total energy consumption of the selected devices; represents the total number of training rounds; represents the weight coefficient for the improvement of the global model accuracy, represents the weight coefficient for the system energy consumption control, represents the weight coefficient for the fairness of device participation; is a fairness index, and:

[0035] (7);

[0036] In formula (7): represents the total number of devices selected to participate in the training; is a positive number used to prevent division-by-zero errors; represents the mean of the participation frequencies of all devices.

[0037] According to the preference of the present invention, in step S4, the participation rate of each group of devices and the weight coefficient of the optimization objective are updated based on the following formula:

[0038] (11);

[0039] (12);

[0040] In formulas (11)-(12): represents the participation rate of the th group in the training round ; represents the participation rate of the th group in the training round ; represents the adjustment value of the participation rate of the th group generated by the current round's strategy; represents the th optimization objective's weight coefficient in the training round , and , respectively corresponding to the weight coefficients representing the improvement of the global model accuracy, the control of the system energy consumption, and the fairness of device participation; represents the learning rate; represents the gradient of the reward function with respect to the weight ; represents the temperature decay term, which gradually decreases as the training round increases; the parameter is used to control the decay speed.

[0041] In another aspect of the present invention, there is provided a device for implementing an optimization method for federated learning device scheduling based on double-layer reinforcement learning, and the device includes:

[0042] A data acquisition module, configured to acquire the current state features of the device , and the current state features of the device include computing power , battery power and network latency ;

[0043] An upper-layer grouping module, configured to divide the devices into groups through a clustering algorithm, and allocate a participation rate to each group of devices through upper-layer reinforcement learning, including: based on the current state features of each device in each group of devices and the change in the global model accuracy and the total energy consumption of the devices, constructing the state of the upper-layer reinforcement learning, and based on this state, allocating a participation rate to each group of devices, and judging whether to optimize and adjust the participation rate of each group of devices based on the feedback of the upper-layer reward function ;

[0044] A lower-layer device selection module, configured to select the devices participating in the federated learning within each group through lower-layer reinforcement learning, including: based on the current state features of each device in each group of devices And based on the changes in the local model accuracy of each device and the device energy consumption, construct the state of the lower-layer reinforcement learning. Based on this state, calculate the scoring mechanism for each device to select the devices participating in federated learning within each group, and measure the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function;

[0045] A federated learning training module, used to construct a device scheduling objective function and initialize the global model parameters , and perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rate of each group of devices and the weight coefficients of the optimization objectives maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein, the optimization objectives include global model accuracy improvement, system energy consumption control, and device participation fairness.

[0046] In another aspect of the present invention, an electronic device is further provided, including:

[0047] At least one processor; and

[0048] A memory, the memory stores instructions, when the instructions are executed by the at least one processor, the at least one processor executes the federated learning device scheduling optimization method based on double-layer reinforcement learning as described above.

[0049] In another aspect of the present invention, a machine-readable storage medium is further provided, which stores executable instructions, and when the instructions are executed, the machine executes the federated learning device scheduling optimization method based on double-layer reinforcement learning as described above.

[0050] Compared with the prior art, the beneficial effects of the present invention are:

[0051] (1) A federated learning device scheduling optimization method based on double-layer reinforcement learning provided by the present invention, on the basis of considering device heterogeneity including problems such as computing power, battery power, and network latency, uses a double-layer reinforcement learning strategy to optimize federated learning device scheduling, aiming to improve the global model performance, reduce device energy consumption, and improve device participation fairness.

[0052] (2) Through the double-layer design of the upper-layer grouping strategy and the lower-layer device selection strategy, combined with the dynamic weight adjustment mechanism, the present invention can achieve the efficiency and robustness of resource scheduling in a heterogeneous device environment. Among them, the upper-layer strategy is responsible for grouping devices according to device status information and assigning appropriate participation rates to each group of devices; the lower-layer strategy selects suitable devices for participating in training within the group according to the device scoring mechanism.

[0053] (3) The present invention synthesizes the multi-objective optimization requirements in the federated learning process, balances the global model accuracy, device energy consumption, and participation fairness, and can adapt to the dynamic changes in device status. Description of the Drawings

[0054] Figure 1 It is a flowchart of the federated learning device scheduling optimization method based on double-layer reinforcement learning according to the present invention.

[0055] Figure 2 It is a comparison test chart of the method Fed-RL of the present invention and the random selection device scheduling method Random Selection on the MNIST dataset. Specific Embodiments

[0056] The present invention will be further described below in conjunction with the drawings and embodiments.

[0057] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0058] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0059] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0060] Aiming at the deficiencies in the prior art, the present invention proposes a federated learning device scheduling optimization method based on double-layer reinforcement learning. By introducing a double-layer design of an upper-layer grouping strategy and a lower-layer device selection strategy, and combining the reinforcement learning mechanism to dynamically adjust the device participation rate and scoring weights, the core problem of resource scheduling in a device heterogeneous environment is effectively solved.

[0061] In the upper-layer grouping strategy, devices are grouped according to state characteristics such as computing power, battery power, and network latency, reducing intra-group heterogeneity and improving the efficiency of grouped resource allocation; the lower-layer device selection strategy then selects the devices suitable for the current training round within each group based on a scoring mechanism, ensuring the accuracy of device selection and the effectiveness of training.

[0062] Meanwhile, by dynamically adjusting the device participation rate and the weight coefficients of each optimization objective through reinforcement learning, the present invention realizes the multi-objective comprehensive optimization of model performance, energy consumption, and participation fairness.

[0063] The following further describes the federated learning device scheduling optimization method and device based on double-layer reinforcement learning of the present invention with specific embodiments.

[0064] Embodiment 1

[0065] Refer Figure 1 , this embodiment provides a federated learning device scheduling optimization method based on double-layer reinforcement learning, which is applied to a federated learning system. The system includes a server and multiple devices participating in federated learning, and the server is communicatively connected to each device. The method includes:

[0066] S1. Obtain the current state characteristics of the device , and the current state characteristics of the device include computing power , battery power , and network latency .

[0067] The method of this embodiment models the device scheduling problem of federated learning in a heterogeneous device environment. Assume that there are devices that may participate in federated learning in the system, that is, it includes devices selected to participate in federated learning, and at the same time includes devices not selected to participate in federated learning. The state of each device can be represented as a feature vector , , , where represents the computing power of the device, such as the performance of the CPU or GPU; represents the remaining battery power of the device, and the range is [0, 1]; represents the network latency of the device, which is used to reflect the communication performance between the device and the server.

[0068] The obtained current state characteristics of the device will be used as the state space data in the subsequent reinforcement learning steps.

[0069] S2. Divide the devices into groups, and allocate a participation rate to each group of devices through upper-layer reinforcement learning, including: based on the current state characteristics of each device in each group of devices, as well as the global model accuracy change and the total device energy consumption, construct the state of the upper-layer reinforcement learning, and allocate the participation rate , and determine whether to optimize and adjust the participation rate of each group of devices based on the feedback of the upper-layer reward function .

[0070] In this embodiment, first, the devices are divided into groups through a clustering algorithm. The clustering algorithm is K-Means clustering, which takes the state feature vector of the devices as the input. To eliminate the scale influence between different metrics, the features of each dimension are standardized before clustering. The K-Means algorithm finally obtains device subsets where the devices within each group have high similarity in terms of computing power, energy consumption status, and communication conditions, which is beneficial to improving the effectiveness of the participation rate allocation strategy and the robustness of scheduling in subsequent upper-layer reinforcement learning.

[0071] Subsequently, the participation rate is allocated to each group of devices through upper-layer reinforcement learning.

[0072] The purpose of upper-layer reinforcement learning is to group the devices based on their resource characteristics, thereby reducing the scheduling complexity brought by device heterogeneity; at the same time, the participation rates of different groups are set so that the global model can achieve accuracy improvement under limited system resources.

[0073] Specifically, the participation rate is allocated to each group of devices through upper-layer reinforcement learning , including:

[0074] Construct the state space: Based on the current state characteristics of each device in each group of devices , including the computing power of the device , battery power , network latency , and the change in the global model accuracy , total device energy consumption construct the state of upper-layer reinforcement learning, where represents the training round; represents the change in the global model accuracy in the current training round , defined as , that is, the difference in the global model accuracy between two consecutive training rounds.

[0075] Construct the action space: Allocate an appropriate participation rate to each group of devices based on the above state .

[0076] Based on the above state information, upper-layer reinforcement learning maps the comprehensive state of each group of devices to the participation rate through the policy function , where the comprehensive state Including the computing power of the devices within the group , battery power , network latency , as well as the current change in the global model accuracy and the total energy consumption of the devices .

[0077] This policy function can be implemented in the form of a neural network, and the parameters are optimized through training so that it can be dynamically adjusted in each round of training to obtain a higher cumulative reward. Specifically:

[0078] (1).

[0079] The goal of policy training is to maximize the cumulative reward function, which comprehensively considers factors such as the improvement of the global model accuracy, the total energy consumption of the devices, and the data quality. The policy parameters are continuously optimized through the policy gradient method to achieve intelligent dynamic allocation of the participation rate.

[0080] Construct the reward function: The upper-layer reward function is constructed based on the change in the global model accuracy , the total energy consumption of the devices and the data quality . Specifically:

[0081] (2);

[0082] In Equation (2): represents the upper-layer reward function; represents the change in the global model accuracy; represents the total energy consumption of the selected devices; represents the scoring weight corresponding to the total energy consumption of the devices; represents the data quality coefficient of the devices in the th group, which is used to preferentially select groups with high data quality to enhance the model performance. This indicator can be comprehensively evaluated by combining factors such as the sample size of the data within the group, the conformity of the data with the task objective distribution (i.e., distribution representativeness), and the annotation integrity.

[0083] Judge whether it is necessary to optimize and adjust the participation rate of each group of devices according to the feedback of the upper-layer reward function to ensure the effectiveness of the group selection and the participation rate .

[0084] S3. Select the devices participating in federated learning within each group through lower-layer reinforcement learning, including: Based on the current state characteristics of each device in each group Construct the state of the lower-layer reinforcement learning based on the local model accuracy change of each device and the device energy consumption. Calculate the scoring mechanism for each device based on this state to select the devices participating in federated learning within each group, and measure the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function.

[0085] In this embodiment, the purpose of the lower-layer reinforcement learning is to select the devices suitable for participating in federated learning training within the group to achieve the resource effectiveness and fairness of scheduling.

[0086] Specifically, selecting the devices participating in federated learning within each group through the lower-layer reinforcement learning includes:

[0087] Construct the state space: Based on the current state characteristics of each device in each group , including the computing power of the device , battery power , network latency , and the change in the local model accuracy of each device , device energy consumption Construct the state of the lower-layer reinforcement learning.

[0088] Construct the action space: Calculate the scoring mechanism for each device based on the above state to select the devices participating in federated learning within each group; the scoring mechanism is:

[0089] (3);

[0090] In formula (3): represents the scoring result of device in the training round ; represents the maximum value of the device computing power in the system; represents the full charge state value of the device; , , , are the scoring weight coefficients of device computing performance, energy consumption status, communication quality, and fair participation respectively, used to adjust the influence of the above factor scores; is the time decay function, representing the time decay term set for considering the historical participation of the device, and:

[0091] (4);

[0092] In formula (4): represents the participation frequency of device in the training round , that is, the frequency of this device being selected in the current training process; represents the time decay coefficient, controlling the decay speed; Indicates the time interval between the current training round and the device last participated in training.

[0093] Construct the reward function: The lower-layer reward function comprehensively considers the contribution of the device to the local model accuracy, energy consumption, and network stability. Specifically:

[0094] (5);

[0095] In formula (5): represents the lower-layer reward function; represents the accuracy change of the local model in the training round ; represents the device 's network stability; represents the device in the training round 's energy consumption; represents the scoring weight corresponding to the local model accuracy change; represents the scoring weight corresponding to the device energy consumption.

[0096] Measure the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function. The lower-layer reward function calculates the performance of the device under the policy execution based on the local accuracy change and energy consumption of the device in each round of training, and provides a value feedback signal for policy learning. This feedback can be used to optimize and update the lower-layer reinforcement learning policy, ensuring that the system selects more cost-effective devices to participate in training in subsequent rounds, thereby improving the overall performance and resource utilization efficiency of the global model.

[0097] S4. Construct the device scheduling objective function and initialize the global model parameters , and perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rate of each group of devices and the weight coefficient of the optimization objective

[0098] maximize the device scheduling objective function to determine the optimal device scheduling strategy; among them, the optimization objectives include global model accuracy improvement, system energy consumption control, and device participation fairness.

[0099] In this embodiment, the device scheduling problem is modeled as a multi-objective optimization problem, and the constructed device scheduling objective function is:

[0100] (6);

[0101] In formula (6): represents the device scheduling objective function; represents the training round the increment of the global model accuracy in; represents in the training round the total energy consumption of the selected devices in; represents the total number of training rounds; represents the weight coefficient for the improvement of the global model accuracy, represents the weight coefficient for system energy consumption control, represents the weight coefficient for device participation fairness; fairness index : inverse frequency weighting, used to prevent resource-rich devices from frequently dominating the training, making the system tend to pay more attention to low-frequency devices, that is, used to measure whether the data of low-frequency participating devices is fully utilized, and:

[0102] (7);

[0103] In formula (7): represents the total number of selected devices participating in the training; is a small positive number, used to prevent division by zero errors and ensure numerical stability; represents the mean of the participation frequencies of all devices.

[0104] The purpose of this method is to achieve the maximization of the device scheduling objective function by dynamically adjusting the grouping participation rate of devices and the weight coefficients of each optimization objective in device selection , so as to find an optimal device scheduling strategy that achieves an overall optimal balance in terms of comprehensive performance, resources, and fairness. Specifically, during the federated learning training process, this method optimizes the device scheduling strategy through reinforcement learning. The core task of reinforcement learning is to dynamically adjust the participation rate of device groups and the weight coefficients of optimization objectives to optimize the device scheduling objective function .

[0105] Reinforcement learning, also known as re-inforcement learning, evaluation learning, or enhanced learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an intelligent agent achieving maximization or a specific goal through learning strategies in the interaction process with the environment. The common model of reinforcement learning is the standard Markov decision process.

[0106] The modeling of reinforcement learning includes:

[0107] 1) Construct the state space:

[0108] The state S consists of the current state characteristics of the device and the global model performance indicators, including:

[0109] (8);

[0110] In formula (8): Indicates the current status characteristics of the device; They are global model accuracy, total device energy consumption, and fairness indicators, which serve as global performance feedback.

[0111] 2) Constructing the action space:

[0112] action Indicates the group participation rate and the weight coefficients of each optimization objective , , Adjustment value:

[0113] (9);

[0114] In formula (9): Indicates The adjustment value of the group device participation rate reflects the dynamic update of the scheduling priority of the group devices in the current round by the upper-layer strategy; They respectively represent the adjustment amounts of the weight coefficients corresponding to the improvement of global model accuracy, system energy consumption control, and device participation fairness in the device scheduling objective function.

[0115] 3) Constructing reward function:

[0116] Reward Value The increment of the equipment scheduling objective function Hooks, that is:

[0117] (10);

[0118] In formula (10): Represents the increment of the equipment scheduling objective function; Indicates the change in global model accuracy; Indicates the change in total energy consumption of the equipment; Indicates the change in fairness.

[0119] The execution process of reinforcement learning is as follows:

[0120] a. Status update:

[0121] The server collects the current status characteristics of each device and global model performance indicators 、 、 。

[0122] b. Policy optimization:

[0123] The policy is the policy function for the reinforcement learning agent to generate actions in the current state, which is used to guide how to adjust the grouping participation rate and the weight coefficients of each optimization objective . The reinforcement learning generates actions through the policy gradient method , adjusts the grouping participation rate and the weight coefficients of each optimization objective .

[0124] c. Execute actions:

[0125] Update the participation rate according to the action output and the weight coefficients of each optimization objective :

[0126] (11);

[0127] (12);

[0128] In formulas (11)-(12): represents the participation rate of the th group in the training round ; represents the participation rate of the th group in the training round ; represents the adjustment value of the participation rate of the th group generated by the current round policy; represents the weight coefficient of the th optimization objective in the training round , and , respectively corresponding to the weight coefficients representing the improvement of the global model accuracy, the control of system energy consumption, and the fairness of device participation; represents the learning rate; represents the gradient of the reward function with respect to the weight ; represents the temperature decay term, which gradually decreases as the training round increases. The purpose of temperature decay is to give a larger weight adjustment amplitude at the initial stage of optimization to quickly find a better solution; as the training progresses, gradually reduce the adjustment amplitude to make the weight update tend to be stable and avoid over-adjustment; the parameter controls the decay speed, and a larger value will make the decay faster and the weight adjustment converge faster; a smaller A value that results in slower attenuation, making the weight update more flexible.

[0129] d. Model training and feedback:

[0130] Use the updated participation rate and the weight coefficients of the optimization objective to complete device selection and training, and the reinforcement learning updates the scheduling policy according to the feedback.

[0131] In this embodiment, the devices participating in the federated learning training are determined through reinforcement learning, and the overall training process of the federated learning is as follows:

[0132] a. Global model initialization:

[0133] The server generates initial global model parameters , and sends them to all participating devices. The initial global model parameters can be generated by random initialization or preset parameters.

[0134] b. Device local model training:

[0135] Each device uses its local dataset to train the received global model , and the optimization objective is to minimize the local loss function:

[0136] (13);

[0137] In formula (13): represents the local loss function, such as cross-entropy loss; represents the model parameters obtained after local training of device .

[0138] The local dataset can be in the form of an image dataset, a text dataset, or a sensor sequence dataset, etc., and specifically varies according to different application tasks. For example, in an image classification task, the local dataset can be an image sample set collected at the device end, and the model can use a multi-layer perceptron (MLP) or a convolutional neural network (CNN) as the local model structure to identify the target category of the image sample. In a text sentiment analysis or instruction recognition task, the local dataset can be the local user corpus, and the model can use a lightweight RNN or a Transformer sub-network to extract features from the text and predict the sentiment label. The method of this embodiment can be widely applied to device environments with edge data perception and training capabilities, such as mobile terminals, intelligent cameras, in-vehicle terminals, etc.

[0139] The trained local model parameters are .

[0140] c. Upload the local model parameters :

[0141] The device uploads the trained local model parameters and the size of the local dataset to the server .

[0142] d. Global model aggregation:

[0143] The server aggregates the local model parameters of all devices into a new global model by weighted averaging: (14);

[0144] (14);

[0145] In Equation (14): represents the updated global model parameters; represents the total number of devices participating in the training; represents the total amount of data, that is, the sum of the local data sample numbers of all devices participating in this round of training; , respectively represent the local dataset sizes of device and device .

[0146] d. Iterative training

[0147] The aggregated global model parameters are distributed to each device again, and the above steps are repeated until the global model converges or reaches the preset number of training rounds

[0148] This method (Fed-RL) and the existing random selection device scheduling method (Random Selection) were compared in terms of performance in federated learning on the MNIST dataset, focusing on analyzing the differences in convergence speed and final model accuracy. The comparison results are as Figure 2 shown. Within the first 30 rounds, the accuracy of Fed-RL is significantly higher than that of Random Selection. Within 60 rounds, Fed-RL has approached 99% accuracy, while Random Selection is still below 96%. Fed-RL finally converges to nearly 99%, which is more stable and efficient than Random Selection. Due to the random device selection strategy, Random Selection has low training efficiency, with a final accuracy below 97% and slow convergence

[0149] In summary, through the designed double-layer reinforcement learning mechanism, the present invention improves the global model training performance of federated learning while taking into account the participation unfairness, and at the same time reduces the participation energy consumption of devices. In practice, this method can be widely applied to multi-device distributed scenarios such as the Internet of Things, edge computing, and intelligent terminals, and is particularly suitable for environments where device resources are highly heterogeneous and dynamically changing. The present invention not only provides an effective solution to the device scheduling optimization problem of federated learning, but also provides a new research idea and technical implementation solution for multi-objective optimization and dynamic resource scheduling problems.

[0150] Embodiment 2

[0151] This embodiment provides a device for implementing a method for optimizing the device scheduling of federated learning based on double-layer reinforcement learning. The device includes:

[0152] A data acquisition module, configured to acquire the current state characteristics of a device , where the current state characteristics of the device include computing power , battery power and network latency ;

[0153] An upper-layer grouping module, configured to divide devices into groups through a clustering algorithm, and allocate participation rates to each group of devices through upper-layer reinforcement learning , including: constructing the state of upper-layer reinforcement learning based on the current state characteristics of each device in each group of devices , the change in global model accuracy, and the total device energy consumption, allocating participation rates to each group of devices based on this state , and determining whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-layer reward function ;

[0154] A lower-layer device selection module, configured to select devices participating in federated learning within each group through lower-layer reinforcement learning, including: constructing the state of lower-layer reinforcement learning based on the current state characteristics of each device in each group of devices , the change in the local model accuracy of each device, and the device energy consumption, calculating a scoring mechanism for each device based on this state to select devices participating in federated learning within each group, and measuring the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function;

[0155] A federated learning training module, configured to construct a device scheduling objective function and initialize global model parameters , and perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rates of each group of devices and the weight coefficients of the optimization objectives Maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein the optimization objectives include improving the global model accuracy, controlling the system energy consumption, and ensuring fairness in device participation.

[0156] Embodiment 3

[0157] This embodiment also provides an electronic device, including: at least one processor; and a memory that stores instructions, which when executed by the at least one processor, cause the at least one processor to execute the federated learning device scheduling optimization method based on double-layer reinforcement learning as described above.

[0158] In this embodiment, the electronic device may include, but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and so on.

[0159] Embodiment 4

[0160] This embodiment also provides a machine-readable storage medium that stores executable instructions, which when executed cause the machine to execute the federated learning device scheduling optimization method based on double-layer reinforcement learning as described above.

[0161] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device is caused to read and execute the instructions stored in the readable storage medium.

[0162] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0163] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD ROM, CD R, CD RW, DVDROM, DVD RAM, DVD RW, DVD RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or the cloud via a communication network.

[0164] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD ROM, optical storage, etc.) that contain computer-usable program code.

[0165] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0168] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for optimizing the device scheduling of federated learning based on double-layer reinforcement learning, characterized in that, The method includes: S1. Obtain the current state characteristics of the device , the current state characteristics of the device include computing power , battery power and network latency ; S2. Divide the devices into groups, and assign participation rates to each group of devices through upper-layer reinforcement learning , including: constructing the state of upper-layer reinforcement learning based on the current state features of each device in each group of devices as well as the change in the global model accuracy and the total energy consumption of the devices, and assigning participation rates to each group of devices based on this state , and judging whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-layer reward function ; S3. Select the devices participating in federated learning within each group through lower-layer reinforcement learning, including: based on the current state characteristics of each device in each group of devices and the local model accuracy changes and device energy consumption of each device to construct the state of lower-layer reinforcement learning, calculate the scoring mechanism for each device based on this state to select the devices participating in federated learning within each group, and measure the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function; S4. Construct the device scheduling objective function and initialize the global model parameters , and perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rate of each group of devices and the weight coefficient of the optimization objective maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein, the optimization objectives include global model accuracy improvement, system energy consumption control, and device participation fairness.

2. The federated learning device scheduling optimization method based on double-layer reinforcement learning according to claim 1, characterized in that In S2, based on the current state characteristics of each device in each group of devices and the global model accuracy change and the total device energy consumption, construct the state of the upper-layer reinforcement learning, and based on this state, allocate the participation rate for each group of devices , specifically, the upper-layer reinforcement learning passes through the policy function to map the comprehensive state of each group of devices to the participation rate , that is: (1); Among them, the comprehensive status of each group of devices includes the computing power of the devices within the group , battery power , network latency , as well as the current change in the global model accuracy and the total energy consumption of the devices .

3. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, wherein, In S2, based on the change in the global model accuracy , the total energy consumption of the device and the data quality construct an upper-layer reward function, that is: (2); In formula (2): represents the upper-layer reward function; represents the change in the global model accuracy; represents the total energy consumption of the selected device; represents the scoring weight corresponding to the total energy consumption of the device; represents the data quality coefficient of the nth group of devices.

4. The method for optimizing the device scheduling of federated learning based on double-layer reinforcement learning according to claim 1, characterized in that, In S3, the scoring mechanism is: (3); In formula (3): represents the device in the training round scoring result; represents the maximum value of the device computing power in the system; represents the full charge state value of the device; and and and are the scoring weight coefficients of the device computing performance, energy consumption status, communication quality, and fair participation respectively; is a time decay function, representing a time decay term set considering the historical participation of the device, and: (4); In formula (4): represents the device in the training round participation frequency; represents the time decay coefficient, used to control the decay speed; represents the current training round distance from the device the time interval since the last participation in training.

5. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, characterized in that, In S3, the lower-level reward function is: (5); In formula (5): represents the lower-level reward function; represents the change in the accuracy of the local model in the training round ; represents the network stability of the device ; represents the device in the training round in the energy consumption; represents the scoring weight corresponding to the change in the accuracy of the local model; represents the scoring weight corresponding to the energy consumption of the device.

6. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, wherein, In S4, the constructed device scheduling objective function is: (6); In formula (6): represents the device scheduling objective function; represents the training round and is the increment of the global model accuracy in it; represents the total energy consumption of the selected devices in the training round ; represents the total number of training rounds; represents the weight coefficient for the improvement of the global model accuracy, represents the weight coefficient for the system energy consumption control, and represents the weight coefficient for the fairness of device participation; is the fairness index, and: (7); In formula (7): represents the total number of devices selected to participate in the training; is a positive number used to prevent division-by-zero errors; represents the average of the participation frequencies of all devices.

7. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 6, wherein In S4, the participation rate of each group of devices and the weight coefficient of the optimization objective are updated based on the following formula: (11); (12); In formulas (11)-(12): represents the participation rate of the th group in the training round ; represents the participation rate of the th group in the training round ; represents the adjustment value of the participation rate of the th group generated by the current round strategy; represents the th weight coefficient of the optimization objective in the training round , and , corresponding to the weight coefficients representing the improvement of the global model accuracy, the control of system energy consumption, and the fairness of device participation respectively; represents the learning rate; represents the gradient of the reward function with respect to the weight ; represents the temperature decay term, which gradually decreases as the training round increases; the parameter is used to control the decay speed.

8. An apparatus for implementing an optimization method for device scheduling in federated learning based on double-layer reinforcement learning, characterized in that, The device includes: A data acquisition module for acquiring the current state characteristics of a device , the current state characteristics of the device including computing power , battery power and network latency ; The upper-level grouping module is used to divide devices into groups through a clustering algorithm and assign participation rates to each group of devices through upper-level reinforcement learning , including: constructing the state of upper-level reinforcement learning based on the current state characteristics of each device in each group of devices , the change in the global model accuracy, and the total energy consumption of the devices, and assigning participation rates to each group of devices based on this state , and judging whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-level reward function ; The lower-layer device selection module is used to select the devices participating in federated learning within each group through lower-layer reinforcement learning, including: based on the current state characteristics of each device in each group of devices and the local model accuracy change and device energy consumption of each device to construct the state of lower-layer reinforcement learning, calculate the scoring mechanism of each device based on this state to select the devices participating in federated learning within each group, and measure the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function; Federated learning training module, used to construct an objective function for device scheduling and initialize global model parameters , perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rate of each group of devices and the weight coefficient of the optimization objective maximize the objective function for device scheduling to determine the optimal device scheduling strategy; wherein, the optimization objectives include improvement of global model accuracy, system energy consumption control, and fairness of device participation.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory, the memory stores instructions, when the instructions are executed by the at least one processor, the at least one processor is caused to execute the federated learning device scheduling optimization method based on double-layer reinforcement learning according to any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that, Executable instructions are stored on the machine-readable storage medium, and when the instructions are executed, the machine is caused to execute the federated learning device scheduling optimization method based on double-layer reinforcement learning according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Equipment scheduling method in multi-task federated learning system

    CN115033382A

  • Federal learning terminal selection and resource scheduling method based on dynamic priority

    CN119299398A

  • Hierarchical user training management system and method oriented to non-independent identically distributed data

    CN113672684A

  • Method and system for federated learning, electronic device, and computer readable medium

    US20220374776A1