Federal learning equipment scheduling optimization method and device based on double-layer reinforcement learning

By adopting the equipment scheduling optimization method of double-layer reinforcement learning in federated learning, the problem of unreasonable resource scheduling in the heterogeneous environment of equipment is solved, and comprehensive optimization of global model performance improvement, energy consumption control and fairness of equipment participation is achieved.

CN120068994AActive Publication Date: 2025-05-30QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Patent Information

Application Number
CN202510512002.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-30
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the heterogeneous environment of equipment, traditional federated learning scheduling algorithms fail to effectively utilize the computing power of high-performance devices, resulting in low-performance devices becoming a system bottleneck, affecting the training efficiency and model performance of the global model, and at the same time there are problems of uneven energy consumption and unfair equipment participation.

Method used

The equipment scheduling optimization method based on double-layer reinforcement learning is adopted, and the equipment is grouped and the participation rate is allocated through the upper-layer reinforcement learning. The equipment scheduling strategy is optimized to improve the performance of the global model, control energy consumption and improve the fairness of equipment participation.

Benefits of technology

It realizes efficient resource scheduling in heterogeneous equipment environments, improves global model performance, reduces equipment energy consumption, and improves the fairness of equipment participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068994A_ABST
    Figure CN120068994A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of federated learning, and particularly relates to a federated learning equipment scheduling optimization method and device based on double-layer reinforcement learning, and the method comprises the steps: obtaining the current state characteristics of equipment; the devices are divided into # imgabs0 # groups, and the participation rate is distributed for each group of devices through upper-layer reinforcement learning; equipment participating in federated learning in each group is selected through lower-layer reinforcement learning; and constructing an equipment scheduling objective function, initializing global model parameters, carrying out federated learning training on the equipment selected based on lower-layer reinforcement learning, and in the training process, maximizing the equipment scheduling objective function by adjusting the participation rate of each group of equipment and the score weight of an optimization objective so as to determine an optimal scheduling strategy of the equipment. According to the method, federated learning equipment scheduling is optimized by using a double-layer reinforcement learning strategy, so that the global model performance is improved, the equipment energy consumption is reduced, and the equipment participation fairness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of federated learning, and more specifically, relates to a method and device for optimizing device scheduling in federated learning based on double-layer reinforcement learning. Background Art

[0002] Traditional centralized training methods not only bring huge communication overhead but also have a serious risk of data privacy leakage because a large amount of raw data needs to be uploaded to the server for unified training. To solve such problems, federated learning, as a distributed machine learning method that can protect data privacy, has emerged.

[0003] Federated learning allows devices to complete model training locally and achieve global aggregation by uploading model parameters instead of raw data, which not only effectively reduces the risk of data privacy leakage but also significantly reduces the bandwidth requirements for data transmission. Due to its unique advantages in privacy protection and communication efficiency, federated learning has been widely applied in scenarios such as the Internet of Things, mobile terminals, and edge computing in recent years.

[0004] However, in an environment where there are significant differences in the resource states of devices such as computing power, battery power, and network latency, how to design an efficient resource scheduling strategy for device heterogeneity to make full use of the computing power of high-performance devices while avoiding low-performance devices from becoming a system bottleneck, thereby improving the global training efficiency and model performance, has become a key problem restricting the development of federated learning technology.

[0005] Traditional federated learning scheduling algorithms usually adopt a unified training strategy and do not design an optimization mechanism for device heterogeneity and dynamics. This leads to the following problems: on the one hand, the computing power and network bandwidth of high-performance devices are often not fully utilized, while low-performance devices may become a system bottleneck, dragging down the training efficiency of the global model. On the other hand, there is a lack of pertinence in energy consumption control, and some devices may be forced to withdraw from training due to excessive power consumption or insufficient battery power, affecting the convergence stability of the global model.

[0006] For example, Chinese patent document CN115033382A discloses a device scheduling method in a multi-task federated learning system, including: constructing a system model of multi-task federated learning; establishing an optimization problem with the goal of minimizing the time of the multi-task federated learning process; scheduling devices to participate in the training process of the federated learning task; transforming the device scheduling process into a multi-armed bandit and matching process; designing a device scheduling algorithm. The present invention schedules the most suitable device for each task in federated learning, thereby minimizing the latency of the multi-task federated learning process.

[0007] In addition, the unfairness problem of device scheduling cannot be ignored. In a heterogeneous environment, devices that participate frequently dominate model training. However, due to insufficient participation, the local data of devices with low participation cannot be effectively learned by the global model, resulting in a decline in the generalization performance of the model. More importantly, device states usually change dynamically over time, such as battery power consumption and network condition fluctuations. Existing scheduling strategies lack flexibility and are difficult to adapt to the dynamically changing device environment, further limiting the practical application of federated learning.

[0008] For example, Chinese patent document CN119299398A discloses a method for federated learning terminal selection and resource scheduling based on dynamic priority, including: obtaining the basic model structure and resource information parameters exchanged between the federated learning central controller and users; calculating the priority value of each user according to the resource information parameters, and sorting the users from largest to smallest according to the priority value; based on the user priority order, selecting a subset of users that meet the training delay requirements; for the subset of users, calculating the minimum bandwidth allocation ratio required to complete the training task, considering the data volume, computing power, and network conditions of the users; using the binary search method to allocate bandwidth resources to the subset of users, and the binary search method adjusts the bandwidth allocation through iteration; determining whether the total energy consumption of the subset of users exceeds a given energy constraint threshold, and if it exceeds the threshold, deleting the user with the lowest priority from the subset of users.

[0009] Based on this, the present invention designs a federated learning device scheduling optimization method based on double-layer reinforcement learning, aiming to solve the problems of the decline in the performance of the global model, uneven energy consumption distribution, and insufficient fairness of device participation caused by unreasonable device resource scheduling during the federated learning training process in a heterogeneous device environment. Summary of the Invention

[0010] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a federated learning device scheduling optimization method based on double-layer reinforcement learning, which is used to optimize the federated learning process in a heterogeneous device environment and achieve comprehensive optimization of the performance of the global model, energy consumption efficiency, and fairness of device participation.

[0011] The present invention also discloses a device loaded with the federated learning device scheduling optimization method based on double-layer reinforcement learning.

[0012] The detailed technical solution of the present invention is as follows: A federated learning device scheduling optimization method based on double-layer reinforcement learning, the method comprising: S1. Obtain the current state characteristics of the device , the current state characteristics of the device including computing power , battery power and network latency ; S2. Divide the devices into groups, and allocate participation rates to each group of devices through upper-layer reinforcement learning , including: based on the current state characteristics of each device in each group of devices and the change in the global model accuracy and the total energy consumption of the devices, construct the state of the upper-layer reinforcement learning, and allocate participation rates to each group of devices based on this state , and determine whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-layer reward function ; S3. Select the devices participating in federated learning within each group through lower-layer reinforcement learning, including: based on the current state characteristics of each device in each group of devices and the change in the local model accuracy of each device and the energy consumption of the device, construct the state of the lower-layer reinforcement learning, calculate the scoring mechanism of each device based on this state to select the devices participating in federated learning within each group, and measure the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function; S4. Construct a device scheduling objective function and initialize the global model parameters , and perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rates of each group of devices and the weight coefficients of the optimization objectives maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein, the optimization objectives include the improvement of the global model accuracy, the control of the system energy consumption, and the fairness of device participation.

[0013] Preferably according to the present invention, in S2, based on the current state characteristics of each device in each group of devices and the change in the global model accuracy and the total energy consumption of the devices, construct the state of the upper-layer reinforcement learning, and allocate participation rates to each group of devices based on this state , specifically, the upper-layer reinforcement learning maps the comprehensive state of each group of devices to the participation rate through the policy function , that is: (1); (1); wherein, the comprehensive state of each group of devices includes the computing power of the devices in the group , battery power , network latency , and the current change in the global model accuracy and the total energy consumption of the devices .

[0014] Preferably according to the present invention, in S2, based on the change in the global model accuracy , the total energy consumption of the devices and data quality Construct the upper-layer reward function, that is: (2); In formula (2): represents the upper-layer reward function; represents the change in the global model accuracy; represents the total energy consumption of the selected device; represents the scoring weight corresponding to the total energy consumption of the device; represents the data quality coefficient of the

[0015] According to the preference of the present invention, in S3, the scoring mechanism is: (3); In formula (3): represents the device in the training round scoring result; represents the maximum value of the computing power of the devices in the system; represents the full charge state value of the device; , , , are the scoring weight coefficients of the computing performance, energy consumption status, communication quality, and fair participation of the device respectively; is a time decay function, representing a time decay term set considering the historical participation of the device, and: (4); In formula (4): represents the device in the training round participation frequency; represents the time decay coefficient, used to control the decay speed; represents the time interval between the current training round and the device last participation in training.

[0016] According to the preference of the present invention, in S3, the lower-layer reward function is: (5); In formula (5): represents the lower-layer reward function; represents the change in the accuracy of the local model in the training round , represents the device network stability; represents the device in the training round The energy consumption in represents the scoring weight corresponding to the change in the local model accuracy; represents the scoring weight corresponding to the device energy consumption.

[0017] Preferably according to the present invention, in the S4, the constructed device scheduling objective function is: (6); In formula (6): represents the device scheduling objective function; represents the training round the increment of the global model accuracy in; represents in the training round the total energy consumption of the selected devices in; represents the total number of training rounds; represents the weight coefficient for the improvement of the global model accuracy, represents the weight coefficient for the system energy consumption control, represents the weight coefficient for the fairness of device participation; is the fairness index, and: (7); In formula (7): represents the total number of devices selected to participate in the training; is a positive number used to prevent division by zero error; represents the mean of the participation frequencies of all devices.

[0018] Preferably according to the present invention, in the S4, the participation rate of each group of devices and the weight coefficient of the optimization objective are updated based on the following formula: (11); (12); In formulas (11)-(12): represents the th group in the training round the participation rate in; represents the th group in the training round the participation rate in; represents the adjustment value of the participation rate of the th group generated by the current round strategy; represents the th optimization objective in the training round the weight coefficient in, and respectively correspond to the weight coefficients representing the improvement of the global model accuracy, the system energy consumption control, and the fairness of device participation; represents the learning rate; Denote the gradient of the reward function with respect to the weights ; Denote the temperature decay term, which gradually decreases as the number of training rounds increases; the parameter is used to control the decay rate.

[0019] In another aspect of the present invention, there is provided an apparatus for implementing an optimization method for device scheduling in federated learning based on double-layer reinforcement learning, the apparatus comprising: A data acquisition module, configured to acquire the current state features of the devices , where the current state features of the devices include computing power , battery power and network latency ; An upper-layer grouping module, configured to divide the devices into groups through a clustering algorithm, and allocate participation rates to each group of devices through upper-layer reinforcement learning , including: constructing the state of upper-layer reinforcement learning based on the current state features of each device in each group of devices as well as the change in global model accuracy and the total energy consumption of the devices, allocating participation rates to each group of devices based on this state , and determining whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-layer reward function ; A lower-layer device selection module, configured to select the devices participating in federated learning within each group through lower-layer reinforcement learning, including: constructing the state of lower-layer reinforcement learning based on the current state features of each device in each group of devices as well as the change in the local model accuracy of each device and the energy consumption of the devices, calculating a scoring mechanism for each device based on this state to select the devices participating in federated learning within each group, and measuring the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function; A federated learning training module, configured to construct a device scheduling objective function and initialize the global model parameters , and perform federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, maximize the device scheduling objective function by adjusting the participation rates of each group of devices and the weight coefficients of the optimization objectives to determine the optimal device scheduling strategy; wherein, the optimization objectives include improving the global model accuracy, controlling the system energy consumption, and ensuring fairness in device participation.

[0020] In another aspect of the present invention, there is also provided an electronic device, comprising: At least one processor; and A memory that stores instructions which, when executed by the at least one processor, cause the at least one processor to execute the method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning as described above.

[0021] In another aspect of the present invention, there is also provided a machine-readable storage medium storing executable instructions which, when executed, cause the machine to execute the method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning as described above.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning provided by the present invention, on the basis of considering issues such as device heterogeneity including computing power, battery power, and network latency, uses a double-layer reinforcement learning strategy to optimize the scheduling of federated learning devices, aiming to improve the global model performance, reduce device energy consumption, and enhance device participation fairness.

[0023] (2) Through the double-layer design of the upper-layer grouping strategy and the lower-layer device selection strategy, combined with the dynamic weight adjustment mechanism, the present invention can achieve the efficiency and robustness of resource scheduling in a heterogeneous device environment. Among them, the upper-layer strategy is responsible for grouping devices according to device status information and assigning an appropriate participation rate to each group of devices; the lower-layer strategy selects suitable devices for participating in training within the group according to the device scoring mechanism.

[0024] (3) The present invention comprehensively considers the multi-objective optimization requirements in the federated learning process, balances the global model accuracy, device energy consumption, and participation fairness, and can adapt to the dynamic changes of device status. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flowchart of the method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to the present invention.

[0026] Figure 2 is a comparison test chart of the method Fed-RL of the present invention and the random selection device scheduling method Random Selection on the MNIST dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The present invention will be further described below in conjunction with the drawings and embodiments.

[0028] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0029] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0030] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0031] Aiming at the deficiencies in the prior art, the present invention proposes a method for optimizing the device scheduling of federated learning based on double-layer reinforcement learning. By introducing a double-layer design of an upper-layer grouping strategy and a lower-layer device selection strategy, and combining the reinforcement learning mechanism to dynamically adjust the device participation rate and scoring weights, the core problem of resource scheduling in a heterogeneous device environment is effectively solved.

[0032] In the upper-layer grouping strategy, devices are grouped according to state features such as computing power, battery power, and network latency, reducing intra-group heterogeneity and improving the efficiency of grouped resource allocation; the lower-layer device selection strategy selects the devices suitable for the current training round in each group based on a scoring mechanism, ensuring the accuracy of device selection and the effectiveness of training.

[0033] At the same time, by dynamically adjusting the device participation rate and the weight coefficients of each optimization objective through reinforcement learning, the present invention realizes the multi-objective comprehensive optimization of model performance, energy consumption, and participation fairness.

[0034] The following further illustrates the method and device for optimizing the device scheduling of federated learning based on double-layer reinforcement learning of the present invention in combination with specific embodiments.

[0035] Embodiment 1 Refer Figure 1 , this embodiment provides a method for optimizing the device scheduling of federated learning based on double-layer reinforcement learning, which is applied to a federated learning system. The system includes a server and multiple devices participating in federated learning, and the server is communicatively connected to each device. The method includes: S1. Obtain the current state features of the device , the current state features of the device include computing power , battery power and network latency .

[0036] The method of this embodiment models the device scheduling problem of federated learning in a heterogeneous device environment. Assume that there are Devices may participate in federated learning, including those selected to participate in federated learning and those not selected. Each device 's state can be represented as a feature vector , , , where represents the computing power of the device, such as the performance of the CPU or GPU; represents the remaining battery power of the device, ranging from [0, 1]; represents the network latency of the device, which is used to reflect the communication performance between the device and the server.

[0037] The currently obtained state features of the device will be used as the state space data in subsequent reinforcement learning steps.

[0038] S2. Divide the devices into groups, and assign participation rates to each group of devices through upper-level reinforcement learning , including: based on the current state features of each device in each group and the change in global model accuracy and the total energy consumption of the devices, construct the state of the upper-level reinforcement learning, and assign participation rates to each group of devices based on this state , and judge whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-level reward function .

[0039] In this embodiment, first, the devices are divided into groups through a clustering algorithm. The clustering algorithm is K-Means clustering, which takes the state feature vector of the device as the input. To eliminate the scale influence between different metrics, the features of each dimension are standardized before clustering. The K-Means algorithm minimizes the sum of the squared Euclidean distances of the state differences of the devices within the group, and finally obtains subsets of devices , and the devices within each group have high similarity in terms of computing power, energy consumption status, and communication conditions, which is beneficial to improving the effectiveness of the participation rate allocation strategy and the robustness of the scheduling in subsequent upper-level reinforcement learning.

[0040] Subsequently, the participation rates are assigned to each group of devices through upper-level reinforcement learning.

[0041] The purpose of the upper-level reinforcement learning is to group the devices based on their resource characteristics, so as to reduce the scheduling complexity caused by device heterogeneity; at the same time, set the participation rates of different groups, so that the global model can achieve accuracy improvement under limited system resources.

[0042] Specifically, the participation rate is allocated to each group of devices through upper-layer reinforcement learning , including: Construct the state space: Based on the current state characteristics of each device in each group of devices , including the computing power of the device , battery power , network latency , and the change in global model accuracy , total device energy consumption Construct the state of the upper-layer reinforcement learning, where represents the training round; represents the change in the global model accuracy in the current training round , defined as , that is, the difference in the global model accuracy between two consecutive training rounds.

[0043] Construct the action space: Allocate an appropriate participation rate to each group of devices based on the above state .

[0044] Based on the above state information, the upper-layer reinforcement learning maps the comprehensive state of each group of devices to the participation rate through the policy function , where the comprehensive state of each group of devices includes the computing power of the devices within the group , battery power , network latency , and the current change in global model accuracy and total device energy consumption .

[0045] This policy function can be implemented in the form of a neural network, and the parameters are optimized through training to dynamically adjust in each round of training to obtain a higher cumulative reward. Specifically: (1).

[0046] The policy training objective is to maximize the cumulative reward function, which comprehensively considers factors such as the improvement of the global model accuracy, total device energy consumption, and data quality. The policy parameters are continuously optimized through the policy gradient method to achieve intelligent dynamic allocation of the participation rate.

[0047] Construct the reward function: The upper-layer reward function is constructed based on the change in global model accuracy , total device energy consumption and data quality , specifically: (2); In formula (2): represents the upper-level reward function; represents the change in the global model accuracy; represents the total energy consumption of the selected devices; represents the scoring weight corresponding to the total energy consumption of the devices; represents the data quality coefficient of the nth group of devices, which is used to preferentially select groups with high data quality to enhance the model performance. This index can be comprehensively evaluated by combining factors such as the sample size of the data within the group, the conformity of the data with the task objective distribution (i.e., distribution representativeness), and annotation integrity.

[0048] Judge whether it is necessary to optimize and adjust the participation rate of each group of devices according to the feedback of the upper-level reward function , to ensure the effectiveness of the grouping and participation rate selection.

[0049] S3. Select the devices participating in federated learning within each group through lower-level reinforcement learning, including: based on the current state characteristics of each device in each group of devices and the change in the local model accuracy of each device and the device energy consumption, construct the state of the lower-level reinforcement learning. Based on this state, calculate the scoring mechanism for each device to select the devices participating in federated learning within each group, and measure the actual effect of the current device selection strategy based on the feedback of the lower-level reward function.

[0050] In this embodiment, the purpose of the lower-level reinforcement learning is to select suitable devices for participating in federated learning training within the group to achieve the resource effectiveness and fairness of the scheduling.

[0051] Specifically, select the devices participating in federated learning within each group through lower-level reinforcement learning, including: Construct the state space: based on the current state characteristics of each device in each group of devices , including the computing power of the device , battery power , network latency , and the change in the local model accuracy of each device , device energy consumption construct the state of the lower-level reinforcement learning.

[0052] Construct the action space: calculate the scoring mechanism for each device based on the above state to select the devices participating in federated learning within each group; the scoring mechanism is: (3); In formula (3): represents the device in the training round n; Represents the maximum computing power of the devices in the system; Represents the full charge state value of the device; , , , Are the scoring weight coefficients for the computing performance, energy consumption status, communication quality, and fair participation of the device respectively, used to adjust the influence of the scores of the above factors; Is a time decay function, representing the time decay term set considering the historical participation of the device, and: (4); In formula (4): Represents the device In the training round The participation frequency, that is, the frequency at which the device is selected in the current training process; Represents the time decay coefficient, controlling the decay speed; Represents the time interval between the current training round and the device Last participation in training.

[0053] Construct the reward function: The lower-layer reward function comprehensively considers the contribution of the device to the local model accuracy, energy consumption, and network stability. Specifically: (5); In formula (5): Represents the lower-layer reward function; Represents the accuracy change of the local model in the training round ; Represents the device Network stability; Represents the device In the training round Energy consumption; Represents the scoring weight corresponding to the local model accuracy change; Represents the scoring weight corresponding to the device energy consumption.

[0054] Based on the feedback of the lower-layer reward function to measure the actual effect of the current device selection strategy. The lower-layer reward function calculates the performance of the device under the strategy execution based on the local accuracy change and energy consumption of the device in each round of training, and provides a value feedback signal for policy learning. This feedback can be used for the optimization and update of the lower-layer reinforcement learning strategy to ensure that the system selects more cost-effective devices to participate in training in subsequent rounds, thereby improving the overall performance and resource utilization efficiency of the global model.

[0055] S4. Construct the device scheduling objective function and initialize the global model parameters , conduct federated learning training based on the devices selected by the underlying reinforcement learning. During the training process, by adjusting the participation rate of each group of devices and the weight coefficients of the optimization objectives maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein, the optimization objectives include global model accuracy improvement, system energy consumption control, and device participation fairness.

[0056] The performance of the global model of federated learning is significantly affected by the heterogeneity of device states. Insufficient device performance may drag down the global model training, while overusing some devices will lead to resource waste and increased system energy consumption. Therefore, a reasonable device scheduling strategy is needed to achieve a balance among performance, energy consumption, and fairness.

[0057] In this embodiment, the device scheduling problem is modeled as a multi-objective optimization problem, and the constructed device scheduling objective function is: (6); In formula (6): represents the device scheduling objective function; represents the increment of the global model accuracy in the training round; represents the total energy consumption of the selected devices in the training round ; represents the total number of training rounds; represents the weight coefficient for global model accuracy improvement, represents the weight coefficient for system energy consumption control, represents the weight coefficient for device participation fairness; the fairness index : inverse frequency weighting, which is used to prevent resource-rich devices from frequently dominating the training and make the system tend to pay more attention to low-frequency devices, that is, to measure whether the data of low-frequency participating devices is fully utilized, and: (7); In formula (7): represents the total number of selected devices participating in the training; is a small positive number used to prevent division by zero errors and ensure numerical stability; represents the mean of the participation frequencies of all devices.

[0058] This method aims to achieve the maximization of the device scheduling objective function by dynamically adjusting the grouped participation rate of devices and the weight coefficients of each optimization objective in device selection , , in order to find an optimal device scheduling strategy that achieves an overall optimal balance in terms of comprehensive performance, resources, and fairness. Specifically, during the federated learning training process, this method optimizes the device scheduling strategy through reinforcement learning. The core task of reinforcement learning is to dynamically adjust the participation rate of device groups. and the weight coefficient of the optimization objective , to optimize the equipment scheduling objective function .

[0059] Reinforcement learning, also known as reinforcement learning, evaluation learning or enhanced learning, is one of the paradigms and methodologies of machine learning. It is used to describe and solve the problem of maximizing or achieving specific goals through learning strategies in the process of interaction between intelligent agents and the environment. The common model of reinforcement learning is the standard Markov decision process.

[0060] Reinforcement learning modeling includes: 1) Construct state space: The state S consists of the current state characteristics of the device and the global model performance indicators, including: (8); In formula (8): Indicates the current status characteristics of the device; They are global model accuracy, total device energy consumption, and fairness indicators, which serve as global performance feedback.

[0061] 2) Constructing the action space: action Indicates the group participation rate and the weight coefficients of each optimization objective , , Adjustment value: (9); In formula (9): Indicates The adjustment value of the group device participation rate reflects the dynamic update of the scheduling priority of the group devices in the current round by the upper-layer strategy; They respectively represent the adjustment amounts of the weight coefficients corresponding to the improvement of global model accuracy, system energy consumption control, and device participation fairness in the device scheduling objective function.

[0062] 3) Constructing reward function: Reward Value The increment of the equipment scheduling objective function Hooks, that is: (10); In formula (10): represents the increment of the device scheduling objective function; represents the change in the global model accuracy; represents the change in the total device energy consumption; represents the change in fairness.

[0063] The execution process of reinforcement learning is as follows: a. State update: The server collects the current state features of each device and the global model performance metrics 、 、 .

[0064] b. Policy optimization: The policy is the policy function for the reinforcement learning agent to generate actions in the current state, which is used to guide how to adjust the grouping participation rate and the weight coefficients of each optimization objective . Reinforcement learning generates actions through the policy gradient method , adjusts the grouping participation rate and the weight coefficients of each optimization objective .

[0065] c. Execute actions: According to the action output, update the participation rate and the weight coefficients of each optimization objective : (11); (12); In equations (11)-(12): represents the th group's participation rate in the training round represents the th group's participation rate in the training round represents the adjustment value of the th group's participation rate generated by the current round's policy; represents the weight coefficient of the th optimization objective in the training round, and represents the learning rate; represents the gradient of the reward function with respect to the weight ; Represents the temperature decay term, which gradually decreases as the number of training rounds increases. The purpose of temperature decay is to give a larger weight adjustment amplitude in the initial stage of optimization to quickly find a better solution; as the training progresses, the adjustment amplitude is gradually reduced to make the weight update tend to be stable and avoid over-adjustment; the parameter controls the decay speed. A larger value will make the decay faster and the weight adjustment converge faster; a smaller value will result in a slower decay and make the weight update more flexible.

[0066] d. Model training and feedback: Use the updated participation rate and the weight coefficients of the optimization objective to complete device selection and training, and the reinforcement learning updates the scheduling policy according to the feedback.

[0067] In this embodiment, the devices participating in the federated learning training are determined through reinforcement learning, and the overall training process of the federated learning is as follows: a. Global model initialization: The server generates initial global model parameters and sends them to all participating devices. The initial global model parameters can be generated by random initialization or preset parameters.

[0068] b. Device local model training: Each device uses its local dataset to train the received global model with the optimization objective of minimizing the local loss function: (13); In formula (13): represents the local loss function, such as cross-entropy loss; represents the model parameters obtained after local training of device .

[0069] The local dataset can be in the form of an image dataset, a text dataset, or a sensor sequence dataset, etc., which specifically varies according to different application tasks. For example, in an image classification task, the local dataset can be an image sample set collected at the device end, and the model can use a multi-layer perceptron (MLP) or a convolutional neural network (CNN) as the local model structure to identify the target category of the image samples. In a text sentiment analysis or instruction recognition task, the local dataset It can be local user corpus, and the model can use lightweight RNN or Transformer sub-network to extract features and predict sentiment labels for text. The method of this embodiment can be widely applied to device environments with edge data perception and training capabilities, such as mobile terminals, smart cameras, and vehicle-mounted terminals.

[0070] The parameters of the trained local model are .

[0071] c. Upload local model parameters : The device will train the local model parameters and local datasets Size Upload to the server.

[0072] d. Global model aggregation: The server uses weighted averaging to calculate the local model parameters of all devices. Aggregate into a new global model: (14); In formula (14): represents the updated global model parameters; Indicates the total number of devices participating in the training; Indicates the total amount of data, that is, the sum of the number of local data samples of all devices participating in this round of training; , Respectively represent devices and equipment The size of the local dataset.

[0073] d. Iterative training The aggregated global model parameters Distribute to each device again and repeat the above steps until the global model converges or reaches the preset training rounds.

[0074] The performance of this method (Fed-RL) and the existing random selection device scheduling method (Random Selection) in federated learning was compared on the MNIST dataset, focusing on analyzing the differences in convergence speed and final model accuracy. The comparison results are shown in Figure 2. Figure 2As shown. Within the first 30 rounds, the accuracy of Fed-RL is significantly higher than that of Random Selection. Within 60 rounds, Fed-RL has approached an accuracy of nearly 99%, while Random Selection is still below 96%. Fed-RL finally converges to nearly 99%, which is more stable and efficient than Random Selection. Due to the random device selection strategy, Random Selection has low training efficiency, with a final accuracy below 97% and slow convergence.

[0075] In summary, through the designed double-layer reinforcement learning mechanism, the present invention improves the global model training performance of federated learning while taking into account participation unfairness, and at the same time reduces the participation energy consumption of devices. In practice, this method can be widely applied to multi-device distributed scenarios such as the Internet of Things, edge computing, and intelligent terminals, especially suitable for environments with highly heterogeneous and dynamically changing device resources. The present invention not only provides an effective solution to the device scheduling optimization problem of federated learning, but also provides new research ideas and technical implementation solutions for multi-objective optimization and dynamic resource scheduling problems.

[0076] Embodiment 2 This embodiment provides a device for implementing a method for optimizing federated learning device scheduling based on double-layer reinforcement learning. The device includes: A data acquisition module for acquiring the current state characteristics of devices , where the current state characteristics of the devices include computing power , battery power and network latency ; An upper-layer grouping module for dividing devices into groups through a clustering algorithm and allocating participation rates for each group of devices through upper-layer reinforcement learning , including: constructing the state of upper-layer reinforcement learning based on the current state characteristics of each device in each group of devices , the change in global model accuracy, and the total device energy consumption, allocating participation rates for each group of devices based on this state , and judging whether to optimize and adjust the participation rates of each group of devices based on the feedback of the upper-layer reward function ; A lower-layer device selection module for selecting devices participating in federated learning within each group through lower-layer reinforcement learning, including: constructing the state of lower-layer reinforcement learning based on the current state characteristics of each device in each group of devices , the change in the local model accuracy of each device, and the device energy consumption, calculating a scoring mechanism for each device based on this state to select devices participating in federated learning within each group, and measuring the actual effect of the current device selection strategy based on the feedback of the lower-layer reward function; A federated learning training module for constructing an objective function for device scheduling and initializing global model parameters , performing federated learning training based on the devices selected by the lower-layer reinforcement learning. During the training process, by adjusting the participation rate of each group of devices and the weight coefficients of the optimization objectives maximize the objective function for device scheduling to determine the optimal device scheduling strategy; wherein, the optimization objectives include improving the global model accuracy, controlling the system energy consumption, and the fairness of device participation.

[0077] Example 3 This embodiment also provides an electronic device, including: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to execute the above-mentioned method for optimizing federated learning device scheduling based on double-layer reinforcement learning.

[0078] In this embodiment, the electronic device may include, but is not limited to: a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile computing device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable computing device, a consumer electronic device, etc.

[0079] Example 4 This embodiment also provides a machine-readable storage medium storing executable instructions that, when executed, cause the machine to execute the above-mentioned method for optimizing federated learning device scheduling based on double-layer reinforcement learning.

[0080] Specifically, a system or device equipped with a readable storage medium may be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device is caused to read and execute the instructions stored in the readable storage medium.

[0081] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0082] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD ROM, CD R, CD RW, DVD ROM, DVD RAM, DVD RW, DVD RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code may be downloaded from a server computer or a cloud via a communication network.

[0083] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD ROM, optical storage, etc.) that contain computer-usable program code.

[0084] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0087] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning, characterized in that: The method comprises: S1. Get the current status characteristics of the device , the current state characteristics of the device Including computing power , Battery level and network latency ; S2. Divide the equipment into groups, and assign participation rates to each group of devices through upper-level reinforcement learning , including: Based on the current status characteristics of each device in each group of devices As well as the change in global model accuracy and the total energy consumption of the device, the state of the upper-level reinforcement learning is constructed, and the participation rate is assigned to each group of devices based on this state. , and based on the feedback of the upper-level reward function, determine whether to optimize and adjust the participation rate of each group of devices ; S3, select the devices in each group to participate in federated learning through lower-level reinforcement learning, including: based on the current state characteristics of each device in each group of devices The local model accuracy changes and device energy consumption of each device are used to construct the state of the lower-level reinforcement learning. Based on this state, the scoring mechanism for each device is calculated to select the devices participating in federated learning in each group, and the actual effect of the current device selection strategy is measured based on the feedback of the lower-level reward function. S4. Construct the equipment scheduling objective function and initialize the global model parameters , based on the devices selected by the lower-level reinforcement learning, the federated learning training is performed. During the training process, the participation rate of each group of devices is adjusted And the weight coefficient of the optimization target Maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein the optimization objectives include improving the accuracy of the global model, controlling system energy consumption, and ensuring fairness of device participation.

2. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, characterized in that: In S2, based on the current status characteristics of each device in each group of devices As well as the change in global model accuracy and the total energy consumption of the device, the state of the upper-level reinforcement learning is constructed, and the participation rate is assigned to each group of devices based on this state. Specifically, the upper layer reinforcement learning is done through the strategy function The comprehensive status of each group of devices Mapped to participation rate ,Right now: (1); Among them, the comprehensive status of each group of equipment Including the computing power of the devices in the group , Battery level , network delay , and the current global model accuracy change and total energy consumption of the equipment .

3. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, characterized in that: In S2, based on the global model accuracy change , Total energy consumption of equipment and data quality Construct the upper reward function, namely: (2); In formula (2): Represents the upper layer reward function; Indicates the change in global model accuracy; Indicates the total energy consumption of the selected equipment; Indicates the scoring weight corresponding to the total energy consumption of the equipment; Indicates Data quality factor of the group device.

4. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, characterized in that: In S3, the scoring mechanism is: (3); In formula (3): Indicates the device In the training round The scoring results in ; Indicates the maximum computing capacity of the device in the system; Indicates the full power status value of the device; , , , They are the scoring weight coefficients for device computing performance, energy consumption status, communication quality, and fair participation; is the time decay function, which represents the time decay term set to consider the historical participation of the device, and: (4); In formula (4): Indicates the device In the training round Frequency of participation in Represents the time attenuation coefficient, which is used to control the attenuation speed; Indicates the current training round distance from the device The time since last training session.

5. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, characterized in that: In S3, the lower layer reward function is: (5); In formula (5): represents the lower layer reward function; Indicates the local model in the training round The change in precision, Indicates the device Network stability; Indicates the device In the training round Energy consumption in Indicates the scoring weight corresponding to the change in local model accuracy; Indicates the scoring weight corresponding to the energy consumption of the equipment.

6. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 1, characterized in that: In S4, the equipment scheduling objective function constructed is: (6); In formula (6): represents the equipment scheduling objective function; Indicates the training round The increment of global model accuracy in ; Indicates the number of training rounds Total energy consumption of selected equipment; Indicates the total number of training rounds; Represents the weight coefficient for improving the accuracy of the global model, Represents the weight coefficient of system energy consumption control, The weight coefficient representing the fairness of device participation; is a fairness indicator, and: (7); In formula (7): Indicates the total number of devices selected for training; A positive number to prevent division by zero errors; Indicates the mean of the participation frequencies of all devices.

7. The method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning according to claim 6, characterized in that: In S4, the participation rate of each group of devices and the weight coefficient of the optimization objective Update based on the following formula: (11); (12); In formula (11)-(12): Indicates Group in training round participation rate in Indicates Group in training round participation rate in Indicates the first Adjustment value for group participation rate; Indicates The optimization goal is in the training round The weight coefficient in , and , which correspond to the weight coefficients representing the improvement of global model accuracy, system energy consumption control, and device participation fairness respectively; represents the learning rate; Represents the reward function with respect to weights The gradient of Represents the temperature decay term, as the training rounds Increase gradually decrease; parameter Used to control the decay speed.

8. A device for implementing a method for optimizing the scheduling of federated learning devices based on double-layer reinforcement learning, characterized in that: The device comprises: Data acquisition module, used to obtain the current status characteristics of the device , the current state characteristics of the device Including computing power , Battery level and network latency ; The upper layer grouping module is used to group devices into groups, and assign participation rates to each group of devices through upper-level reinforcement learning , including: Based on the current status characteristics of each device in each group of devices As well as the change in global model accuracy and the total energy consumption of the device, the state of the upper-level reinforcement learning is constructed, and the participation rate is assigned to each group of devices based on this state. , and based on the feedback of the upper-level reward function, determine whether to optimize and adjust the participation rate of each group of devices ; The lower-level device selection module is used to select the devices in each group that participate in federated learning through lower-level reinforcement learning, including: based on the current state characteristics of each device in each group of devices The local model accuracy changes and device energy consumption of each device are used to construct the state of the lower-level reinforcement learning. Based on this state, the scoring mechanism for each device is calculated to select the devices participating in federated learning in each group, and the actual effect of the current device selection strategy is measured based on the feedback of the lower-level reward function. Federated learning training module, used to build the device scheduling objective function and initialize the global model parameters , based on the devices selected by the lower-level reinforcement learning, the federated learning training is performed. During the training process, the participation rate of each group of devices is adjusted And the weight coefficient of the optimization target Maximize the device scheduling objective function to determine the optimal device scheduling strategy; wherein the optimization objectives include improving the accuracy of the global model, controlling system energy consumption, and ensuring fairness of device participation.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the federated learning device scheduling optimization method based on double-layer reinforcement learning as described in any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores executable instructions, which, when executed, enable the machine to execute the federated learning device scheduling optimization method based on double-layer reinforcement learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hierarchical user training management system and method oriented to non-independent identically distributed data

    CN113672684A

  • Equipment scheduling method in multi-task federated learning system

    CN115033382A

  • High-timeliness federal edge learning scheduling strategy design method

    CN116341679A

  • Method for accelerating convergence of global federated learning model and federated learning system

    CN116416508A

  • Federal learning method and system based on deep reinforcement learning

    CN116486192A

Cited By

  • Intelligent optical fiber distribution cooperative scheduling method, device, equipment and medium

    CN120263715A

  • Deep learning optimization system and method for intelligent equipment

    CN120524976A