Federal learning resource allocation optimization method and system for edge intelligent network
By applying game models and two-layer deep reinforcement learning methods in a federated learning environment, the resource allocation optimization problem of dynamic and heterogeneous multi-mobile devices is solved, efficient resource allocation and profit maximization are achieved, and the overall efficiency of edge intelligent networks is improved.
Patent Information
- Application Number
- CN202510056464.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-30
AI Technical Summary
In the federated learning environment of dynamic and heterogeneous multimobile devices, there are challenges in resource allocation optimization, including device heterogeneity, energy consumption limitations, network bandwidth fluctuations, and insufficient incentive mechanisms, resulting in reduced training efficiency, waste of resources and insufficient user participation enthusiasm.
A game model is adopted and combined with two-layer deep reinforcement learning methods, a federated learning resource allocation optimization method for edge intelligent networks is designed. Through the Stackelberg game model, the competitive relationship between edge servers and mobile devices is modeled as pricing strategies and contribution decisions. DDPG and MADDPG algorithms are used to learn the optimal pricing strategies and contribution strategies respectively to achieve optimization of resource allocation.
It has achieved efficient resource allocation optimization for devices with limited resources, improved the overall efficiency of multi-device collaboration in edge intelligent networks, maximized the profits of edge servers and mobile devices, reduced energy consumption and latency, and improved resource utilization and user participation enthusiasm.
Smart Images

Figure CN120066765A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of edge computing and artificial intelligence, and relates to an optimization method and system for federated learning resource allocation for an edge intelligent network. Background Art
[0002] With the rapid popularization of the Internet of Things, 5G networks, and intelligent devices, edge computing, as a technology that sinks computing power to the network edge, has gradually received wide attention. By executing computing tasks at a location close to the data source, edge computing can not only effectively reduce network latency but also reduce the load on the central server, meeting real-time requirements. At the same time, the development of artificial intelligence (AI) technology has made it a trend to deploy AI models to edge nodes for inference and decision-making. By deploying and executing artificial intelligence algorithms in the edge computing environment for real-time processing of local data, reducing data transmission volume and response time, edge intelligence not only utilizes the low latency and high real-time performance of edge computing but also combines the powerful analysis and reasoning capabilities of AI, thus being widely applied in multiple application scenarios such as intelligent manufacturing, intelligent transportation, and smart cities.
[0003] However, edge devices usually face challenges such as limited computing resources, limited battery capacity, and network bandwidth fluctuations, which pose greater constraints and challenges to the execution of machine learning tasks in the edge intelligent network. To address data privacy and bandwidth limitation issues, Federated Learning (FL) has emerged. Federated Learning realizes distributed collaborative learning in the edge intelligent environment by locally training models on edge devices and aggregating parameters, thereby protecting user privacy and saving network resources. However, the practical application of Federated Learning faces many problems, especially when allocating resources on mobile devices with limited resources, which is particularly complex.
[0004] First, due to the heterogeneity of edge devices, i.e., there are significant differences in computing power, energy consumption level, storage capacity, and network connection status among different devices, the complexity of resource allocation increases greatly, making the unified resource allocation strategy unable to meet the needs of all devices, resulting in a decline in training efficiency and resource waste. Second, due to limited energy, mobile devices are prone to reducing their device performance due to excessive energy consumption when participating in federated learning. This situation not only affects the performance of individual devices but may also lead to training convergence problems in the entire federated learning system. In addition, during the process of participating in federated learning, devices may reduce their enthusiasm for participation due to issues such as energy consumption, thereby affecting the overall cooperation efficiency of the system. Current incentive mechanisms are mostly simple integral or reward mechanisms, often unable to fully consider the contribution differences of devices and actual resource consumption, resulting in insufficient user participation enthusiasm. How to design a reasonable incentive mechanism to ensure that devices are willing to contribute computing resources is another major challenge faced by federated learning.
[0005] To address the above problems, a large amount of research has been conducted in the fields of incentive mechanisms and resource allocation for federated learning in recent years, but there are still the following challenges: (1) Traditional resource allocation methods rely on the static scheduling of the central server, ignoring the dynamic characteristics of edge devices, such as energy consumption, computing power, and network status, resulting in low system efficiency, a long training process, and even the situation where devices drop out midway due to insufficient power. (2) The deficiencies of device heterogeneity and incentive mechanisms disrupt the balance of resource utilization. Traditional methods are difficult to reflect the actual contributions of different devices, especially devices with tight resources have low enthusiasm under high consumption. (3) There is a trade-off between the model convergence speed and system overhead. Frequent parameter updates and communications will increase the energy consumption and bandwidth burden. Especially when the network is unstable, the problems of delay and energy consumption are more prominent. Therefore, it is still a difficult problem to achieve optimal resource allocation in a federated learning environment with dynamic and heterogeneous multi-mobile devices. Summary of the Invention
[0006] To solve the problem of optimal resource allocation in a federated learning environment with dynamic and heterogeneous multi-mobile devices, the present invention discloses a method and system for optimizing federated learning resource allocation for an edge intelligent network, comprehensively considering performance indicators such as energy consumption, convergence, and system profit, aiming to achieve efficient resource allocation optimization for resource-limited devices and improve the overall efficiency of multi-device cooperation in the edge intelligent network. The present invention considers the competitive relationship between the edge server and mobile devices, adopts a game model to maximize the profit associated with energy consumption and delay, and solves the optimal solution through a two-layer deep reinforcement learning method to achieve efficient resource allocation in the federated learning system.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] A method for optimizing the resource allocation of federated learning for edge intelligent networks, comprising the following steps:
[0009] Step 1: Considering the competition relationship between the edge server and mobile devices, design a two-stage Stackelberg game based on the incentive mechanism; the edge server, as the game leader, decides the pricing strategy of mobile devices in each communication round; the mobile devices, as the followers, make corresponding actions according to the server's pricing strategy and their own resource status after receiving the server's strategy, respectively improving the utility of each role;
[0010] Step 2: Establish a resource allocation model for multi-mobile-device heterogeneous federated learning with the goal of maximizing the profit associated with energy consumption and latency. The model includes: a resource allocation optimization model for the edge server involving utility and pricing, and a resource allocation optimization model for mobile devices involving heterogeneous parameter selection and communication resource allocation;
[0011] Step 3: Solve the optimal solution based on a two-layer deep reinforcement learning method. Model the game problem as an MDP model, and propose a two-layer deep reinforcement learning solution to find the optimal solution. The edge server, as an agent, uses the DDPG algorithm to learn the pricing strategy through interaction with the environment. The algorithm architecture includes an actor network and a critic network; mobile devices apply the MADDPG algorithm to make corresponding decision actions, and each mobile device is regarded as an agent with its own actor and critic networks;
[0012] Step 4: Perform resource allocation for heterogeneous and dynamic devices according to the overall optimal solution.
[0013] Further, the specific process of Step 1 is as follows: In the federated learning of multiple mobile devices, a set of distributed mobile devices train the model locally and upload the updated model parameters to the edge server for aggregation; at the beginning of each round of communication, the edge server first decides the pricing strategy of mobile devices. The edge server gives a pricing for the unit contribution of each mobile device which characterizes the contribution of the model parameters trained by the device in each round to the improvement of the overall system model performance; mobile devices formulate their own behavioral strategies according to this decision, and each mobile device uses weight quantization to transmit the locally trained model parameters from high precision to low precision.
[0014] Further, the resource allocation optimization model for the edge server involving utility and pricing in Step 2 is:
[0015]
[0016] The resource allocation optimization model for mobile devices involving heterogeneous parameter selection and communication resource allocation is as follows:
[0017]
[0018] Where represents the quantization strategy, is the local iteration strategy, λ 1 and λ 2 are cost-related weight coefficients, k is the number of communication rounds, b n represents the bandwidth allocated to mobile device n, B max is the upper limit of the total wireless bandwidth, T k,global is the total delay of one global iteration in the k-th round, is the total GPU delay of device n in the k-th round.
[0019] Furthermore, the specific process of step 3 is as follows:
[0020] Construct the federated learning game problem into an MDP model Then propose a two-layer deep reinforcement learning scheme to find the optimal solution. When the state s k is sensed in the k-th round of communication, the agent takes the corresponding action Then the environment is transferred to the next state s k+1 , and the agent obtains the reward γ ∈ (0, 1] is the reward discount factor; decouple the federated learning game algorithm into edge server pricing decisions and mobile device contribution decisions:
[0021] The edge server pricing decision includes:
[0022] In each round of communication, the edge server makes pricing decisions first; first, the state set s k observed by the server in the k-th communication round includes the pricing strategies {p k-l , …, p k-1} of the previous l rounds, the historical contribution strategies {u k-l , …, u k-1} of the mobile devices, and the channel gain information {g k-l , …, {g k-1}; secondly, the edge server, as an agent, uses the DDPG algorithm to learn the pricing strategy through interaction with the environment. The algorithm architecture includes an actor network and a critic network; the actor network consists of an actor network π parameterized by θ π and a target actor network π' parameterized by θ π' , takes the state of the edge server as input, and outputs the action pricing strategy; the critic network is composed of θQ Parameterized critic networks \(Q\) and \(\theta\) Q' Composed of a parameterized target critic network \(Q'\), which takes the state and action as inputs and outputs the \(Q\)-value of the edge server;
[0023] At the beginning of learning, the edge server initializes the actor network, the critic network, and the replay buffer In the \(k\)-th communication round, the edge server first determines the pricing strategy \(p\) by inputting the observed current state \(s\) k into the actor network \(\pi\) parameterized by \(\theta\) π ; after the action is completed, the edge server obtains the reward \(r\) k , and the state transfers to \(s'\) k ; store \((s, p, r, s')\) into the replay buffer k+1 ; then perform mini-batch sampling to update the critic, that is, minimize the loss function by gradient descent, expressed as: k , \(p\) k , \(r\) k , \(s'\) k+1 ) After that, update the actor policy through the sampled policy gradient, expressed as:
[0024]
[0025] where
[0026] \(y\) k \(= r\) k +\(\gamma Q'\{s'\), \(\pi'(s'|\theta')\})\) (13) k+1 , \(\pi'(s'|\theta')\) k+1 |\(\theta'\) π′ )|\(\theta'\) Q′ )
[0027] Then, update the actor policy through the sampled policy gradient, expressed as:
[0028]
[0029] Finally, the edge server updates the parameters of the target actor network and the target critic network, expressed as:
[0030] \(\theta'\) π′ \(= \eta\theta'\) π +(1 - \(\eta\))\(\theta\) π′ (14)
[0031] \(\theta\) Q′ \(= \eta\theta\) Q +(1 - \(\eta\))\(\theta\) Q′ (15)
[0032] where \(\eta \ll 1\);
[0033] The mobile device contribution decision includes:
[0034] According to the pricing decision made by the edge server and the time-varying environmental state, the mobile device uses the multi-agent deep reinforcement learning algorithm MADDPG to learn the optimal contribution strategy and bandwidth resource allocation; each mobile device is regarded as an agent, including an actor network and a critic network, and does not know the non-shareable information of other agents; among them, the actor network consists of The parameterized actor network μ n and The parameterized target actor network μ n ', taking the observation of device n as the input and outputting the action, μ = {μ 1 , …, μ N} and μ' = {μ 1 ', …, μ N '} are the policy set and target policy set of all agents; the critic network consists of the critic network and the target critic network ', taking the state and action as the input and outputting the Q value of device n;
[0035] Before the learning starts, the mobile device receives the initial global observation state o k , including the price decision made by the edge server in the current round; each device selects the action according to the current policy in the k-th communication round. After the joint action is completed, the reward is obtained and transferred to the next state o k+1 ; (o k , a k , r k , o k+1 ) is stored in the replay buffer . After that, mini-batch sampling is performed to update the critic, that is, the loss function is minimized by the gradient descent method, which is expressed as:
[0036]
[0037] where
[0038]
[0039] Then, the actor policy is updated by the sampled policy gradient, which is expressed as:
[0040]
[0041] Finally, each mobile device updates the parameters of the target network, which is expressed as:
[0042]
[0043] Among them, η << 1.
[0044] The present invention also provides a federated learning resource allocation optimization system for an edge intelligent network, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the federated learning resource allocation optimization method for the edge intelligent network.
[0045] The beneficial effects of the present invention are as follows:
[0046] (1) Aiming at the fact that the static resource allocation method cannot reflect the resource competition relationship between the edge server and the mobile device in federated learning, the present invention proposes a two-stage Stackelberg game model based on an incentive mechanism, which can dynamically adjust the resource allocation scheme according to the feedback of the device, fully consider the computing power, energy consumption, and network status of the mobile device, and improve the enthusiasm of the edge server and the participating devices.
[0047] (2) The present invention fully considers the heterogeneity of mobile devices, balances system computing and communication, and the server and mobile devices can be adjusted in real time according to the current resource status, avoiding excessive consumption or waste of resources, and ensuring low energy consumption and high profit during the federated learning process.
[0048] (3) The present invention applies the relationship between the number of communication rounds and convergence, as well as the contribution of the model parameters selected by the device in each round of training to the improvement of the overall system model performance, conducts convergence analysis, obtains a closed-form expression of the convergence bound, and clarifies the relationship between the contributions of the participants and the optimization variables. Further, a dynamic solution based on two-layer deep reinforcement learning is adopted, enabling the achievement of game equilibrium in complex edge scenarios where information is partially non-sharable, thereby maximizing the profits of both parties. Description of the Drawings
[0049] Figure 1 It is an example diagram of a federated learning incentive mechanism resource allocation model based on mobile devices.
[0050] Figure 2 It is the process of a resource allocation optimization method based on two-layer deep reinforcement learning.
[0051] Figure 3 It is a comparison diagram of system costs before learning convergence in the embodiment.
[0052] Figure 4 It is a comparison diagram of energy consumption before learning convergence in the embodiment. Detailed Embodiments
[0053] The technical solution provided by the present invention will be described in detail below in conjunction with specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0054] When the present invention is specifically implemented, without loss of generality, it is assumed that federated learning uses the classical convolutional neural network (CNN) LeNet model, which consists of multiple convolutional layers, pooling layers, and fully connected layers. The transmission power P of each device n = 0.2W, and the channel gain g of the mobile device n is modeled as Rayleigh fading with an average path loss of 10 -3 The bandwidth is assumed to be B max = 100MHz. For GPU computing, considering the different computing capabilities of different devices, assume the frequency of the GPU memory module the frequency of the GPU core module the GPU core voltage The network in the DRL scheme adopted by the present invention is a four-layer fully connected neural network, the number of neurons in the hidden layer is 256, the number of mini-batches is 256, and the learning rate is 0.001. The Softmax and ReLU functions are used to activate the actor and critic networks in the edge server, and the tanh and ReLU activation functions are used in the mobile device. For soft target updates, η = 0.001 is used. In this example, the entire game is divided into two stages: in the first stage, the server publishes an incentive policy; in the second stage, the mobile device selects the optimal participation policy according to the server's policy and its own state. The specific implementation steps of this method are as follows:
[0055] A federated learning resource allocation optimization method for an edge intelligent network provided by the present invention includes the following steps:
[0056] Step 1: Considering the competitive relationship between the edge server and the mobile device, a two-stage Stackelberg game is designed based on the incentive mechanism.
[0057] In this game process, the interaction between the server and the mobile device is described by the Stackelberg game model as Figure 1 shown, ensuring that the game reaches an equilibrium state. In the equilibrium state, both the pricing strategy of the server and the response strategy of the device can reach the optimal, that is, the utility functions of both parties are maximized.
[0058] In the federated learning of multiple mobile devices, the set A distributed mobile device trains a model locally and uploads the updated model parameters to the edge server for aggregation, which is called the aggregator. At the beginning of each round of communication, the aggregator first decides the pricing strategy for the mobile devices, and the mobile devices formulate their own behavior strategies based on this decision. The aggregator, that is, the edge server, acts as the leader in the game and determines the unit contribution for each device to make a pricing decision which characterizes the contribution of the model parameters trained by the device in each round to the improvement of the overall system model performance. This incentive mechanism can be modeled as a Stackelberg game. When formulating the strategy, the server comprehensively considers multiple factors, including the computing power of the mobile devices, the current network status, and the contribution degree of participating in federated learning, etc. Secondly, as the follower of the game, after receiving the pricing strategy from the server, the mobile device makes a response based on the pricing strategy and its own resource status, such as computing resources and network status, which can improve the utility of each role respectively and maximize the profit of the participants. The goal of both parties is to maximize their own profit.
[0059] For the edge server, it needs to consider the overall improvement of the system performance by the mobile devices and the price it offers; for the mobile device, it not only needs to consider the price income given by the server, but also considers the resources consumed by the device during the computing and communication processes. In this way, the device can decide how much resources to invest in the participation process according to the actual situation. Considering the limited communication resources and the competition among multiple mobile devices, each mobile device uses weight quantization to quantize the local trained model parameters from high precision to low precision for transmission, thus reducing the communication energy consumption. The convergence speed of the learning model is characterized by controlling the number of local iterations and quantization noise. Devices with more local training iterations and larger transmission bit widths of model parameters will bring a greater improvement in global performance, that is, fewer communication rounds K, while resulting in more energy consumption and delay in a single round of communication. Therefore, when the model converges to the required accuracy, the number of communication rounds k required is inversely proportional to the utility of each device in each round. Given the pricing strategy of the edge server for the unit contribution of each device in the k-th communication round and the contribution strategy of each mobile device the utilities of the edge server and the mobile devices can be calculated. The utility functions of the server and the mobile devices in the k-th communication round are the goals of both parties. The action goal of the aggregator, that is, the edge server, is to find the optimal pricing strategy that maximizes the utility function, and then the device selects its own contribution parameters according to the pricing strategy of the edge server to maximize the profit.
[0060] Step 2: Establish a resource allocation model for multi-mobile-device heterogeneous federated learning with the goal of maximizing the profit associated with energy consumption and delay.
[0061] In the present invention, the orthogonal frequency division multiple access (OFDMA) protocol is adopted to upload the local training results of mobile devices to the edge server, and the data stored on the mobile devices is used for distributed training without updating to the edge server. The upper limit of the total wireless bandwidth is B max b n represents the bandwidth allocated to mobile device n. The delay for uploading the locally updated model parameters from device n to the aggregator is expressed as:
[0062]
[0063] where d n is the data size of device n for uploading quantization parameters. The mobile device of the present invention applies a heterogeneous weight quantization function Q(·) to convert the parameters from 32 bits to q n bits, and the reduction ratio is expressed as N 0 is the Gaussian noise power, is the channel gain of n to the aggregator in the k-th communication round, modeled as Rayleigh fading with an average path loss of 10 -3 Given the transmission power P n = 0.2 W, the communication energy consumption of mobile device n can be calculated as:
[0064]
[0065] Advanced mobile devices have promoted the development of federated learning systems that use GPUs instead of CPUs for local training. Due to the high-speed computing power of GPUs and their advantages in processing data-intensive tasks, the computing model of the present invention is based on GPUs instead of CPUs. During the federated learning process, each mobile device iteratively trains the local model using its own data samples and then updates the model parameters to the edge server for aggregation. Respectively, and give the number of cycles required for mobile device n to process data acquisition in the GPU memory module and computing in the core module. Therefore, the computing time for one round of training of the mobile device is expressed as:
[0066]
[0067] where and represent the frequencies of the GPU memory module and the core module of device n respectively. ρ n is the number of local iterations of n in one round of iteration, and other irrelevant computing delays are represented by In addition, given the GPU core voltage of device n as the computing power of mobile device n can be calculated as:
[0068]
[0069] where is the total power consumption of other irrelevant components. and are related hardware constant coefficients. Therefore, the computing energy consumption of mobile device n in one round can be calculated as:
[0070]
[0071] The total GPU latency of device n in the k-th round is given by and the total latency of one global iteration in the k-th round is expressed as:
[0072]
[0073] In the k-th round, the total GPU energy consumption of mobile device n is Therefore, the total energy consumption in the k-th round is expressed as:
[0074]
[0075] Given the pricing strategy for the unit contribution of each device by the edge server in the k-th communication round and the contribution strategy of each mobile device the utility function of the aggregator in the k-th communication round can be calculated as:
[0076]
[0077] Then the utility function of the mobile device in the k-th communication round is expressed as:
[0078]
[0079] where λ 1 and λ 2 are cost-related weight coefficients.
[0080] In the present invention, the action goal of the aggregator, i.e., the edge server, is to find the optimal pricing strategy that maximizes the utility function, and then the device selects its own contribution parameters according to the pricing strategy of the edge server to achieve profit maximization. Meanwhile, an attempt is made to explore the optimal bandwidth resource allocation for each mobile device Therefore, the goals of the edge server and the mobile device in the k-th round of communication can be formulated as an optimization problem:
[0081] For the edge server:
[0082]
[0083] For the mobile device:
[0084]
[0085] Among them represents the quantization strategy which is the local iteration strategy
[0086] Step 3: Solve the optimal solution based on a two-layer deep reinforcement learning method
[0087] In this step of the present invention, the game problem is modeled as an MDP model and a two-layer deep reinforcement learning (DRL) scheme is proposed to find the optimal solution. The method process is as Figure 2 shown. When the state s k is sensed in the k-th round of communication, the agent takes the corresponding action and then transfers the environment to the next state s k+1 , and the agent obtains the reward γ∈(0,1] is the reward discount factor. The goal of deep reinforcement learning is to find the best policy π that maps states to actions to maximize the expected cumulative discounted reward. In the present invention, the federated learning game algorithm based on DRL is decoupled into pricing decision and contribution decision, and the profits of the edge server and the mobile device are maximized simultaneously
[0088] This step is divided into two sub-steps: pricing decision and contribution decision
[0089] Sub-step 3-1: Edge server pricing decision
[0090] In each round of communication, the edge server is the leader with the priority decision-making authority and makes the pricing decision first. First, the state set s k observed by the server in the k-th communication round includes the pricing strategies {p k-l ,…,p k-1} of the historical l rounds the historical contribution strategies {u k-l ,…,u k-1} of the mobile device, and the channel gain information g k-l ,…,g k-1}. Secondly, the edge server, as an agent, uses the DDPG algorithm to learn the pricing strategy through interaction with the environment. The algorithm architecture includes an actor network and a critic network. The actor network consists of an actor network π parameterized by θ π and a target actor network π' parameterized by θ π' . Taking the state of the edge server as the input, it outputs the action pricing strategy. The critic network consists of a critic network Q parameterized by θ Q and θ Q'It consists of a parameterized target critic network Q', which takes the state and action as inputs and outputs the Q-value of the edge server.
[0091] At the beginning of learning, the edge server initializes the actor network, the critic network, and the replay buffer. In the k-th communication round, the edge server first determines the pricing strategy p by inputting the observed current state s k into the actor network π parameterized by θ π . After the action is completed, the server obtains the reward r k , and the state transfers to s k . Store (s k+1 , p k , r k , s k ) in the replay buffer k+1 . After that, mini-batch sampling is performed to update the critic, that is, the loss function is minimized by gradient descent, which is expressed as:
[0092]
[0093] where
[0094] y k = r k + γQ'(s k+1 , π'(s k+1 |θ π′ )|θ Q′ ) (13)
[0095] Then, the actor policy is updated by the sampled policy gradient, which is expressed as:
[0096]
[0097] Finally, the edge server updates the parameters of the target actor network and the target critic network, which is expressed as:
[0098] θ π' = θ π + (1 - η)θ π' (14)
[0099] θ Q' = ηθ Q + (1 - η)θ Q' (15)
[0100] where η << 1 (let's assume η = 0.001), that is, the soft update method is adopted instead of directly copying the parameters.
[0101] Sub-step 3-2: Mobile device contribution decision.
[0102] According to the pricing decisions made by the edge server and the time-varying environmental state, the mobile device adopts the multi-agent deep reinforcement learning algorithm MADDPG to learn the optimal contribution strategy and bandwidth resource allocation. Each mobile device is regarded as an agent, including an actor network and a critic network, and does not know the non-shareable information of other agents, which conforms to the practical significance of privacy protection. Among them, the actor network consists of the parameterized actor network μ n and the parameterized target actor network μ n '. Taking the observation of device n as the input and outputting the action, μ = {μ 1 , …, μ N} and μ' = {μ 1 ', …, μ N '} are the policy set and the target policy set of all agents. The critic network consists of the critic network and the target critic network . Taking the state and action as the input, it outputs the Q value of device n.
[0103] Before the learning starts, the mobile device receives the initial global observation state o k , including the price decision made by the edge server in the current round. Each device selects an action according to the current policy in the k-th communication round . After the joint action is completed, it obtains the reward and transfers to the next state o k+1 . Store (o k , a k , r k , o k+1 ) into the replay buffer . Then, perform mini-batch sampling to update the critic, that is, minimize the loss function by gradient descent method, which is expressed as:
[0104]
[0105] where
[0106]
[0107] Next, update the actor policy through the sampled policy gradient, which is expressed as:
[0108]
[0109] Finally, each mobile device updates the parameters of the target network, which is expressed as:
[0110]
[0111] Similarly, η << 1 (let's assume η = 0.001), that is, the soft update method is adopted.
[0112] Step 4: Allocate resources for heterogeneous and dynamic devices according to the overall optimal solution.
[0113] The most important feature of deep reinforcement learning in solving optimization problems is flexible reward design. When the reward function is reasonably designed, high-efficiency learning performance can be obtained. The goal of the present invention is to maximize the profit of the federated learning system and improve the energy consumption savings of mobile devices under the guarantee of converging to a certain accuracy. Considering the resource competition relationship between the edge server and mobile devices, the incentive mechanism game model proposed by the present invention can dynamically adjust the resource allocation scheme according to the feedback of heterogeneous devices, fully considering the computing power, energy consumption and network status of mobile devices, and improving the enthusiasm of the edge server and participating devices. As the deep reinforcement learning algorithm continuously optimizes the strategy in each round of iteration, the system can gradually approach the optimal solution, thereby realizing the efficient allocation of global resources and the maximization of system performance.
[0114] In real federated learning, the network and mobile users have strong dynamics and heterogeneity, including changes in computing and bandwidth resources. For specific situations, it is necessary to design a resource allocation scheme suitable for the current situation for edge servers and mobile devices with corresponding state information changes according to Step 3. According to the optimal resource allocation scheme obtained in Step 3, by reasonably allocating limited computing and communication resources, the federated learning system can achieve resource allocation that maximizes profit. Based on this embodiment for experiments, the results show that compared with the benchmark method, each intelligent device can select the optimal solution, and can reduce the system cost (as Figure 3 shown) and save energy consumption (as Figure 4 shown).
[0115] The present invention also provides a federated learning resource allocation optimization system for an edge intelligent network, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the federated learning resource allocation optimization method for the edge intelligent network.
[0116] It should be noted that the above content only illustrates the technical idea of the present invention and cannot limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A method for optimizing the allocation of federated learning resources for edge intelligent networks, characterized in that: The steps include: Step 1: Considering the competitive relationship between edge servers and mobile devices, a two-stage Stackelberg game is designed based on the incentive mechanism. The edge server, as the game leader, determines the pricing strategy of mobile devices in each communication round. After receiving the server's strategy, the mobile device acting as a follower takes corresponding actions based on the server's pricing strategy and its own resource status, respectively improving the utility of each role; Step 2: Establish a multi-mobile device heterogeneous federated learning resource allocation model with the goal of maximizing the profit associated with energy consumption and latency. The model includes: a resource allocation optimization model involving utility and pricing for edge servers, and a resource allocation optimization model involving heterogeneous parameter selection and communication resource allocation for mobile devices; Step 3: Solve the optimal solution based on the two-layer deep reinforcement learning method. Model the game problem as an MDP model and propose a two-layer deep reinforcement learning solution to find the optimal solution. The edge server, as an intelligent agent, uses the DDPG algorithm to learn the pricing strategy through interaction with the environment. The algorithm architecture includes an actor network and a critic network. The mobile device uses the MADDPG algorithm to make corresponding decision actions. Each mobile device is regarded as an intelligent agent with its own actor and critic network. Step 4: Allocate resources for heterogeneous and dynamic devices based on the overall optimal solution.
2. The method for optimizing the allocation of federated learning resources for edge intelligent networks according to claim 1, characterized in that: The step 1 specifically includes the following process: In the federated learning of multiple mobile devices, the collection The distributed mobile devices train the model locally and upload the updated model parameters to the edge server for aggregation; at the beginning of each round of communication, the edge server prioritizes the pricing strategy of the mobile device, and the edge server contributes a unit of Give pricing Characterizes the contribution of the model parameters trained by the device in each round to the improvement of the overall system model performance; The mobile device formulates its own behavior strategy based on the decision. Each mobile device uses weight quantization to transfer the local training model parameters from high precision to low precision.
3. The method for optimizing the allocation of federated learning resources for edge intelligent networks according to claim 1, characterized in that: The resource allocation optimization model of the edge server involving utility and pricing in step 2 is: The resource allocation optimization model for mobile devices involving heterogeneous parameter selection and communication resource allocation is: in represents a quantitative strategy, is the local iteration strategy, λ1 and λ2 are the weight coefficients related to the cost, k is the number of communication rounds, and b n represents the bandwidth allocated to mobile device n, B max is the upper limit of the total wireless bandwidth, T k,global is the total delay of a global iteration in round k, is the total GPU latency of device n in round k.
4. The method for optimizing the allocation of federated learning resources for edge intelligent networks according to claim 1, characterized in that: The step 3 specifically includes the following process: Constructing the federated learning game problem as an MDP model Then a two-layer deep reinforcement learning scheme is proposed to find the optimal solution. In the kth round of communication, the state s is perceived. k When , the agent takes corresponding actions Then the environment is transferred to the next state s k+1 , the agent gets reward γ∈(0,1] is the reward discount factor; the federated learning game algorithm is decoupled into edge server pricing decision and mobile device contribution decision: The edge server pricing decisions include: In each round of communication, the edge server makes a pricing decision first; first, the state set s observed by the server in the kth communication round k Including the pricing strategy of the historical l rounds {p k-l ,…,p k-1 }, Historical contribution strategy of mobile devices k-l ,…,u k-1 }, and the channel gain information {g k-l ,…,g k-1 }; Secondly, the edge server, as an intelligent agent, uses the DDPG algorithm to learn pricing strategies through interaction with the environment. The algorithm architecture includes an actor network and a critic network; the actor network consists of θ π Parameterized actor networks π and θ π′ The parameterized target actor network π' takes the state of the edge server as input and outputs an action pricing policy; the critic network consists of θ Q Parameterized critic network Q and θ Q′ The parameterized target critic network Q' takes the state and action as input and outputs the Q value of the edge server; At the beginning of learning, the edge server initializes the actor network, critic network and replay buffer In the kth communication round, the edge server first transforms the observed current state s k Input to the π Determine the pricing strategy p in the parameterized actor network π k ; After the action is completed, the edge server receives a reward r k , the state is transferred to s k+1 ; will (s k ,p k ,r k ,s k+1 ) is stored in the replay buffer Then, small batch sampling is performed to update the critic, that is, the loss function is minimized by the gradient descent method, which is expressed as: in y k =r k +γQ'(s k+1 ,π'(s k+1 |θ π′ )|θ Q′ ) (13) Next, the actor policy is updated by the sampled policy gradient, expressed as: Finally, the edge server updates the parameters of the target actor network and the target critic network, expressed as: i π′ =eth π +(1-η)θ π′ (14) i Q′ =eth Q +(1-η)θ Q′ (15) Among them, η<<1; The mobile device contribution decision includes: According to the pricing decisions made by the edge server and the time-varying environmental state, the mobile device uses the multi-agent deep reinforcement learning algorithm MADDPG to learn the optimal contribution strategy and bandwidth resource allocation; each mobile device is regarded as an agent, including an actor network and a critic network, and does not know the non-shared information of other agents; the actor network consists of Parameterized actor network μ n and Parameterized target actor network μ n ', taking the observation of device n as input and outputting actions, μ = {μ1,…,μ N } and μ'={μ1',…,μ N '} is the strategy set and target strategy set of all agents; the critic network is composed of the critic network and target critic network Composition, taking state and action as input, outputs the Q value of device n; Before learning begins, the mobile device receives the initial global observation state o k , including the price decision made by the edge server in the current round; each device chooses an action according to the current strategy in the kth communication round In joint action After completion, get reward Transition to next state o k+1 ; will (o k ,a k ,r k ,o k+1 ) is stored in the replay buffer Then, small batch sampling is performed to update the critic, that is, the loss function is minimized by the gradient descent method, which is expressed as: in Next, the actor policy is updated by the sampled policy gradient, expressed as: Finally, each mobile device updates the parameters of the target network, expressed as: Among them, η<<1.
5. A federated learning resource allocation optimization system for edge intelligent networks, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method for optimizing federated learning resource allocation for edge intelligent networks as described in any one of claims 1-4.
Citation Information
Cited By
Multi-network multi-granularity resource optimization method based on federal deep reinforcement learning
CN121099375A
Multi-network multi-granularity resource optimization method based on federated deep reinforcement learning
CN121099375B
Interruption risk-oriented federated learning method and system, product and medium
CN121365751A
A federated learning method, system, product and medium for interruption risk
CN121365751B
Federal distributed processing method and system based on traffic data elements
CN121765017A