Interference-assisted resource optimization method for asynchronous federated learning system
By introducing interference-assisted resource optimization methods into the asynchronous federated learning system, combining the two-step resource allocation strategies of DDPG and Transformer networks, the resource configuration of terminal devices is optimized, and the problem of eavesdropping attacks and low resource utilization is solved, and the low cost and efficient training of asynchronous federated learning is achieved.
Patent Information
- Application Number
- CN202311697288.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
The existing asynchronous federated learning system is vulnerable to eavesdropping attacks during global model broadcasting, and has low resource utilization, resulting in high computing and communication costs.
The interference assisted resource optimization method for asynchronous federated learning systems is adopted. By building a dynamic model of delay and energy consumption, combining the two-step resource allocation strategy and interference assist technology of DDPG and Transformer networks, the CPU frequency, transmission power and interference power of the terminal equipment are optimized to ensure that the downlink data transmission rate is greater than the given confidentiality rate.
In a dynamically changing wireless environment, minimize the delay and energy consumption required for asynchronous federated learning, improve the robustness and confidentiality rate of the system, and reduce training time and energy costs.
Smart Images

Figure CN120151844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of network resource allocation, specifically an interference-assisted resource optimization method for asynchronous federated learning systems. Background Art
[0002] In existing asynchronous federated learning, multiple updates are required to achieve high-precision federated learning results, which brings huge computational and communication costs. In addition, although federated learning protects user privacy by exchanging model parameters, since the wireless channel is open access, malicious eavesdroppers may intercept the model parameters transmitted by the server and infer the original data by analyzing the differences in these parameters. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art that cannot cope with eavesdropping attacks that may occur during the global model broadcast process and have low resource utilization, the present invention proposes an interference-assisted resource optimization method for asynchronous federated learning systems. For the first-come, first-served mode, that is, the local models that arrive at the server first are selected to participate in the aggregation of the asynchronous federated learning process, analyze the dynamic modeling of delay and energy consumption in the actual wireless communication scenario, and based on the two-step resource allocation strategy and interference-assisted technology of DDPG, minimize the delay and energy consumption required for asynchronous federated learning on the basis of ensuring physical layer security.
[0004] The present invention is realized through the following technical solutions:
[0005] The present invention relates to an interference-assisted resource optimization method for asynchronous federated learning systems. By constructing a dynamic model of delay and energy consumption for the asynchronous federated learning system in the first-come, first-served mode, using a two-step algorithm that combines Deep Deterministic Policy Gradient (DDPG) and Transformer network deployed at the base station equipped with an edge server, for the state information collected in the current environment during each global model aggregation, make resource allocation decisions including the scheduling strategy of uplink and downlink terminal devices and the CPU frequency, transmit power, and interference power of the terminal devices, and use interference-assisted technology to make the downlink data transmission rate greater than the given secrecy rate to counter eavesdropping attacks.
[0006] The present invention relates to a system for implementing the above method, including: a base station equipped with an edge server and multiple terminal devices configured with full-duplex antennas. Among them, the base station equipped with an edge server generates a global model by weighted aggregation of the collected local models, generates a resource allocation decision according to the collected state information, and sends back the resource allocation decision and the global model together; the terminal devices perform local training according to the resource allocation decision, local data set, and downloaded global model, and upload the local models, and some terminal devices send interference signals when the base station equipped with an edge server broadcasts the global model. Technical effects
[0007] The present invention combines the improved DDPG with the device scheduling interference assistance technology of the Transformer encoder module, and is applicable to the resource allocation strategy of the first-come, first-served asynchronous federated learning system. Compared with the prior art, the present invention combines the interference assistance technology with the asynchronous federated system, and uses the scheduled device to send an auxiliary interference signal to counter eavesdropping attacks. The present invention jointly optimizes the uplink and downlink terminal device scheduling, the CPU frequency of the terminal device, the transmission power of the terminal device, and the interference power of the terminal device in a dynamically changing wireless environment to reduce the time and energy cost consumption of the first-come, first-served mode asynchronous federated learning training. The two-step algorithm proposed by the present invention has stronger robustness compared with the benchmark algorithm. Brief description of the drawings
[0008] Figure 1 It is a flowchart of the present invention;
[0009] Figure 2 It is a schematic diagram of the application scenario of the embodiment;
[0010] Figure 3 It is a schematic diagram of the iteration of the asynchronous federated learning task;
[0011] Figure 4 It is a schematic diagram of the two-step algorithm combining DDPG and the Transformer network;
[0012] Figure 5 It is a schematic diagram of the DDPG network architecture;
[0013] Figures 6 - 8 It is a schematic diagram of the effect of the embodiment. Detailed implementation manners
[0014] As Figure 2 shown, this embodiment relates to a resource optimization scenario for an asynchronous federated learning system, including: a base station equipped with an edge server, an eavesdropper, and multiple terminal devices configured with full-duplex antennas Among them: the size of the local dataset owned by the terminal device k is D k , when the selected terminal device for aggregation receives the global model from the base station equipped with the edge server, some of the terminal devices are selected to send interference signals to the eavesdropper to resist eavesdropping attacks by increasing the downlink data transmission rate.
[0015] As Figure 1 shown, this embodiment relates to a physical layer security resource optimization method for an asynchronous federated learning system based on the above scenario. By constructing and training an asynchronous federated learning system, an asynchronous federated learning task is completed through multiple rounds of iteration, including:
[0016] Step 1: Asynchronous federated learning performs dynamic modeling of latency and energy consumption in a first-come, first-served manner and initializes it, specifically including:
[0017] Step 1.1: Construct the local training latency and energy consumption model: When a local training requires M rounds of iteration, the computing latency of a local training The total energy consumed by the terminal device k to complete the local model training where: C k is the number of CPU cycles required for the terminal device k to execute one bit of data, is the CPU frequency of the terminal device in round l, that is, the CPU frequency of the terminal device in a local training is fixed at the CPU frequency configuration when the terminal device starts training, ζ k is the effective capacitance coefficient based on the device chip.
[0018] Step 1.2: Construct the local model upload latency and energy consumption model: The communication latency for uploading the local model from the terminal device k to the base station carrying the edge server Communication energy consumption where: The system bandwidth B is evenly distributed to all terminal devices, S is the size of the local model parameters, is the transmission power of the model upload of the terminal device k, is the Gaussian noise of the base station carrying the edge server, is the channel gain from the terminal device k to the base station carrying the edge server, remains unchanged during a local training and is fixed at the value when the terminal device starts training. Then the latency generated by the asynchronous federated learning system in the local model training and upload phases in the l-th round
[0019] In the asynchronous federated learning system, terminal devices often cannot complete local training or model upload in one global iteration. Introduce as the computing and communication latency required for the terminal device k to complete the remaining local training task volume in the l-th round. When the terminal device k participates in aggregation during the l - 1-th round, its required latency is reset to the computing and communication latency required for one local task and otherwise its required latency is where: The latency not completed in the l - 1-th round The global latency T of the previous round l-1 .
[0020] Step 1.3: Construct the global model aggregation latency and energy consumption model: When the processing rate used by the base station carrying the edge server remains unchanged, the latency required for one global model aggregation is fixed at T agg; Compared with the terminal device, the base station equipped with an edge server has sufficient computing power, so the energy consumption during aggregation can be ignored.
[0021] Step 1.4, construct the global model broadcast delay and energy consumption model: To enhance the anti-eavesdropping ability of the system, the terminal devices participating in the aggregation can send interference signals to improve the downlink data transmission rate. Under the cooperative interference, the data transmission rate of the global model broadcast from the base station equipped with an edge server to the terminal device k When the size of the global model is equal to the size of the local model, the delay of the global model broadcast from the base station equipped with an edge server to the terminal device in the l-th round The energy consumption required for the terminal device k to send interference signals in the l-th round Where: Are the channel gains from the base station equipped with an edge server to the terminal device k, from the base station equipped with an edge server to the eavesdropper, and from the terminal device k to the eavesdropper, respectively. p s Is the server transmission power, And Are the background Gaussian noises of the terminal device k and the eavesdropper, respectively. Indicates whether the terminal device k sends interference signals to the eavesdropper in the l-th round. When the terminal device k sends interference signals with Power, then Otherwise The set Π of terminal devices sending interference l .
[0022] Step 1.5, consider the computing energy consumption and communication energy consumption of the terminal devices during aggregation. The total delay and total energy consumption required to complete L rounds of asynchronous federated learning iteration only need to sum the delay T l For each round and sum the energy consumption E l For each round:
[0023] Step 1.6, the base station equipped with an edge server initializes the global machine learning model ω 0 , specifies the training termination condition, and distributes ω 0 To each terminal device.
[0024] Step 2, local training and uploading, specifically including:
[0025] Step 2.1, the terminal device k uses the global model Used for resource allocation decision-making in the Round to perform local training and generate a new local model Where: The round when the terminal device k starts training If the terminal device k participates in the l-th round of aggregation, otherwise it does not participate.
[0026] Step 2.2 The terminal device k sends the updated local model and the local status information s l to the base station equipped with the edge server.
[0027] Step 3: Global model aggregation: The base station equipped with the edge server receives a specified number of local models and local status information on a first-come, first-served basis. After weighting the model parameters according to the freshness of the local models at the base station equipped with the edge server, weighted aggregation is performed according to the size of the local data volume. While obtaining the global model of this round, a resource allocation decision is generated based on the collected status information, specifically including:
[0028] Step 3.1) Based on the minimum latency and energy consumption obtained in Step 1, construct an optimization objective problem for the resource allocation scheme: To jointly evaluate the latency and energy consumption of the system while ensuring model accuracy, the optimization objective is set as Decision variable where: Θ = ξ 1 T g + ξ 2 E g is the total consumption of an asynchronous federated learning task, ξ 1 and ξ 2 are the weight coefficients of latency and energy consumption respectively, and satisfy ξ 1 + ξ 2 = 1, |Φ l | is the number of terminal devices participating in the aggregation in the l-th round, Υ 1 is the corresponding weight coefficient. To prevent malicious eavesdropping, it is required that the minimum downlink data transmission rate during the training process should not be less than the given target rate
[0029] Step 3.2 Adopt a two-step algorithm combining Deep Deterministic Policy Gradient (DDPG) and Transformer network as Figure 4 shown to jointly optimize the decision variable Ω l in Step 3 into and two parts, specifically including:
[0030] Step 3.2.1 Convert the problem to be optimized into a Markov decision process: The DDPG agent generates an action based on the current state space information s l i.e., the solution of the first part of the decision variable On this basis, simplify the optimization objective to a linear programming problem, and then obtain the final decision variable. Then, from the decision variable Ωl After acting on the environment, the next state space information s is obtained l+1 and the reward r of the current decision l , where: the state space information statistically aggregates the terminal devices in the last 4 rounds of the k; the reward is to avoid the algorithm converging to the extreme. On the basis of the optimization objective, the reward punishes the terminal devices that have not participated in the aggregation for multiple rounds, Υ 2 is the penalty coefficient, and the operator ⊙ takes 1 when all the values are 0.
[0031] Step 3.2.2 constructs a DDPG network architecture combined with Transformer, as Figure 4 shown, specifically including: an Actor online network and an Actor target network with the same structure, a Critic online network and a Critic target network with the same structure, where: the Critic online network splices the output result of the encoder layer and the action output by the Actor online network, and obtains the state-action value function Q(s l ,a l ) through a linear layer.
[0032] Both the described Actor online network and Actor target network include: a Transformer encoder layer and a linear layer, where: the Transformer encoder layer is composed of 3 identical layers;
[0033] Both the described Critic online network and Critic target network include: a Transformer encoder layer, a feature-action mapping layer and a linear layer, where: after splicing the output result of the Transformer encoder layer and the action output by the Actor online network passing through the feature-action mapping layer, the state-action value function Q(s l ,a l ) is obtained through a linear layer.
[0034] The described encoder layer includes: a multi-head self-attention mechanism, a feed-forward neural network, a normalization layer and a residual connection, where: the multi-head self-attention mechanism converts the mapped state space information into multiple groups of independent query (Query) vectors, key (Key) vectors and value (Value) vectors, and makes each group of query vectors, key vectors and value vectors pass through the scaled dot-product attention in parallel, calculates the corresponding attention output respectively, splices the attention output and reduces the dimension by using a linear transformation to generate the final output.
[0035] The described feed-forward neural network is composed of two linear layers and an activation function.
[0036] Step 3.2.3 solves the decision variables, specifically including:
[0037] i) DDPG outputs decision variables: In each decision-making process, the DDPG network generates new CPU frequency and transmission power configuration schemes for all terminal devices. For the CPU frequency and transmission power of the unaggregated terminal devices, the original resource configuration data is used to overwrite the new scheme. According to the revised CPU frequency and transmission power configuration schemes, the terminal devices are sorted in ascending order of delay from small to large.
[0038] ii) Determine the decision When the number of aggregated terminal devices is n, according to the first-come-first-served principle, the first n terminal devices are the aggregated terminal devices Φ l (n) in this round. The delay of the nth terminal device is the delay generated during the local model training and uploading phases of the system, and the corresponding energy consumption is obtained.
[0039] iii) Simplify the objective problem: According to the determined decision variables and the derivable results, the original objective problem is simplified to: minimizing the sum of interference powers with the following constraints, which is a simple linear programming problem.
[0040] iv) Generate the decision and and substitute them into the optimal solution of the above linear programming problem: The aggregated terminal devices are sorted in descending order of channel gain h ke , i.e., h 1e > h 2e > … > h ie > … h ne . For the terminal device (p max is the transmission power upper limit of the terminal device) that satisfies the condition, the interference power transmitted is The interference power transmitted by the terminal device is The interference power transmitted by the terminal device max is p , i.e., The terminal device does not transmit interference, i.e., i According to q and When there is no terminal device satisfying the condition, step 4.3.5 is skipped.
[0041] v) Calculate the reward r l (n). According to the current decision, calculate the reward r l (n) when the number of aggregated terminal devices is n.
[0042] vi) Output the final decision. Repeat steps 4.3.2-4.3.5, traverse the number of possible terminal devices participating in the aggregation from 0 to K, and select the one that makes the profit r l The decision when (n) is maximum is output as the final decision.
[0043] Step 3.2.4 DDPG network update: Set the final decision variable Ω l Act on the environment to obtain the state space information s at the next moment l+1 and the benefit r of the current action l , and s l , a l , r l ,s l+1 Put it into the experience replay pool used to train the DDPG network until the number of data in the experience replay pool reaches the preset threshold, then randomly select small batches of data from the experience replay pool and use soft update to update the Actor online network, Actor target network, Critic online network, and Critic target network respectively.
[0044] Step 4: Global model broadcast: The base station equipped with the edge server distributes the global model parameters ω of this round to the terminal devices participating in the aggregation l At the same time, some of the terminal devices participating in the aggregation send interference signals to the eavesdropper, return to step 3 and repeat until the training termination condition is met, then the training is completed.
[0045] The global model parameters of this round are in: For Model The staleness, staleness function It is a monotonically decreasing function that aims to reduce the weight of outdated local models and increase the weight of the latest local models during aggregation.
[0046] Step 5: The trained DDPG agent is deployed on the base station equipped with the edge server. The DDPG agent obtains the state space information s during the current global aggregation. l , using the trained DDPG Actor online network to generate action decisions a l , use the two-step algorithm to calculate the final decision variable Ω l , apply it to the next round of asynchronous federated learning iteration, and obtain the state space information s at the next moment l+1 ; Repeat this process until the federated learning task training is completed.
[0047] After specific actual experiments, a scenario is realized in a circular area with a diameter of 200 meters. The base station deploying the edge server is located at the center position. Ten users and eavesdroppers are randomly distributed within this circular area, but it is necessary to ensure that the data transmission rate between the base station carrying the edge server and the users is greater than the data transmission rate between the eavesdroppers and the base station carrying the edge server. The channel gain follows Rayleigh fading with a mean of 10 -L(d) / 20 , where \(L(d)=20\lg(d)+20\times\lg(2450)+32.45\), and \(d\) is the distance between the base station carrying the edge server and the user. The federated learning task uses a convolutional neural network model for local training on the MNIST dataset and the CIFAR-10 dataset. The size of the CNN model is 0.2 MB.
[0048] To prove the superiority of the present invention, five feasible benchmark schemes are adopted: 1) LAFL: using traditional fully connected layers in the actor-critic network of the DDPG algorithm; 2) DAFL: only using the traditional DDPG algorithm; 3) RAFL: random resource allocation; 4) NAFL: not using cooperative interference; 5) SFL: all terminal devices participate in each global aggregation.
[0049] As Figure 6 shown, it is the change of the federated learning model accuracy and the total task consumption \(\Theta\) based on the MNIST dataset after smoothing processing for different schemes. Compared with the benchmark schemes, the proposed TAFL scheme can achieve higher accuracy with less energy consumption and latency.
[0050] As Figure 7 shown, it is the influence of the latency-energy consumption ratio on the model accuracy and the total consumption incurred in the present invention (the default value in the experiments of the present invention is 1:4). It can be found that the resource consumption required to achieve the same accuracy increases with the increase of the latency-energy consumption ratio. This is because numerically, time is several to dozens of times that of energy consumption. Therefore, the larger the time proportion, the greater the total consumption. In practical applications, the ratio can be set by considering the costs of latency and energy consumption in combination.
[0051] As Figure 8 shown, it is the comparison of the robustness of different schemes. By setting a more difficult-to-achieve scenario (higher given secrecy rate and fewer terminal devices), the percentage of different schemes achieving the target requirements during the training process is tested. When the comparison schemes need multiple rounds of exploration to meet the task requirements (or cannot meet them), only the proposed TAFL scheme can always complete the task with 100%.
[0052] Table 1 Algorithm Bandwidth 1MHz Bandwidth 1.5MHz Bandwidth 2MHz TAFL 7.5Mbits / s 11.3Mbits / s 15.7Mbits / s NAFL 3.2Mbits / s 4.8Mbits / s 6.4Mbits / s
[0053] Table 2 Dataset TAFL RAFL SFL MNIST IID 543 1193 1895 MNIST Non-IID 2059 4638 11530 CIFAR-10 IID 448 958 1574 CIFAR-10 Non-IID 2791 5469 11943
[0054] As shown in Table 1, it is the maximum downlink data transmission rate that can be achieved with or without using interference-assisted technology under different channel bandwidths. It can be found that using cooperative interference can effectively improve the system secrecy rate.
[0055] As shown in Table 2, it is the total consumption required by the proposed TAFL scheme and the benchmark scheme when the specified accuracy (97% for MNIST and 59% for CIFAR-10) is first reached under different datasets and data distributions. Compared with the synchronous federated algorithm, it can be reduced by about 70%-80%, and compared with the resource random allocation algorithm, it can be reduced by about 50%-55%.
[0056] Compared with the prior art, the present method adopts a joint uplink and downlink terminal device scheduling, terminal device CPU frequency, terminal device transmission power, and terminal device interference power allocation strategy, reducing the latency and energy required by the first-come, first-served asynchronous federated learning. When allocating resources, it adopts a two-step DDPG algorithm based on a combined transformer encoder, which has strong robustness. When broadcasting the global model, it uses physical layer security technology and adopts cooperative interference, effectively improving the system secrecy rate.
[0057] Those skilled in the art can make local adjustments to the above specific implementation in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation, and all implementation solutions within its scope are subject to the present invention.
Claims
1. A resource optimization method for an asynchronous federated learning system, characterized in that, by constructing a dynamic model of delay and energy consumption for an asynchronous federated learning system in the first-come, first-served mode, using a two-step algorithm combining deep deterministic policy gradient and Transformer network deployed at a base station equipped with an edge server to make, for the state information collected in the current environment at each global model aggregation, resource allocation decisions including the uplink and downlink terminal device scheduling strategies and the CPU frequency, transmission power, and interference power of the terminal devices, and using interference-aided technology to make the downlink data transmission rate greater than a given secrecy rate to combat eavesdropping attacks.
2. The resource optimization method for an asynchronous federated learning system according to claim 1, characterized in that specifically it includes: Step 1: Perform dynamic modeling of delay and energy consumption for asynchronous federated learning in a first-come, first-served manner and initialize it; Step 2: Local training and uploading, specifically including: Step 3: Global model aggregation: The base station equipped with an edge server receives a specified number of local models and local state information according to the first-come, first-served principle. After weighting the model parameters according to the freshness of the local models at the base station equipped with an edge server, weighted aggregation is performed according to the size of the local data volume to obtain the global model of this round. At the same time, according to the collected state information, resource allocation decisions are generated; Step 4, Global model broadcast: The base station equipped with an edge server distributes the global model parameters ω of this round to the terminal devices participating in aggregation. l Meanwhile, some of the terminal devices participating in aggregation send interference signals to the eavesdropper, return to Step 3 and repeat until the training termination condition is met, then the training is completed. The global model parameters of this round Wherein: is the model staleness, and the staleness function is a monotonically decreasing function, aiming to reduce the weight of the stale local model and increase the weight of the latest local model during aggregation; Step 5: Deploy the trained DDPG agent on the base station equipped with an edge server. The DDPG agent generates an action decision a based on the state space information s obtained during the current global aggregation l , and uses the trained Actor online network of DDPG to generate an action decision a l , calculates the final decision variable Ω using a two-step algorithm l , applies it to the next round of asynchronous federated learning iteration, and obtains the state space information s at the next moment l+1 ; Continuously repeat this process until the federated learning task training ends.
3. The resource optimization method for an asynchronous federated learning system according to claim 2, characterized in that the said Step 1 specifically includes: Step 1.
1. Construct the local training delay and energy consumption model: When a local training requires M rounds of iteration, the computing delay of a local training The total energy consumed by the terminal device k to complete the local model training where: C k is the number of CPU cycles required for the terminal device k to execute one bit of data, is the CPU frequency of the terminal device in rounds, that is, the CPU frequency of the terminal device in a local training is fixed to the CPU frequency configuration when the terminal device starts training, ζ k is the effective capacitance coefficient based on the device chip; Step 1.
2. Construct the local model upload delay and energy consumption model: The communication delay of the local model uploaded from the terminal device k to the base station equipped with the edge server Communication energy consumption where: The system bandwidth B is evenly allocated to all terminal devices, S is the size of the local model parameters, is the transmission power of the model upload of the terminal device k, is the Gaussian noise of the base station equipped with the edge server, is the channel gain from the terminal device k to the base station equipped with the edge server, remains unchanged during one local training and is fixed at the value when the terminal device starts training; then the delay generated in the local model training and upload phase of the l-th round of the asynchronous federated learning system Step 1.
3. Construct the global model aggregation delay and energy consumption model: When the processing rate used by the base station equipped with an edge server remains unchanged, the delay required for one global model aggregation is fixed at T agg ; Compared with the terminal device, the base station equipped with an edge server has sufficient computing power, so the energy consumption during aggregation is ignored; Step 1.
4. Construct the global model broadcast delay and energy consumption model: To enhance the anti-eavesdropping ability of the system, the participating terminal devices send interference signals to improve the downlink data transmission rate. Under the cooperative interference effect, the base station equipped with an edge server broadcasts the data transmission rate of the global model to the terminal device k. When the size of the global model is equal to the size of the local model, the delay of the global model broadcast from the base station equipped with an edge server to the terminal device in the l-th round The energy consumption required for the terminal device k to send interference signals in the l-th round Where: Are the channel gains from the base station equipped with an edge server to the terminal device k, from the base station equipped with an edge server to the eavesdropper, and from the terminal device k to the eavesdropper respectively, and p s Is the server transmit power, And Are the background Gaussian noises of the terminal device k and the eavesdropper respectively, Indicates whether the terminal device k sends interference signals to the eavesdropper in the l-th round. When the terminal device k sends interference signals with Power, then Otherwise The set Π of terminal devices sending interference l ; Step 1.
5. When considering the aggregation of terminal devices, calculate their computing energy consumption and communication energy consumption. The total delay and total energy consumption required to complete L rounds of asynchronous federated learning iterations only need to sum the delay T l for each round and the energy consumption E l for each round: Step 1.6: The base station equipped with an edge server initializes the global machine learning model ω 0 , specifies the training termination condition, and distributes ω 0 to each terminal device.
4. The resource optimization method for an asynchronous federated learning system according to claim 3, characterized in that in In an asynchronous federated learning system, end devices often cannot complete local training or model uploading in one global iteration. Introduce The computing and communication delay required for end device k to complete the remaining local training task volume in the l-th round. When end device k participates in aggregation during the (l - 1)-th round, its required delay is reset to the sum of the computing and communication delays required for one local task Otherwise, its required delay is Where: the delay not completed in the (l - 1)-th round The global delay T of the previous round l-1 .
5. The resource optimization method for an asynchronous federated learning system according to claim 2, characterized in that the said Step 2 specifically includes: Step 2.
1. The terminal device k, according to the global model used for resource allocation decision in the round, performs local training to generate a new local model where: the round when the terminal device k starts training is that the terminal device k participated in the l-th aggregation, otherwise it did not participate. Step 2.2 The terminal device k sends the updated local model and the local status information s l to the base station equipped with the edge server.
6. The resource optimization method for an asynchronous federated learning system according to claim 2, characterized in that the said Step 3 specifically includes: Step 3.1) Based on the minimum latency and energy consumption obtained in Step 1, construct an optimization objective problem for the resource allocation scheme: To jointly evaluate the latency and energy consumption of the system while ensuring the model accuracy, the optimization objective is set as Decision variable where: Θ = ξ 1 T g + ξ 2 E g is the total consumption of one asynchronous federated learning task, ξ 1 and ξ 2 are the weight coefficients of latency and energy consumption respectively, and satisfy ξ 1 + ξ 2 = 1, |Φ l | is the number of terminal devices participating in aggregation in l rounds, Υ 1 is the corresponding weight coefficient; To prevent malicious eavesdropping, it is required that the minimum downlink data transmission rate should not be less than the given target rate Step 3.2 uses a two-step algorithm combining Deep Deterministic Policy Gradient (DDPG) and Transformer network to jointly optimize the decision variable Ω in Step 3 l into and two parts for joint optimization.
7. The resource optimization method for an asynchronous federated learning system according to claim 6, characterized in that the said Step 3.2 specifically includes: Step 3.2.1 Convert the problem to be optimized into a Markov decision process: The DDPG agent generates an action, i.e., the solution of the first part of the decision variables, based on the current state space information s l On this basis, simplify the optimization objective into a linear programming problem, and then obtain the final decision variables. Then, the decision variables Ω Act on the environment to obtain the next state space information s l And the reward r of the current decision l+1 l l , where: The state space information Statistics the aggregation situation of the terminal device k in the recent 4 rounds; the reward Is to avoid the algorithm converging to the extreme. The reward punishes the terminal devices that have not participated in the aggregation for multiple rounds on the basis of the optimization objective. Υ 2 Is the penalty coefficient, and the operator ⊙ takes 1 when all the values are 0; Step 3.2.2 Construct a DDPG network architecture combined with Transformer, as shown in Figure 4, which specifically includes: an Actor online network and an Actor target network with the same structure, and a Critic online network and a Critic target network with the same structure. Among them: the Critic online network splices the output result of the encoder layer with the action output by the Actor online network, and then obtains the state-action value function Q(s l , a l ) through a linear layer; Step 3.2.3 Solve the decision variables, specifically including: i) DDPG outputs decision variables: For each decision, the DDPG network generates a new CPU frequency and transmission power configuration plan for all terminal devices. The CPU frequencies and transmission powers of the terminal devices not participating in the aggregation continue to use the original resource configuration data to overwrite the new plan; According to the corrected CPU frequency and transmission power configuration plan, the terminal devices are sorted in ascending order of delay ii) Determine the decision When the number of participating aggregation terminal devices is n, according to the first-come, first-served principle, the first n ones are the aggregation terminal devices Φ l (n) in this round. The delay of the nth terminal device is the delay generated by the system in the local model training and uploading phases, and the corresponding energy consumption is obtained. iii) Simplify the target problem: According to the determined decision variables and the derivable results, the original target problem is simplified to: find the sum of interference powers with as the constraints, which is a simple linear programming problem with the minimum value; iv) Generate a decision and Substitute into the optimal solution of the above linear programming problem: Arrange the terminal devices participating in the aggregation in descending order according to the channel gain h ke That is, h 1e >h 2e >…>h ie >…h ne , satisfying (p max is the upper limit of the transmission power of the terminal device) The interference power transmitted by the terminal device is The interference power transmitted by the terminal device is p max , that is, q i =p max , The terminal device does not transmit interference, that is According to q i , the corresponding and When there is no terminal device that meets the conditions , then skip step 4.3.5; v) Calculate the revenue r l (n); Calculate the revenue r when the number of participating aggregation terminal devices is n according to the current decision l (n); vi) Output the final decision; repeat steps 4.3.2 - 4.3.5, traverse the number of possible participating end devices n from 0 to K, and select the decision when the revenue r l (n) is the largest as the final decision for output; Step 3.2.4 DDPG network update: Apply the final decision variable Ω l to the environment to obtain the state space information s at the next moment l+1 and the reward r of the current action l , and put s l , a l , r l , s l+1 into the experience replay pool for training the DDPG network. After the number of data in the experience replay pool reaches the preset threshold, randomly select a small batch of data from the experience replay pool and update the Actor online network, Actor target network, Critic online network, and Critic target network respectively in a soft update manner.
8. The resource optimization method for an asynchronous federated learning system according to claim 2, characterized in that both the Actor online network and the Actor target network include: a Transformer encoder layer and a linear layer, where: the Transformer encoder layer is composed of 3 identical layers; The described Critic online network and Critic target network both include: a Transformer encoder layer, a feature-action mapping layer, and a linear layer, where: after concatenating the output result of the Transformer encoder layer with the action output by the Actor online network that has passed through the feature-action mapping layer, the state-action value function Q(s l , a l ) the said encoder layer includes: a multi-head self-attention mechanism, a feed-forward neural network, a normalization layer, and a residual connection, where: the multi-head self-attention mechanism converts the mapped state space information into multiple groups of independent query (Query) vectors, key (Key) vectors, and value (Value) vectors, enables each group of query vectors, key vectors, and value vectors to pass through scaled dot-product attention in parallel, calculates the corresponding attention outputs respectively, splices the attention outputs and generates the final output after dimensionality reduction using a linear transformation; the said feed-forward neural network consists of two linear layers and an activation function.
9. A resource optimization system for an asynchronous federated learning system implementing the method according to any one of claims 1-8, characterized in that, it includes: a base station equipped with an edge server and multiple terminal devices configured with full-duplex antennas, wherein the base station equipped with the edge server generates a global model by weighted aggregation based on the collected local models, and generates a resource allocation decision based on the collected status information, and sends back the resource allocation decision and the global model together; The terminal device performs local training according to the resource allocation decision, the local data set and the downloaded global model, and uploads the local model. Some terminal devices send interference signals when the base station equipped with the edge server broadcasts the global model.
Citation Information
Cited By
Federal learning aggregation method, medium and system
CN120822582A
Multi-dimensional resource management joint optimization method based on wireless edge network
CN120835006A
Method, device, system and equipment for enhancing park signal by combining 5G ultra-dense networking with federated learning and storage medium
CN121099354A