A Federated Learning Method for Wireless Charging Based on Three-Stage Stackelberg Game Theory
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-08
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]为了克服现有技术的不足,本发明提供了一种基于三阶段Stackelberg博弈的无线充电联邦学习方法,该方法解决了联邦学习中终端设备的能源限制和个体自私问题,激励所有角色参与系统并确保联邦学习任务的成功完成
Smart Images

Figure CN118657234B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communication technology, specifically relating to a wireless charging federated learning method based on a three-stage Stackelberg game. Background Technology
[0002] The large-scale deployment of IoT devices and smart applications has generated massive amounts of data, demonstrating the enormous potential for training effective machine learning models. Deep reinforcement learning, as a powerful machine learning method, has achieved remarkable results in many tasks such as image recognition, natural language processing, and intelligent recommendation. However, centralized data training models pose a serious risk of privacy breaches due to the privacy implications of the data. Furthermore, the transmission of raw data consumes significant communication resources. Federated Learning (FL), as a distributed machine learning paradigm, offers a promising solution to these problems. It allows End Equipment Workers (EWs) to participate in the learning process and train the model while maintaining the data locally, significantly reducing the risk of privacy breaches. Since only the model parameters are offloaded to cloud servers, the data privacy of the EWs is maintained. Simultaneously, the size of the model parameters is typically much smaller than the original data, further reducing the size of data transmission.
[0003] However, training complex tasks quickly depletes the energy of the Energy Controller (EW), posing a challenge to obtaining a high-performance Functional Rendering (FL) model. Thanks to advancements in radio frequency technology, Wireless Power Transfer (WPT) can transfer energy to the EW to address the energy constraint problem. Specifically, after the Base Station (BS) distributes the FL model as a cloud server, the EW can use a dedicated frequency to obtain energy from the Charging Service Provider (CSP) via WPT for training and transmitting the model. However, without incentives, the EW is unwilling to use its own computing resources to train the FL model for the BS. The CSP incurs costs when transferring energy, while the BS desires a better FL model for less compensation. From an economic perspective, considering the self-interest of all parties, without incentives, they have no motivation to participate in the FL process. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this invention provides a wireless charging federated learning method based on a three-stage Stackelberg game. This method addresses the energy constraints and individual selfishness issues of terminal devices in federated learning, incentivizing all roles to participate in the system and ensuring the successful completion of the federated learning task. The method first designs a framework of base station-terminal device-wireless charging service provider. The base station issues the federated learning task, aiming to obtain a better federated learning model at a lower cost. The terminal device trains a local federated learning model, hoping to obtain more rewards with less energy consumption. The charging service provider transfers energy to the terminal device via wireless charging during model training and uploading, while charging a fee. Then, utility formulas are designed for each of the three roles. To obtain the optimal strategies for all roles, the proposed game problem is analyzed using backward induction, and the unique existence of Stackelberg equilibrium and Nash equilibrium is proved. Finally, the Lagrange subgradient method is used to obtain the approximate optimal solution for the base station. The method proposed in this invention can effectively incentivize all roles to participate in the framework, thereby solving the energy constraint problem of terminal devices in federated learning, and simultaneously addressing the individual rationality problem of all roles.
[0005] The technical solution adopted by this invention to solve its technical problem is as follows:
[0006] Step 1: Construct utility functions for BS, EW, and CSP based on their respective parameters, and define the three-stage Stackelberg game problem;
[0007] Step 2: Use backward induction to find the optimal solutions for BS, EW, and CSP; prove that there is a unique Stackelberg game equilibrium between CSP and EW, a unique Nash equilibrium between EW, and a unique Stackelberg game equilibrium between EW and BS.
[0008] Step 3: Approximate the optimal solution of BS using the Lagrange subgradient method.
[0009] Furthermore, the parameters of the BS include the BS satisfaction parameter, the BS trade-off parameter for the delay in completing the federated learning task, and the reward parameter paid by the BS to the EW; the parameters of the EW include the EW's federated learning contribution parameter, network effect satisfaction parameter, and unit bid parameter for purchasing energy; the parameters of the CSP include the energy selling parameter, energy cost parameter, and energy conversion coefficient parameter.
[0010] Furthermore, in step 1, the utility function of the constructed BS is expressed as:
[0011]
[0012] Where: the first term on the right is the definition of BS's satisfaction with the local EW model, η1,η2>0 are BS's satisfaction parameters, which depend on BS's requirements for the accuracy of the global federated learning model and the local model; the second term is BS's trade-off on global aggregation latency, which depends on the EW with the slowest local training, λ>0 is BS's trade-off parameter on the completion latency of the federated learning task, T n The third term represents the total latency of EW n in a full round of FL; the fourth term is the sum of the federated learning task rewards paid by BS to all EWs, I n δ is the parameter representing the compensation paid by BS to EW n. n It represents the contribution of EW n to the federated learning model; EW n represents the terminal device n; N represents the total number of all EWs;
[0013] The utility function of BS must satisfy the constraint that the payment strategy of BS to EW is within the set range;
[0014] Furthermore, in step 1, the utility function of the constructed EW is expressed as:
[0015]
[0016] Where: the first item on the right is the payment BS made to EW n, I n The first term is the unit price of compensation; the second term is the satisfaction that EW obtains from network effects, ζ. n It is the network satisfaction weight value of EW n, φ mn This is the social relationship value between EWn and EWm, which depends on the data similarity between EWm and EWn, as well as the historical degree of cooperation between EWm and EWn. This information is stored in the BS (Background Object Model). m≠n, Denotes the set of EW; δ m The third term represents the federated learning contribution of EWm; the fourth term represents the energy resources purchased by EWn from the CSP, where s n EWn is the unit price offered for purchasing energy; E n This represents the total energy consumption of EWn in a complete round of global FL.
[0017] The utility function of EW must satisfy the constraint that the payment strategy of EW to CSP is within the set range.
[0018] Furthermore, in step 1, the utility function of the constructed CSP is expressed as:
[0019]
[0020] Where: the first item on the right is the energy payment received by the CSP from the EW; the second item is the CSP's energy resource cost; due to energy transmission losses, the CSP calculates the cost energy value as E. n / μ n μ n ∈[0,1] represents the energy conversion coefficient transmitted from CSP to EW; here, a≥0 and b≥0 are energy cost parameters;
[0021] The utility function of the CSP must satisfy the constraint that the energy strategy sold by the CSP to the EW is within a set range.
[0022] Furthermore, the three-stage Stackelberg game problem is expressed as:
[0023] In the first phase, BS acts as the leader and EW as the follower; BS controls the reward vector paid to EW. To maximize one's own utility formula In the second phase, EW engages in a game, with EW n following BS's payoff unit price strategy I. n Other EW's federated learning contribution decisions determine their own s n To maximize its own utility, the energy unit price vector of all EWs is defined as follows: In the third phase, EW and CSP engage in a game to ultimately determine the energy vector ε for CSP across all EW wireless transmissions, where ε = {E1,...,E...} n ,...,E N} T Therefore, the three-stage Stackelberg game Ω is represented as:
[0024] Ω={(BS,EW,CSP),(I n ,s n E n ),(U BS U n U CSP )}.
[0025] Furthermore, in step 2, the method of backward induction to find the optimal solutions for BS, EW, and CSP refers to first finding the optimal solution for CSP, then finding the optimal solution for EW, and finally finding the optimal solution for BS.
[0026] Furthermore, in step 2, there exists a unique Stackelberg game equilibrium between CSP and EW:
[0027] when At that time, there exists a unique Stackelberg game equilibrium between CSP and EW, where ε is the policy solution of CSP, and ε *It is the optimal strategy solution for CSP. and These are the optimal policy sets for BS and EW, respectively.
[0028] Furthermore, in step 2, there exists a unique Nash equilibrium between EW:
[0029] When there is one and only one Nash equilibrium strategy in EW, i.e. At this point, there exists a utility function. in It is the optimal strategy for other EW (Enhanced Warp) strategies. This represents the optimal federated learning contribution strategy for EW. This represents EW's optimal energy purchase price strategy.
[0030] Furthermore, in step 2, there exists a unique Stackelberg game equilibrium between BS and EW:
[0031] when At that time, there exists a unique Stackelberg game equilibrium between BS and EW, where: This represents the set of reward price strategies for BS. This represents the set of optimal reward price strategies for BS. Let ε represent the optimal policy set for all EWs. * This represents the optimal strategy set for CSP.
[0032] Furthermore, in step 3, the Lagrange subgradient method is: introducing dual variables to transform solving the original problem into solving the dual problem.
[0033] The beneficial effects of this invention are as follows:
[0034] (1) This invention solves the problem of limited energy of terminal devices in federated learning, enabling terminal devices to fully participate in the federated learning system, thereby accelerating the training of federated learning.
[0035] (2) The introduction of wireless charging service providers complicates the interaction between base stations, terminal devices, and wireless charging service providers. This invention analyzes their interactions using a three-stage Stackelberg game, fully considering the individual rationality of all roles. This ensures that all roles fully participate in the wireless charging federated learning system. Attached Figure Description
[0036] Figure 1 This is a flowchart of the method of the present invention;
[0037] Figure 2 A schematic diagram of wireless charging federated learning provided in an embodiment of the present invention;
[0038] Figure 3 A schematic diagram of a three-stage Stackelberg game provided for embodiments of the present invention;
[0039] Figure 4 Convergence performance diagram of the three-stage Stackelberg game provided in the embodiments of the present invention;
[0040] Figure 5 The convergence performance diagram of the Lagrange subgradient method provided in the embodiments of the present invention is shown.
[0041] Figure 6 The utility variation diagram of each method under different EW numbers provided in the embodiments of the present invention;
[0042] Figure 7 This is a social welfare analysis diagram of various methods under different EW quantities provided in the embodiments of the present invention;
[0043] Figure 8 The global loss variation diagram of the WRN (Wide Res-Net) model on the CIFAR-10 dataset provided in this embodiment of the invention;
[0044] Figure 9 The graph shows the global accuracy variation of the WRN model provided in this embodiment of the invention on the CIFAR-10 dataset. Detailed Implementation
[0045] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0046] The large-scale deployment of IoT devices and smart applications has generated massive amounts of data, demonstrating the enormous potential for training effective machine learning models. As a powerful machine learning method, deep reinforcement learning has achieved remarkable results in many tasks such as image recognition, natural language processing, and intelligent recommendation. However, centralized data training models pose a serious risk of privacy breaches due to the privacy implications of the data. Furthermore, the transmission of raw data consumes significant communication resources. FL, as a distributed machine learning paradigm, offers a promising solution to these problems. It allows the EW (External Web Controller) to participate in the learning process and train the model while maintaining the data locally, significantly reducing the risk of privacy breaches. Since only the model parameters are offloaded to the cloud server, the data privacy of the EW is maintained. Simultaneously, the size of the model parameters is typically much smaller than the raw data, which also reduces the size of data transmission.
[0047] However, training complex tasks quickly depletes the energy of the EW (Extended Flow) server, posing a challenge to obtaining a high-performance FL (Flexible Flow) model. Thanks to advancements in radio frequency (RF) technology, the WPT (Cloud Power Transfer) can transfer energy to the EW to address the energy constraint problem. Specifically, after the BS (Base Station) distributes the FL model as a cloud server, the EW can use a dedicated frequency to obtain energy from the CSP (Cloud Service Provider) via the WPT for training and model transfer. However, without incentives, the EW is unwilling to use its own computing resources to train the FL model for the BS. The CSP incurs costs when transferring energy, while the BS desires a better FL model for less compensation. From an economic perspective, considering the self-interest of all parties, without incentives, they have no motivation to participate in the FL process.
[0048] To address the aforementioned issues, this invention employs a three-stage Stackelberg game to incentivize EW, BS, and CSP to participate in the FL system, and uses backward induction to analyze their optimal solutions. Furthermore, the optimal solution for BS is obtained using the Lagrange subgradient method.
[0049] Figure 1 This diagram illustrates a federated learning method for wireless charging based on a three-stage Stackelberg game, as provided in an embodiment of the present invention. It consists of a Base Set (BS), an Edge Builder (EW), and a Concurrent Spread Provider (CSP). Let N = {1, ..., n, ..., N} represent the set of EWs. The BS needs to train a single FL task, and each EW has its own private dataset. The CSP has ample energy. To incentivize the EW to train a well-developed local model, the BS pays corresponding rewards based on its contribution. However, due to limited energy, the EW must purchase energy from the CSP to complete local training and model transfer. Throughout the learning process, the EW can continue to collect energy from the CSP on a dedicated frequency band outside the communication band.
[0050] The energy consumption of the federated learning process in EW mainly consists of computation and communication. The computation delay for one iteration of EW n in one training round is expressed as:
[0051]
[0052] Where c n The number of CPU cycles required to train with 1 bit of sample data, d n f is the size of the dataset owned by EW n. n The computing resources owned by EW n.
[0053] After training, the EW uploads the model parameters over the wireless channel using orthogonal frequency division multiple access. The transmission rate from EW n to BS is expressed as:
[0054] R n =W n log2(1+pn h n / W n σ) (2)
[0055] Among them W n This is the bandwidth allocated by BS to EW n, p n It is the transmission power of EW n, h n σ represents the channel gain between EWn and BS, and σ represents the Gaussian noise power.
[0056] Then, the transmission delay between EW n and BS can be calculated as follows:
[0057]
[0058] Where g n This represents the size of the EW transmission model.
[0059] Ignoring the aggregation latency of the federated model and the latency of the distributed global model, the total latency of EWn in a full round of FL can be expressed as:
[0060]
[0061] Where τ n This represents the number of iterations of EW n.
[0062] Accordingly, the computational energy consumption of one local iteration of EW n is:
[0063]
[0064] Where ρ n It is the chipset capacitance coefficient of EW n.
[0065] Furthermore, the communication energy consumption of EW n can be expressed as:
[0066]
[0067] Therefore, the total energy consumption of EWn in a complete FL cycle can be expressed as:
[0068]
[0069] EW can obtain energy from CSP during model training and uploading. Therefore, the maximum energy that EWn can obtain is expressed as:
[0070]
[0071] Where 0≤μ n ≤1 is the energy conversion coefficient. This is the transmission power allocated by the CSP to EW n, h csp,nIt is the channel gain between CSP and EW n.
[0072] Therefore, the energy that EW can purchase satisfies the following constraints:
[0073]
[0074] Next, the utility functions for EW, BS, and CSP are given.
[0075] For EWn, its utility function is the reward obtained from BS and the satisfaction derived from network effects, minus the compensation given to CSP. Before giving the utility formula, we first present a measure of EWn's federated learning contribution. The number of iterations τ n With local model accuracy θ n The relationship between them is τ n =τ0log[1 / (1-θ)] n )].here, It is a constant, representing the lower bound of local iterations when the model reaches convergence ∈ 0. Furthermore, It is a constant that depends on the loss function of the FL model. The upper limit of accuracy achievable by EWn during local training is... Parameter ι、 Let and k represent the minimum error, learning rate, and decay rate required to achieve maximum accuracy, respectively. It can be seen that EW requires more iterations and a larger dataset to obtain better local model parameters. Therefore, this invention uses the number of iterations τ. n and dataset size d n As a measure of FL contribution to EW n, that is:
[0076]
[0077] in These are control parameters.
[0078] Therefore, the utility of EW n is expressed as:
[0079]
[0080] The first item is the compensation paid by BS to EW n, I n The first term is the unit price of compensation; the second term is the satisfaction that EW obtains from network effects, ζ. n It is the network satisfaction weight value of EW n, φ mn It is EW n and EW m The social relationship value between EWm and EWn depends on the data similarity between EWm and EWn and the historical closeness of their cooperation; this information is stored in BS; the third item is the energy resources that EWn purchases from CSP, where s n It is the unit price for purchasing energy in EW n.
[0081] For BS, a trade-off between global model quality and global aggregation latency needs to be considered, and it is defined as follows:
[0082]
[0083] The first term is the definition of BS's satisfaction with the local EW models, where η1,η2>0 depends on BS's requirements for the accuracy of the global federated learning model and the local models. The second term is BS's trade-off regarding global aggregation latency, mainly depending on the EW with the slowest local training, where λ>0 determines BS's latency trade-off for the federated learning task. The third term is the sum of the federated learning task rewards paid by BS to all EWs.
[0084] For CSP, its utility function is the energy fee collected from EW minus the energy cost, that is:
[0085]
[0086] The first item is the energy payment received by the CSP from the EW, and the second item is the CSP's energy resource cost. Due to energy transmission losses, the CSP calculates the cost energy value as E. n / μ n Here, a≥0 and b≥0 are energy cost parameters.
[0087] Figure 3 The game process between three entities is described. In the joint optimization problem, this invention considers maximizing the utility functions of BS, EW, and CSP. To determine the optimal strategy, the solution is divided into three stages: In the first stage, BS acts as the leader and EW as the follower; BS controls the reward vector paid to EW. To maximize one's own utility formula, where Subsequently, in the second phase, EW engages in a game, with EW n making decisions based on BS's reward system. n The decision of other EWs determines their optimal energy unit price s n To maximize its own utility, the energy unit price vector of all EWs is defined as follows: Although the variable parameter that determines the model quality is the number of iterations τ n But if EW n wants to control τ n It controls the energy E purchased from the CSP n τ can then be determinedn Since the transmission energy consumption is constant, purchasing more energy allows for more iterations. Finally, in the third stage, the EW and CSP engage in a game to determine the CSP's energy vector ε for all EW wireless transmissions, where ε = {E1,...,E...} n ,...,E N} T Therefore, the three-stage Stackelberg game Ω can be represented as:
[0088] Ω={(BS,EW,CSP),(I n ,s n E n ),(U BS U n U CSP (14) The goal is to maximize the utility of all three parties. The optimization problems for BS, EW, and CSP are given by the following equation:
[0089]
[0090]
[0091]
[0092] C1, C2, and C3 indicate that BS's payment strategy, EW's pricing strategy, and CSP's energy sales strategy are all within a certain range.
[0093] The proposed game is analyzed using backward induction. First, the optimal energy selling strategy of CSP is solved. Then, the existence of a unique Nash equilibrium is proved, and the optimal energy unit price strategy of EW is given. Finally, the optimal payment strategy of BS is given.
[0094] First, we define the Stackelberg game equilibrium between CSP and EW, and then we give the optimal solution for CSP.
[0095] Definition 1: When At that time, there exists a unique Stackelberg game equilibrium between CSP and EW, where and These are the optimal strategies for BS and EW, respectively.
[0096] Theorem 1: For a given unit price strategy s of EW n n CSP's optimal energy sales strategy It can be represented as:
[0097]
[0098] Proof: Based on the utility of CSP, the first-order partial derivative can be obtained as follows:
[0099]
[0100] Based on formula (19), U can be obtained. CSP Regarding E n The second derivative:
[0101]
[0102] A second-order partial derivative less than 0 means that U CSP Relative to E n The utility function is convex upwards. Therefore, it can be seen that making U... CSP E with first derivative equal to 0 n The optimal value is obtained from formula (18).
[0103] For constraint C3, according to formula (8), the maximum sales energy of CSP can be determined. With T n Related, and T n The control variable τ of EWn n Therefore, constraint C3 is satisfied when solving for the optimal strategy of EW.
[0104] Next, we prove the existence and uniqueness of the EW Nash equilibrium. Then, we solve for the optimal solution of EW. Finally, we give the s that satisfy C3. n constraint.
[0105] Definition 2. There exists one and only one Nash equilibrium strategy in EW, that is At this point, there exists a utility function. in It is the optimal strategy for other EWs.
[0106] Note: From formulas (7), (10) and (18), it can be seen that δ n With s n Therefore, EW n only needs to control s n This allows control of δ n .
[0107] When a Nash equilibrium exists, neither party can increase its utility by unilaterally changing its strategy. Before proving the Nash equilibrium, an assumption needs to be made.
[0108] Assumption 1: First, we give the assumption of network effects, namely... in
[0109] Note. Assumption 1 ensures that all EW increases by s. nAll have negative marginal utility. If this assumption does not hold, then all EW can increase infinitely by s. n To gain more from network effects without considering I n Therefore, this assumption is made to ensure s n It is bounded and more in line with reality. Based on assumption 1, we can obtain the following Theorem 2.
[0110] Theorem 2: Considering a dynamic strategy with a fixed number of EWs, for each EW, its utility function satisfies formula (11). Therefore, there exists a unique Nash equilibrium point for each EW.
[0111] Proof: First, we need to substitute the optimal solution of CSP. We will rewrite the utility function of EW and then use the properties of the Hessian matrix to prove the existence of a unique Nash equilibrium.
[0112] According to formulas (7) and (18), when EW n purchases energy E n When determined, the number of iterations τ n for:
[0113]
[0114] Then, by rewriting the utility function of EW n and substituting formulas (18) and (21) into formula (13), we can obtain:
[0115]
[0116] The Hessian matrix H can be used to determine whether a Nash equilibrium exists. H is derived from U. n right The second-order partial derivatives are obtained. First, EWn is calculated for equation (22) with respect to s. n The first-order partial derivatives of are:
[0117]
[0118] Based on formula (23), then U n For s n The second-order partial derivatives can be obtained as follows:
[0119]
[0120] Similarly, based on formula (23), We can obtain:
[0121]
[0122] Therefore, the Hessian matrix can be represented as:
[0123]
[0124] in and Based on assumption 1, we can obtain -H can be proven to be strictly diagonally dominant positive definite, and correspondingly, H is strictly convex diagonally. Therefore, it can be proven that there exists a unique Nash equilibrium energy-based bidding strategy.
[0125] Next, we will give the optimal solution for EW.
[0126] Theorem 3. For a given BS unit payment strategy I n The optimal closed-form solution of EW n It can be represented as:
[0127]
[0128] Proof: According to It can be known that U n For s n It is convex upwards, therefore U n The first partial derivative of s is equal to 0 n To achieve the optimal value, let each EW... To obtain the optimal energy unit price, i.e., formula (27).
[0129] Finally, s is given. n A constraint that makes the calculated Satisfying C3. According to the maximum transferable energy constraint formula (9), combined with formulas (4), (7), (8) and (23), we can obtain s. n The constraints are:
[0130]
[0131] when When , it means that the energy gained during EW training is greater than or equal to the energy consumed, therefore formula (9) must be satisfied. Conversely, EW n needs to consider the following constraints:
[0132]
[0133] For the sake of simplicity, use represent
[0134] when When, the value of the closed-form solution It must satisfy constraint formula (29) and C2, that is, if The value is greater than the constraint value and Then the value of EW n can only take the following values: and The smaller one. Accordingly, when When, closed-form solution Only C2 needs to be satisfied.
[0135] Therefore, the optimal solution can be obtained. for:
[0136]
[0137] The optimal solution for BS is given next, followed by the definition of the Stackelberg game equilibrium between BS and EW.
[0138] Definition 3. When At that time, there exists a unique Stackelberg game equilibrium between BS and EW, where and ε * These are the optimal strategies for EW and CSP, respectively.
[0139] According to formula (29), we know with I n Irrelevant. Therefore, in different The methods for obtaining the optimal solution of BS are different. In formula (30) The solutions for the two cases are similar, therefore the more complex solution is chosen. To solve the subsequent problems, the simpler solutions are similar. Next, based on... Sub-case discussion of optimal unit payment strategy
[0140] (1)
[0141] exist In the case of s n with I n It is irrelevant, therefore, according to formulas (10) and (21), δ n with I n If it is irrelevant, then we can also know T n with I n Irrelevant. Therefore, U BS Rewritten as:
[0142]
[0143]
[0144]
[0145] in because It is the minimum value, that is Therefore, we can obtain C4.
[0146] It can be seen that U BS Compared to I n Monotonically decreasing, therefore the optimal unit payment price for:
[0147]
[0148] (2)
[0149] Similar to case (1), U BS Compared to I n Monotonically decreasing, optimal unit payment strategy Calculated by the following formula:
[0150]
[0151] (3)
[0152] First, substitute the closed-form solution of EW. We will rewrite the utility formula of BS and then give the solution of BS in this subcase.
[0153] According to formulas (21), (27) and achievable Then, the data contribution δ of EW n n It can be rewritten as:
[0154]
[0155] in
[0156] Next, δ n Substituting into formula (12), U BS It can be rewritten as:
[0157]
[0158] To make formula (35) easier to handle, the following substitutions are made:
[0159]
[0160] Therefore, the utility function of BS is transformed into:
[0161]
[0162]
[0163]
[0164]
[0165] According to formula (36), t is greater than all T. n Therefore, we can obtain C5. Then, combining... From (27), we can obtain C6.
[0166] It can be seen that U BS For I n It is convex upwards. As a method for solving constrained mathematical optimization problems, the Karush-Kuhn-Tucker (KKT) approach is suitable for solving the formula, then U... BS The Lagrange function can be expressed as:
[0167]
[0168] Where α n ,β n ,γ n and ψ n , is a Lagrange multiplier associated with C1, C5 and C6.
[0169] Calculation formula (36) for I n The first-order partial derivatives are:
[0170]
[0171] make We can obtain:
[0172]
[0173] To express formula (38) more concisely, we use the symbols C, D, and F, which are represented as follows:
[0174]
[0175]
[0176]
[0177] Since directly eliminating the Lagrange multiplier is difficult to solve convex optimization problems, the Lagrange dual problem of the primal problem is solved. Derive the Lagrange dual function v(α) n ,β n ,γ n ,ψ n ):
[0178]
[0179]
[0180] Since formula (36) is convex upwards, it is easy to find feasible points (I) that satisfy constraints C1, C5, and C6. n Since the dual function v is not differentiable, the Lagrange subgradient method is used to solve the function v. The variable α n ,β n ,γ n and ψ n Update using the gradient ascent method with the following equation:
[0181]
[0182]
[0183]
[0184]
[0185] Where q k It is the step size of the k-th Lagrange iteration.
[0186] when and Once determined, I is updated using formula (38). n Formula (36) updates t. Variable I n ,t and Lagrange multipliers and Continuously update until convergence or the maximum number of iterations is reached.
[0187] Simulation experiments demonstrate that the deep reinforcement learning-based method proposed in this invention significantly reduces user energy consumption compared to other benchmark methods in various scenarios. The technical effectiveness of the above solution is then verified using specific experimental data.
[0188] Example:
[0189] In the experiment, this invention considers an ESP system with a set of RSUs and vehicles traveling on the road. Both the RSUs and VWs possess computing resources, with the RSUs having greater computing resources than the VWs. In each time slot, multiple VUs with computing tasks and VWs possessing computing resources are randomly distributed within the coverage area of the RSUs.
[0190] In the experiment, this invention considers a wireless charging federated learning method that includes BS, CSP and EW.
[0191] The embodiments of this invention compare the performance of the proposed method with other benchmark methods in a multi-user scenario: Through simulation experiments, the following benchmark testing methods are used to verify the rationality and effectiveness of the proposed method:
[0192] (1) No CSP Utility Method (NCUM): NCUM does not consider the utility of CSP, meaning that CSP will fully meet the energy needs of EW. The game exists only between BS and EW.
[0193] (2) Resource Preference Method (RPM): The unit price is determined based on the computing resources and datasets owned by the EW, that is, the better the type of EW, the higher the reward.
[0194] (3) Uniform Method (UM): The unit payment price is uniform for all EWs.
[0195] (4) Random Method (RM): The unit price of all EWs is random.
[0196] Without loss of generality, the following are the main assumptions used in the simulation: the training data size d of EW n Set to [25, 35] MB, calculate the number of CPU cycles c required to compute 1 bit of data. n Set to [10, 30] cycles. Computational resource f n Set to [1.5, 2.5] GHz, transmission power p n The bandwidth is set to [0.1, 0.2]W. The channel gain model for EW and BS / CSP is a path loss model, i.e., 32.44 + 20lg(dist), where dist is the distance between EW and BS / CSP, distributed between [0, 150]m. The communication bandwidth W... n Set to 1MHz, noise power spectral density Set to -174dBm / Hz
[45] . Power transferred from CSP to EW Set it to 0.5W.
[0197] Figure 4 and Figure 5 The convergence performance of the proposed method is described. The number of EWs is set to 20. Figure 4 As shown, the utilities of BS, CSP, and EW change as the game progresses and converge after approximately five rounds. Figure 5The convergence performance of the Lagrange subgradient method for finding the optimal solution of Black-Scholes is described. It can be seen that the method converges after approximately 15 rounds. Therefore, the good convergence performance of the proposed method is verified.
[0198] Figure 6 The utility of BS, CSP, and EW under different numbers of EWs is shown. When the number of EWs increases from 12 to 32, the energy required to be purchased increases, thus increasing the utility of CSP. Then, as the number of EWs increases, the marginal utility of BS decreases. This is because the marginal improvement in the accuracy of the FL model decreases as the number of iterations of EWs increases. Finally, the average utility of EWs gradually decreases. This is because, as can be seen from Equation (12), the satisfaction of BS does not increase linearly, leading to more intense competition when the number of EWs increases. In addition, as shown in 7, compared with the other four benchmarks, the proposed method can improve social welfare by 23.87%, 19.09%, 35.37%, and 51.86%. First, NCUM does not consider the utility of CSP, and the energy purchased by EWs is not the optimal solution for CSP, so its social welfare is less than that of the proposed method. In addition, the social welfare of RPM increases first and then decreases, mainly because the marginal satisfaction of BS decreases as the number of EWs increases. However, the unit payment price of RPM increases linearly with the resources owned by EWs. Finally, due to unreasonable incentives, the social welfare of UM and RM is low.
[0199] Figure 8 and Figure 9 A comparison of FL performance under different methods is presented. The superiority and rationale of the proposed method are evaluated using a WRN model on the CIFAR-10 dataset. The number of exponents (EW) is set to 20 in both FL tasks. The WRN model is trained on the CIFAR-10 dataset, which consists of a training set of 50,000 color images and a test set of 10,000 images. Figure 8 and Figure 9 It can be seen that the convergence performance of this method is only slightly lower than that of NCUM. However, NCUM achieves better FL performance at the expense of CSP utility. It does not adequately incentivize CSP participation in the wireless power supply FL framework, therefore the NCUM method cannot solve the energy-constrained problem of EW.
Claims
1. A wireless charging federated learning method based on a three-stage Stackelberg game, characterized in that, Design a framework for base station-terminal device-wireless charging service provider, where the base station issues federated learning tasks; The terminal device trains a locally federated learning model; the charging service provider transfers energy to the terminal device via wireless charging transmission during the model training and uploading process, and charges a fee in return. Throughout the learning process, the EW continues to harvest energy from the CSP on a dedicated frequency band outside the communication band; The energy consumption of the EW's federated learning process mainly consists of computing and communication. The computation latency for one iteration in a training round is expressed as: in The number of CPU cycles required to train with 1 bit of sample data. for The size of the dataset possessed, for The computing resources available; After training is complete, EW uploads the model parameters over a wireless channel using orthogonal frequency division multiple access. The transmission rate to the BS is expressed as: in BS is allocated bandwidth yes Transmission power, yes Channel gain between BS and BS Indicates Gaussian noise power; Then, The transmission delay between the BS and the base station can be calculated as follows: in The size of the EW transmission model; Ignoring the aggregation latency of the federated model and the latency of the distributed global model; therefore, The total delay in a full round of FL is expressed as: in express The number of iterations; Accordingly, The computational energy consumption of one local iteration is: in yes The chipset capacitance coefficient; Furthermore, Communication energy consumption is expressed as: Therefore, in a complete round of FL The total energy consumption is expressed as: EW can obtain energy from CSP during model training and uploading; therefore, The maximum energy obtained is expressed as: in It is the energy conversion coefficient. It is CSP Allocated transmission power, Is CSP and Channel gain between; Therefore, the energy that EW can purchase satisfies the following constraints: Includes the following steps: Step 1: Construct utility functions for BS, EW, and CSP based on their parameters, and define the three-stage Stackelberg game problem; the constructed utility function of BS is expressed as: Where: the first item on the right is BS's definition of satisfaction with the EW local model, The first parameter represents the satisfaction level of the BS (Browser Support Builder), which depends on the BS's requirements for the accuracy of the global model and the local model in federated learning; the second parameter represents the BS's trade-off regarding global aggregation latency, which depends on the EW (Execution Workflow) which is the slowest in local training. For the trade-off parameter between BS and the delay in completing federated learning tasks, express The total latency throughout a full round of FL; the third item is the sum of the federated learning task rewards paid by BS to all EWs. It was BS that paid to The reward parameter, yes The contribution of the federated learning model; Indicates terminal device ; This represents the total number of all EWs; The utility function of BS must satisfy the constraint that the payment strategy of BS to EW is within the set range; The utility function of the constructed EW is expressed as: Among them: the first item on the right is the payment made by BS to... The reward, The first is the unit price of compensation; the second is the satisfaction that EW obtains from network effects. yes The weight value of network satisfaction, yes and The social relationship value between them depends on and Data similarity and historical and The degree of closeness of cooperation, this information is stored in BS, , Represents the set of EW; express The third item is EW's contribution to federal learning; Energy resources purchased from CSP, of which yes The unit price for purchasing energy; Indicates the current complete round of global FL Total energy consumption; The utility function of EW must satisfy the constraint that the payment strategy of EW to CSP is within the set range; The utility function of the constructed CSP is expressed as: Where: the first item on the right is the energy payment received by the CSP from the EW; the second item is the CSP's energy resource cost; due to energy transmission losses, the CSP calculates the cost energy value as follows: , This represents the energy conversion factor from CSP to EW; here, and It is an energy cost parameter; The utility function of the CSP must satisfy the constraint that the energy strategy sold by the CSP to the EW is within a set range. The three-stage Stackelberg game problem is expressed as: In the first phase, BS acts as the leader and EW as the follower; BS controls the reward vector paid to EW. To maximize one's own utility formula In the second phase, EW engages in a game of strategy. According to BS's compensation unit price strategy Other EW's federated learning contribution decisions determine their own To maximize its own utility, the energy unit price vector of all EWs is defined as follows: In the third phase, EW and CSP engage in a game to ultimately determine the energy vector of CSP for all EW wireless transmissions. ,in Therefore, the three-stage Stackelberg game Represented as: Step 2: Use backward induction to find the optimal solutions for BS, EW, and CSP; prove that there is a unique Stackelberg game equilibrium between CSP and EW, a unique Nash equilibrium between EW, and a unique Stackelberg game equilibrium between EW and BS. Step 3: Approximate the optimal solution of BS using the Lagrange subgradient method.
2. The wireless charging federated learning method based on a three-stage Stackelberg game as described in claim 1, characterized in that, The parameters of the BS include the BS satisfaction parameter, the BS trade-off parameter for the delay in completing the federated learning task, and the reward parameter paid by the BS to the EW; the parameters of the EW include the EW's federated learning contribution parameter, network effect satisfaction parameter, and unit price parameter for purchasing energy; the parameters of the CSP include the energy selling parameter, energy cost parameter, and energy conversion coefficient parameter.
3. The wireless charging federated learning method based on a three-stage Stackelberg game as described in claim 1, characterized in that, In step 2, the reverse induction method for finding the optimal solutions for BS, EW, and CSP refers to first finding the optimal solution for CSP, then finding the optimal solution for EW, and finally finding the optimal solution for BS.
4. The wireless charging federated learning method based on a three-stage Stackelberg game as described in claim 3, characterized in that, In step 2, there exists a unique Stackelberg game equilibrium between CSP and EW: when At that time, there exists a unique Stackelberg game equilibrium between CSP and EW, where It is a CSP strategy solution. It is the optimal strategy solution for CSP. and These are the optimal policy sets for BS and EW, respectively.
5. The wireless charging federated learning method based on a three-stage Stackelberg game as described in claim 4, characterized in that, In step 2, there exists a unique Nash equilibrium between EW: When there is one and only one Nash equilibrium strategy in EW, i.e. At this point, a utility function exists. ,in It is the optimal strategy for other EWs. This represents the optimal federated learning contribution strategy for EW. This represents EW's optimal energy purchase price strategy.
6. The wireless charging federated learning method based on a three-stage Stackelberg game as described in claim 5, characterized in that, In step 2, there exists a unique Stackelberg game equilibrium between BS and EW: when At that time, there exists a unique Stackelberg game equilibrium between BS and EW, where: This represents the set of reward price strategies for BS. This represents the set of optimal reward price strategies for BS. This represents the optimal policy set for all EWs. This represents the optimal strategy set for CSP.
7. The wireless charging federated learning method based on a three-stage Stackelberg game as described in claim 6, characterized in that, In step 3, the Lagrange subgradient method is: introducing dual variables to transform solving the original problem into solving the dual problem.