Sampling rate and power optimization method based on single-cluster air federated edge learning system

By optimizing the transmit power and data sampling rate of edge devices in a single-cluster aerial federated edge learning system, the problems of energy consumption and data heterogeneity are solved, and efficient model training under energy-constrained conditions is achieved.

CN121367557APending Publication Date: 2026-01-20CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511398359.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively balance the impact of energy consumption of edge devices and data heterogeneity on model training, resulting in high device energy consumption and poor model training performance.

Method used

By combining the Lyapunov drift-penalty algorithm with a non-convex optimization method, the transmit power of edge devices and the sampling rate of training data are optimized, thus achieving an effective solution to the resource allocation problem.

Benefits of technology

Under energy-constrained conditions, it significantly improves model convergence performance and energy utilization efficiency, reduces equipment energy consumption, and enhances the accuracy and stability of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121367557A_ABST
    Figure CN121367557A_ABST
Patent Text Reader

Abstract

The invention discloses a sampling rate and power optimization method based on a single-cluster air federated edge learning system. According to the method, an optimization problem with the optimization objective of minimizing the difference between expected global loss and optimal loss after T times of communication is established, and an original optimization problem is converted into an approximate optimization problem containing an explicit formula. In an optimization problem solving process, long-term energy constraint is converted into queue stability constraint based on a Lyapunov optimization method, and an EARA algorithm is provided to realize online joint optimization of a data sampling rate and a transmission power scaling factor. Experimental results show that the EARA algorithm can significantly improve the convergence performance of the model under a relatively low energy budget.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of communication, in particular to a sampling rate and power optimization method based on a single-cluster air-federated edge learning system. BACKGROUND

[0002] In recent years, with the wide popularity of smart home, industrial Internet of Things and smart office and the like, a large number of edge devices (such as smart phones, sensors, smart home appliances and the like) continuously collect and generate massive data. However, these devices usually rely on battery power supply and face limited energy resources, and the collected data has significant heterogeneity in type, distribution and quality. Existing researches mainly focus on adopting power control strategies to reduce the communication errors due to poor channel quality of devices and minimize the mean square error as an optimization target. These methods fail to take into account the influence of device energy consumption and data heterogeneity on model training. Therefore, the application proposes a single-cluster air-federated edge learning model training system, which comprehensively considers the local training of edge devices and the model transmission process, constructs an optimization problem under the long-time energy constraint, and aims to narrow the gap between the global expected loss and the optimal loss. To solve this problem, the application proposes an energy-aware joint optimization of the transmission power and training data sampling rate of edge devices, so as to maximize the system training energy under the energy constraint. SUMMARY

[0003] The application aims to overcome the defects in the prior art and proposes a sampling rate and power optimization method for a single-cluster air-federated edge learning system, which adopts a Lyapunov drift plus penalty algorithm combined with a non-convex optimization method to effectively solve the resource allocation problem.

[0004] The technical scheme adopted by the application is a sampling rate and power optimization method based on a single-cluster air-federated edge learning system, which comprises the following steps:

[0005] An air-federated edge learning system model is constructed, and an optimization problem is established to minimize the gap between the expected global loss and the optimal loss after T times of communication, which is represented as:

[0006]

[0007] indicates the sampling rate of the data sample for label m in the edge device k, indicates a power control coefficient, a constant C1 and a constant B. indicates an edge device index set, K indicates the number of edge devices, and M indicates the number of sample labels indicates the total amount of data samples with a value of m, p mdenotes the global data sample distribution with label m, each device transmits d total analog symbols corresponding to its model size denotes the noise power, denotes the parameters of the local model, denotes the transmission power budget of edge device k, the total communication rounds are T, and the index set of the communication rounds is denoted as denotes the total energy consumption of edge device k in the t-th communication round, H k denotes the energy budget.

[0008] where constraint C1 limits the transmit power of each edge device to not exceed its maximum transmit power; constraint C2 requires that the total energy consumption of each edge device for local training and wireless communication within T rounds cannot exceed its energy budget H k ; constraint C3 specifies the value range of the data sampling rate;

[0009] Based on the Lyapunov optimization method and the upper limit of the gap between the expected global loss and the optimal loss in the convergence analysis, the optimization problem is converted into a new optimization problem;

[0010] The converted new optimization problem is decomposed into two sub-problems, which are iteratively solved by alternating optimization; online joint optimization of the data sampling rate and the transmission power scaling factor is realized.

[0011] Further, the aerial federated edge learning system includes K edge devices and 1 edge server, and the edge devices and the edge server are connected through a wireless network to collaboratively train a shared global deep learning model.

[0012] In the aerial federated edge learning system, after sampling, the training data subset obtained by the edge device k is and the new data distribution is calculated by the following formula respectively:

[0013]

[0014] Based on the data distribution Q k , the loss function of the edge device k is:

[0015]

[0016] denotes the expectation based on the data distribution Q k , and denotes the indicator function, f(w,x i ) denotes the loss function of the global deep learning model, x i denotes the feature of the i-th data sample of the edge device, and y idenotes a data sample label, denotes the label y i denotes the expectation under the condition m, denotes the model parameter vector of the edge device k, denotes a d-dimensional real number space, each device transmits a total of d analog symbols corresponding to the model size thereof.

[0017] An energy consumption model is also constructed in the air federal edge learning system model, and the total energy consumption of the edge device k in the tth communication round is represented as:

[0018]

[0019] N F denotes the number of floating point operations required to process one data sample, f k denotes the CPU frequency of the edge device k, n F denotes the number of floating point operations per cycle, E denotes the number of iterations of local training, denotes the power control coefficient, denotes the model parameter vector, denotes the computational energy consumption of the edge device k in the tth communication round, denotes the communication energy consumption of the edge device k in the tth communication round.

[0020] Further, the optimization problem is converted into a new optimization problem, which is represented as:

[0021]

[0022] denotes the Lyapunov drift, V≥0 is a weight parameter, denotes the aggregation weight of, denotes the sampling rate of the edge device k for the data sample of the class j.

[0023] The application also provides a single-cluster air federal edge learning system based on the single-cluster air federal edge learning system, comprising K edge devices and 1 edge server, the edge devices and the edge server are connected through a wireless network, the edge devices and the edge server store a computer program, and the computer program is executed by a processor to realize the steps of the sampling rate and power optimization method of the single-cluster air federal edge learning system.

[0024] The application finally provides a computer readable storage medium storing executable instructions for being executed by a processor to realize the sampling rate and power optimization method of the single-cluster air federal edge learning system.

[0025] The application has the following beneficial effects:

[0026] To realize the effective utilization of system resources, the present application mainly studies the joint communication and computing resource optimization problem of energy-constrained edge devices participating in federated learning training in the data heterogeneous environment. Specifically, to deeply analyze the influence of energy constraints and data heterogeneity on the performance of federated learning, the present application theoretically derives the relationship between global loss and energy consumption, and transforms the original optimization problem into an approximate optimization problem containing an explicit formula. In the optimization problem solving process, the present application transforms the long-term energy constraint into a queue stability constraint based on the Lyapunov optimization method, and proposes an EARA algorithm to realize the online joint optimization of data sampling rate and transmission power scaling factor. Experimental results show that the proposed EARA algorithm can significantly improve the model convergence performance under lower energy budget. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 System model of federated learning in the air.

[0028] Figure 2 Comparison between the upper bound of the loss function derived on the MNIST dataset and the actual training loss of the EARA algorithm.

[0029] Figure 3 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm.

[0030] Figure 4 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm. Figure 5 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm.

[0031] Figure 6 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm. Figure 7 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm.

[0032] Figure 8 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm. Figure 9 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm.

[0033] Figure 10 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm. Figure 11 Comparison between the upper bound of the loss function derived on the CIFAR-10 dataset and the actual training loss of the EARA algorithm. DETAILED DESCRIPTION

[0034] In order to make the purposes, technical solutions and advantages of the present application clearer, the specific embodiments of the present application will be described in further detail below.

[0035] The example of the present application provides a sampling rate and power optimization method of a single-cluster air-federated edge learning system, which comprises the following steps:

[0036] Step 1: Construct a system model, wherein the air-federated edge learning model comprises a sample-level data sampling model and a communication model based on air computing, and an energy consumption model is further designed.

[0037] The air-federated edge learning system is composed of K edge devices and 1 edge server, and the devices are connected through a wireless network to collaboratively train a shared global deep learning model. Each device owns a local data set wherein k,i x k,i represents the features of the i-th data sample of the device k, y represents the label of the i-th data sample of the device k, represents the total amount of local data samples owned by the edge device k. All local data sets constitute a global data set

[0038] M classification task is one of the most common and widely used tasks in the field of machine learning, such as image classification, text classification, etc., which has a clear application background in smart home, industrial Internet of Things, etc. At the same time, the M classification task can well reflect the data heterogeneity challenge in federated learning, and provide clear evaluation indicators (such as classification accuracy) for algorithm performance. Therefore, the present application assumes an M classification task as the learning task of the air-federated edge learning system model. Data sample label For any edge device k, the total amount of data samples with label value m wherein is an indicator function. Due to the differences in the environment of the edge device, user behavior and data collection method, the data samples collected by different devices may have significant heterogeneity, and data distribution is a key indicator for measuring data heterogeneity. In the present application, the local data distribution vector of the edge device k is represented as wherein represents the distribution of data samples with label m in the device k, and is defined as follows:

[0039] and

[0040] In order to simulate the data heterogeneity environment, the present application assumes that each edge device has a different data distribution, i.e. The core goal of learning task is to optimize the model parameter vector by training on the dataset, so as to minimize the value of the loss function, and then improve the estimation performance of the model parameters. The loss function of the model defined in the application is f(w,x i ), the loss function F k of the edge device k participating in the training is represented as

[0041]

[0042] In the formula, w is the model parameter vector of the edge device k; represents a d-dimensional real number space, and each device sends a total of d analog symbols corresponding to the model size thereof. Federated edge learning needs to continuously aggregate the local model updates of participating devices to obtain the optimal global model. Therefore, the global loss function F can be defined as:

[0043]

[0044] In the formula, ρ k represents the aggregation weight, and the optimal model parameter vector w * is obtained by minimizing the global loss function F. Therefore, the optimization objective of the over-the-air federated edge learning system training can be represented by the following formula:

[0045]

[0046] In order to solve the above formula, the entire training process needs to consider the performance loss caused by the influence of the local calculation stage and the communication stage of model transmission and aggregation. In addition, due to the energy limitation of edge devices, the application assumes that the total communication round is fixed as T, and the index set thereof is represented as

[0047] In the local training stage, due to the significant difference in data distribution of different devices, if all the data of all devices is used for training, it may cause the optimization direction of each device to be inconsistent, thereby affecting the convergence performance of the global model. In order to solve this problem, the application adopts a sample-level data sampling strategy, that is, each device selects a part of data points from its local data set for training. This method not only can alleviate the influence of data distribution difference, but also can effectively reduce the calculation energy consumption, thereby improving the training efficiency while ensuring the model performance. Since the data sampling process adjusts the data class distribution, the application widely uses the cross-entropy loss function as the loss function of the model, and the loss function F k can be rewritten as:

[0048]

[0049] In the formula, f m(·) represents the prediction probability of the model for the real class m, and the value range is 0 < f m (·) ≤ 1. represents the mathematical expectation of the distribution P k represents the indicator function, represents the expectation under the condition that the label y i = m.

[0050] In the air federal edge learning system, the data sampling rate represents the sampling rate of the data sample for the class m in the edge device k. After sampling, the edge device k obtains the training data subset and the new data distribution can be calculated by the following formula respectively:

[0051]

[0052] Based on the data distribution Q k , the loss function of the edge device k can be rewritten as:

[0053] represents the expectation of the distribution Q k what, represents the indicator function.

[0054] In the local model updating and training stage, each edge device updates its current local model parameter by setting . Wherein, w t-1 represents the latest global model parameter vector before the start of the tth round of local training, represents the model parameter vector of the edge device k at the initial stage of the tth round. The model is broadcasted by the edge server to all edge devices. Based on the data subset the edge device updates w k by iteration. The present application assumes that each edge device performs E iterations for local model training. At the τth (τ ∈ E) iteration, the local model parameter vector update can be represented as:

[0055]

[0056] In the formula, α represents the learning rate, represents the gradient of the loss function with respect to the vector . After the training is completed, all edge devices upload their model parameter vectors ​To the edge server, the edge server can aggregate these parameter vectors to obtain a new round of global model representation as follows:

[0057]

[0058] With traditional orthogonal multiple access communication technology, the edge server needs to successfully decode all local models transmitted by the devices before model aggregation, which can cause large communication delay and bandwidth consumption. As can be seen from the following formula, the edge server is only interested in the weighted sum of local model parameters. Therefore, in order to speed up the learning speed and improve the spectrum utilization, the present application adopts over-the-air computing technology to utilize the superposition characteristics of multiple access fading channels for fast gradient aggregation. In order to simplify the analysis and not lose generality, the present application considers a frequency non-selective block fading channel model, in which the wireless channel remains unchanged in each outer iteration and can change at different iterations. In addition, it is assumed that each edge device knows its own channel state information completely, which allows them to compensate for the phase introduced by the wireless channel. At the same time, the edge server knows the global channel state information to facilitate power control. Specifically, represents the complex channel coefficient from the edge device k to the edge server in the tth round of communication. In order to alleviate the influence of the fading channel and compensate for the phase distortion, the transmission signal scaling factor can be calculated as follows:

[0059]

[0060] In the formula, represents the power control coefficient; is the complex conjugate of the channel coefficient. In addition, symbol-level synchronization at the edge server can be ensured by adopting the timing advance technology in 5G new radio, which adjusts the transmission time of each device according to the received timing advance technology command with reference to the public clock. Then the signal vector r received at the edge server t is expressed as:

[0061]

[0062] In the formula, n t represents the complex Gaussian noise vector received by the edge server. At the same time, the edge device has a transmission power budget Therefore, the transmission signal needs to satisfy the following constraint:

[0063]

[0064] represents the mathematical expectation.

[0065] The edge server processes the received signal by receiving a factor ζ t to obtain new global model aggregation parameters, denoted as:

[0066]

[0067] wherein The global model aggregation parameters w t+1 are rewritten as:

[0068]

[0069] wherein denotes the aggregation weight and satisfies n t = n t / ζ t denotes the equivalent noise.

[0070] Energy consumption is a key factor limiting edge devices to participate in learning tasks. Most edge devices (such as smartphones, Internet of Things devices) rely on battery power, and devices with insufficient energy may not be able to participate in federated learning, resulting in data loss or model bias, thereby affecting model performance. In aerial federated edge learning, the energy consumption of edge devices mainly comes from the computational energy consumption of local model training and the communication energy consumption of model updating. Among them, the computational energy consumption is closely related to the model complexity, the amount of training data, and the device hardware performance (such as CPU / GPU computing power); the communication energy consumption depends on the data volume, transmission power, and network type (such as Wi-Fi, 5G). Therefore, under the condition of limited energy budget, the computational energy consumption and the communication energy consumption need to be considered comprehensively.

[0071] In the aerial federated edge learning system, the computational workload of an edge device k is denoted as wherein E denotes the number of iterations of local training, N F denotes the number of floating-point operations required to process one data sample. The local computing capability of an edge device k is denoted as wherein f k denotes the CPU frequency of the device k, and n F represents the floating-point operations per cycle. Therefore, the training time T of an edge device k is The power consumption of the processor can be represented as wherein κ is the energy loss coefficient, mainly depending on the chip architecture. Therefore, the computational energy consumption of an edge device k in the t-th communication round is defined as:

[0072] ​​​

[0073] In the formula This indicates that at a given CPU frequency f k With the number of iterations E in local training, computational energy consumption is mainly affected by the size of the training dataset. Influence.

[0074] In an airborne federated edge learning system, communication energy consumption can be represented by the square of the L2 norm of the transmitted signal. This is because the communication energy consumption of wireless communication is mainly related to the transmission power, which is directly related to the amplitude (i.e., the L2 norm) of the transmitted signal. Specifically, the transmission power is proportional to the square of the L2 norm of the signal. Therefore, with a fixed communication time, the communication energy consumption of edge device k in the t-th communication round... Defined as:

[0075]

[0076] The total energy consumption of edge device k in the t-th communication round can be expressed as:

[0077]

[0078] In an airborne federated edge learning framework, model accuracy is a core metric for evaluating system performance. However, edge devices are typically constrained by energy resources, which can significantly impact the quality of local training and the reliability of model updates, thus limiting the overall model's convergence performance. To achieve an effective trade-off between energy efficiency and model accuracy, this chapter proposes optimization from two dimensions: data sampling and transmission power control. Specifically, during the local training phase, the data sampling rate is introduced as a variable to dynamically adjust the distribution and volume of data participating in training, balancing computational energy consumption and model accuracy. During the model update phase, the transmission power scaling factor is optimized to adaptively adjust the transmission power, thereby achieving an optimal balance between communication energy consumption and model update reliability. In each communication round, this invention jointly optimizes the sampling rate on the device. and uplink transmission power scaling factor This minimizes the upper bound of the optimal gap in the loss function. Therefore, the optimization objective of this invention is to minimize the gap between the expected global loss and the optimal loss after T communications. The formulaic optimization problem expression is:

[0079]

[0080] Among them, constraint C1 restricts the transmit power of each edge device from exceeding its maximum transmit power; constraint C2 requires that the total energy consumption of each edge device for local training and wireless communication within T rounds cannot exceed its energy budget H. k Constraint C3 specifies the range of values ​​for the data sampling rate.

[0081] Note that the problem contains two different non-convex components. Specifically, the first term involves the ratio of the normalized resource allocation variables and, combined with the norm calculation, makes it a highly non-linear non-convex expression. The second term is a fractional term with respect to the total transmit power, containing the square of the denominator, which makes it a non-convex function. Moreover, the constraint C1 contains a quadratic term which is also a non-convex constraint, further increasing the non-convexity of the problem. Therefore, the problem cannot be directly solved by standard convex optimization methods (e.g., gradient descent or convex optimization tools). Due to the long-term energy constraint C2, the offline optimization of the problem requires optimal partitioning of the energy budget for all devices in each round. Therefore, the present invention implements an online optimization algorithm that does not require any prior knowledge, but only relies on the current channel state information and energy state information.

[0082] Step 2: Transform the optimization problem in Step 1 based on the Lyapunov optimization method and the upper bound of the gap between the expected global loss and the optimal loss in the convergence analysis.

[0083] Step 1 proposes a single-cluster aerial federated edge learning model training system that considers both local training and model transmission on edge devices. Under the long-term energy constraint, an optimization problem is constructed to narrow the gap between the expected global loss and the optimal loss. Based on the Lyapunov optimization method and the upper bound of the gap between the expected global loss and the optimal loss in the convergence analysis, the optimization problem in Step 1 is transformed.

[0084] Based on the Lyapunov framework, the time-average inequality of energy can be transformed into a queue stability constraint. The energy state of the edge device is defined as a virtual queue which can be represented as:

[0085]

[0086] where is the queue length of the edge device k in round t that deviates from the long-term energy budget, and Therefore, by controlling the stability of , it is ensured that the aerial federated edge learning system will not terminate the training due to energy consumption constraints during the training process.

[0087] To optimize the objective function while stabilizing the system state, a drift-plus-penalty method is introduced. The core idea is to minimize the following equation at each round i.e.:

[0088]

[0089] where is expressed as Lyapunov drift, whose magnitude reflects the trend of energy state change. V≥0 is a weight parameter that can gradually minimize the objective function J while ensuring the stability of the energy queue. V emphasizes the objective function of the optimization model, and vice versa. Specifically, V→∞ means maximizing the objective function only without considering energy consumption. When V=0, only the stability of the energy queue is ensured. Then, the objective function in the problem is decoupled into a function of each round , that is:

[0090]

[0091] denotes the sampling rate of the data sample of the category j in the edge device k.

[0092] The online optimization sub-problem can be expressed as:

[0093]

[0094] Step 3: decompose the transformed deterministic optimization problem obtained in step 2 into two sub-problems, and solve them iteratively by alternating optimization; the present application proposes an energy-aware alternating resource allocation algorithm EARA, which maximizes the system training performance under the condition of energy limitation by jointly optimizing the transmission power and the training data sampling rate of the edge device.

[0095] To solve the problem , first fix the sampling rate , and obtain the sub-problem expression about the transmission power scaling factor :

[0096]

[0097] where is the equivalent maximum transmission power. denotes the queue length of the energy consumption deviation of the edge device k from the long-term energy budget in the round t. From the above formula , it can be seen that the second term of the objective function has the form of the square of the inverse, and by calculating the Hessian matrix of the formula to be greater than 0, it can be known that the function is convex with respect to as a whole, but since is a linear sum of multiple , due to the existence of the square term, the problem may still have non-convexity in the high-dimensional optimization space, so that the standard convex optimization solver cannot be directly applied. Therefore, the problem ​The problem is non-convex in general. To solve this problem, the idea of inverse convex optimization is adopted to transform the non-convex problem into an equivalent convex optimization problem. Specifically, an auxiliary variable The objective function is then rewritten as follows:

[0098]

[0099] The problem is equivalent to:

[0100]

[0101] To handle the fractional term in the above equation, an auxiliary variable u t is introduced: t 2 and a constraint:

[0102]

[0103] which is non-convex, is linearized using Taylor expansion as follows:

[0104]

[0105] where denotes the result of the previous iteration. Then, the problem P3.1 can be converted to:

[0106]

[0107] C1 is a specific constant, The uplink transmission latency of device k in round t.

[0108] The problem P3.2 is convex and can be solved directly using the standard convex optimization tool CVX.

[0109] Based on the power scaling factor The problem can be restated as:

[0110]

[0111] A denotes the energy consumption required to compute a single data sample, and B is a specific constant. p m denotes the global sample distribution with label m.

[0112] Since The objective function contains a non-convex fractional term, and the problem is non-convex in general. The invention adopts the block coordinate descent method to optimize each variable block Specifically, in each iteration, the objective function is minimized with respect to a single variable block, while the other variable blocks remain unchanged. The subproblem is as follows:​

[0113]

[0114] To solve the non-convexity caused by the fractional term, the present application further uses SCA technique to construct a convex approximation function for the fractional term, and obtains the approximate solution of the original problem by minimizing the convex approximation function. Specifically, let The fractional term in the above equation can be rewritten as Use the first-order Taylor expansion to convexify, and define the value of the last iteration as Approximate the fractional term as

[0115]

[0116] Process the absolute value term in the objective function, and let

[0117]

[0118] Since is not a convex function, use linear relaxation to approximate it:

[0119]

[0120] In the above equation, is the value of the last iteration. At this time, the fractional term is converted into a linear term. Introduce an auxiliary variable u k Relax the quadratic term, i.e.

[0121]

[0122] At this time, the problem P4.1 can be converted into

[0123]

[0124] The auxiliary variable is denoted as The value of the tth iteration is The value of the tth iteration of the sampling rate is denoted as

[0125] The problem P4.2 is a convex optimization problem, which can be directly solved by using the standard convex optimization solver CVX. The present application sorts out the proposed EARA algorithm, as shown in Table 1.

[0126]

[0127]

[0128] Figures 4-11It can be seen that the EARA test accuracy is obviously better than other comparative algorithms. The EARA algorithm dynamically adjusts the transmission power and data sampling rate in each round, ensures the continuity and stability of the training process, and avoids premature depletion or waste of energy. At the same time, the alternating optimization strategy of EARA explicitly models the communication distortion and data heterogeneity, and dynamically balances the device sampling rate and transmission power under the energy constraint. Thus, the online joint optimization of data sampling rate and transmission power scaling factor is realized.

[0129] The above-mentioned are specific embodiments of the present application and the technical principles used. If changes are made in accordance with the concept of the present application, and the resulting functions still do not exceed the spirit covered by the specification and drawings, they should still be within the scope of protection of the present application.

Claims

1. A sampling rate and power optimization method based on a single-cluster aerial federal edge learning system, characterized in that, The method comprises the following steps: An aerial federated edge learning system model is constructed, and an optimization problem is established to minimize the gap between an expected global loss and an optimal loss after T communications, which is represented as: denotes the sampling rate of data samples for label m in edge device k, denotes the power control coefficient, C1 and B are constants, denotes the edge device index set, K denotes the number of edge devices, and M denotes the number of sample labels denotes the total amount of data samples with label m, p m denotes the global data sample distribution for label m, each device transmits a total of d analog symbols corresponding to its model size, denotes the noise power, denotes the parameters of the local model, denotes the transmission power budget of edge device k, the total communication round is T, and the index set of the communication round is denoted as denotes the total energy consumption of edge device k in the tth communication round, H k denotes the energy budget; wherein constraint C1 limits the transmit power of each edge device not to exceed its maximum transmit power; constraint C2 requires that each edge device cannot exceed its energy budget H in T rounds of local training and wireless communication k ; and constraint C3 specifies the range of data sampling rate Based on the Lyapunov optimization method and the upper limit of the gap between the expected global loss and the optimal loss in the convergence analysis, the optimization problem is converted into a new optimization problem. The converted new optimization problem is decomposed into two sub-problems, and the online joint optimization of the data sampling rate and the transmission power scaling factor is realized by iterative solution through the alternating optimization method.

2. The sampling rate and power optimization method based on a single-cluster aerial federal edge learning system according to claim 1, characterized in that: The aerial federated edge learning system comprises K edge devices and one edge server, and the edge devices and the edge server are connected through a wireless network to collaboratively train a shared global deep learning model.

3. The sampling rate and power optimization method based on a single-cluster air-federated edge learning system according to claim 2, characterized in that: In the air federation edge learning system, after sampling, the training data subset obtained by the edge device k and the new data distribution are calculated by the following formula respectively: Based on the data distribution Q k The loss function of the edge device k is: denotes the expectation based on the data distribution Q k , denotes the indicator function, f(w, x i ) denotes the loss function of the global deep learning model, x i denotes the feature of the i-th data sample of the edge device, y i denotes the label of the data sample, denotes the expectation under the condition that the label y i = m, is the model parameter vector of the edge device k, denotes the d-dimensional real space, each device transmits a total of d analog symbols corresponding to the model size thereof.

4. The sampling rate and power optimization method based on a single-cluster air-federated edge learning system according to claim 3, characterized in that: An energy consumption model is further constructed in the aerial federated edge learning system model, and the total energy consumption of the edge device k in the tth communication round is represented as: N F denotes the floating point operations required to process one data sample, f k denotes the CPU frequency of edge device k, n F denotes the floating point numbers operated per cycle, E denotes the number of iterations of local training, denotes the power control coefficient, denotes the model parameter vector, denotes the computation energy consumption of edge device k in the t-th communication round, denotes the communication energy consumption of edge device k in the t-th communication round.

5. The sampling rate and power optimization method based on single-cluster aerial federal edge learning system according to any one of claims 1-4, characterized in that: The optimization problem is converted into a new optimization problem, and the new optimization problem is represented as: denotes the Lyapunov drift, V > 0 is a weight parameter, denotes the aggregated weight, denotes the sampling rate of data samples of class j in edge device k.

6. The sampling rate and power optimization method based on the single-cluster aerial federated edge learning system according to claim 5, characterized in that: The two sub-problems are: fixing the sampling rate first obtaining the sub-problem of the transmission power scaling factor ​ is the equivalent maximum transmit power, denotes the complex channel coefficient from edge device k to the edge server in the t-th round of communication, denotes the queue length of the energy consumption deviation from the long-term energy budget of edge device k in round t; Based on power scaling factor Problem Rephrased as: A represents the energy consumption required for calculating a single data sample, and B is a constant.

7. The sampling rate and power optimization method based on a single-cluster aerial federal edge learning system according to claim 6, characterized in that: Based on the sub-problem Introducing auxiliary variables The objective function is then rewritten as follows: Then the sub-problem is equivalent to: Introduce the auxiliary variable u t = (z t ) 2 and one constraint: Linearize it using Taylor expansion as follows: where denotes the result of the previous iteration; then, the problem P3.1 is transformed into: C1 is a constant, is the uplink transmission latency for device k in round t; Problem P3.2 is convex, and is solved by using a standard convex optimization tool CVX.

8. The sampling rate and power optimization method based on a single-cluster aerial federal edge learning system according to claim 6, characterized in that: Based on sub-problems Each variable block is optimized using a block coordinate descent approach At each iteration, the objective function is minimized with respect to a single variable block while holding the other variable blocks constant, the sub-problem is as follows: Then the fractional term is constructed a convex approximation function using SCA technique, and the approximate solution of the original problem is obtained by minimizing the convex approximation function; specifically, let where the fractional term is rewritten as The first-order Taylor expansion is used for convexification, and the value of the last iteration is defined as The fractional term is approximated: The absolute value term in the objective function is processed, and the following equation is obtained: Due to Not a convex function, approximate with linear relaxation: wherein is the value of the previous iteration; an auxiliary variable u is introduced k Relaxing the quadratic term, i.e.: At this time, problem P4.1 is converted into: denotes an auxiliary variable the value of the t-th iteration, denotes the sampling rate the value of the t-th iteration; Problem P4.2 is a convex optimization problem, and is directly solved by using a standard convex optimization solver CVX. 9.A single-cluster based federated learning system in the air edge, characterized in that: The aerial federated edge learning system comprises K edge devices and one edge server, and the edge devices and the edge server are connected through a wireless network, and the edge devices and the edge server store computer programs, and the computer programs are executed by a processor to realize the steps of the sampling rate and power optimization method based on the single-cluster aerial federated edge learning system according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The executable instructions are stored for being executed by a processor to realize the sampling rate and power optimization method based on the single-cluster aerial federated edge learning system according to any one of claims 1 to 8.