Efficient federated learning method for constructing electromagnetic atlas through cooperation of multiple unmanned aerial vehicles
By employing an efficient federated learning method for constructing electromagnetic maps through multi-UAV collaboration, combining physical neural networks and generative adversarial networks, and utilizing a multi-agent deep deterministic policy gradient algorithm, this approach addresses the limitations of energy and computing power for UAVs in electromagnetic map construction. It achieves fast and accurate electromagnetic map construction while reducing latency and energy consumption.
Patent Information
- Application Number
- CN202511100365.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-21
AI Technical Summary
Unmanned aerial vehicles (UAVs) face challenges in constructing electromagnetic maps due to limited energy and computing power, as well as insufficient coverage and flexibility of traditional ground communication facilities, making it difficult to quickly collect and process massive amounts of terminal data.
An efficient federated learning method for constructing electromagnetic maps through multi-UAV collaboration is adopted. By optimizing the communication and computing resources of UAVs and edge nodes, combining physical neural networks and generative adversarial networks, and using a multi-agent deep deterministic policy gradient algorithm for model training and decision-making, a fast and accurate electromagnetic map construction is achieved.
The most accurate electromagnetic spectrum model is obtained in the shortest time with the least energy consumption, which improves the physical consistency and interpretability of the model, enhances the prediction and analysis capabilities of electromagnetic signal transmission, and reduces latency and energy consumption.
Smart Images

Figure CN120996087A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of unmanned aerial vehicles (UAVs) and multi-agent reinforcement learning, specifically involving an efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps. Background Technology
[0002] With the digitalization of various industries, the number of wireless devices is growing exponentially, making spectrum resources increasingly scarce and difficult to manage. Electromagnetic mapping (EMG) collects a certain amount of electromagnetic data to construct a multi-dimensional spectrum information database for the entire region, presenting it clearly and intuitively to facilitate electromagnetic environment analysis and spectrum resource management. EMG offers numerous conveniences in many fields and holds an indispensable position.
[0003] With the rapid development of drone technology, drones are increasingly widely used in many fields, including but not limited to environmental monitoring, logistics delivery, disaster relief, and security monitoring. The flexibility and maneuverability of drones make them an ideal platform for collecting and processing large amounts of field data. Some researchers have studied scenarios where rechargeable drones with limited battery power collect sensor network data and proposed an algorithm based on hierarchical reinforcement learning to jointly optimize the drone's flight path and charging process. Others have proposed a method for constructing urban electromagnetic environment maps based on drones by combining measured data with geographic information system (GIS) information.
[0004] However, due to the limitations of traditional terrestrial communication facilities in terms of coverage and flexibility, the key lies in how to quickly collect network data in the face of a massive number of widely distributed terminals. Furthermore, given the limited energy and computing power of drones, effectively utilizing these resources to support complex data processing tasks presents a challenge. Summary of the Invention
[0005] To address the aforementioned technical challenges, this application provides an efficient federated learning method for constructing electromagnetic spectrum models through multi-UAV collaboration. By optimizing the communication and computing resources of UAVs and edge nodes, the most accurate electromagnetic spectrum model is obtained in the shortest time with minimal energy consumption. The model is expanded using GANs, with each edge computing node uploading its local GAN network generator G to the UAV for sharing. When an edge computing node needs to train a federated learning model, the UAV provides the corresponding generator to produce additional electromagnetic spectrum models, reducing the time required for additional model generation. By incorporating physical laws or prior knowledge into the prediction of the electromagnetic spectrum using PINN, the physical consistency and interpretability of the model are enhanced. The entire model training process is modeled as a partially observable Markov decision process, and solved using a multi-agent deep deterministic policy gradient algorithm, enabling UAVs to make effective decisions even in dynamic environments.
[0006] To achieve the above objectives, this application employs the following technical solution:
[0007] This application discloses an efficient federated learning method for constructing electromagnetic maps collaboratively using multiple drones. Several drones are deployed in a network scenario, and point-to-point communication occurs between the drones and user terminals. The federated learning method specifically includes the following steps:
[0008] Step 1: All user terminals are divided into K clusters. Each cluster contains an edge computing node. Each user terminal must be bound to an edge computing node k. Each user terminal collects environmental information data and physical data, and uses the environmental information data and physical data to perform partial local computing training to obtain a local model.
[0009] Step 2. Drone Scene Modeling and Analysis: In a network scenario, multiple drones are deployed. Within one model aggregation task cycle, the drones' flight trajectories are planned to acquire models from all user terminals. After completing the task, the drones return to the starting point to recharge and then proceed to the next round of model aggregation. Specifically:
[0010] Step 2.1, Edge computing node model parameter communication model analysis: The user terminal offloads the remaining computing tasks from Step 1 to the edge computing node, and transmits the local model at the same time. The edge computing node model parameters are first transmitted to the edge computing node for aggregation, and then wait for the drone to acquire them. Physical neural networks and generative adversarial networks (GANs) will be introduced into federated learning to expand the loss function.
[0011] Step 2.2, UAV Motion Model Modeling: Construct a UAV motion model and determine the feasibility and safety conditions for UAV flight;
[0012] Step 2.3, UAV channel model analysis: Based on the path loss between the UAV and the edge computing node and the connection probability of the line-of-sight link, the average path loss between the UAV and the edge computing node is obtained;
[0013] Step 2.4: Aggregate drone models to obtain the energy consumption of drone model aggregation;
[0014] Step 2.5: Model aggregation between drones to obtain the total energy consumption for transmission between drones;
[0015] Step 2.6: Determine whether the convergence condition of the UAV inter-drone model gathered in Step 2.5 is met. If convergence is achieved, use the gathered UAV inter-drone model to make predictions and obtain the electromagnetic spectrum map. If convergence is not achieved, the UAV will send the gathered UAV inter-drone model to each terminal to update the local model and start the next round of training and model gathering of federated learning.
[0016] Step 3: Formulation of optimization problem: Based on Step 1 and Step 2, create optimization problem to improve the accuracy of electromagnetic spectrum;
[0017] Step 4: Each UAV solves the optimization problem created in Step 3 using Multi-Agent Deep Deterministic Policy Gradient (MADDPG), modeling the solution process as a partially observable Markov decision process.
[0018] A further improvement of this application is that, in step 1, Q user terminals are randomly distributed throughout the scene, and each user terminal... There is a D q A local dataset consisting of data samples Where m q and n q These represent the data features and their corresponding labels, respectively. The total dataset is... Local model training specifically involves calculating the loss L of the HATA model based on the physical data collected from the user terminal. hata :
[0019]
[0020] Where f is the signal frequency, d is the transmission distance, and h is the signal frequency. te h represents the effective height of the base station antenna. re To determine the effective height of the receiving antenna, α(h) re ) is the correction factor.
[0021] A further improvement in this application is that, in step 2.1, a physical neural network and a generative adversarial network (GAN) are introduced into federated learning to expand the loss function, specifically:
[0022] Step 2.2.1: Based on the loss L obtained in Step 1.1 hata Calculate the received signal power P r :
[0023] P r =P t -L hata
[0024] Among them, P t Transmission power;
[0025] Step 2.2.2: The received signal power P calculated in step 2.2.1 r The physical neural network (PINN) model uses environmental information data collected from user terminals as input features to serve as the ground truth for training. For training, the local loss function of user terminal q is:
[0026]
[0027] in, Let i represent the sample loss function for the i-th data sample. For the predicted value of the i-th data sample, The amount of data selected for training;
[0028] global loss function It is the weighted average of the local loss functions of all user terminals, expressed as:
[0029]
[0030] The weights of the joint model in round t+1 are aggregated using the local model weights ω submitted by the participating nodes and their data volume percentages, and are expressed as follows:
[0031]
[0032] in, D represents the weights of the local model submitted by terminal q in round t+1. q / D represents the percentage of data owned by terminal q;
[0033] Step 2.2.3: Train a Generative Adversarial Network (GAN) at each edge computing node to augment the model. The GAN model consists of a discriminator D and a generator G. The loss function of the GAN model is:
[0034]
[0035] in, This indicates that the input model is the local model, σ is random noise, and the generator after training transforms the random noise σ into data with an approximate distribution to the samples. The optimization objective of the generator is to minimize... G The optimization objective of the discriminator is max(D,G). D L(D,G);
[0036] Step 2.2.4, the time representation of model augmentation using the anti-generative network (GAN) model is as follows:
[0037]
[0038] Among them, D ω f represents the amount of data representing the aggregated model parameters. GAN This represents the number of CPU cycles required to process a unit of data in a Generative Adversarial Network (GAN) model task. The computational power allocated to the edge computing node k for the Generative Adversarial Network (GAN) model task is represented by: The energy consumed in this process is expressed as:
[0039]
[0040] in, The energy consumption per CPU cycle for edge computing node k;
[0041] Step 2.2.5: Take the contents of step 2.2.2... And the loss function L of the generative adversarial network (GAN) model in step 2.2.3 GAN The loss function is incorporated into the local model loss function to obtain the final loss function, thus accelerating the convergence of federated learning.
[0042] L=ρ1L data +ρ2L PINN +ρ3L GAN
[0043] Where ρ1, ρ2, ρ3 are weighting coefficients and ρ1 + ρ2 + ρ3 = 1, L data This represents the loss of the local model.
[0044] A further improvement in this application is that, in step 3, based on the local model training results of steps 1 and 2 and the analysis of the UAV scene modeling in step 2, an optimization problem is created, specifically:
[0045] The local model computation time for edge computing nodes is:
[0046]
[0047] The time for the edge computing node aggregation model is:
[0048]
[0049] Using the Taylor approximation method, the above equation can be rewritten as follows:
[0050]
[0051] The total time for distributed federated learning is the sum of the local model computation time on the edge computing nodes and the aggregated model computation time.
[0052]
[0053] Transmission power consumption
[0054]
[0055] The time function T(u) taken by the drone u:
[0056]
[0057] The energy consumption function E(u) of the drone u:
[0058]
[0059] Based on this, the optimization problem can be formulated as follows:
[0060]
[0061] Constraint C1 states that the received signal-to-noise ratio of the edge computing node must be greater than a threshold. Constraint C2 states that the received power of the UAV during flight must be greater than the received threshold. Constraints C3 and C4 are constraints on the access of resource blocks by the terminal / edge computing node / UAV. Constraint C5 states that the allocated wireless resource blocks should not exceed the available wireless resources. Constraint C6 is a constraint on the terminal's transmit power. Constraint C7 states that the terminal must offload its model to at most one edge computing node without violating the computing capabilities of the edge computing node. Constraint C8 states that the overall computing power of the edge computing node should be less than or equal to the total computing power. Constraint C9 states that the time for the UAV to converge a model in one round should not exceed the UAV's flight time. Constraint C10 states that the energy consumption of the UAV to complete a data acquisition task in one round should be less than the UAV's own battery power. Constraints C11 and C12 are to ensure the feasibility and safety of UAV flight; therefore, the distance between the UAV and obstacles, and between UAVs, must be maintained above a safe distance.
[0062] The beneficial effects of this application are:
[0063] This application models the multi-UAV sensing and computation problem as a partially observable Markov decision process, and models the time, energy consumption, and accuracy of each communication and computation stage. By optimizing the communication and computational resources of UAVs and edge nodes, the most accurate electromagnetic spectrum is obtained in the shortest time with the least energy consumption.
[0064] This application proposes a hybrid knowledge embedding architecture combining PINN and GAN. PINN enhances the physical plausibility of the generated data through its physical conservation constraints, selects different transmission formulas for different terrains, and accelerates model convergence using physical knowledge constraints. GAN's adversarial training mechanism improves the detailed reconstruction capability of electromagnetic field strength distribution. GAN learns the patterns of local model parameters and then expands the model to improve its performance. Prior knowledge of electromagnetic signal transmission and a weighted sum of GAN's loss are introduced into the neural network's loss function. This combination of methods effectively improves the electromagnetic spectrum model's ability to predict and analyze complex physical phenomena.
[0065] This application utilizes sensing, computation, and communication models to establish action and state spaces, and designs corresponding reward functions based on execution time, energy consumption, model accuracy, and flight penalties. The MADDPG algorithm from multi-agent reinforcement learning is then used to solve the problem, ultimately yielding the model convergence strategy for the UAV network.
[0066] Compared with traditional methods, the proposed algorithm has advantages in terms of latency, energy consumption and accuracy, and the constructed spectrum map is more in line with the actual situation. Attached Figure Description
[0067] Figure 1 This is a schematic diagram of the service scenario deployed by the present invention.
[0068] Figure 2 This is a flowchart of the algorithm of the present invention.
[0069] Figure 3 This is a comparison of the convergence of the loss function of this application (GAN+PINN) with three other methods.
[0070] Figure 4 This is a comparison chart of the mean squared error of this application with three other federated learning algorithms.
[0071] Figure 5 This is a comparison chart of the predicted electromagnetic spectrum and the actual situation in this application.
[0072] Figure 6 This is a comparison of the latency and energy consumption of this application with three other algorithms for the next model aggregation.
[0073] Figure 7 This is a comparison chart of the model accuracy under different numbers of UAVs in this application.
[0074] Figure 8 This is the flight trajectory diagram of the drone in this application.
[0075] Figure 9 This is a graph showing the energy consumption and latency balance of drones under different algorithms.
[0076] Figure 10 This is a schematic diagram illustrating how the latency of the next model aggregation varies with the number of terminal devices under different algorithms.
[0077] Figure 11 This is a schematic diagram illustrating how the energy consumption of the model aggregation changes with the number of terminal devices under different algorithms. Detailed Implementation
[0078] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0079] like Figure 1As shown, this application presents an efficient federated learning method for constructing electromagnetic maps collaboratively using multiple drones. Several drones are deployed in a network scenario, and point-to-point communication occurs between the drones and user terminals. K edge computing nodes are deployed to pre-collect data from user terminals within the acquisition range and await the drones to perform model aggregation tasks. The federated learning method specifically includes the following steps:
[0080] Step 1: All user terminals are divided into K clusters, and each cluster contains an Edge Computing Node (ECN). Each user terminal must be bound to an edge computing node k. Each user terminal collects environmental information data and physical data, and uses the environmental information data and physical data to perform partial local computation training to obtain a local model.
[0081] In step 1, Q user terminals are randomly distributed throughout the scene, and each user terminal... There is a D q A local dataset consisting of data samples Where m q and n q Let represent the data characteristics and their corresponding labels, respectively. Generally, it is assumed that the data from each user terminal follows an IID distribution. Therefore, the total dataset in the system is . Local model training specifically involves calculating the loss L of the HATA model based on the physical data collected from the user terminal. hata :
[0082]
[0083] Where f is the signal frequency, d is the transmission distance, and h is the signal frequency. te h represents the effective height of the base station antenna. re To determine the effective height of the receiving antenna, α(h) re ) is a correction factor, determined by the buildings in the propagation environment.
[0084] Step 2. Drone Scene Modeling and Analysis: In a network scenario, multiple drones are deployed. Within one model aggregation task cycle, the drones' flight trajectories are planned to acquire models from all user terminals. After completing the task, the drones return to the starting point to recharge and then proceed to the next round of model aggregation. Specifically:
[0085] Step 2.1, Edge Computing Node Model Parameter Communication Model Analysis: Assume user terminal q processes... Each local model is partially computed using a sample of data, and each computation point requires f. q The remaining local model is sent to edge computing nodes for computation, specifically:
[0086] Establish a link between the user terminal q and the edge computing node. Distributed federated learning update transmission between the edge computing node and the user terminal q uses... A resource block set, where edge computing nodes and user terminals q use a single block. To communicate: Access to a single block cannot exceed that of a single user terminal. For all user terminals, the allocated resource blocks shall not exceed the available radio resources, that is: User terminal q will offload the local model to at most one edge computing node, that is... in, Let be the variable associated with the user terminal q and the edge computing node. If the user terminal q is associated with the edge computing node k, then... Otherwise, it is 0. The computing capabilities of edge computing nodes must not be violated: The remaining local model is sent to edge computing nodes for computation, specifically including the following steps:
[0087] Step 2.1.1: After the user terminal q completes part of the local task computation, the edge computing node processes the remaining local model. The time for each calculation point is:
[0088]
[0089] in, For the data that has already been calculated on the user terminal q, f q The number of CPU cycles required to process a single data sample, c q,k For edge computing nodes The computational power allocated to the learning task for user terminal q, with a value ranging from: c min ≤c q,k ≤c max The overall computing power of edge computing node k is equal to or less than the total computing power:
[0090]
[0091] in, The variable associated between the user terminal q and the edge computing node;
[0092] Step 2.1.2: The time it takes for the user terminal q to train the local model, thus obtaining the energy consumption of the user terminal during the training of the local model.
[0093]
[0094] Where, δq (f q ) 2 The energy consumption of the CPU cycle of the user terminal q;
[0095] Step 2.1.3: The user terminal q updates the local model using the federated learning method. The time required to update the edge computing nodes in step 2.1.1 is:
[0096]
[0097] in, The available computing power for edge computing nodes to perform learning tasks, where α is the local accuracy α of the w-th prediction metric:
[0098]
[0099] in, This represents the w-th predicted value of the local model, i.e., the predicted electromagnetic signal strength. Let w be the true value of the local model, which is the electromagnetic signal strength in the environment. W is the number of prediction indicators for which information is collected for each user terminal to train the local model and make predictions. w is the number of types of prediction values.
[0100] Step 2.1.4, each user terminal q uses Each element of the gradient vector used to transmit the parameters of the federated learning model is quantized using bits. The local model has m model parameters, and each user terminal q needs to send a total of bits during each federated learning session. The transmission time from user terminal q to edge computing node k, where model parameters are uploaded, is expressed as:
[0101]
[0102] Energy transferred is
[0103]
[0104] Where, λ q,j λ is a binary variable; when user terminal q is allocated to resource block j, λ q,j =1, otherwise λ q,j =0,m q For data features, The number of bits for the parameter. The transmission rate between the user terminal q and the edge computing node: γ is the resource block bandwidth transmitted between the user terminal q and the edge computing node. q Signal-to-interference-plus-noise ratio (SIR) when user terminal q uses resource block j: p qFor the user terminal's transmit power, g q,j For the wireless channel gain of the terminal, This is the cumulative interference excluding the user terminal q, and the user terminal's transmit power p. q In addition to adhering to certain ranges, its summation is subject to the following constraints.
[0105]
[0106] In this step, physical neural networks and generative adversarial networks (GANs) will be introduced into federated learning to expand the loss function, specifically:
[0107] Step 2.2.1: Based on the loss L obtained in Step 1.1 hata Calculate the received signal power P r :
[0108] P r =P t -L hata
[0109] Among them, P t Transmission power;
[0110] Step 2.2.2: The received signal power P calculated in step 2.2.1 r The physical neural network (PINN) model uses environmental information data collected from user terminals as input features to serve as the ground truth for training. For training, the local loss function of user terminal q is:
[0111]
[0112] in, Let i represent the sample loss function for the i-th data sample. For the predicted value of the i-th data sample, The amount of data selected for training;
[0113] global loss function It is the weighted average of the local loss functions of all user terminals. To obtain the optimal parameters, edge computing nodes and terminals coordinate to minimize the global risk by optimizing the transmission model weights, which can be expressed as:
[0114]
[0115] The weights of the joint model in round t+1 are aggregated using the local model weights ω submitted by the participating nodes and their data volume percentages, and are expressed as follows:
[0116]
[0117] in, D represents the weights of the local model submitted by terminal q in round t+1. q / D represents the percentage of data owned by terminal q;
[0118] Step 2.2.3: Train a Generative Adversarial Network (GAN) at each edge computing node to augment the model. The GAN model consists of a discriminator D and a generator G. The loss function of the GAN model is:
[0119]
[0120] in, This indicates that the input model is the local model, and σ represents random noise. After training, the generator transforms the random noise σ into data that approximates the distribution of the samples. The optimization objectives of the generator and the discriminator are different; the optimization objective of the generator is to minimize... G The optimization objective of the discriminator is max(D,G). D L(D,G);
[0121] Step 2.2.4, the time representation of model augmentation using the anti-generative network (GAN) model is as follows:
[0122]
[0123] Among them, D ω f represents the amount of data representing the aggregated model parameters. GAN This represents the number of CPU cycles required to process a unit of data in a Generative Adversarial Network (GAN) model task. The computational power allocated to the edge computing node k for the Generative Adversarial Network (GAN) model task is represented by: The energy consumed in this process is expressed as:
[0124]
[0125] in, The energy consumption per CPU cycle for edge computing node k;
[0126] Step 2.2.5: Take the contents of step 2.2.2... And the loss function L of the generative adversarial network (GAN) model in step 2.2.3 GAN The loss function is incorporated into the local model loss function to obtain the final loss function, thus accelerating the convergence of federated learning.
[0127] L=ρ1L data +ρ2L PINN +ρ3L GAN
[0128] Where ρ1, ρ2, ρ3 are weighting coefficients and ρ1 + ρ2 + ρ3 = 1, Ldata This represents the loss of the local model.
[0129] Step 2.2, UAV motion model modeling: Construct a UAV motion model and determine the feasibility and safety conditions for UAV flight.
[0130] The kinematic model of the UAV is modeled as follows, neglecting wind and air resistance during flight:
[0131]
[0132] Among them, (x u (t),y u (t),z u (t) represents the position coordinates of the UAV u at time t. u v u (t) represents the velocity of the drone u at time t. Let be the yaw angle of the UAV at time t. Let be the pitch angle of the UAV at time t. Let u be the horizontal component of the acceleration of the UAV at time t. Let be the vertical component of the acceleration of the UAV u at time t. Let x be the derivative of u. The velocity, pitch angle, and yaw angle of the UAV at time t satisfy the following constraints:
[0133] v min ≤v u (t)≤v max
[0134]
[0135] Among them, v min v max These are the minimum and maximum speeds of the drone u;
[0136] The drone u searches for the optimal trajectory to reach the edge computing node in a 3D scene. In this environment, there are multiple selectable edge computing nodes k. The distance from drone u to edge computing node k is:
[0137]
[0138] Among them, l k =(x k ,y k ,z k () represents the coordinates of the edge computing node;
[0139] Due to the complexity of the environment, drones do not fly at the same horizontal altitude. During drone flight path planning, drones not only need to avoid environmental obstacles but also collisions between themselves. Obstacles in the environment are denoted by ∈ = {∈1, ∈2, ..., ∈...}. ob ,…,∈ Ob} represents the coordinates of the center of the ob-th obstacle. It is its height.
[0140] Therefore, in order to achieve the feasibility and safety of drone flight, the following conditions need to be met:
[0141]
[0142] Where, d min It is the minimum safe distance between the drone and the obstacle, and between two drones, when the distance d between the drone and the edge computing node. u,k ≤d th When the drone is considered to have reached the vicinity of the edge computing node that needs to acquire data, d th It is the target distance threshold.
[0143] Step 2.3, UAV Channel Model Analysis: Communication between the UAV and edge computing nodes reuses the 5G communication frequency band, potentially leading to co-channel interference between links. Because the simplified Line of Sight (LoS) channel model ignores the effects of multipath fading and shadowing, this channel model is inaccurate in areas with many obstacles. There are both line-of-sight (LoS) links and non-line-of-sight (NLoS) links between the UAV and edge computing nodes. The path losses for the LoS and NLoS links are:
[0144]
[0145] Where, d u,k Let η be the distance from the drone u to the edge computing node k. LoS and η NLoS These are the average path overhead losses for LoS and NLoS links, respectively. c Where c is the carrier frequency and c is the speed of light;
[0146] The connection probability of the line-of-sight link between the drone and the edge computing node is expressed as:
[0147]
[0148] Among them, both ρ4 and ρ5 are related to the environment, h u It refers to the drone's flight altitude;
[0149] The probability of a non-line-of-sight (NLoS) link connection between a drone and an edge computing node is: Γ NLoS =1-Γ LoS ;
[0150] This application only considers the average path loss ζ ave (dB) is:
[0151]
[0152] in, Let u be the angle between the drone and the edge computing node k.
[0153] Step 2.4: When the drone is performing model convergence, it first needs to fly towards the target location in a certain direction and distance, which generates flight time. and energy consumption
[0154]
[0155] in, It is flight power;
[0156] The channel gain representation from edge computing node k to drone u:
[0157]
[0158] Where, d u,k Let ρ6 be the distance from the drone u to the edge computing node k.
[0159] To ensure that the drone can properly converge the model, the receiving power at the receiver needs to be greater than the threshold P. th Therefore, the received power of the drone u is P. k -ζ ave (dB) yields the achievable convergence rate of the drone model:
[0160]
[0161] in, It is the communication bandwidth of the drone u to acquire the edge computing node model. It is the cumulative interference from other edge computing nodes on the UAV's u-communication link, p u It is the drone's transmit power, g k′,u The channel gain from the interference edge computing node k′ to the UAV u is expressed as:
[0162]
[0163] Where, d u,k′It is the distance from the interference edge computing node k′ to the UAV u at time t;
[0164] Because the research scenario requires acquiring a large number of terminal training model parameters, these parameters need to be obtained during flight. Each edge computing node uses... Bits are used to quantize each element of its gradient vector, since the neural network model that integrates the physical neural network and the generative adversarial network has m k Each edge computing node needs to send a total of [number] parameters during each iteration. Bit-to-edge computing node, drone model convergence time Represented as:
[0165]
[0166] in, λ represents the convergence rate of the drone model. k,j λ ∈{0,1} is a binary variable representing the resource block allocated to the edge computing node, i.e., when the edge computing node k is allocated to resource block j, λ k,j =1, otherwise λ k,j =0,β u,k ∈{0,1} is a binary variable representing whether the edge computing node k is within the acquisition range of the drone u. If it is within β u,k =1; otherwise β u,k =0, to obtain the energy consumption of the drone model. for:
[0167]
[0168] in, This refers to the unit energy cost for data acquisition.
[0169] Step 2.5: Model aggregation between UAVs to obtain the total energy consumption for transmission between UAVs; after all UAVs receive the transmitted model, this application considers both device-to-device (D2D) communication and local computation, and transmission commands will be executed in parallel under each communication mode. This application assumes that the server has a total computing resource C. total , making C u If the computing resources allocated to the UAV u represent the resources available for use, then this application is subject to the following constraints:
[0170]
[0171] Assuming each drone is equipped with a CPU as a computing device, after the drone completes a round of model aggregation, in order to obtain a more complete global model, the drones will perform collaborative computing on the model in a D2D manner.
[0172] Assume that the communication radius of the UAV is r, and the two UAVs can communicate only within a distance of r. To avoid interference between D2D communications, different bandwidths are allocated to them. Using N u =
[0173] {n|n∈U,d n,u <r} represents the set of potential communicable UAVs of UAV u. The transmission rate between UAV u and UAV n is:
[0174]
[0175] Among them, is the channel bandwidth allocated by the UAV to D2D communication, is the transmission power from UAV u to UAV n, is the channel gain between UAV u and UAV n, and d n,u 2 Let n be the energy consumption per CPU cycle of the drone. Then, the total energy consumption for transmission using D2D communication is:
[0186] Step 2.6: Determine whether the convergence condition of the UAV inter-drone model gathered in Step 2.5 is met. If convergence is achieved, use the gathered UAV inter-drone model to make predictions and obtain the electromagnetic spectrum map. If convergence is not achieved, the UAV will send the gathered UAV inter-drone model to each terminal to update the local model and start the next round of training and model gathering of federated learning.
[0187] Step 3, Formulation of the optimization problem: Based on the local model training results from Steps 1 and 2 and the analysis of the UAV scene modeling from Step 2, an optimization problem is created, specifically:
[0188] The local model computation time for edge computing nodes is:
[0189]
[0190] The time for the edge computing node aggregation model is:
[0191]
[0192] Using the Taylor approximation method, the above equation can be rewritten as follows:
[0193]
[0194] The total time for distributed federated learning is the sum of the local model computation time on the edge computing nodes and the aggregated model computation time.
[0195]
[0196] Transmission power consumption
[0197]
[0198] The time function T(u) taken by the drone u:
[0199]
[0200] The energy consumption function E(u) of the drone u:
[0201]
[0202] Based on this, the optimization problem can be formulated as follows:
[0203]
[0204] Constraint C1 states that the received signal-to-noise ratio of the edge computing node must be greater than a threshold. Constraint C2 states that the received power of the UAV during flight must be greater than the received threshold. Constraints C3 and C4 are constraints on the access of terminals / edge computing nodes / UAVs to resource blocks, and the access permission of a specific resource block must not exceed that of a single terminal / edge computing node / UAV. Constraint C5 states that the allocated wireless resource blocks must not exceed the available wireless resources. Constraint C6 is a constraint on the terminal's transmit power. Constraint C7 states that, without violating the computing capabilities of the edge computing nodes, the terminal must offload its model to at most one edge computing node. Constraint C8 states that the overall computing power of the edge computing nodes should be less than or equal to the total computing power. Constraint C9 states that the time for a UAV to converge a model in one round must not exceed the UAV's flight time. Constraint C10 states that the energy consumption of the UAV to complete a data acquisition task in one round must be less than the UAV's own battery power. Constraints C11 and C12 are to ensure the feasibility and safety of UAV flight; therefore, the distance between the UAV and obstacles, and between UAVs, must be maintained above a safe distance.
[0205] Step 4: Each UAV solves the optimization problem created in Step 3 using Multi-Agent Deep Deterministic Policy Gradient (MADDPG). The solution process is modeled as a partially observable Markov decision process, defined as a tuple. Each agent has a current system state. The agent's policy π(s) u,t Select an action As a result of an action, the agent receives a reward and a next state from the environment; all agents share the same reward function. Value function of an agent
[0206] If the discount factor for future returns is γ∈[0,1), then the discount feedback return is γQ. u (s u,t+1 ,a u,t+1 ), where state, action, and reward are defined as:
[0207] 1) State Space
[0208] Considering the UAV's 3D geographic location, the number of model bits required for federated learning aggregation, the UAV's energy consumption and maximum energy, and the time consumption are represented as state S. t ={u[t],b[t],E[t],E max Let u[t] be the three-dimensional coordinates of the UAV at time t, b[t] be the number of bits that still need to be transmitted for model aggregation at time t, and E[t] be the energy consumption at time t, including E Tran andE(u),Emax Let T be the maximum energy of the drone, and T be the current time consumed, including T. FL and T(u);
[0209] 2) Action Space
[0210] At each time t, the drone performs flight, data transmission, and computation actions. Since each action affects the task cost, it is reasonable to consider them together. The action space is defined as A. t ={x d ,y d ,z d N u ,T bit}, where x d ,y d ,z d N is the flight distance of the drone in three-dimensional coordinates. u These are the task proportions and selection numbers for D2D communication by UAVs, T bit Let b[t] be the number of bits transmitted by the drone at that moment;
[0211] 3) Reward function
[0212] The drone needs to converge models within a certain time and energy consumption range to make the final model more accurate. Therefore, the final reward function is related to time, energy consumption, accuracy, and consists of four parts, where φ i ,i={1a,1b,2,3,4a,4b} are parameters. When the model convergence is completed, the lower the drone's energy consumption, the greater the reward. A negative sign is added before the energy consumption calculation to indicate the energy consumption reward function, which is expressed as:
[0213] r1(s u,t ,a u,t )=-φ 1a E(t)-φ 1b D sum
[0214] Where D sum It is a normalization function relative to the number of transmitted bits b[t], expressed as:
[0215]
[0216] ρ7 and ρ8 are normalization parameters. As the flight time increases, the penalty terms related to energy consumption increase, which helps guide the UAV to complete the mission as soon as possible and stop energy consumption.
[0217] The less time a drone takes to complete a mission, the greater the reward. The time-related reward function is defined as follows:
[0218] r2(s u,t ,a u,t )=-φ2T
[0219] The higher the accuracy of the model amassed by federated learning, the larger the accuracy-related reward function, defined as:
[0220] r3(s u,t ,a u,t )=φ3α(t)
[0221] To prevent the drone from flying out of the target area and causing accidents, and to prevent the drone from failing to acquire enough bits, a penalty function is defined to ensure that the drone can fly safely and effectively and converge the model, as shown in the following expression:
[0222]
[0223] in, This indicates that the drone has flown out of the target area, is outside the permitted range, or has encountered obstacles, or failed to maintain a safe distance. Otherwise... Indicates the number of bits that could not be retrieved;
[0224] The reward function for the entire problem can be expressed as the sum of the above rewards:
[0225]
[0226] The MADDPG algorithm framework consists of: each drone equipped with an actor network, a critic network, and a corresponding target network. The drone u observes the environment and obtains its current state value s. u,t The state values are then input into the actor network. An actor is a policy network, and its output is a policy probability distribution π(s). u,t ), which represents the probability that the drone will achieve the optimal solution by performing each action in the current state. The drone will determine the probability based on π(s). u,t To select the current action a u,t After performing the action, the current reward r is obtained. u (t), and arrive at the next state s. u,t+1 And add the interaction to the experience replay area.
[0227] The experience replay region samples data to update the critic network, enabling it to better predict the Q-value of a specific action in a given state. The trained critic network is then used to guide the actor network updates, ensuring that the policies generated by the actors achieve the largest possible Q-value within the critic network. The experience replay region collects interaction data from all drones. During training, both the critic and actor networks have access to this overall interaction data, representing centralized training. Individual drones, however, can only select their current action through their own actor network, representing distributed execution.
[0228] To accelerate algorithm convergence, this application introduces a priority experience replay mechanism. This mechanism selects experience samples that contribute significantly to network learning, assigns them higher sampling weights to make them easier to sample, and accelerates the training process. Samples with larger TD errors are considered as having made greater contributions. The TD error is calculated as follows:
[0229] δ m =r u (t)+γQ′ u (s t+1 ,a t+1 )-Q u (s t ,a t )
[0230] The probability of each sample being selected from the experience pool is:
[0231]
[0232] Where M represents the total number of samples in a Mini-batch, ∈ is a positive constant used to prevent the probability from being 0, and μ is used to adjust the priority of high-contribution samples.
[0233] Critics network training: The loss function of the critic network is defined as...
[0234] L u,Q (ω)=E[(r u (t)+γQ′ u (s t+1 ,a t+1 )-Q u (s t ,a t )) 2 ]
[0235] Among them, Q′ u It is the Q-value output by the target critic network, Q u It is the Q-value of the current critic network, r u It was obtained from the experience replay area in st Execute a in state t Rewards for actions.
[0236] Actor network training: The actor network is updated via policy gradient, and the loss function is as follows.
[0237]
[0238] To ensure stable training, each actor and critic network in MADDPG corresponds to its target network. After a certain number of training epochs, the target network is updated using the current actor and critic networks. The target network provides a relatively stable reference point. The target actor and critic networks update their network parameters using a soft update method.
[0239] ω Q′ =εω Q +(1-ε)ω Q′
[0240] ω π′ =εω π +(1-ε)ω π′
[0241] To verify the performance of this application, three other federated learning algorithms were selected as comparison schemes. The first federated learning algorithm proposes a three-layer architecture, using the HierFAVG algorithm to periodically aggregate models between edge servers and cloud servers, reducing the communication frequency with the cloud while maintaining efficient model training. The second federated learning algorithm, based on the model-agnostic meta-learning framework, proposes a personalized FedAvg algorithm variant (Per-FedAvg) to solve the personalized FL problem. This algorithm fine-tunes the initial shared model through multiple rounds of communication between the server and users, and performs multiple gradient updates locally. The third federated learning algorithm proposes a novel FL structure that avoids vehicles directly participating in model aggregation, reducing the impact of vehicle mobility on the learning process.
[0242] The effects of GAN and PINN on the convergence of the loss function are as follows: Figure 3 As shown, adding only PINN results in a slight decrease in the final loss function convergence value, while the convergence speed is slightly improved, reaching convergence in approximately 5000 epochs. Adding only GAN results in a significant decrease in the loss function convergence value, while the convergence speed is also accelerated. Adding both GAN and PINN further improves the model performance, ultimately reaching convergence in approximately 1500 epochs, with the convergence value being 46.86% higher than the original neural network.
[0243] Figure 4The impact of GAN and physical neural networks on the mean squared error (MSE), i.e., the impact on model accuracy, is shown. Compared with the original neural network model, adding only PINN reduces the MSE by 2.07, while adding only GAN reduces the MSE by 5.95. The model that incorporates both PINN and GAN performs best, with an MSE reduction of 7.1. This demonstrates that the proposed hybrid knowledge embedding architecture of PINN and GAN significantly improves model performance.
[0244] Figure 5 This section compares the electromagnetic spectrum predicted by several models with the actual electromagnetic spectrum. It shows that the signal intensity distribution predicted by the models differs from the electromagnetic spectrum of the actual environment (e.g., ...). Figure 5 (a) is largely the same, accurately predicting the distribution trend, but differing in details. After adding only PINN to the model (e.g.) Figure 5 (b) The trend transition of signal strength is not smooth enough, which is caused by missing data in some ranges; after only adding the GAN model (such as... Figure 5 (c) Due to the divergence of data generated by GANs, some predicted values are either too high or too low; after simultaneously introducing GAN and PINN (such as... Figure 5 (d) not only supplements the missing electromagnetic data in some areas, but also introduces physical knowledge to reduce the divergence of the prediction results. The model's prediction is more accurate and closest to the actual situation.
[0245] Figure 6 This paper compares the performance of the proposed MADDPG method with three other algorithms. Latency and energy consumption are two crucial constraints during UAV flight, determining whether the UAV can complete a mission effectively and completely. When there are 200 terminal devices in the TE layer, the proposed algorithm outperforms the other algorithms in both energy consumption and latency. The proposed algorithm achieves an energy consumption of 868J and a latency of 18.5s for a single model aggregation, representing improvements of 46.84%, 34.86%, and 58.71% in time efficiency, respectively, and 43.81%, 0.08%, and 51.01% in energy efficiency, respectively. This effectively improves the UAV's endurance and reduces the time required to complete the mission.
[0246] Figure 7This demonstrates the impact of the number of drones on model accuracy. After a certain training period, the model's accuracy continuously increases until convergence. As the number of drones increases, the training speed gradually accelerates, and the convergence speed increases. When the number of drones is 5 or 6, convergence is near at t=2000s, with a slight improvement in accuracy at convergence, and all accuracies reaching over 95%. However, when the number of drones is sufficient, further increasing the number does not yield significant improvements, as shown by the almost overlapping curves for u=5 and u=6 in the figure. Furthermore, the more drones there are, the higher the cost. Figure 7 As can be seen from the data, considering both accuracy and cost, using 5 drones yields the best training results.
[0247] Figure 8 The diagram illustrates the flight paths of four drones. Starting from a central charging station, they traverse all edge nodes of four regional sections according to a planned route, avoiding obstacles along the way. At each node, they collect and aggregate model data. Finally, they return to the starting point, completing the model aggregation among the multiple drones. The path planning algorithm enables the drones to traverse all nodes to acquire model data with minimal energy and time consumption, while avoiding obstacles to prevent collisions.
[0248] Figure 9 This paper illustrates the balance between energy consumption and latency for the UAV in this task. It shows that these two requirements are mutually restrictive: lower latency leads to higher energy costs, and vice versa. In practical applications, an appropriate operating point should be selected based on the UAV's own attributes and specific operational needs to achieve a proper balance between energy consumption and latency. Figure 9 As can be seen, the optimal operating point for the algorithm in this application is approximately 500 J in energy consumption and 25 s in latency. Although these two figures are contradictory, there are also overall differences between them under different algorithms. It can be seen that the method proposed in this application has the lowest overall cost in terms of energy consumption and latency, with energy consumption in the range of [400, 800] J and latency in the range of [5, 45] s. Its curve is also superior to the other three algorithms.
[0249] For different numbers of TEs, the latency and energy consumption of a single cloud aggregation in the four algorithms vary with the number of TEs as follows: Figure 10 and Figure 11 As shown. From Figure 10 and Figure 11As can be seen, with the increase of the number of training times (TEs), the latency and energy consumption of the four algorithm models all increased to varying degrees. However, the algorithm in this application outperformed the other algorithms in both latency and energy consumption at different TE counts. With 250 TEs, it consumed only about 20.3 seconds of latency and 1154.8 J of energy, representing improvements of 60.09%, 30.53%, and 45.53% in training time, respectively, and improvements of 50.08%, 11.97%, and 40.53% in energy consumption, respectively. Figure 11 In this algorithm, as the number of training elements (TEs) increases, the aggregation latency grows slowly. This is because the more TEs there are, the greater the deviation of the data from the original distribution, requiring more training time. The GAN framework in this paper extends the model to compensate for this, resulting in the lowest latency and a slow growth rate. Figure 11 In this context, the increase in the number of training elements (TEs) involved in model training directly leads to an increase in the energy consumption of federated learning. Figure 10 and Figure 11 This demonstrates that the performance of the algorithm in this paper is less affected by the number of TEs, and it has good applicability to learning scenarios with a large number of variable TEs.
Claims
1. A multi-UAV collaborative method for constructing an efficient federated learning model of electromagnetic spectrum, wherein several UAVs are deployed in a network scenario, and point-to-point communication is performed between the UAVs and the user terminal, characterized in that: The federated learning method specifically includes the following steps: Step 1: All user terminals are divided into K clusters, and each cluster contains an edge computing node (ECN). Each user terminal must be bound to an edge computing node k. Each user terminal collects environmental information data and physical data, and uses the environmental information data and physical data to perform partial local computing training to obtain a local model. Step 2. Drone Scene Modeling and Analysis: In a network scenario, multiple drones are deployed. Within one model aggregation task cycle, the drones' flight trajectories are planned to acquire models from all user terminals. After completing the task, the drones return to the starting point to recharge and then proceed to the next round of model aggregation. Specifically: Step 2.1, Edge computing node model parameter communication model analysis: The user terminal offloads the remaining computing tasks from Step 1 to the edge computing node, and transmits the local model at the same time. The edge computing node model parameters are first transmitted to the edge computing node for aggregation, and then wait for the drone to acquire them. Physical neural networks and generative adversarial networks (GANs) will be introduced into federated learning to expand the loss function. Step 2.2, UAV Motion Model Modeling: Construct a UAV motion model and determine the feasibility and safety conditions for UAV flight; Step 2.3, UAV channel model analysis: Based on the path loss between the UAV and the edge computing node and the connection probability of the line-of-sight link, the average path loss between the UAV and the edge computing node is obtained; Step 2.4: Assemble the drone models and obtain the energy consumption of drone model aggregation; Step 2.5: Model aggregation between drones to obtain the total energy consumption for transmission between drones; Step 2.6: Determine whether the convergence condition of the UAV inter-drone model gathered in Step 2.5 is met. If convergence is achieved, use the gathered UAV inter-drone model to make predictions and obtain the electromagnetic spectrum map. If convergence is not achieved, the UAV will send the gathered UAV inter-drone model to each terminal to update the local model and start the next round of training and model gathering of federated learning. Step 3: Formulation of optimization problem: Based on Step 1 and Step 2, create optimization problem to improve the accuracy of electromagnetic spectrum; Step 4: Each UAV solves the optimization problem created in Step 3 using Multi-Agent Deep Deterministic Policy Gradient (MADDPG), modeling the solution process as a partially observable Markov decision process.
2. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 1, characterized in that: In step 1, Q user terminals are randomly distributed throughout the scene, and each user terminal... There is a D q A local dataset consisting of data samples Where m q and n q These represent the data features and their corresponding labels, respectively. The total dataset is... Local model training specifically involves calculating the loss L of the HATA model based on the physical data collected from the user terminal. hata : Where f is the signal frequency, d is the transmission distance, and h is the signal frequency. te h represents the effective height of the base station antenna. re To determine the effective height of the receiving antenna, α(h) re ) is the correction factor.
3. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 2, characterized in that: In step 2.1, suppose the user terminal q processes... A portion of the local model computation is completed using sample data, while the remaining local model is sent to edge computing nodes for computation. Specifically, a link is established between the user terminal q and the edge computing node, and the distributed federated learning update transmission between the edge computing node and the user terminal q uses... A resource block set, where edge computing nodes and user terminals q use a single block. To communicate: Access to a single block cannot exceed that of a single user terminal. For all user terminals, the allocated resource blocks shall not exceed the available radio resources, i.e.: User terminal q can offload the local model to at most one edge computing node. in, Let be the variable associated with the user terminal q and the edge computing node. If the user terminal q is associated with the edge computing node k, then... Otherwise, it is 0. The computing capabilities of edge computing nodes must not be violated: The remaining local model is sent to edge computing nodes for computation, specifically including the following steps: Step 2.1.1: After the user terminal q completes part of the local task computation, the edge computing node processes the remaining local model. The time for each calculation point is: in, For the data that has already been calculated on the user terminal q, f q The number of CPU cycles required to process a single data sample, c q,k For edge computing nodes The computational power allocated to the learning task for user terminal q, with a value ranging from: c min ≤c q,k ≤c max The overall computing power of edge computing node k is equal to or less than the total computing power: in, The variable associated between the user terminal q and the edge computing node; Step 2.1.2: The time it takes for the user terminal q to train the local model, thus obtaining the energy consumption of the user terminal during the training of the local model. Where, δ q (f q ) 2 The energy consumption of the CPU cycle of the user terminal q; Step 2.1.3: The user terminal q updates the local model using the federated learning method. The time required to update the edge computing nodes in step 2.1.1 is: in, The available computing power for edge computing nodes to perform learning tasks, where α is the local accuracy α of the w-th prediction metric: in, This represents the w-th predicted value of the local model, i.e., the predicted electromagnetic signal strength. Let w be the true value of the local model, which is the electromagnetic signal strength in the environment. W is the number of prediction indicators for which information is collected for each user terminal to train the local model and make predictions. w is the number of types of prediction values. Step 2.1.4, each user terminal q uses Each element of the gradient vector used to transmit the parameters of the federated learning model is quantized using bits. The local model has m model parameters, and each user terminal q needs to send a total of bits during each federated learning session. The transmission time from user terminal q to edge computing node k, where model parameters are uploaded, is expressed as: Energy transferred is Where, λ q,j λ is a binary variable; when user terminal q is allocated to resource block j, λ q,j =1, otherwise λ q,j =0,m q For data features, The number of bits for the parameter. The transmission rate between the user terminal q and the edge computing node: γ is the resource block bandwidth transmitted between the user terminal q and the edge computing node. q Signal-to-interference-plus-noise ratio (SIR) when user terminal q uses resource block j: p q For the user terminal's transmit power, g q,j For the wireless channel gain of the terminal, This is the cumulative interference excluding the user terminal q, and the user terminal's transmit power p. q In addition to adhering to certain ranges, its summation is subject to the following constraints.
4. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 3, characterized in that: In step 2.1, the physical neural network and generative adversarial network (GAN) are introduced into federated learning to expand the loss function, specifically as follows: Step 2.2.1: Based on the loss L obtained in Step 1.1 hata Calculate the received signal power P r : P r =P t -L hata Among them, P t Transmission power; Step 2.2.2: The received signal power P calculated in step 2.2.1 r The physical neural network (PINN) model uses environmental information data collected from user terminals as input features to serve as the ground truth for training. For training, the local loss function of user terminal q is: in, Let i represent the sample loss function for the i-th data sample. For the predicted value of the i-th data sample, The amount of data selected for training; global loss function The weighted average of the local loss functions of all user terminals is expressed as: The joint model weights in round t+1 are represented by aggregating the local model weights ω submitted by participating nodes and their data volume percentages: in, D represents the weights of the local model submitted by terminal q in round t+1. q / D represents the percentage of data owned by terminal q; Step 2.2.3: Train a Generative Adversarial Network (GAN) at each edge computing node to augment the model. The GAN model consists of a discriminator D and a generator G. The loss function of the GAN model is: in, This indicates that the input model is the local model, and σ represents random noise. After training, the generator transforms the random noise σ into data that approximates the distribution of the samples. The optimization objective of the generator is to minimize... G The optimization objective of the discriminator is to max out L(D,G). D L(D,G); Step 2.2.4, the time representation of model augmentation using the anti-generative network GAN model is as follows: Among them, D ω f represents the amount of data representing the aggregated model parameters. GAN This represents the number of CPU cycles required to process a unit of data in a Generative Adversarial Network (GAN) model task. The computational power allocated to the edge computing node k for the Generative Adversarial Network (GAN) model task is represented by: The energy consumed in this process is expressed as: in, The energy consumption per CPU cycle for edge computing node k; Step 2.2.5: Take the contents of step 2.2.2... And the loss function L of the generative adversarial network (GAN) model in step 2.2.3 GAN The loss function is incorporated into the local model loss function to obtain the final loss function, thus accelerating the convergence of federated learning. L=ρ1L data +p2L PINN +p3L GAN Where ρ1, ρ2, ρ3 are weighting coefficients and ρ1 + ρ2 + ρ3 = 1, L data This is the loss for the local model.
5. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 4, characterized in that: In step 2.2, under the condition that wind force and air resistance are not considered in the environment, the kinematic model of the UAV is modeled as follows: Among them, (x u (t),y u (t),z u (t) represents the position coordinates of the UAV u at time t. u v u (t) represents the velocity of the drone u at time t. Let be the yaw angle of the UAV at time t. Let be the pitch angle of the UAV at time t. Let u be the horizontal component of the acceleration of the UAV at time t. Let be the vertical component of the acceleration of the UAV u at time t. Let x be the derivative of u. The velocity, pitch angle, and yaw angle of the UAV at time t satisfy the following constraints: v min ≤v u (t)≤v max Among them, v min v max These are the minimum and maximum speeds of the drone u; The distance from drone u to edge computing node k is: Among them, l k =(x k ,y k ,z k ) represents the coordinates of the edge computing node; Obstacles in the environment are represented by ∈ = {∈1, ∈2, ..., ∈ ob ,…,∈ Ob } represents the coordinates of the center of the ob-th obstacle. It is its height; The feasibility and safety of drone flights need to meet the following conditions: Where, d min It is the minimum safe distance between the drone and the obstacle, and between two drones, when the distance d between the drone and the edge computing node. u,k ≤d th When the drone is considered to have reached the vicinity of the edge computing node that needs to acquire data, d th It is the target distance threshold.
6. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 5, characterized in that: In step 2.3, there are line-of-sight (LoS) links and non-line-of-sight (NLoS) links between the UAV and the edge computing nodes. The path losses for the LoS and NLoS links are as follows: Where, d u,k Let η be the distance from the drone u to the edge computing node k. LoS and η NLoS These are the average path overhead losses for LoS and NLoS links, respectively. c Where c is the carrier frequency and c is the speed of light; The connection probability of the line-of-sight link between the drone and the edge computing node is expressed as: Where ρ4 and ρ5 are parameters, h u It refers to the drone's flight altitude; The probability of a non-line-of-sight (NLoS) link connection between a drone and an edge computing node is: Γ NLoS =1-Γ LoS ; Average path loss ζ ave (dB) is: in, 7. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 6, characterized in that: In step 2.4, the flight time generated by the UAV during model convergence. and energy consumption in, It is flight power; The channel gain representation from edge computing node k to drone u: Where, d u,k ρ6 is the distance from the drone u to the edge computing node k, and ρ6 is a parameter; The received power of the UAV u is P k -ζ ave (dB) yields the achievable convergence rate of the drone model: in, It is the communication bandwidth of the drone u to acquire the edge computing node model. It is the cumulative interference from other edge computing nodes on the UAV's u-communication link, p u It is the drone's transmit power, g k′,u The channel gain from the interference edge computing node k′ to the UAV u is expressed as: Where, d u,k′ It is the distance from the interference edge computing node k′ to the UAV u at time t; To acquire model parameters during flight, each edge computing node uses... Bits are used to quantize each element of its gradient vector, since the neural network model that integrates the physical neural network and the generative adversarial network has m k Each edge computing node needs to send a total of [number] parameters during each iteration. Bit-to-edge computing node, drone model convergence time Represented as: in, For the convergence rate of the drone model, ζ k,j λ ∈{0,1} is a binary variable representing the resource block allocated to the edge computing node, i.e., when the edge computing node k is allocated to resource block j, λ k,j =1, otherwise λ k,j =0,β u,k ∈{0,1} is a binary variable representing whether the edge computing node k is within the acquisition range of the drone u. If it is, β u,k =1; otherwise, β u,k =0, to obtain the energy consumption of the drone model. for: in, This refers to the unit energy consumption cost for data acquisition.
8. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 7, characterized in that: In step 2.5, assuming the communication radius of the drone is r, and communication between two drones is only possible within a distance r, the transmission rate between drone u and drone n is: in, It is the channel bandwidth allocated to D2D communication by the drone. The transmission power from drone u to drone n. Let d be the channel gain between UAV u and UAV n. n,u This represents the distance between drone u and drone n; The transmission time from drone u to drone n is: Where, λ u,j λ ∈{0,1} is a binary variable representing the resource block allocated to the drone, i.e., when drone u is allocated to resource block j, λ u,j =1, otherwise λ u,j =0,β u,k Whether there is communication between the drone and the edge computing node. For the communication rate, the energy consumption for data transmission in D2D communication is: The global model convergence time after the drones are converged is: Where α is the model accuracy, the total transmission time calculated using D2D communication is: The energy consumed by D2D communication is: Among them, f n δ represents the local CPU frequency of drone n. n (f n ) 2 Let n be the energy consumption per CPU cycle of the drone. Then, the total energy consumption for transmission using D2D communication is:
9. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 1, characterized in that: In step 3, based on the local model training results from steps 1 and 2 and the analysis of the UAV scene modeling from step 2, an optimization problem is created, specifically: The local model computation time for edge computing nodes is: The time for the edge computing node aggregation model is: Using the Taylor approximation method, the above equation can be rewritten as follows: The total time for distributed federated learning is the sum of the local model computation time on the edge computing nodes and the aggregated model computation time. Transmission power consumption The time function T(u) taken by the drone u: The energy consumption function E(u) of the drone u: Based on this, the optimization problem can be formulated as follows: Constraint C1 states that the received signal-to-noise ratio of the edge computing node must be greater than a threshold. Constraint C2 states that the received power of the UAV during flight must be greater than the received threshold. Constraints C3 and C4 are constraints on the access of resource blocks by the terminal / edge computing node / UAV. Constraint C5 states that the allocated wireless resource blocks should not exceed the available wireless resources. Constraint C6 is a constraint on the terminal's transmit power. Constraint C7 states that the terminal must offload its model to at most one edge computing node without violating the computing capabilities of the edge computing node. Constraint C8 states that the overall computing power of the edge computing node should be less than or equal to the total computing power. Constraint C9 states that the time for the UAV to converge a model in one round should not exceed the UAV's flight time. Constraint C10 states that the energy consumption of the UAV to complete a round of data acquisition should be less than the UAV's own battery power. Constraints C11 and C12 are to ensure the feasibility and safety of UAV flight; therefore, the distance between the UAV and obstacles, and between UAVs, must be maintained above a safe distance.
10. The efficient federated learning method for multi-UAV collaborative construction of electromagnetic maps according to claim 9, characterized in that: In step 4, the solution process of the optimization problem is modeled as a partially observable Markov decision process, defined as a tuple. Each agent has a current system state. The agent's policy π(s) u,t Select an action As a result of an action, the agent receives a reward and a next state from the environment; all agents share the same reward function. Value function of an agent If the discount factor for future returns is γ∈[0,1), then the discount feedback return is γQ. u (s u,t+1 ,a u,t+1 ), where state, action, and reward are defined as: 1) State Space Considering the UAV's 3D geographic location, the number of model bits required for federated learning aggregation, the UAV's energy consumption and maximum energy, and the time consumption are represented as state S. t ={u[t],b[t],E[t],E max Let u[t] be the three-dimensional coordinates of the UAV at time t, b[t] be the number of bits that still need to be transmitted for model aggregation at time t, and E[t] be the energy consumption at time t, including E Tran andE(u),E max Let T be the maximum energy of the drone, and T be the current time consumed, including T. FL and T(u); 2) Action Space At each time t, the action space is defined as A. t ={x d ,y d ,z d N u ,T bit }, where x d ,y d ,z d N is the flight distance of the drone in three-dimensional coordinates. u These are the task proportions and selection numbers for D2D communication by UAVs, T bit Let b[t] be the number of bits transmitted by the drone at that moment; 3) Reward function The reward function is related to time, energy consumption, and accuracy, and consists of four parts, where φ i ,i={1a,1b,2,3,4a,4b} are parameters. When the model convergence is completed, the lower the drone's energy consumption, the greater the reward. A negative sign is added before the energy consumption calculation to indicate the energy consumption reward function, which is expressed as: r1(s u,t ,am u,t )z-φ 1a E(t)-φ 1b D sum Where D sum It is a normalization function relative to the number of transmitted bits b[t], expressed as: ρ7 and ρ8 are normalization parameters; The less time a drone takes to complete a mission, the greater the reward. The time-related reward function is defined as follows: r2(s u,t ,a u,t )=-φ2T The higher the accuracy of the model amassed by federated learning, the larger the accuracy-related reward function, defined as: r3(s u,t ,a u,t )=φ3α(t) Define a penalty function that enables the drone to fly safely and efficiently and converge the model, as shown in the following expression: in, This indicates that the drone U has flown out of the target area; otherwise... Indicates the number of bits that could not be retrieved; The reward function for the entire problem can be expressed as the sum of the above rewards: