Federal learning security optimization method based on cooperative beam forming and node selection in multi-unmanned aerial vehicle network

By introducing a federated learning security optimization method that combines cooperative beamforming and node selection into multi-UAV networks, the problem of coordinating security, system cost, and model accuracy in multi-UAV cooperative networks is solved, achieving efficient learning and security defense in complex environments.

CN121690348APending Publication Date: 2026-03-17BEIJING FORESTRY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610016690.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies cannot collaboratively optimize security, overall system cost, and model accuracy in multi-drone cooperative networks. Traditional optimization methods struggle to handle high-dimensional hybrid action spaces and dynamic network topologies, resulting in insufficient security, uneven energy consumption, and slow model convergence.

Method used

We employ a federated learning-based safety optimization method based on cooperative beamforming and node selection. Through dynamic role partitioning, weighted combination of weight coefficients, and deep reinforcement learning algorithms, we construct a joint optimization multi-objective function. We combine graph attention and multi-head attention mechanisms to handle the mixed action space, achieving a balance between safety, cost, and accuracy.

Benefits of technology

It achieves a dynamic balance between multi-UAV collaborative defense and system effectiveness, improves physical layer security, balances federated learning efficiency and overall system cost, improves model convergence speed and accuracy, and solves the problems of intelligent decision-making and efficient solution in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121690348A_ABST
    Figure CN121690348A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning security optimization method based on cooperative beam forming and node selection in a multi-unmanned aerial vehicle network, and aims to solve the problem that security, system cost and accuracy cannot be collaboratively optimized in the prior art. The method comprises the following steps: S1, establishing a multi-unmanned aerial vehicle federated learning system containing dynamic role division and a channel model; s2, carrying out quantitative modeling on the comprehensive cost, the communication security and the model accuracy of the system; s3, constructing a joint optimization objective function taking the transmitting power, the training round number and the interference parameters as variables, and establishing a Markov decision process; s4, designing a deep reinforcement learning algorithm based on SAC to solve the problem; and S5, integrating a graph attention mechanism and a multi-head attention mechanism in the solving network so as to process a dynamic topology and hybrid action space. According to the method, by establishing a unified optimization framework, multi-index collaborative optimization is realized, and the system cost is effectively reduced while the safety and the accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of wireless communication, distributed machine learning and physical layer security, and more specifically, to a federated learning security optimization method and system based on cooperative beamforming and node selection in a multi-UAV cooperative network. Background Technology

[0002] With the rapid development of IoT technology, drones, due to their high mobility, flexible deployment, and line-of-sight transmission advantages, are widely used in mobile edge computing and data collection. To protect user privacy while utilizing distributed data, federated learning, as an emerging distributed machine learning paradigm, has been introduced into drone networks. This allows drones to train models locally and interact only with model parameters, thus avoiding the direct transmission of raw data. Although drone-assisted federated learning has broad application prospects, it still faces significant technical challenges in the practical deployment of multi-drone networks. First, there is the issue of communication security. Due to the broadcast nature of wireless channels, drones are highly susceptible to interception by malicious aerial eavesdroppers when transmitting model parameters over the air interface, leading to model inversion and privacy leaks. Existing upper-layer encryption or differential privacy technologies often come with huge computational overhead or sacrifice model accuracy, making them unsuitable for resource-constrained drone networks. While physical layer security technologies are considered an effective alternative, current research is mostly limited to single-drone jamming. Due to the power and size limitations of individual drones, their jamming capabilities are limited and they struggle to cope with complex eavesdropping environments. Second, there is the issue of system cost and efficiency. Drones have limited onboard battery capacity, and their energy consumption includes flight propulsion, hovering, computation, and communication energy. Balancing these energy consumptions with training latency while ensuring security is a complex resource scheduling problem. Existing research often focuses only on optimizing a single aspect and lacks a comprehensive system-wide approach. Considering the overall cost; finally, there is the issue of model accuracy. Fading and noise in the wireless channel can interfere with the aggregation of model parameters, directly affecting the convergence performance of federated learning. Furthermore, the selection of UAV nodes is crucial to the model's convergence speed and accuracy; random or blind selection strategies often lead to slow convergence. In summary, existing technologies lack security defense mechanisms for multi-UAV collaborative scenarios and fail to collaboratively optimize the three mutually constraining objectives of security, overall system cost, and model accuracy within a unified framework. Moreover, related joint optimization problems typically exhibit high non-convexity, strong coupling, and mixed action spaces containing continuous and discrete variables. Traditional convex optimization methods struggle to handle dynamically changing network topologies, resulting in low solution efficiency and a tendency to get trapped in local optima. Therefore, there is an urgent need for an optimization method that can leverage the advantages of multi-UAV collaboration to achieve the optimal balance between security, cost, and accuracy through intelligent decision-making. Summary of the Invention

[0003] To address the shortcomings of existing technologies in collaboratively optimizing security, overall system cost, and model accuracy in multi-UAV federated learning networks, as well as the difficulty of traditional optimization methods in handling high-dimensional hybrid action spaces and dynamic network topologies, this invention provides a federated learning security optimization method based on cooperative beamforming and node selection in multi-UAV networks. This method solves the technical problem of achieving a comprehensive balance of multiple performance indicators for collaborative defense and efficient learning among multi-UAVs under complex electromagnetic environments and resource-constrained conditions.

[0004] To achieve the aforementioned objectives, the present invention employs a federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network, comprising: S1. Establish a wireless communication system model consisting of a base station, multiple drones and a potential eavesdropping drone. Introduce a dynamic role division mechanism to divide the drone swarm into a learning group that performs federated learning tasks and an interference group that performs cooperative interference tasks. Mathematically represent the channels of each communication link. S2. Based on the system model, quantitative modeling is performed on the three core performance indicators of system overall cost, model accuracy and communication security. S3. The three indicators of quantified system cost, safety and accuracy are weighted and combined through weight coefficients to construct a joint optimization multi-objective function with specific system parameters as optimization variables, and it is established as a Markov decision process. S4. Design a deep reinforcement learning algorithm based on the maximum entropy principle for Soft Actor-Critic (SAC) to solve the joint optimization multi-objective function; S5. In the policy network of the solution algorithm, graph attention and multi-head attention mechanisms are integrated to capture the dynamic spatial topological features of the UAV swarm and process the mixed action space.

[0005] Furthermore, the mathematical characterization of the channel described in S1 specifically involves: modeling the air-to-ground channel from the learning group UAV to the base station as a probabilistic channel model that includes line-of-sight transmission probability and non-line-of-sight transmission probability; modeling the channel from the jamming group UAV to the eavesdropper as a line-of-sight dominated air-to-air channel model; and using the array factor to characterize the interference gain of the jamming group's cooperative beamforming on the eavesdropper.

[0006] Furthermore, the quantitative modeling of core performance indicators described in S2 includes: establishing a weighted system comprehensive cost model that takes into account all UAV flight time, communication time, flight propulsion energy consumption, interference energy consumption, and computing energy consumption as a cost indicator; establishing a physical layer security rate model that uses cooperative beamforming to transmit artificial noise to interfere with eavesdroppers as a security indicator; and establishing an accuracy cost model determined by the convergence error theoretical bound of federated learning in noisy channels and the related terms of the optimization variables as an accuracy indicator.

[0007] Furthermore, the specific system parameters described in S3 include at least the number of local training rounds and signal transmission power of the learning group UAV, and the three-dimensional target interference position and cooperative interference excitation current weight of the interference group UAV.

[0008] Furthermore, the SAC-based deep reinforcement learning algorithm described in S4 is specifically executed as follows: the agent interacts with the environment to collect experience data and stores it in the experience replay pool; the critic network is updated by minimizing the Bellman error; and the actor network policy is updated by maximizing the cumulative reward in combination with the maximum entropy term.

[0009] Furthermore, the integrated graph attention and multi-head attention mechanism described in S5 is specifically manifested as follows: a graph attention network is used as a state encoder to dynamically calculate the attention weights of neighboring nodes based on the relative distance between UAVs to aggregate spatial features; a multi-head actor network structure is adopted to output discrete training round actions and continuous power and position actions through parallel decision heads.

[0010] The beneficial effects of this invention are as follows: 1. Achieving a dynamic balance between multi-UAV collaborative defense and system effectiveness: This invention proposes a collaborative beamforming framework based on dynamic role division and scoring node selection. It intelligently selects nodes by comprehensively evaluating learning gain, cost, and location dispersion, and utilizes directional beams generated by collaborative interference to overcome the defense bottleneck caused by the power limitations of single UAVs. This mechanism significantly enhances physical layer security while effectively balancing federated learning efficiency and overall system cost.

[0011] 2. A theory-driven accuracy optimization metric was constructed: Addressing the difficulty in directly optimizing the accuracy of federated learning models, this invention extracts an "accuracy cost" term directly related to the optimization variables from the upper bound of the convergence error through rigorous theoretical derivation. This provides a solid mathematical foundation for the design of the reward function, rather than relying solely on empirical inspiration, thereby effectively accelerating model convergence.

[0012] 3. Achieves intelligent decision-making and efficient solution in complex environments: To address the challenges of highly dynamic changes in the topology of multi-UAV networks and the mixed action space, this invention designs a Graph Attention Multi-Head Actor-Soft Actor-Critic (GAMA-SAC) algorithm. This algorithm effectively extracts spatial cooperation features through a graph attention network and uses a multi-head attention mechanism to process mixed actions in parallel. It avoids the problem that traditional optimization algorithms are prone to getting trapped in local optima or having low solution efficiency in non-convex, high-dimensional problems, and has extremely strong robustness and real-time performance. Attached Figure Description

[0013] Figure 1 A flowchart of a federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network.

[0014] Figure 2 This is a scenario model diagram of a multi-UAV collaborative federated learning system in an embodiment of the present invention, illustrating the spatial relationships between the base station, the learning group UAVs, the jamming group UAVs, and the eavesdropper. Detailed Implementation

[0015] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0016] like Figure 1 As shown, in one embodiment of the present invention, a federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network is provided, the method comprising: S1. Establish a wireless communication system model consisting of a base station, multiple drones, and a potential eavesdropper. Introduce a dynamic role division mechanism to divide the drone swarm into a learning group and an interference group, and mathematically represent the channels of each communication link. S2. Based on the system model, quantitative modeling is performed on the three core performance indicators of system overall cost, model accuracy and communication security. S3. The three indicators of quantified system cost, safety and accuracy are weighted and combined through weight coefficients to construct a joint optimization multi-objective function with specific system parameters as optimization variables, and it is established as a Markov decision process. S4. Design a SAC deep reinforcement learning algorithm based on the maximum entropy principle to solve the joint optimization multi-objective function; S5. In the policy network of the solution algorithm, graph attention and multi-head attention mechanisms are integrated to capture the dynamic spatial topological features of the UAV swarm and process the mixed action space.

[0017] To clearly define the technical environment and physical premises in which this invention is applied, system and channel modeling is performed in S1. Specifically, this invention considers a three-dimensional spatial scenario including a ground base station, a cluster of N drones, and a potential aerial mobile eavesdropper. In each round of federated learning communication, the system performs dynamic role allocation: Q drones are selected from the N drones according to a specific strategy as learning drones. One unit is responsible for performing local training and uploading model parameters; the remaining NQ units automatically become jamming drones. They are responsible for coordinating the emission of artificial noise to interfere with eavesdroppers.

[0018] For the communication link in this scenario, this invention provides the following mathematical representation: First, considering the complex ground environment, the air-to-ground channel from the learning UAV to the base station is modeled as a probabilistic channel model including line-of-sight and non-line-of-sight transmissions. For the ... The probability of a learning drone establishing a line-of-sight link with a base station. Depends on its elevation angle The higher the elevation angle, the higher the line-of-sight probability. The corresponding average path loss... The calculation is a weighted sum of line-of-sight and non-line-of-sight path losses. Secondly, for the interference channel from the jamming drone swarm to the eavesdropper, cooperative beamforming is used to treat the jamming drone swarm as a distributed antenna array, and an array factor is introduced. To characterize the directional gain of the interference beam in space.

[0019]

[0020] in, It is the excitation current weight of the j-th interfering UAV. Represents the imaginary unit. Is related to wavelength The relevant phase constant. Furthermore... and These are the elevation and azimuth angles, respectively. The array factor describes the directional gain of the jamming beam in space, and its definition involves the three-dimensional coordinates of each jamming UAV and the weight of its emitted excitation current. By optimizing the current weights and positions of each jamming node, the jamming energy can form a high-gain main lobe in the direction of the eavesdropper, thereby maximizing the jamming effect. Finally, the air-to-air channel from the learning UAV to the eavesdropper is modeled as a line-of-sight dominated channel, where the path loss is mainly determined by the Euclidean distance between the UAV and the eavesdropper.

[0021] To translate performance objectives from the physical world into mathematical language that can be processed by algorithms, S2 requires quantitative modeling of the system's core metrics. This process mainly includes defining the drone selection score, overall system cost, communication security, and the cost of model accuracy. S21. Establish a drone selection and scoring model. To achieve intelligent node role allocation, this invention defines a comprehensive scoring index. To evaluate the first The potential value of a drone. This score is composed of three weighted components: first, the learning gain, defined as the average loss value of the drone's local dataset under the current model, i.e.: , A larger loss value indicates a greater potential contribution of the data to model optimization; secondly, the learning cost quantifies the computational and communication resource consumption required for the UAV to participate in training, expressed as: , Lower cost is better; finally, there is the dispersion index, defined as the average distance between the drone and other nodes, i.e.: , The efficiency of cooperative beamforming largely depends on the compactness of the geometry of the jamming drone swarm. A drone already in the center, close on average to the other drones, if selected as a jammer, will help form a more compact jamming array, thus reducing the total movement cost required for all jammers to reach the optimal jamming position. Therefore, a higher index means that it is more difficult for that node to form a compact beam array if it is used as a jamming node; conversely, the closer it is to the center, the more suitable it is for cooperative jamming. The final total score is calculated as follows: The system then selects the optimal learning node based on this information.

[0022] S22. Establish a comprehensive system cost model. This invention constructs a comprehensive cost function that includes both time and energy dimensions. The time cost is comprised of the total latency of each round of federated learning. The total latency is determined by the longest time taken for all learning drones to complete local training and parameter upload, i.e.: . Energy cost This covers the computing, communication, and hovering energy consumption of the learning group's UAVs, as well as the jamming group's UAVs' jamming launch and flight propulsion energy consumption. The overall cost is expressed as a weighted sum of the two: .

[0023] S23. Establish a physical layer security rate model. To quantify communication security, define the security rate. To learn the difference between the legal channel capacity from the drone to the base station and the eavesdropping channel capacity to the eavesdropper, Shannon's formula is used, and its expression is:

[0024] in, This represents the signal-to-interference-plus-noise ratio at the base station. The signal-to-interference-plus-noise ratio at the eavesdropper's location. Specifically, in calculating... At that time, the denominator included the strong interference power generated by the jamming group's UAVs through cooperative beamforming. The interference power is optimized by adjusting the current weights of the interference nodes. The beam gain is formed in the direction of the eavesdropper, thereby effectively suppressing the eavesdropping signal-to-noise ratio.

[0025] S24. Establish a model accuracy cost model. Since the final accuracy of federated learning is difficult to directly use as a real-time term of the optimization function, this invention, based on convergence analysis under noisy channels, derives a theoretical bound on the model convergence error and extracts the part directly related to the optimization variables as the accuracy cost. This cost consists of a communication error term and a local update drift term, and the specific formula is as follows: . The first term in the formula reflects the transmission power. The first term reflects the impact on suppressing uplink noise and reducing aggregation error; the second term reflects the number of local training rounds. The impact of model drift caused by non-independent and identically distributed data. Minimize This theoretically guarantees that the model converges to a higher accuracy. After quantifying the core performance metrics, S3 aims to construct joint optimization problems and model the solution process as a Markov decision process.

[0026] S31, Construct a joint optimization multi-objective function. To achieve synergistic balance among the multiple objectives, this invention introduces three non-negative weighting factors. These correspond to the overall system cost, the cost of model accuracy, and the cost of communication security, respectively. Therefore, a joint optimization objective function is constructed with the goal of minimizing the weighted sum:

[0027] in, It is the set of optimization variables. The objective function aims to minimize a combined cost consisting of three weighted components. This refers to system cost (including energy consumption and latency); The cost of precision in federated learning; It is the system's average security rate.

[0028] Constraints This ensures the communication security of all learning drones. and The range of values ​​for transmit power and number of local training rounds were restricted respectively. It is to ensure the synchronization of the jamming task and the learning task in time, including the movement time of the jamming drone. It is determined by its maximum speed in the horizontal and vertical directions. The regulations stipulate that drones must be operated within a pre-defined safe geographical space. and This ensured that the mission energy consumption of both the learning drone and the jamming drone did not exceed their currently available battery power. Finally, It is a safety distance constraint to avoid collisions between drones.

[0029] S32. Establish a Markov decision process. Considering that the above optimization problem is a complex problem that is non-convex, highly coupled, and contains mixed variables, traditional convex optimization methods are difficult to directly handle dynamic topological changes. Therefore, this invention reconstructs it as a Markov decision process, defining the following triples. Among them, the state space The aim is to enable intelligent agents to fully perceive their environment and their state. It includes current global and local information, specifically spatial topology information consisting of the three-dimensional position coordinates of all drones, base station locations, and predicted eavesdropper locations; resource status information consisting of the remaining energy level of each drone to prevent nodes from going offline due to energy depletion; channel feature information consisting of the channel gain estimated from the current time slot; and a dynamic graph adjacency matrix consisting of the drone positional relationships, which is used for subsequent graph neural network extraction of topological features; and action space to address the mixed variable characteristics of the problem. Defined as a hybrid action space It consists of two parts: discrete actions and continuous actions. The discrete action part is the number of globally uniform local training rounds. The continuous action section includes the transmit power of each learning drone. Increment of target flight position for each interfering drone and the excitation current weight of cooperative interference Finally, in the reward function In terms of design, in order to guide the agent toward minimizing the objective function In terms of direction learning, this invention will use the reward function Designed for total cost The negative value of the total cost It is a weighted sum that combines all optimization objectives:

[0030] in, This is a hard constraint penalty. We introduced a "safety shield" mechanism, which penalizes actions that would cause the drone spacing to be less than a certain value. If an action is deemed inappropriate, it will be prevented, and the agent will receive a severe penalty. By maximizing this reward function, the agent is guided to learn an optimal strategy that simultaneously balances communication security, system efficiency, learning performance, and flight safety.

[0031] For the highly coupled and complex optimization problem constructed in S3, direct solution is very difficult. Therefore, S4 aims to design an efficient deep reinforcement learning algorithm to solve this multi-objective joint optimization problem. Given that this problem has a high-dimensional state space and complex action constraints, this invention proposes a SAC algorithm framework based on the maximum entropy principle. This framework adopts a soft actor-critic structure, aiming to maximize the expected cumulative reward while maximizing the policy entropy, thereby enhancing the agent's exploration ability in complex environments and preventing premature convergence to local optima.

[0032] S4. SAC Deep Reinforcement Learning Solution Algorithm Based on the Maximum Entropy Principle. The specific solution process first establishes an experience replay pool. Used to store experience tuples generated by the interaction between the agent and the environment. During the training phase, the agent follows the current policy network. Sampling action Observe the new state after performing the action. And receive a reward The transformed data is then stored in the replay pool. To break the correlation between data and improve sample utilization, the algorithm randomly samples batches of data from the replay pool to update network parameters.

[0033] SAC consists of two core networks: parameters are... The actor network is used for the generation strategy, and the parameters are... The critic network is used to evaluate the value of actions, i.e., the Q-value. In the network update step, the critic network is first updated with the goal of minimizing the soft Bellman residual, i.e., minimizing the loss function through gradient descent. , The target value is: . The design of this target value incorporates an entropy regularization term. This allows the Q-value to reflect not only the expected cumulative reward but also the entropy of future states, thus encouraging the agent to explore action spaces with higher uncertainty. A double-Q network mechanism is employed here to mitigate the overestimation problem.

[0034] The actor network is then updated with the goal of maximizing the weighted sum of the Q-value and policy entropy, and the corresponding loss function is defined as: . To handle gradient backpropagation in a continuous action space, this invention employs a reparameterization technique, transforming the random action sampling process into a deterministic function transformation. This allows the policy network parameters to be directly optimized via gradient descent. Furthermore, the algorithm introduces an automatic entropy adjustment mechanism to dynamically adjust the temperature parameter. In the early stages of training, high entropy is maintained to encourage exploration, while entropy is reduced in the later stages of training to ensure the stability and accuracy of the strategy.

[0035] To enable the deep reinforcement learning algorithm in S4 to effectively handle the dynamically changing UAV network topology and complex hybrid action space in the scenario of this invention, S5 further integrates graph attention mechanism and multi-head output mechanism in the policy network and value network.

[0036] S5. Network Optimization Integrating Graph Attention and Multi-Head Attention Mechanisms. Addressing the dynamic topology problem of constantly changing relative positions in a drone swarm during flight and the potential for nodes to join or leave at any time, this invention employs a graph attention network at the agent's state encoding layer. Specifically, each drone is treated as a node in a graph. Its input original feature vector This includes the node's location, energy, and channel state. The graph attention network layer computes the node's... Its neighboring nodes Attention coefficient between To dynamically aggregate neighborhood information, the calculation formula is as follows: ,in For a shared linear transformation matrix, For attention vectors, This represents the concatenation operation. The normalized coefficient is used as a weight to perform a weighted summation of the features of neighboring nodes, thereby generating a new high-dimensional feature vector that incorporates local cooperative relationships. This mechanism allows the agent to focus on key neighbors that have the greatest impact on the current task, without being limited by changes in the total number of nodes.

[0037] To address the hybrid nature of the action space, which contains both discrete and continuous variables with distinct physical meanings, this invention designs a multi-head output structure within the actor network. The global feature vector, encoded by a graph attention network, is input into a shared fully connected layer and then splits into two independent decision heads: a global discrete decision head and a role-specific continuous decision head. The global discrete decision head uses a Softmax activation function to output a probability distribution, which is used to select a globally uniform number of local training rounds. This action applies to all learning nodes; the role-specific continuous decision head is further subdivided into learning subheads and interference subheads, which output the transmit power of the learning drone respectively. And the displacement of the interference drone and complex disturbance weights For continuous actions, the network output mean and standard deviation are sampled using a Gaussian distribution to maintain exploratory nature, and the Tanh function is used to map the output to a legal physical constraint range. This multi-head decoupling design avoids interference between action variables of different dimensions and properties within the same network layer, significantly improving the propagation efficiency of policy gradients and training stability.

[0038] Through the close coordination of the above steps, this invention constructs a complete closed-loop optimization system. First, at the physical layer, cooperative beamforming technology is used to overcome the limitations of single-machine defense. Then, through rigorous mathematical modeling, security, energy consumption, and accuracy are mapped into a calculable cost function. Finally, with the help of the improved GAMA-SAC algorithm, the optimal resource scheduling and interference strategies are found in complex dynamic environments, thereby achieving the best balance between security, energy consumption, and model accuracy in the UAV federated learning system.

[0039] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit the scope of protection of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; however, the above modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of protection of the technical solutions of the embodiments of this application.

[0040] In summary, compared with the prior art, the present invention has the following advantages: By establishing a unified optimization framework based on dynamic role division, this invention achieves coordinated optimization of three mutually constraining core indicators in multi-UAV cooperative networks: overall system cost, communication security, and model accuracy. It breaks through the limitations of existing technologies that rely solely on single-UAV defense or isolated research on a single indicator. It can significantly improve the physical layer security capacity by utilizing cooperative beamforming while effectively balancing federated learning efficiency and system resource consumption. To address the difficulty in directly optimizing the accuracy of federated learning models, this invention does not rely on experience or simulation methods. Instead, it derives and extracts a convergence error upper bound related to communication noise, transmission power, and the number of local training rounds through rigorous theoretical analysis of the convergence of federated learning in noisy channels. This term is used as an accuracy cost indicator, and quantitative optimization of accuracy is achieved based on solid mathematical evidence, which greatly improves the reliability of reward function design and the stability of model convergence. To address the challenges of solving problems in the highly dynamic topology and mixed action space of multi-UAV networks, this invention designs a GAMA-SAC deep reinforcement learning algorithm based on the maximum entropy principle. It innovatively integrates a graph attention network to capture spatial cooperation features and utilizes a multi-head attention mechanism to achieve parallel decoupling control of discrete and continuous mixed variables. This solves the bottleneck of traditional convex optimization algorithms in handling dynamic topology and non-convex constraints, and balances the real-time performance of the solution with global optimization capabilities. In the system modeling and mechanism design phase, this invention is fully adapted to the multi-UAV collaborative combat scenario. It introduces a UAV selection and scoring mechanism to intelligently classify learning and interference roles. The system model covers dynamic graph structure, base station and mobile eavesdropper. The security model uses multi-UAV collaborative beamforming to construct a virtual antenna array to overcome the power limitation of a single UAV. Moreover, the optimization variables are all actual adjustable flight and communication parameters, which makes the method flexibly adaptable to complex and ever-changing electromagnetic environments. It has strong practicality and applicability in resource-constrained multi-UAV federated learning scenarios.

Claims

1. A federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network, characterized in that, The method comprises the following steps: S1, a wireless communication system model composed of a base station, multiple unmanned aerial vehicles and a potential eavesdropping unmanned aerial vehicle is established, a dynamic role division mechanism is introduced to divide the unmanned aerial vehicle group into a learning group and an interference group, and the channels of each communication link are mathematically characterized; S2, based on the system model described in S1, the three core performance indicators of system comprehensive cost, model accuracy and communication security are quantitatively modeled; S3, the three indicators of system cost, security and accuracy after quantification are combined by weight coefficients, a joint optimization multi-objective function with specific system parameters as optimization variables is constructed, and it is established as a Markov decision process; S4, a soft actor-critic deep reinforcement learning algorithm based on the maximum entropy principle is designed to solve the joint optimization multi-objective function; S5, in the strategy network of the solving algorithm, graph attention and multi-head attention mechanisms are integrated to capture the dynamic spatial topological features of the unmanned aerial vehicle group and process the mixed action space.

2. The federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network according to claim 1, characterized in that, In S1, the channel is mathematically characterized as follows: the air-to-ground channel from the learning group unmanned aerial vehicle to the base station is modeled as a probabilistic channel model containing line-of-sight transmission probability and non-line-of-sight transmission probability, the channel from the interference group unmanned aerial vehicle to the eavesdropping unmanned aerial vehicle is modeled as an air-to-air channel model dominated by line-of-sight, and the interference gain of the interference group cooperative beamforming to the eavesdropper is characterized by an array factor.

3. The federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network according to claim 1, characterized in that, In S2, the quantitative modeling of the core performance indicators includes: S21, a weighted system comprehensive cost model considering the flight time, communication time, flight energy consumption, interference energy consumption and calculation energy consumption of all unmanned aerial vehicles is established as the cost indicator; S22, a physical layer secrecy rate model for transmitting artificial noise through cooperative beamforming to interfere with the eavesdropper is established as the security indicator; S23, an accuracy cost model determined by the related items of the optimization variables in the upper bound of the convergence error theory of federated learning in a noisy channel is established as the accuracy indicator.

4. The method of claim 3, wherein, The accuracy cost model determined by the accuracy indicator in S23 is derived based on the learning gain acceleration factor, the gradient heterogeneity upper bound and the communication channel noise variance, which can map the abstract model accuracy to a calculable mathematical expression.

5. The federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network according to claim 1, characterized in that, The specific system parameters in S3 include at least the local training round number, signal transmission power of the learning group unmanned aerial vehicle, and three-dimensional target interference position, cooperative interference excitation current weight of the interference group unmanned aerial vehicle.

6. The federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network according to claim 1, characterized in that, The deep reinforcement learning algorithm based on SAC in S4 has the following specific execution process: the agent interacts with the environment to collect experience data into an experience replay pool, updates the critic network by minimizing the Bellman error, and updates the actor network policy by maximizing the cumulative reward combined with the maximum entropy term.

7. The federated learning security optimization method based on cooperative beamforming and node selection in a multi-UAV network according to claim 1, characterized in that, The integration of graph attention and multi-head attention mechanisms in S5 is specifically manifested as follows: a graph attention network is used as a state encoder to dynamically calculate the attention weights of neighbor nodes according to the relative distance between unmanned aerial vehicles to aggregate spatial features; a multi-head actor network structure is used to output discrete training round number actions and continuous power and position actions through parallel decision heads.