Unmanned aerial vehicle cluster environment adaptive optimization method based on machine learning

Through the machine learning-based multi-agent reinforcement learning and federated learning framework, combined with adaptive communication and fault tolerance mechanisms, the problems of collaborative efficiency and robustness of drone clusters in complex environments are solved, and efficient and reliable task execution is achieved.

CN120669755AActive Publication Date: 2025-09-19XIAN BAOTONG DEFENSE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511156618.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-09-19
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing drone swarm control methods are difficult to achieve efficient collaboration under complex and changeable environmental conditions. They have problems such as extensive environmental modeling, insufficient robustness, and low collaboration efficiency. In particular, the mission failure rate is high when electromagnetic interference and communication interruption occur.

Method used

It adopts a multi-agent reinforcement learning framework and federated learning mechanism based on machine learning, combined with graph attention mechanism and multimodal data fusion, to achieve dynamic environment adaptation and fault tolerance, enhance anti-interference ability through adaptive communication protocol and optical communication link switching, and use LSTM network for fault detection and compensation.

Benefits of technology

It improves the efficiency and reliability of drone clusters in completing tasks in complex environments, reduces energy consumption and communication delays, enhances the ability to resist electromagnetic interference, and ensures mission success rate and positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669755A_ABST
    Figure CN120669755A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle cluster environment adaptive optimization method based on machine learning, which relates to the technical field of unmanned aerial vehicle cluster cooperative control and comprises the steps of basic framework construction, training parameter optimization, dynamic strategy adjustment, anti-interference communication enhancement and fault tolerance self-reconstruction. Coupling the optimized federated learning model with a multi-modal data fusion module, performing self-attention mechanism fusion on sensor data after time synchronization, constructing a high-dimensional state vector containing an agent state and an environment feature, inputting the high-dimensional state vector into a reinforcement learning framework to generate a joint action decision, and realizing fault detection through an LSTM network. And performing fault compensation by using redundant sensor data and the multi-modal fusion model. Through a dynamic graph attention mechanism and multi-target federated learning, the cluster can dynamically adjust a strategy to cope with complex environments such as electromagnetic interference and obstacle change, and a multi-agent collaborative decision and residual error compensation fault self-healing mechanism ensures that the cluster can still complete a task when a node fails or communication is interrupted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cooperative control of unmanned aerial vehicle (UAV) clusters, and in particular to a method for adaptively optimizing an UAV cluster environment based on machine learning. Background Art

[0002] Existing drone swarm control methods usually rely on preset rules or fixed strategies, which make it difficult to cope with complex and changing environmental conditions (such as electromagnetic interference, dynamic changes in obstacles, communication interruptions, etc.).

[0003] The existing technology has the following core defects: Extensive environmental modeling: Traditional methods rely on manually designed environmental features, which make it difficult to capture continuous spatial features such as dynamic electromagnetic spectrum and airflow disturbances, resulting in decision lag (average delay of more than 200ms).

[0004] Low collaboration efficiency: The centralized control architecture has a single point of failure (the mission failure rate exceeds 70% when the leading drone loses contact), while the distributed approach has high communication overhead, with a single node transmitting up to 500KB of data per second, resulting in excessive energy consumption.

[0005] Insufficient robustness: Under strong electromagnetic interference of -20dBm, the traditional communication protocol has a bit error rate of over 30%, and lacks a dynamic compensation mechanism when the sensor fails. When monocular vision fails, the positioning error increases by 5 times.

[0006] In view of this, the present invention is proposed to solve the above technical problems. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for adaptive optimization of UAV cluster environment based on machine learning, so as to solve the technical problems of existing UAV cluster optimization methods such as extensive environmental modeling, low collaborative efficiency and insufficient robustness.

[0008] The purpose of the present invention is to provide a method for adaptive optimization of a UAV cluster environment based on machine learning, comprising the following steps: S1. Basic framework construction: Model the interaction between agents and the environment through the graph attention mechanism, generate action strategies based on the hierarchical actor-critic network architecture, and form a multi-agent reinforcement learning framework; S2. Training parameter optimization: Each drone uses the environmental data collected by local multimodal sensors to conduct distributed training through a federated learning mechanism to establish a multi-objective federated learning model, providing efficient model support for real-time decision-making; S3, dynamic strategy adjustment, couples the optimized federated learning model with the multimodal data fusion module, fuses the time-synchronized sensor data through a self-attention mechanism, constructs a high-dimensional state vector containing the agent state and environmental characteristics, and inputs it into the reinforcement learning framework to generate joint action decisions; S4, anti-interference communication enhancement: During the action decision execution process, the spectrum sensing module monitors the electromagnetic environment in real time. When the energy density in the 2.4GHz band is detected to be greater than −80dBm / MHz and lasts for more than 200ms, the three-layer communication adaptive mechanism is triggered; S5, fault-tolerant self-reconstruction, for possible sensor failure or node loss scenarios, fault detection is achieved through the LSTM network, and fault compensation is performed using redundant sensor data and multimodal fusion models.

[0009] Furthermore, the multi-agent reinforcement learning framework adopts a shared experience pool and distributed training, and the agents achieve collaborative optimization through joint action decision-making.

[0010] Furthermore, the federated learning mechanism dynamically adjusts the local training rounds and convergence accuracy through a successive convex approximation algorithm to minimize the total energy consumption of the cluster.

[0011] Furthermore, multimodal data fusion adopts the self-attention mechanism to construct the environment state vector to support the decision-making of the reinforcement learning model.

[0012] Furthermore, the anti-interference communication enhancement in step S4 adopts an adaptive communication protocol, including dynamic adjustment of OLSR parameters and optical communication link switching, to cope with electromagnetic interference.

[0013] Furthermore, fault-tolerant self-reconfiguration adopts fuzzy adaptive control and distributed consensus algorithm to achieve sensor fault detection and cluster topology self-reconfiguration.

[0014] Furthermore, in the multi-agent reinforcement learning framework, a dynamic interaction graph consisting of environment nodes and agent nodes is constructed through the graph attention mechanism, and the agent collaboration weights are calculated in real time. The collaboration weight calculation formula is: ,in, is the collaborative attention weight between agents, is the state vector of agent i at time t, is the state vector of agent j at time t, leakyReLU is the activation function, Softmax j To perform Softmax normalization weight on the j dimension, is the weight matrix of the attention mechanism.

[0015] The federated learning mechanism achieves energy-precision balance through a multi-objective optimization function. The optimization function is: + - A, and , represent the local training round and convergence accuracy threshold respectively, 、 、 is the weight coefficient, and A are the objective function terms.

[0016] Furthermore, the optical communication link switching condition is: when the energy density of the 2.4GHz band is greater than −80dBm / MHz and the duration is greater than 200ms, the 850nm visible light communication module is enabled, and the OLSR protocol Hello interval is adjusted to: T hello =T0˙ ,in, is the link connectivity probability, β is the adaptive coefficient of 0.5-1.0, T hello The interval for sending Hello messages is adjusted dynamically, and T0 is the Hello interval.

[0017] By adopting the above technical solution, the present invention has the following beneficial effects: Environmental adaptability: Through the dynamic graph attention mechanism and multi-objective federated learning, the cluster can dynamically adjust its strategy to cope with complex environments such as electromagnetic interference and obstacle changes, shortening task completion time by 40% and reducing energy consumption by 35%.

[0018] Communication efficiency: Federated learning reduces data transmission by 92%, and the communication success rate of visible light communication remains above 98% under strong interference, with latency reduced by 53% and energy consumption reduced by 63%.

[0019] Mission reliability: Multi-agent collaborative decision-making and residual compensation fault self-healing mechanism ensure that the cluster can still complete the task when a node fails or communication is interrupted. The positioning accuracy only drops by 15% when a sensor fails, and the cluster's self-reconfiguration success rate reaches 99.2%. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are part of this application and are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but do not constitute an undue limitation of the present invention. Obviously, the drawings described below are only some embodiments. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without inventive effort. In the accompanying drawings: Figure 1 This is a framework diagram of the UAV cluster environment adaptive optimization method based on machine learning provided in this embodiment of the present application.

[0021] It should be noted that these drawings and textual descriptions are not intended to limit the conceptual scope of the present invention in any way, but rather to illustrate the concept of the present invention for those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0022] The present application provides a method for adaptively optimizing a drone cluster environment based on machine learning, comprising the following steps: S1. Basic Framework Construction: The graph attention mechanism is used to model the interaction between agents and between agents and the environment. Action strategies are generated based on a hierarchical actor-critic network architecture to form a multi-agent reinforcement learning framework. Each node in the drone swarm acts as an independent agent, and a dynamic interaction graph containing environmental nodes (obstacles, interference sources) and agent nodes is constructed. The collaborative weights between agents are calculated in real time through the graph attention mechanism (GAT). Actor network: It is divided into a global strategy layer and an individual execution layer. The global strategy layer outputs the cluster-level task allocation matrix ϵ ,The individual execution layer outputs flight control commands (speed, heading angle, obstacle avoidance acceleration) based on the global decision and local state; Critic Network: This introduces a contrastive learning mechanism that improves the accuracy of value function estimation by using positive and negative sample pairs, reducing training loss by 40% compared to traditional methods. Agents perform distributed training using a shared experience pool, dynamically adjusting their strategies to adapt to environmental changes.

[0023] S2. Training parameter optimization: Each UAV uses the environmental data collected by local multimodal sensors to conduct distributed training through the federated learning mechanism, establish a multi-objective federated learning model, provide efficient model support for real-time decision-making, and solve the optimal local training rounds through the Lagrange multiplier method. and convergence accuracy , while ensuring global model accuracy while reducing communication energy consumption and latency. An asynchronous aggregation mechanism is introduced, with the leading drone setting a dynamic aggregation window (100ms-500ms), allowing nodes with weak computing power to prioritize uploading lightweight model parameters (transmitting only gradient symbols, reducing data volume by 90%). S3, dynamic strategy adjustment, couples the optimized federated learning model with the multimodal data fusion module, fuses the time-synchronized sensor data through a self-attention mechanism, constructs a high-dimensional state vector containing the agent state and environmental characteristics, and inputs it into the reinforcement learning framework to generate joint action decisions; S4, anti-interference communication enhancement: During the action decision execution process, the spectrum sensing module monitors the electromagnetic environment in real time. When the energy density in the 2.4GHz band is detected to be greater than −80dBm / MHz and lasts for more than 200ms, the three-layer communication adaptive mechanism is triggered. The three-layer communication adaptive mechanism is as follows: Physical layer: Enable 850nm wavelength visible light communication module (communication distance 500m, bit error rate <10 -6 ); Data link layer: Time division multiple access (TDMA) is used instead of contention access to reduce the probability of collision; Network layer: Predict link status based on dynamic Bayesian networks and optimize the Hello message interval of the OLSR protocol; S5, fault-tolerant self-reconstruction, for possible sensor failure or node loss scenarios, fault detection is achieved through the LSTM network, and fault compensation is performed using redundant sensor data and multimodal fusion models.

[0024] The multi-agent reinforcement learning framework adopts a shared experience pool and distributed training, and the agents achieve collaborative optimization through joint action decision-making.

[0025] The federated learning mechanism dynamically adjusts the local training rounds and convergence accuracy through a successive convex approximation algorithm to minimize the total energy consumption of the cluster.

[0026] Multimodal data fusion utilizes a self-attention mechanism to construct an environmental state vector to support the reinforcement learning model's decision-making. The drone is equipped with multimodal sensors (LiDAR, cameras, spectrum sensing modules, etc.) to collect environmental data in real time. After time synchronization (accuracy <1ms), the self-attention mechanism is used to fuse this data, constructing an 18-dimensional environmental state vector (including position, speed, battery level, sensor status, and environmental characteristics). This is then input into the reinforcement learning model to generate joint action decisions, enabling coordinated optimization of path planning, task allocation, and obstacle avoidance. Single-frame data processing takes less than 15ms.

[0027] The anti-interference communication enhancement in step S4 adopts an adaptive communication protocol, including dynamic adjustment of OLSR parameters and optical communication link switching, to cope with electromagnetic interference.

[0028] Fault-tolerant self-reconfiguration uses fuzzy adaptive control and a distributed consensus algorithm to implement sensor fault detection and cluster topology self-reconfiguration. It also builds a long short-term memory (LSTM) prediction model. When the inertial measurement unit (IMU) data residual exceeds three times the standard deviation, fault isolation is triggered (detection delay <50ms). The self-reconfiguration algorithm uses an improved distributed consensus protocol: , where α is the inertia coefficient, is the adjacency matrix weight, and the smooth reconstruction of the cluster topology (reconstruction time < 200ms) is achieved by dynamically adjusting α (0.3-0.7).

[0029] In the multi-agent reinforcement learning framework, a dynamic interaction graph consisting of environment nodes and agent nodes is constructed through the graph attention mechanism, and the collaborative weights of the agents are calculated in real time. The collaborative weight calculation formula is: ,in, is the collaborative attention weight between agents, is the state vector of agent i at time t, is the state vector of agent j at time t, leakyReLU is the activation function, Softmax j To perform Softmax normalization weight on the j dimension, is the weight matrix of the attention mechanism.

[0030] The federated learning mechanism achieves energy-precision balance through a multi-objective optimization function. The optimization function is: + - A, and , represent the local training round and convergence accuracy threshold respectively, 、 、 is the weight coefficient, and A are the objective function terms.

[0031] The optical communication link switching condition is: when the energy density of the 2.4GHz band is greater than −80dBm / MHz and the duration is greater than 200ms, the 850nm visible light communication module is enabled and the OLSR protocol Hello interval is adjusted to: T hello =T0˙ ,in, is the link connectivity probability, β is the adaptive coefficient of 0.5-1.0, T hello The interval for sending Hello messages is adjusted dynamically, and T0 is the Hello interval.

[0032] 1. System deployment The drone swarm consists of k follower drones and one leading drone, with the following hardware configuration: Table 1 Hardware parameters of drone cluster ; 2. Training phase (1) Initialization State space: Define an 18-dimensional state vector S, which includes position (x, y, z), velocity (v_x, v_y, v_z), quaternion attitude (q_1, q_2, q_3, q_4), remaining battery E, sensor health status (3D one-hot encoding), and environmental characteristics (6 dimensions such as obstacle distance and interference intensity). Action space: Contains 12 discrete actions, including 4 types of mission assignment (reconnaissance, transportation, relay, and return) and 8 types of flight control (forward, backward, left, right, translation, lifting, and turning). (2) Online learning At each time step t, agent i collects the local state , generating actions through the DGAT-MARL network , and get rewards after execution (The reward function includes five dimensions, including task completion, energy consumption, and safety, and adopts a sparse reward + intrinsic incentive mechanism.) Experience tuple ( , , , ) are stored in the local experience pool. When the capacity reaches 1000, old experience training is mixed in at a ratio of 1:10 to avoid overfitting. (3) Federated Learning Aggregation The leading drone initiates parameter aggregation every 5 seconds, and each node dynamically adjusts the local training rounds based on the remaining power (When the battery level is >50% =20, otherwise =5), and reduces the communication volume by using gradient compression technology (only transmitting the gradients with the top 20% of absolute values), reducing the amount of single aggregation data from 500KB to 40KB.

[0033] 3. Task execution phase (1) Environmental perception and decision-making Multimodal sensor data is fused through a self-attention mechanism to generate a 128-dimensional environment state vector, which is input into the DGAT-MARL network to output a joint action, including the task allocation matrix T and individual control instructions. The kinematic model converts this into a PWM signal to drive the motor, and the decision delay is controlled within 80ms. (2) Implementation of anti-interference communication The spectrum analyzer scans the 2-6 GHz band at 100 Hz. If interference power exceeds the threshold for three consecutive cycles, the visible light communication module is initialized (generating synchronization signals and configuring optical intensity modulation parameters, which takes less than 50 ms). The OLSR protocol also dynamically adjusts the Hello interval based on link quality (default 1 second, shortened to 0.2 seconds in the presence of interference). Network coding technology is also enabled to increase throughput by 30%.

[0034] (3) Fault handling mechanism When the LSTM model predicts a residual error greater than 3σ for two consecutive cycles, a fault signal is sent to adjacent nodes, initiating sensor data fusion compensation. (If GPS fails, positioning is performed using visual SLAM and barometer fusion, maintaining an accuracy of ±50 cm.) When a node fails, adjacent nodes recalculate neighbor weights using a distributed consensus algorithm, completing topology reconstruction within three communication cycles (each 100 ms).

[0035] This specific embodiment is merely an explanation of the invention and is not a limitation of the invention. After reading this specification, those skilled in the art may make non-creative modifications to this embodiment as needed. However, as long as they are within the scope of protection of the invention, they are protected by patent law.

Claims

1. A self-adaptive optimization method for UAV cluster environment based on machine learning, characterized by: The following steps are involved: S1. Basic framework construction: Model the interaction between agents and the environment through the graph attention mechanism, generate action strategies based on the hierarchical actor-critic network architecture, and form a multi-agent reinforcement learning framework; S2. Training parameter optimization: Each drone uses the environmental data collected by local multimodal sensors to conduct distributed training through a federated learning mechanism to establish a multi-objective federated learning model, providing efficient model support for real-time decision-making; S3, dynamic strategy adjustment, couples the optimized federated learning model with the multimodal data fusion module, fuses the time-synchronized sensor data through a self-attention mechanism, constructs a high-dimensional state vector containing the agent state and environmental characteristics, and inputs it into the reinforcement learning framework to generate joint action decisions; S4, anti-interference communication enhancement: During the action decision execution process, the spectrum sensing module monitors the electromagnetic environment in real time. When the energy density in the 2.4GHz band is detected to be greater than −80dBm / MHz and lasts for more than 200ms, the three-layer communication adaptive mechanism is triggered; S5, fault-tolerant self-reconstruction, for possible sensor failure or node loss scenarios, fault detection is achieved through the LSTM network, and fault compensation is performed using redundant sensor data and multimodal fusion models.

2. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 1 is characterized in that: The multi-agent reinforcement learning framework adopts a shared experience pool and distributed training, and the agents achieve collaborative optimization through joint action decision-making.

3. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 2 is characterized in that: The federated learning mechanism dynamically adjusts local training rounds and convergence accuracy through a successive convex approximation algorithm to minimize the total energy consumption of the cluster.

4. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 3 is characterized in that: The multimodal data fusion adopts a self-attention mechanism to construct an environment state vector to support the decision-making of the reinforcement learning model.

5. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 1, characterized in that: The anti-interference communication enhancement in step S4 adopts an adaptive communication protocol, including dynamic adjustment of OLSR parameters and optical communication link switching, to cope with electromagnetic interference.

6. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 1, characterized in that: The fault-tolerant self-reconfiguration adopts fuzzy adaptive control and distributed consensus algorithm to realize sensor fault detection and cluster topology self-reconfiguration.

7. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 2 is characterized in that: In the multi-agent reinforcement learning framework, a dynamic interaction graph containing environment nodes and agent nodes is constructed through the graph attention mechanism, and the agent collaboration weight is calculated in real time. The collaboration weight calculation formula is: ,in, is the collaborative attention weight between agents, is the state vector of agent i at time t, is the state vector of agent j at time t, leakyReLU is the activation function, Softmax j To perform Softmax normalization weight on the j dimension, is the weight matrix of the attention mechanism.

8. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 3 is characterized in that: The federated learning mechanism achieves energy-precision balance through a multi-objective optimization function, which is: + - A, and , represent the local training round and convergence accuracy threshold respectively, 、 、 is the weight coefficient, and A are the objective function terms.

9. The method for adaptive optimization of UAV cluster environment based on machine learning according to claim 5, characterized in that: The optical communication link switching condition is: when the energy density of the 2.4GHz band is greater than −80dBm / MHz and the duration is greater than 200ms, the 850nm visible light communication module is enabled, and the OLSR protocol Hello interval is adjusted to: T hello =T0˙ ,in, is the link connectivity probability, β is the adaptive coefficient of 0.5-1.0, T hello The interval for sending Hello messages is adjusted dynamically, and T0 is the Hello interval.

Citation Information

Patent Citations

  • Multi-agent confrontation method and system based on dynamic graph neural network

    CN113627596A

  • Federal learning method and system

    CN116861239A

  • Multi-agent cluster consistency cooperative control method based on behavior prediction

    CN117319232A

  • Multi-rotor unmanned aerial vehicle cluster training neural network model by applying federated learning framework

    CN117873171A

  • Extensible deep reinforcement learning multi-unmanned aerial vehicle path planning cooperation method

    CN117930864A

Cited By

  • Navigation obstacle avoidance method and device, electronic equipment, storage medium and program product

    CN121254846A

  • Complex mine radar adaptive anti-interference detection method based on reinforcement learning

    CN121477158A

  • Multi-agent reinforcement learning method for task allocation of unmanned aerial vehicle cluster

    CN121581137A

  • Toughness recovery method, system and device for unmanned aerial vehicle cluster network, and storage medium

    CN122205481A