Multi-unmanned aerial vehicle cooperative efficient communication method and system based on low earth orbit satellite

Through the multi-UAV collaborative communication method based on low-orbit satellites, using reinforcement learning for dynamic resource allocation and security verification, the problem of insufficient resource management and security in multi-UAV systems is solved, and the reliability and security of efficient communication and task execution are achieved.

CN120454814APending Publication Date: 2025-08-08BEIJING TH SMART AVIATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510468746.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In multi-UAV systems, the prior art is difficult to achieve efficient resource allocation management and control, and the lack of an effective security verification mechanism, resulting in insufficient performance and security of the communication system.

Method used

A multi-UAV collaborative communication method based on low-orbit satellites is adopted, and a dynamic resource allocation is allocated using reinforcement learning, and a security verification mechanism is added during the training process to build a framework structure of state space, action space and reward space, and combine the simulation environment of satellite orbit, drone dynamics and channel model to optimize communication decisions.

Benefits of technology

It realizes efficient communication and task execution efficiency, ensures policy reliability and system security and robustness, reduces system interference and spectrum conflicts, and improves communication performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120454814A_ABST
    Figure CN120454814A_ABST
Patent Text Reader

Abstract

The invention provides a multi-unmanned aerial vehicle cooperative efficient communication method based on a low-orbit satellite, and the method is characterized in that the method comprises the following steps: obtaining all related parameters in the communication of multiple unmanned aerial vehicles and the low-orbit satellite, and defining the frame structures of a state space, an action space and a reward space according to the related parameters, and obtaining a first frame structure; a simulation tool is used for constructing a simulation environment integrating a satellite orbit model, a multi-unmanned aerial vehicle dynamics model and a channel model, a first virtual simulation environment is obtained, and safety verification is added in the training process to filter dangerous actions; and verifying the performance of the optimal model I in the virtual simulation environment I, and migrating the optimal model I to an actual system for testing to complete a real-time communication decision between the multiple unmanned aerial vehicles and the low-orbit satellite. Dynamic resource allocation is carried out by adopting reinforcement learning, so that high communication and task execution efficiency is realized; and a security verification mechanism is added in the decision training process, so that the strategy reliability and the system security and robustness are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of low-orbit satellite communication technology, and in particular to an efficient communication method and system for multi-UAV collaboration based on low-orbit satellites. Background Art

[0002] As a new type of aircraft, drones are increasingly used in aerial photography, agriculture, plant protection, micro selfies, express delivery, disaster relief, wildlife observation, infectious disease monitoring, surveying and mapping, news reporting, power inspection, disaster relief, film and television shooting, etc., and they play an irreplaceable role in life; network communication technology based on low-orbit satellite technology brings more advantages to the practical application of drones, and can provide a larger range of Internet service platforms and more accurate positioning systems; on the other hand, complex environments and multi-terminal interconnected communication platforms require efficient communication modules to ensure their system performance and system control and management capabilities; and when multiple drone systems work together, efficient resource allocation management and control is one of the goals of intelligent optimization of drones equipped with low-orbit satellite communications. Summary of the Invention

[0003] In view of the defects in drone communication, the present invention provides an efficient communication method and system for multi-drone collaboration based on low-orbit satellites, adopts reinforcement learning methods for dynamic resource allocation, achieves efficient communication and task execution efficiency, and adds a security verification mechanism during the reinforcement learning training process to ensure the reliability of the strategy and the security and robustness of the system.

[0004] In a first aspect, the present invention provides an efficient communication method for multi-UAV collaboration based on low-orbit satellites, characterized by comprising:

[0005] Obtain relevant parameters in the communication between multiple UAVs and low-orbit satellites, and use them to define the framework structure of the state space, action space, and reward space to obtain framework structure 1;

[0006] A simulation environment is constructed using a simulation tool that integrates a satellite orbit model, a multi-UAV dynamics model, and a channel model to obtain a first virtual simulation environment. A deep reinforcement learning model constructed using the first framework is trained and hyperparameters are adjusted to obtain the first optimal model. Safety verification is incorporated into the training process to filter out dangerous actions.

[0007] Verify the performance of the optimal model in the virtual simulation environment and migrate it to the actual system for testing to complete real-time communication decisions between multiple UAVs and low-orbit satellites.

[0008] Furthermore, the state space uses environmental parameters, drone status, network load, and historical information as main reference parameters; the action space uses resource allocation actions, collaborative control actions, and mode switching as main reference parameters; and the reward function is composed of communication performance rewards, negative rewards, and multi-objective trade-offs.

[0009] Furthermore, the efficient communication method for multi-UAV collaboration based on low-orbit satellites is characterized in that the multi-objective trade-off includes a communication performance reward based on throughput and latency, a negative reward based on energy consumption and spectrum conflict, and a multi-objective trade-off with a weighted balance of multiple reference targets.

[0010] Furthermore, the multi-UAV dynamics model adopts a relative motion model, a cooperative dynamics equation and an external interference model.

[0011] Furthermore, the multi-UAV dynamics model also adopts multi-machine coupling effect modeling, including communication and control coupling, and downwash airflow interference modeling.

[0012] Furthermore, the satellite orbit model adopts a multi-satellite constellation modeling and orbit visualization module.

[0013] The second aspect is an efficient communication system for multi-UAV collaboration based on low-orbit satellites, including the following modules:

[0014] The communication input module obtains the relevant parameters of the communication between multiple UAVs and low-orbit satellites, and uses them to define the framework structure of the state space, action space and reward space, obtaining framework structure 1;

[0015] The communication model training module uses simulation tools to build a simulation environment that integrates a satellite orbit model, a multi-UAV dynamics model, and a channel model to obtain a virtual simulation environment. The module then trains a deep reinforcement learning model based on the framework structure and adjusts hyperparameters to obtain the optimal model. Safety verification is incorporated into the training process to filter out dangerous actions.

[0016] The communication model decision module verifies the performance of the optimal model in the virtual simulation environment and migrates it to the actual system for testing, completing real-time communication decisions between multiple UAVs and low-orbit satellites.

[0017] The above embodiment has the following advantages or beneficial effects:

[0018] (1) Dynamic resource allocation to achieve efficient communication and task execution;

[0019] (2) A security verification mechanism is added to the decision-making training process to ensure the reliability of the strategy and the security and robustness of the system.

[0020] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0022] Figure 1 A flowchart of an efficient communication method for multi-UAV collaboration based on low-orbit satellites provided by the present invention;

[0023] Figure 2 This is a module flow chart of an efficient communication system for multi-UAV collaboration based on low-orbit satellites provided by the present invention. DETAILED DESCRIPTION

[0024] To make the objectives, technical solutions, and advantages of this application more clearly understood, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] Low-Earth Orbit Satellite Network Description: The Low-Earth Orbit Satellite Network consists of a large number of small communications satellites, orbiting at an altitude of approximately 500-1200 kilometers. Compared to traditional communications satellites, it offers advantages such as wide coverage, low transmission latency (20-40 milliseconds), and high communication speeds (over 100 Mbps). It can provide high-speed and stable communication services for drones worldwide. Drones based on the Low-Earth Orbit Satellite Network can communicate with the Low-Earth Orbit Satellite Network.

[0026] DRL (Deep Reinforcement Learning) deep reinforcement learning.

[0027] GAT (Graph Attention Network) graph neural network.

[0028] SNR (Signal-to-Noise Ratio) refers to the ratio of signal to noise in the system.

[0029] Throughput refers to the amount of data (measured in bits, bytes, packets, etc.) successfully transmitted per unit time for a network, device, port, virtual circuit or other facility.

[0030] Latency refers to the time it takes for a message or packet to be transmitted from one end of the network to the other.

[0031] The packet loss rate refers to the percentage of lost packets to transmitted packets during network transmission.

[0032] Doppler shift refers to the phase and frequency changes caused by the propagation path difference when a mobile station moves in a certain direction at a constant rate.

[0033] Atmospheric attenuation refers to the energy attenuation phenomenon that occurs when electromagnetic waves propagate in the atmosphere.

[0034] The atmospheric attenuation coefficient refers to the energy loss per unit received power when it travels per unit distance in the atmosphere.

[0035] A frequency band refers to a continuous frequency range in the radio spectrum, which is usually used for frequency division for different communication systems or applications.

[0036] Bandwidth refers to the number of bits that can be transmitted per second on a communication link.

[0037] Transmitting power refers to the operating power of the transmitting antenna of a wireless product, which determines the strength and distance of the wireless signal. The greater the power, the stronger the signal.

[0038] A satellite access node refers to the exchange point in satellite communications where users enter and exit the satellite communication network.

[0039] Downwash, also known as rotor downwash, occurs when a drone is in a hovering state. Rotating rotors cause airflow to flow from above to below the rotors, generating induced velocity. The combined velocity of the induced velocity and the blade element's circular motion creates the blade element's true velocity, causing air to flow in the direction opposite to the pull. Under these conditions, the pressure on the rotor's upper surface is low, while the pressure on its lower surface is high, generating lift between the upper and lower surfaces. Since the drone maintains a constant altitude during flight, the air with the high pressure on the lower surface flows downward, forming a downwash.

[0040] Example 1

[0041] like Figure 1 As shown in FIG, an efficient communication method for multi-UAV collaboration based on low-orbit satellites includes the following steps:

[0042] Step S01: Obtain relevant parameters in the communication between multiple UAVs and low-orbit satellites, and use them to define the framework structure of the state space, action space, and reward space to obtain framework structure 1;

[0043] The state space uses environmental parameters, drone status, network load, and historical information as primary reference parameters; the action space uses resource allocation actions, collaborative control actions, and mode switching as primary reference parameters; and the reward function is composed of communication performance rewards, negative rewards, and multi-objective trade-offs. Environmental parameters include interference source distribution, atmospheric attenuation coefficient, spectrum occupancy heat map, Doppler shift, and bandwidth; drone status includes position, speed, heading, battery charge, and queue length for data to be transmitted; network load includes spectrum occupancy and the communication coordination status of adjacent drones; and historical information includes the communication quality and switching records of dynamically changing communication links.

[0044] Among them, the drone’s own state parameters are represented by three-dimensional coordinates: (x, y, z)∈R 3 , velocity vector: (v x ,v y ,v z ), heading angle: θ∈[0,360°);

[0045] Interference source distribution represents {(x i ,y i ,P i )}, interference source location (x, y) and transmission power P i ;

[0046] The atmospheric attenuation coefficient is expressed as ɑ rain (dB / km);

[0047] The historical link quality represents the SNR and packet loss rate of the sliding window sampling N moments;

[0048] The switching record is represented by the time consumption and success rate of M satellite switching.

[0049] Finally, we can get the S-dimensional state space, which can be expressed as follows:

[0050]

[0051] The collected data are normalized and then feature fused. Different feature extraction methods are adopted according to the data characteristics of data of different dimensions. The feature vectors are obtained and then spliced. Among them, for sequence data, the LSTM method is used to extract sequence feature data. For spatial data such as radar, neural networks such as CNN are used to extract high-dimensional feature data. For topological structure data, GAT is used to learn the features of drone-satellite topological structure.

[0052] For the modeling of the action space, several aspects of the action space modeling reference include resource allocation actions, collaborative control actions, and mode switching as key actions for adjusting the state of the drone. Among them, the resource allocation actions specifically include satellite access nodes, allocating spectrum blocks (frequency bands and bandwidths), and adjusting transmit power;

[0053] The collaborative control actions include triggering multi-satellite diversity transmission, coordinating the reuse of UAV resources within the cluster, and allocating time slots.

[0054] Mode switching includes direct mode, multi-satellite collaborative mode, and disaster recovery fallback link.

[0055] The specific parameter structure is: Satellite access node in resource allocation action:

[0056] Discrete N satellite IDs;

[0057] Spectrum allocation: Frequency bands and bandwidths include S-band (2 GHz, 20 MHz bandwidth), Ku-band (14 GHz, 50 MHz bandwidth), and Ka-band (30 GHz, 100 MHz bandwidth);

[0058] Power control: balances communication quality and energy consumption. For example, on rainy days, power is increased to compensate for attenuation. ΔP∈[-30%,30%] is based on the current power baseline.

[0059] Modulation and coding scheme: QPSK 1 / 2, 16QAM 3 / 4, 64QAM 5 / 6, etc.

[0060] Data priority scheduling: control instructions (weight 1.0) > sensor data (0.7) > logs (0.3);

[0061] Cluster head election in collaborative control layer actions: binary action: {0 (no election), 1 (election)};

[0062] Relay collaboration: {direct connection to satellite, drone A relays, drone B relays};

[0063] Spectrum reuse: TDMA time slot allocation, time slot numbering (1…N), to avoid co-frequency interference between multiple drones.

[0064] The actions in the action space are parameterized and encoded and stored together in the form of a suitable data structure.

[0065] For example, the following storage method:

[0066]

[0067] The reward and penalty function is closely linked to the system goals, achieving the optimal balance between communication quality, energy consumption and other dimensions. Communication performance rewards include throughput rewards, delay penalties, packet loss rate penalties, etc. Among them, throughput rewards:

[0068] Delay Penalty: ,applying exponential penalties for delays exceeding the threshold;

[0069] Packet loss rate penalty: R loss =-ɑ3·Packet loss rate 2 , impose a secondary penalty on the packet loss rate to strictly control the loss of key data.

[0070] Energy consumption rewards include energy consumption penalties, battery life balance rewards, etc.

[0071] Energy consumption penalty: R energy =-η1·Energy consumption of this step, directly penalizing energy loss.

[0072] Endurance Balance Rewards: The higher the remaining power, the higher the reward.

[0073] In some special scenarios, there are system stability penalties and collaborative scenario penalties. Specifically, there are frequent satellite switching: P handover =-μ1·number of switching times 2 ;

[0074] Mode Oscillation: P oscillation =-μ2·∏(mode switching frequency>f max );

[0075] Communication interruption: P outage = -μ3·interruption duration;

[0076] Spectrum Conflict: P interference =-υ1·Number of conflicting frequency bands;

[0077] Relay Overload: P relay =-υ2·Relay queue length.

[0078] The total reward function integration can be expressed as follows:

[0079]

[0080] Normalize each reward item to prevent a certain reward from dominating the entire learning process.

[0081]

[0082] Step S02: Use a simulation tool to construct a simulation environment that integrates a satellite orbit model, a drone dynamics model, and a channel model to obtain a virtual simulation environment one; wherein, a deep reinforcement learning model composed of the framework structure one is trained, and hyperparameters are adjusted to obtain an optimal model one.

[0083] Building a virtual simulation environment primarily involves modeling three aspects: satellite orbit models, drone dynamics models, and channel models. The simulation tools involved include Gazebo and ROS2 for dynamics simulation, STK, a satellite orbit simulation tool, for high-precision visualization, and MATLAB 5G Toolbox, which supports millimeter-wave channel modeling.

[0084] The multi-UAV dynamics model uses a relative motion model, cooperative dynamics equations, and an external interference model. For formation control and collision avoidance, the relative motion states of the multi-UAVs are defined as follows: Coupling model with communication delay added to cooperative dynamics model: γ ij : cooperative gain coefficient; τ: communication delay

[0085] The multi-UAV dynamics model also adopts multi-machine coupling effect modeling, including communication and control coupling, and downwash interference modeling.

[0086] Aerodynamic interactions between rotors:

[0087]

[0088] r ij : horizontal distance between drones; z ij : vertical height difference; h: attenuation height constant

[0089] The satellite orbit model uses multi-satellite constellation modeling and orbit visualization modules.

[0090] The modular model is implemented and integrated, and the reinforcement learning environment is encapsulated to support distributed training, multiple DRL algorithms, and customized interfaces.

[0091] Collect a large amount of training data, continuously learn and update strategies, and synchronize multiple modules until key indicators reach convergence, including reward function loss, strategy loss, value function loss, etc.

[0092] Incorporate safety judgment into training to filter out potentially dangerous actions.

[0093] Step 03: Verify the performance of the optimal model 1 in the virtual simulation environment 1, and migrate it to the actual system for testing to complete real-time communication decision-making between multiple drones and low-orbit satellites.

[0094] Verify and evaluate various model metrics, including measurement stability, generalization, and real-time performance. In a multi-UAV collaborative scenario, verify spectrum reuse efficiency, with a reduction of co-channel interference exceeding 40%. Adopt a lightweight deployment strategy, compress the model, and deploy it to the UAV edge device management unit.

[0095] An online learning mechanism can be used to continuously learn data streams and further optimize the model.

[0096] Example 2

[0097] like Figure 2 As shown in the figure, the efficient communication system for multi-UAV collaboration based on low-orbit satellites includes the following modules:

[0098] The communication input module obtains the relevant parameters of the communication between multiple UAVs and low-orbit satellites, and uses them to define the framework structure of the state space, action space and reward space, obtaining framework structure 1;

[0099] The communication model training module uses simulation tools to build a simulation environment that integrates a satellite orbit model, a multi-UAV dynamics model, and a channel model to obtain a virtual simulation environment. The module then trains a deep reinforcement learning model based on the framework structure and adjusts hyperparameters to obtain the optimal model. Safety verification is incorporated into the training process to filter out dangerous actions.

[0100] The communication model decision module verifies the performance of the optimal model in the virtual simulation environment and migrates it to the actual system for testing, completing real-time communication decisions between multiple UAVs and low-orbit satellites.

[0101] Through the above scheme, the following advantages or beneficial effects are achieved:

[0102] (1) Dynamic resource allocation to achieve efficient communication and task execution;

[0103] (2) A security verification mechanism is added to the decision-making training process to ensure the reliability of the strategy and the security and robustness of the system.

[0104] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. An efficient communication method for multi-UAV collaboration based on low-orbit satellites, characterized in that: The following steps are involved: Obtain relevant parameters in the communication between multiple UAVs and low-orbit satellites, and use them to define the framework structure of the state space, action space, and reward space to obtain framework structure 1; A simulation environment is constructed using a simulation tool that integrates a satellite orbit model, a multi-UAV dynamics model, and a channel model to obtain a first virtual simulation environment. A deep reinforcement learning model constructed using the first framework is trained and hyperparameters are adjusted to obtain the first optimal model. Safety verification is incorporated into the training process to filter out dangerous actions. Verify the performance of the optimal model in the virtual simulation environment and migrate it to the actual system for testing to complete real-time communication decisions between multiple UAVs and low-orbit satellites.

2. The high-efficiency communication method for multi-UAV collaboration based on low-orbit satellites according to claim 1, characterized in that: The state space uses environmental parameters, drone status, network load, and historical information as main reference parameters; the action space uses resource allocation actions, collaborative control actions, and mode switching as main reference parameters; and the reward function is composed of communication performance rewards, negative rewards, and multi-objective trade-offs.

3. The high-efficiency communication method for multi-UAV collaboration based on low-orbit satellites according to claim 1, characterized in that: The multi-objective trade-off includes a communication performance reward based on throughput and delay, a negative reward based on energy consumption and spectrum conflict, and a multi-objective trade-off based on a weighted balance of multiple reference objectives.

4. The high-efficiency communication method for multi-UAV collaboration based on low-orbit satellites according to claim 1, characterized in that: The multi-UAV dynamics model adopts a relative motion model, a cooperative dynamics equation and an external interference model.

5. The high-efficiency communication method for multi-UAV collaboration based on low-orbit satellites according to claim 4, characterized in that: The multi-UAV dynamics model also adopts multi-machine coupling effect modeling, including communication and control coupling, and downwash interference modeling.

6. The high-efficiency communication method for multi-UAV collaboration based on low-orbit satellites according to claim 1, characterized in that: The satellite orbit model adopts a multi-satellite constellation modeling and orbit visualization module.

7. An efficient communication system for multi-UAV collaboration based on low-orbit satellites, characterized by: Includes the following modules: The communication input module obtains the relevant parameters of the communication between multiple UAVs and low-orbit satellites, and uses them to define the framework structure of the state space, action space and reward space, obtaining framework structure 1; The communication model training module uses simulation tools to build a simulation environment that integrates a satellite orbit model, a multi-UAV dynamics model, and a channel model to obtain a virtual simulation environment. The module then trains a deep reinforcement learning model based on the framework structure and adjusts hyperparameters to obtain the optimal model. Safety verification is incorporated into the training process to filter out dangerous actions. The communication model decision module verifies the performance of the optimal model in the virtual simulation environment and migrates it to the actual system for testing, completing real-time communication decisions between multiple UAVs and low-orbit satellites.