A Multi-UAV Flight Search and Rescue Trajectory Optimization Method for Disaster Area Rescue

By building a UAV communication model and mobile model, combining FSO and RF communication, using distributed Markov quadruple and multimodal deep reinforcement learning to optimize the UAV trajectory, the communication difficulties after the damage to communication facilities in the disaster area are solved, and efficient multi-UAV collaborative communication coverage is achieved.

CN115866574BActive Publication Date: 2025-07-11GUIZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211454058.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-07-11
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

After the disaster occurs, communication facilities are destroyed, resulting in difficulty in emergency communication. The existing drone reinforcement learning methods cannot effectively solve the problem of explosion in state space and action space dimensions of collaborative work of multiple drones, and it is difficult to achieve rapid and long-distance unregulated mobile user communication in disaster areas.

Method used

Build a UAV user communication model, user mobile model and UAV base station communication model, combine FSO and RF communication, Markov quadruple and multimodal deep reinforcement learning can be observed through the distributed part, optimize the drone's flight trajectory, and dynamically optimize the drone's position using the enhanced k-means algorithm and policy gradient algorithm.

Benefits of technology

It realizes fast and long-distance communication coverage in the disaster area, maximizes the total communication throughput and minimizes the coverage, and improves the communication efficiency in the disaster area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115866574B_ABST
    Figure CN115866574B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-UAV flight search and rescue trajectory optimization method for disaster area rescue, including: using UAVs equipped with radio frequency (RF) communication modules and free space optical (FSO) communication modules as relay nodes to connect to remote base stations. Considering RF / FSO communication spectrum resource management and minimizing the number of UAVs, a multi-UAV collaborative trajectory optimization method for disaster area emergency communication is proposed through an enhanced K-means algorithm and an extended multi-modal deep reinforcement learning algorithm. At the same time, based on multi-modal deep reinforcement learning, the UAVs are trained to work collaboratively according to the positions and movement trajectories of mobile users. The present invention improves the utilization rate of RF / FSO communication spectrum resources through the multi-modal deep reinforcement learning algorithm, and at the same time maximizes the total link throughput. Compared with other methods, the present invention also improves the RF allocation efficiency and the bit rate of FSO communication backhaul.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of new generation information technology, and particularly relates to a method for optimizing the flight search and rescue trajectories of multiple unmanned aerial vehicles (UAVs) for disaster area rescue. Background Art

[0002] When disasters such as earthquakes, tsunamis, and mountain floods occur, some communication facilities may be damaged, and emergency communication becomes the top priority. The research on emergency communication after the destruction of communication facilities has attracted the common attention of the academic and engineering circles. Considering the destructiveness of ground roads and the urgency of emergency communication guarantee during disasters, how to quickly and remotely communicate with irregular mobile users in the disaster area has become a difficult point. Free Space Optics (FSO) communication can provide large communication capacity and high-speed long-distance transmission functions by using laser light waves as carriers and the atmosphere as the transmission medium without the need to lay optical fibers. UAVs, on the other hand, have high flexibility. How to use UAVs equipped with FSO communication modules (connected to remote base stations) and radio frequency communication modules (accessed by disaster area users) as tools for disaster area emergency communication has great significance and has been widely applied.

[0003] Introducing UAV assistance in emergency communication has great advantages. A method for deploying the optimal positions of UAVs has been proposed to optimize the uneven access volume and comprehensive coverage in the disaster area. Users in the disaster area often move irregularly for self-rescue, while search and rescue personnel carry out organized rescue, which shows high spatial / temporal dynamic characteristics. Establishing models and access matching mechanisms in a dynamic multi-user and multi-UAV collaborative environment is a difficult point.

[0004] Reinforcement learning, as one of the branches of artificial intelligence, is usually used for decision-making in complex environments and has also made many contributions in the field of UAVs. Reinforcement learning has many advantages. It can make multi-step decisions to obtain the optimal reward in a dynamic environment. However, single-UAV reinforcement learning is mainly used for dynamic environments with low-dimensional action spaces and state spaces and cannot solve the problem of the dimensional explosion of the state space and action space in the collaborative work of multiple UAVs for emergency communication. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for optimizing the flight search and rescue trajectories of multiple UAVs for disaster area rescue to solve the problems existing in the above-mentioned prior art.

[0006] To achieve the above purpose, the present invention provides a method for optimizing the flight search and rescue trajectories of multiple UAVs for disaster area rescue, including:

[0007] Constructing a UAV-user communication model, a user movement model, and a UAV-base station communication model;

[0008] Obtaining the average path loss between the UAV and the user through the UAV-user communication model;

[0009] Obtain the moving speed and moving direction of the user through the user movement model;

[0010] Obtain the backhaul rate between the UAV and the base station under different visibility conditions through the UAV base station communication model;

[0011] Construct a distributed partially observable Markov quadruple based on the average path loss, the moving speed and moving direction of the user, and the backhaul rate between the UAV and the base station;

[0012] Based on the distributed partially observable Markov quadruple, introduce and enhance the k-means algorithm, and through an extended multi-modal deep reinforcement learning method, achieve dynamic optimization of the UAV flight trajectory.

[0013] Optionally, the process of obtaining the average path loss includes: respectively obtaining the path loss in line-of-sight wireless transmission and non-line-of-sight wireless transmission, obtaining the probability of establishing line-of-sight wireless transmission based on environmental parameters, obtaining the probability of establishing non-line-of-sight wireless transmission based on the elevation angle between the UAV and the user, multiplying the path loss by the corresponding transmission probability and then summing to obtain the average path loss.

[0014] Optionally, the user movement model includes: simulating the user's movement speed using the Maxwell-Boltzmann distribution, and simulating the user's movement direction using the Gaussian distribution.

[0015] Optionally, the UAV base station communication model includes: judging the environmental quality function value based on visibility, obtaining the atmospheric attenuation factor based on the wavelength of the transmitted signal, the environmental quality function value, and visibility, calculating the backhaul rate of the free space optical connection between the UAV and the base station under different visibility conditions based on the atmospheric attenuation factor and the straight-line distance between the UAV and the base station. The calculation formula for the backhaul rate is as follows:

[0016]

[0017]

[0018] Where, P t is the FSO transmit power of the UAV, τ t is the optical efficiency of the transmitter and receiver, τ atm is the atmospheric transmission value at the wavelength of the laser transmitter, α is the atmospheric attenuation factor, φ FSO is the FSO beam diameter of the receiver of UAV i. Where θ t is the transmitter divergence, E p = hc / λ is the photon energy, h is Planck's constant, c is the speed of light, λ is the transmission wavelength, N bis the sensitivity for the average number of received photons b. Finally, L represents the straight-line distance between the UAV and the base station.

[0019] Optionally, calculate the signal-to-noise-plus-interference ratio of the communication between the UAV and the user based on the transmission power of the UAV, the average path loss, and the additive white Gaussian noise power.

[0020] Optionally, the distributed partially observable Markov quadruple includes observation data, a state space, an action space, and a reward function; the observation data includes the position and speed of the UAV itself, the position and speed of the ground user, and the position information of the ground base station; the state space is based on the position of the ground base station, the dynamic speed of the UAV, and the dynamic speed of the user; the action space is commission and maintaining a fixed altitude, flying different numbers of unit lengths forward and backward and left and right, constituting a continuous action space; the reward function is based on maximizing the number of connected users and the global traffic throughput.

[0021] Optionally, the reward function r all (t) is expressed as follows:

[0022]

[0023] where Θ mu represents the number of disconnected users, which is calculated by the formula N uav is the number of UAVs, N mu is the total number of users, N mu,j is the N users connected to UAV j; χ, ψ represent discount factors, χ is fixed, and ψ changes with the number of users, Γ ij is the signal-to-noise-plus-interference ratio, and R is the backhaul rate.

[0024] Optionally, the dynamic optimization process of the UAV flight trajectory includes: initializing the UAV position based on the enhanced k-means algorithm; using the neutralized training value network to transform the distributed partially observable Markov quadruple into a Markov process, and transforming the reward function into a global cumulative reward; based on the task objective, constructing a policy using the policy gradient algorithm, the policy is fitted by a neural network to obtain a policy network, taking the maximization of the transformed reward function as the objective function, introducing the importance sampling method, and clipping and limiting the objective function through the KL divergence, obtaining new data through the real-time interaction between the initialized UAV and the environment, and updating the policy network based on the continuously updated objective function to achieve the dynamic optimization of the UAV flight trajectory.

[0025] The technical effects of the present invention are:

[0026] The present invention uses a drone that combines FSO communication and RF communication as a relay base station to provide communication services for ground users. First, an RF communication model between the drone and the users and an FSO communication model between the drone and the ground base station are established. At the same time, a Gaussian random distribution is introduced to model the movement model of users in the disaster area. Subsequently, based on the partially observable Markov process, the action space, state space, and reward function of the drone are established. At the same time, an enhanced k-means algorithm is proposed to initialize the position of the drone. Finally, based on multi-modal deep reinforcement learning, the drone trajectory is dynamically optimized, achieving the effect of minimizing the disaster area covered by the drone and maximizing the total communication throughput. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0028] Figure 1 is the network architecture diagram of the disaster area emergency communication in the embodiment of the present invention;

[0029] Figure 2 is the flow chart of the drone trajectory optimization algorithm based on multi-modal deep reinforcement learning in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and combine the embodiments to detail this application.

[0031] Embodiment 1

[0032] As Figure 1-2 shown, this embodiment provides a multi-drone flight search and rescue trajectory optimization method for disaster area rescue, including:

[0033] S1: Establish an RF communication model between the drone and the ground users, introduce the Maxwell-Boltzmann distribution, model the movement direction and movement speed of the ground mobile users, and finally construct an FSO communication model between the drone and the distant ground base station. The network model is as Figure 1 shown.

[0034] S2: Model the trajectory optimization process of the drone as a distributed partially observable Markov process, and specifically establish the observation space and action space of the drone; the state space of the environment and the overall environmental reward function.

[0035] S3: Based on multi-modal deep reinforcement learning, introduce an enhanced k-means algorithm to initialize the position of the drone to accelerate the convergence speed of the algorithm, and dynamically optimize the flight trajectory of the drone, thereby maximizing the number of ground personnel served and the total throughput.

[0036] Furthermore, the step S1 includes

[0037] S101: Establish a communication model for line-of-sight wireless transmission (LoS) between the unmanned aerial vehicle and the mobile user. Let and be the path losses between the unmanned aerial vehicle i and the mobile user j in LoS and non-line-of-sight wireless transmission (NLoS), respectively:

[0038]

[0039]

[0040] Here, A LoS and A NLoS represent the path losses at the reference distance d ij = 1, d ij is the distance between the unmanned aerial vehicle i and the mobile user j, and γ LoS and γ NLoS represent the path loss exponents of LoS and NLoS transmissions, respectively.

[0041] S1011: The LoS probability between the unmanned aerial vehicle i and the mobile user j can be calculated as:

[0042]

[0043] where α and β are two environmental parameters that depend on the urban environment deployment model. The NloS probability formula for the communication link is:

[0044]

[0045]

[0046] where is the elevation angle between the unmanned aerial vehicle i and the mobile user j, and h i is the vertical height of the unmanned aerial vehicle i. The LoS and NloS probabilities here determine the average path loss between the unmanned aerial vehicle and the user, and the average path loss can be calculated as:

[0047]

[0048] S1012: The signal-to-noise plus interference ratio (SINR) of the wireless communication between the unmanned aerial vehicle i and the mobile user j is:

[0049]

[0050] where σ is the power of additive white Gaussian noise (AWGN), N uav is the number of unmanned aerial vehicles, is the transmission power of the drone, is the total interference power. The drone communicates with the mobile user to ensure a high signal-to-noise ratio. When the SINR of the mobile user with all drones is lower than the threshold δ, the mobile user is considered a disconnected user.

[0051] S1013: It can be deduced from the above formula that the bandwidth that each user can obtain can be expressed as:

[0052]

[0053] where B i is the channel bandwidth of drone i, and N mu,i are the N users connected to drone i.

[0054] S102: The mobile users in the disaster area include rescue workers and disaster victims, and their positions are constantly moving in the disaster area. In this embodiment, the Maxwell-Boltzmann distribution is considered to simulate the speed of the mobile users' movement. The walking speed of people is distributed between 0.9 m / s and 2.2 m / s. The probability density function of the walking speed is as follows:

[0055]

[0056] where v is the average speed of people, and v r.m.s. is the root mean square of the speed v≡|V|. Then, a Gaussian distribution simulation method is used to simulate the movement direction of the users. The probability density function of the walking direction is:

[0057]

[0058] S103: The backhaul rate R of the FSO connection between the drone and the ground base station can be calculated as:

[0059]

[0060]

[0061] where P t is the FSO transmission power of the drone, τ t is the optical efficiency of the transmitter and receiver, τ atm is the atmospheric transmission value at the wavelength of the laser transmitter, α is the atmospheric attenuation factor, and φ FSO is the FSO beam diameter of the receiver of drone i. Where θ t is the transmitter divergence, E p =h c / λ is the photon energy, h is the Planck constant, c is the speed of light, λ is the transmission wavelength, and N b is the sensitivity of the average received photon number b. Finally, L represents the distance between drone i and ground base station k.

[0062]

[0063]

[0064] where d ik represents the distance between the UAV i and the ground base station k, and h k is the height of the ground base station. The unit of α is dB / km; represents the visibility, with the unit of kilometer; Q is the environmental quality function; λ is the wavelength of the transmitted signal. The relationship between different visibilities and Q is as follows:

[0065]

[0066] According to the above formula, the return rate R of the UAV to the ground base station under different visibilities can be calculated.

[0067] Furthermore, the step S2 includes

[0068] S201: Considering the dynamic characteristics of the environment, the distribution of UAVs, and the locality of UAV observations, the UAV trajectory optimization problem is formulated as a distributed partially observable Markov problem. The observation, state, action, and reward of this distributed partially observable Markov game at time t are defined as {O, S, A, R}. The detailed definitions of the distributed partially observable Markov elements are as follows.

[0069] S202: Observation. At time slot t, the agent collects the environmental information within the observation range, which includes the position and speed of the UAV itself, the position and speed of the ground users, and the position information of the ground base station in time slot t. Therefore, the observation o(t) is defined as:

[0070] o(t) = (d i (t), u j (t)) (16)

[0071] Here represents the position of UAV i and the position of ground user j at time slot t.

[0072] S203: The state space is composed of the states of all UAVs, ground users, and ground base stations in the entire environment at time t. The state space s(t) is defined as:

[0073]

[0074] where since the position of the ground base station is invariant, m k = (x k , y k ) defines the position of the ground base station k. In addition and It represents the dynamic speeds of the UAVs and the users.

[0075] S204: An action is the value that guides the agent's actions generated according to the strategy in the Markov game. The action space of UAV i is as follows:

[0076]

[0077] Here, each UAV maintains a fixed altitude in time slot t, and at the same time flies different numbers of unit lengths left - right and front - back, thus forming a continuous action space.

[0078] S205: A reward is a function described by actions and states, which is the reward generated after the agent executes an action in the current state. In the disaster area emergency communication problem, each UAV has the same goal, that is, to maximize the number of connected users and the global traffic throughput. Therefore, the rewards of all UAVs are defined as follows in time slot t:

[0079]

[0080] Here, Θ mu represents the number of disconnected users, which is calculated by the formula that is, the total number of users minus the number of users communicating with all connected UAVs; χ, ψ represent discount factors, χ is fixed, and ψ changes with the number of users.

[0081] Furthermore, the step S3 includes,

[0082] S301: Enhanced k - means algorithm. Considering the urgency of emergency communication and the locality of UAV observations, this embodiment proposes an enhanced k - means algorithm to initialize the UAV trajectories. Considering a set of observations {m1,..., m n}, where each observation m n is a vector in, d = 2 is the coordinate dimension of the mobile users. It is desired to assign n mobile users to k UAVs, so the n observations are divided into S = {S1,..., S k} clusters. Specifically, the following optimization problem needs to be solved:

[0083]

[0084] where c1 defines cluster j, and each centroid in a set of centroids {c1,..., c k} is associated with a cluster. At the same time, the UAVs in the disaster area cannot collect the location information of all mobile users. Therefore, when the distances between the mobile users and all UAVs When it is considered that no UAV can observe this mobile user, this mobile user will not be assigned to any cluster. In the algorithm update, the centroid c i at (T + 1) is updated by the following formula:

[0085]

[0086] When no new centroid is generated, it indicates that the algorithm converges. At this time, the final centroid set {c1,..., c k} is selected as the initial position of the UAV.

[0087] S302: Multimodal reinforcement learning algorithm. In the centralized training stage, the neutralized training value network is used to make the partially observable Markov process become a Markov process. The Markov process at time slot t is defined as <S, A, O, R, P, n, γ>, where S = {O1, O2,..., O i} is the global state space, A i is the action space of UAV i, O i is the local observation state s of UAV i, R is the global cumulative reward, and P(s′|s, A) is the transition probability from S to S′ after all n UAVs execute the action A = (α1, α2,..., α n ) and the ground users change their positions.

[0088] In this embodiment of the multimodal reinforcement learning algorithm, the UAV is modeled using the actor-critic framework. The actor-critic framework is a policy gradient algorithm that has two key components, an actor and a critic. The UAV is regarded as the Actor. The UAV constructs a policy π according to the task objective. The policy π is fitted by a neural network with parameters θ. The Actor selects actions through π. The other component, the critic, is used to evaluate the value of the actions selected by the policy π. Starting from the initial state of the disaster area, the UAV continuously interacts with the environment to obtain a complete sequence τ and stores it in the experience replay area Then the sum of rewards for all stages is as follows:

[0089]

[0090] where is the probability of a sequence τ i occurring. In the Markov game, the current reward is more important. Then where γ represents the discount factor. To find the optimal policy to maximize the cumulative reward obtained by the agent, it is necessary to update the policy network parameters θ by maximizing the reward objective function:

[0091]

[0092] A here θ = Q(s t , a t ) - V(s t ) is the advantage function used to replace R(τ). The Proximal Policy Optimization algorithm is an online learning algorithm. In the Proximal Policy Optimization algorithm, the method of importance sampling is introduced to increase sample efficiency. Formulas (22) and (23) are rewritten as:

[0093]

[0094]

[0095] In the importance sampling method, if the difference in the probability distributions of π θ and π θ′ is too large, it will lead to an increase in the number of required samples, thus reducing sample efficiency. Therefore, the KL(π θ ||π θ′ ) = H(π θ , π θ′ ) - H(π θ ) divergence is introduced in the Proximal Policy Optimization algorithm to clip or limit the objective function J θ :

[0096]

[0097] if KL(π θ ||π θ ′) > KL max , increase β (28)

[0098] if KL(π θ ||π θ ′) < KL min , decrease β (29)

[0099] As described above, the above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-UAV flight search and rescue trajectory optimization method for disaster area rescue, characterized in that It includes the following steps: Construct a UAV-user communication model, a user movement model, and a UAV-base station communication model; Obtain the average path loss between the UAV and the user through the UAV-user communication model; Obtain the movement speed and direction of the user through the user movement model; Obtain the backhaul rate between the UAV and the base station under different visibility conditions through the UAV-base station communication model; Construct a distributed partially observable Markov quadruple based on the average path loss, the movement speed and direction of the user, and the backhaul rate between the UAV and the base station; Based on the distributed partially observable Markov quadruple, introduce and enhance the k-means algorithm, and realize the dynamic optimization of the UAV flight trajectory through an extended multi-modal deep reinforcement learning method; The user movement model includes: simulating the user's movement speed using the Maxwell-Boltzmann distribution and simulating the user's movement direction using the Gaussian distribution; Calculate the signal-to-noise-plus-interference ratio of the communication between the UAV and the user based on the transmission power of the UAV, the average path loss, and the additive white Gaussian noise power; The distributed partially observable Markov quadruple includes observation data, a state space, an action space, and a reward function; the observation data includes the position and speed of the UAV itself, the position and speed of the ground user, and the position information of the ground base station; the state space is based on the position of the ground base station, the dynamic speed of the UAV, and the dynamic speed of the user; the action space is to commission and maintain a fixed altitude and fly different numbers of unit lengths forward and backward and left and right, constituting a continuous action space; the reward function is based on maximizing the number of connected users and the global traffic throughput; Reward function The expression is as follows: Among them, represents the number of disconnected users, which is calculated by the formula and is the number of drones, is the total number of users, is the number of users connected to the drone ; , represents the discount factor, is fixed, while varies with the number of users, is the signal-to-noise plus interference ratio, is the backhaul rate; The dynamic optimization process of the UAV flight trajectory includes: initializing the UAV position based on the enhanced k-means algorithm; using a neutralized training value network to convert the distributed partially observable Markov quadruple into a Markov process, and converting the reward function into a global cumulative reward; based on the task objective, constructing a policy using the policy gradient algorithm, where the policy is fitted by a neural network to obtain a policy network, taking the maximization of the converted reward function as the objective function, introducing the importance sampling method, and clipping and restricting the objective function through the KL divergence, obtaining new data through the real-time interaction between the initialized UAV and the environment, and updating the policy network based on the continuously updated objective function to realize the dynamic optimization of the UAV flight trajectory; The modeling of the UAV utilizes the actor-critic framework, which has two key components, actor and critic; the UAV is regarded as the Actor, and the UAV constructs a policy π according to the task objective, where the policy π is fitted by a neural network and its parameters are θ, and the Actor selects actions through π; the other component, critic, is used to evaluate the value of the actions selected by the policy π; Update the policy network parameters by maximizing the reward objective function to find the optimal policy that enables the agent to obtain the optimal policy with the maximum cumulative reward.

2. The multi-UAV flight search and rescue trajectory optimization method for disaster area rescue according to claim 1, wherein the process of obtaining the average path loss includes: respectively obtaining the path loss in line-of-sight wireless transmission and non-line-of-sight wireless transmission, obtaining the probability of establishing line-of-sight wireless transmission based on environmental parameters, obtaining the probability of establishing non-line-of-sight wireless transmission based on the elevation angle between the UAV and the user, multiplying the path loss by the corresponding transmission probability and then summing them up to obtain the average path loss.

3. The multi-UAV flight search and rescue trajectory optimization method for disaster area rescue according to claim 1, wherein the UAV base station communication model includes: judging the environmental quality function value based on visibility, obtaining the atmospheric attenuation factor based on the wavelength of the transmitted signal, the environmental quality function value, and visibility, calculating the backhaul rate of the free space optical connection between the UAV and the base station under different visibilities based on the atmospheric attenuation factor and the straight-line distance between the UAV and the base station. The calculation formula of the backhaul rate is as follows: wherein, is the FSO transmission power of the UAV, is the optical efficiency of the transmitter and receiver, is the atmospheric transmission value at the wavelength of the laser transmitter, is the atmospheric attenuation factor, is the UAV FSO beam diameter of the receiver, where is the divergence of the transmitter, is the photon energy, is the Planck constant, is the speed of light, is the transmission wavelength, is the average number of received photons sensitivity of, finally, represents the straight-line distance between the UAV and the base station.

Citation Information

Patent Citations

  • Unmanned aerial vehicle trajectory and power joint optimization method based on deep reinforcement learning

    CN111263332A

  • Multi-unmanned aerial vehicle track and intelligent reflecting surface shift joint optimization method and system

    CN113364495A