Urban low-altitude dense traffic flow airspace-oriented unmanned aerial vehicle double-layer conflict resolution method
By building a hybrid integer nonlinear planning model in the UAV system and using a dual deep Q network and attention mechanism, the secondary conflict problem of UAV in ultra-low altitude airspace is solved, and the airspace stability and safety are improved.
Patent Information
- Application Number
- CN202510148987.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing technology is difficult to effectively resolve secondary conflicts of drones in ultra-low altitude airspace, and the existing conflict digestion methods lack real-time quantitative modeling and integrated conflict digestion frameworks before and after takeoff, resulting in increased airspace instability and conflict risk.
A hybrid integer nonlinear planning model based on conflict digestion strategy is adopted, and the initial parameters are optimized in combination with an improved random fractal search algorithm, and after takeoff, dual deep Q network and attention mechanism are used for training, and conflict-free maneuvering actions are output in real time.
Real-time quantitative secondary conflict measurement and effective digestion of drones in ultra-low altitude airspace is achieved, which improves airspace stability and overall safety, and reduces conflict occurrence and training time.
Smart Images

Figure CN119992883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and urban air traffic systems, and in particular to a method for resolving unmanned aerial vehicle (UAV) double-layer conflicts in urban low-altitude dense traffic flow airspace. Background Art
[0002] Unmanned aerial vehicles, also known as drones, have expanded their applications from the military to entertainment, agriculture, logistics, and even construction. These applications have brought many benefits to society, but they also introduce a large amount of air traffic to ultra-low-altitude airspace in urban areas. Without proper traffic management, heavy air traffic in ultra-low-altitude airspace may pose a great risk to property and pedestrians on the ground. Major aviation regulators have proposed the development of traffic management frameworks to ensure that the operation of drones in ultra-low-altitude airspace can be safe and efficient. These developing frameworks have some common key modules, and autonomous conflict resolution is one of them. At present, many studies have been devoted to the development of functional modules for autonomous conflict resolution, but they still face the following three major limitations.
[0003] (1) Most existing research focuses on resolving primary conflicts, that is, simply avoiding obstacles. However, in future ultra-low altitude airspace, traffic throughput is expected to far exceed the current level of manned air traffic, possibly by several orders of magnitude. When air traffic density increases, resolving only primary conflicts may trigger more secondary conflicts, that is, drones may have new conflicts with other drones after avoiding the current obstacle. In addition, this means that drones need to perform more obstacle avoidance operations, thereby increasing the instability of the entire airspace.
[0004] (2) Secondary conflict is an important indicator for measuring airspace stability, but to date, there is no real-time quantitative modeling method for secondary conflict. At the same time, most obstacle avoidance modules currently rely on traditional optimization methods or geometric algorithms, which limits their processing capabilities when facing a large number of obstacles and makes it difficult to effectively reduce secondary conflicts.
[0005] (3) Most of the current conflict resolution methods focus on a single flight phase, such as before or after takeoff. No research has explored how to integrate the two-phase conflict resolution into an integrated framework to improve the overall safety of UAV flight in ultra-low altitude airspace. According to existing research, relying solely on conflict detection and resolution before takeoff, complex weather conditions (such as strong winds) or mechanical limitations will cause the UAV to be unable to strictly follow the planned path, thereby increasing the risk of conflict. If only relying on conflict detection and resolution after takeoff, in order to ensure computational efficiency, learning algorithms are usually required. However, for such complex and highly uncertain ultra-low altitude airspace, the time required for pre-training will increase exponentially with the size of the airspace, and it is difficult to achieve effective conflict resolution. Summary of the invention
[0006] Purpose of the invention: In view of the above problems, the purpose of the present invention is to provide a double-layer conflict resolution method for UAVs in urban low-altitude dense traffic flow airspace.
[0007] Technical solution: The present invention provides a method for resolving UAV double-layer conflicts in urban low-altitude dense traffic airspace, comprising:
[0008] Before the intelligent UAV takes off, a mixed integer nonlinear programming model based on conflict resolution strategy is constructed, and the improved random fractal search algorithm is used to optimize the mixed integer nonlinear programming model to obtain the initial parameters of the intelligent UAV takeoff.
[0009] After the intelligent drone takes off, a dual deep Q network is constructed and trained. The introduced attention mechanism is used to simulate the influence of surrounding neighbors on the intelligent drone. The trained dual deep Q network is used to enable the intelligent drone to output conflict-free maneuvers at each time step until the intelligent drone reaches its final destination.
[0010] Furthermore, the process of constructing a mixed integer nonlinear programming model based on a conflict resolution strategy includes:
[0011] Based on the flight plan and the real-time flight data of intelligent drones, an initial four-dimensional route is constructed for all drones in the airspace;
[0012] Check whether there is a flight conflict on the four-dimensional route, and assign a resolution strategy to each flight conflict detected;
[0013] Define the optimization objective function of the i-th flight conflict as f i (ζ i ), where ζ i is a feasible solution in the search space, in are the decision variables of the conflict resolution strategy, representing the position, take-off time and cruising speed of the intelligent drone respectively. The expression of the objective function is:
[0014] min:f obj =ω risk R CP +ω t_delay T delay +ω t_air T air
[0015] In the formula, ω risk ,ω t_delay ,ω t_air are the weight factors of risk, delay and flight time respectively, and ω risk +ωt_delay +ω t_air =1, R CP Indicates potential conflict risk, T delay represents the total delay time of all drones, T air Represents the total flight time of all drones;
[0016] The constraints for constructing the objective function are expressed as:
[0017]
[0018]
[0019] In the formula, Indicates drone U g The actual departure time, Indicates drone U g The actual flight time, Indicates drone U g Scheduled departure time, Indicates drone U g Scheduled flight time, Indicates drone U g The delay time, N uav Represents the total number of drones; N p Represents the total number of adjacent waypoint pairs, d ij Represents arc J ij The Euclidean distance of Indicates the battery life. represents the number of drones departing from the starting point, represents the number of drones that have reached the destination, Indicates drone U g In Arc J ij Cruising speed on and Respectively represent the minimum and maximum values of the UAV cruising speed;
[0020] The mixed integer nonlinear programming model is composed of the objective function and constraints.
[0021] Furthermore, the process of optimizing the mixed integer nonlinear programming model using the improved random fractal search algorithm includes the following steps:
[0022] In step 101, assuming that a volume particle is regarded as a potential solution, the particle is initialized randomly within the problem condition constraints, which can be expressed as:
[0023] P i =B lower +λ(B upper -B lower )
[0024] Where P i represents the i-th initial particle in the population, B lower and B upper Respectively represent the lower and upper limits of the constraint vector, λ is a random number, and λ∈[0,1];
[0025] Step 102, generating new particles in the search space through Gaussian random walk distribution, expressed as:
[0026]
[0027] Where P i η is composed of particle P i The i-th new particle generated, η is the number of new particles generated, μ P is the mean of the Gaussian distribution of the new particle position, μ P =|P i |, σ is the standard deviation, P best is the particle with the best solution in the current group, and Represents the weight factor for adjusting the degree of development and exploration, which follows a uniform distribution and ranges from [0,1];
[0028] Step 103, all generated particles are evaluated by the fitness function, the solution particles with fitness values greater than the threshold remain unchanged, and the remaining solution particles are updated when the following conditions are met, and the conditions are expressed as:
[0029]
[0030] Where P ai Indicates the probability of a particle being updated, rank(P i ) represents the ranking of the particle after evaluation by the fitness function, n p represents the total number of particles;
[0031] Step 104, record the position of the best particle in the population as X best , calculate the distance value d between the updated solution particle and the best particle Pi , summarizing all distance values into a distance vector D P =[d1 d2…d n ] T ;
[0032] Step 105, calculate the fitness value of each solution particle through the fitness function, and summarize all fitness values into a fitness vector F P =[f1 f2…f n ] T ;
[0033] Step 106, normalize the distance vector and the fitness vector, and calculate the score vector of each solution particle, expressed as:
[0034] S P =ω FDB F P_norm +(1-ω FDB )D P_norm
[0035] In the formula, F P_norm is the normalized value of the fitness vector, D P_norm is the normalized value of the distance vector, ω FDB is the weight parameter;
[0036] Step 107, sort all the score vectors numerically, select the top 5% of the score vectors as the dominant population, and update the probability of each particle being updated according to the following formula:
[0037]
[0038] In the formula, l rate represents the update rate, P dominant represents the probability matrix of particles in the dominant population, N dominant represents the total number of particles in the dominant population;
[0039] Step 108, repeating steps 101 to 107 until the maximum number of iterations is reached, stopping the iteration, and selecting the solution particle with the highest fitness function value as the final solution to the optimization problem.
[0040] Furthermore, the fitness function consists of an objective function and a penalty function, expressed as:
[0041] f value =f obj (x,t,v)+ψ(x,t,v)
[0042] Where ψ(w,v,t) is the penalty function, expressed as:
[0043]
[0044] In the formula, ω fc Represents the weight factor of flight conflict, N fc Represents the total number of flight conflicts in a solution.
[0045] Furthermore, the dual deep Q network includes a policy network and a target network. The policy network and the target network have the same structure, both of which are composed of multi-layer perceptrons and attention mechanisms.
[0046] Furthermore, constructing a dual deep Q network and performing network training on it includes the following steps:
[0047] Step 201, initializing the experience pool, and initializing the parameters θ of the strategy network and the parameters θ' of the target network;
[0048] Step 202, in the simulation environment, the intelligent drone uses a policy network and combines the ε-greedy algorithm to generate actions, and uses a reward function to calculate the reward of the current action;
[0049] Step 203: Collecting experience of intelligent drones Add to the experience pool. When the number of experiences in the experience pool is greater than or equal to the batch size, random batch sampling is performed from the experience pool, and the parameters θ of the policy network are updated according to the loss function; represents the complete observation vector of the intelligent UAV at time t, represents the action of the intelligent drone at time t, r t represents the reward obtained by the intelligent drone at time t;
[0050] Step 204, updating the number of iterations and periodically copying the parameters of the policy network to the parameters of the target network;
[0051] Step 205, repeating steps 202 to 204 until the maximum number of iterations is reached, and then stopping the iteration.
[0052] Furthermore, the trained dual deep Q network is used to enable the intelligent drone to output conflict-free maneuvers at each time step, including:
[0053] Step 301, construct and initialize the simulation environment;
[0054] Step 302, creating a pre-planned path for each drone;
[0055] Step 303: The intelligent drone obtains the current complete observation vector This includes the intelligent drone's own state vector The state vectors of the surrounding K neighboring drones And the state vector related to the static obstacles around the intelligent drone
[0056] Step 304, input the complete observation vector into the dual deep Q network to obtain all actions and corresponding Q values of the intelligent drone under the current observation state;
[0057] Step 305: The intelligent drone selects the action corresponding to the maximum Q value to execute.
[0058] Further, step 304 includes:
[0059] First, use a multi-layer perceptron to transform the complete observation vector of the intelligent drone Abstract a vector with a length of 256, then use the abstract vector representing the state of the intelligent drone itself as the query vector of the attention mechanism, and use the abstract vector representing the state of other drones around the intelligent drone as the key vector and value vector. The weight ω of the value vector is calculated by the key vector and the query vector. The expression is:
[0060]
[0061] In the formula, MLP represents multi-layer perceptron, dim represents the dimension of input and output;
[0062] Calculate the context vector V according to the weight c , represents the influence of the surrounding neighboring drones on the current main drone, expressed as:
[0063] V c =ωV t neighbor
[0064] Where V t neighbor represents the feature vector extracted from the vector representing the surrounding neighbor drones by a multi-layer perceptron, and the expression is:
[0065]
[0066] The context vector V c A link operation is performed with the abstract vector of the static obstacles around the drone, and then it passes through a multi-layer perceptron to obtain the final Q value, and the optimal action for the next step is selected from it, which is expressed as:
[0067]
[0068] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:
[0069] 1. This paper proposes a quantitative method that can measure secondary conflicts in real time and integrate this quantitative index into the reward function of deep reinforcement learning, so that the deep reinforcement learning method can take optimized actions that both reduce secondary conflicts and avoid primary conflicts when performing tasks;
[0070] 2. The present invention introduces a new optimization method in the pre-tactical stage of the flight to provide more reasonable initial parameters, thereby accelerating the convergence speed in the training process and improving the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 This is a flowchart of the UAV double-layer conflict resolution method for urban low-altitude dense traffic flow airspace;
[0072] Figure 2 This is a schematic diagram of the simulated operation airspace;
[0073] Figure 3 This is a grid diagram of the simulated operation airspace;
[0074] Figure 4 It is the network structure diagram of the dual-depth Q network;
[0075] Figure 5 This is the network structure diagram of the attention mechanism;
[0076] Figure 6 A schematic diagram of the environment used in the computer simulation;
[0077] Figure 7 A comparison chart of the number of secondary conflicts. DETAILED DESCRIPTION
[0078] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0079] The method for resolving the dual-layer conflict of UAVs in the urban low-altitude dense traffic airspace described in this embodiment is as follows: Figure 1 As shown, including:
[0080] Before the intelligent UAV takes off, a mixed integer nonlinear programming model based on conflict resolution strategy is constructed, and the improved random fractal search algorithm is used to optimize the mixed integer nonlinear programming model to obtain the initial parameters of the intelligent UAV takeoff.
[0081] After the intelligent drone takes off, a dual deep Q network is constructed and trained. The introduced attention mechanism is used to simulate the influence of surrounding neighbors on the intelligent drone. The trained dual deep Q network is used to enable the intelligent drone to output conflict-free maneuvers at each time step until the intelligent drone reaches its final destination.
[0082] In this embodiment, a double-layer conflict resolution framework is proposed specifically for ultra-low-altitude urban environments, assuming that the sky traffic flow is very dense, and deep reinforcement learning is combined to propose a double-layer conflict resolution framework, which includes two stages of conflict resolution before and during flight for intelligent drones. Before flight, the improved random fractal search algorithm is used to find the optimal take-off time and cruising speed of each drone, and they are used as the initial parameters after take-off; then the dual deep Q network framework is used as the learning framework, so that the intelligent drone can output conflict-free maneuvers at each time step; then the attention mechanism is used to construct the influence of surrounding neighboring drones on the intelligent drone, so that the learned conflict resolution strategy can adapt to any number of neighboring drones. In this embodiment, pre-flight conflict resolution and real-time conflict resolution are combined, which not only greatly reduces the occurrence of air collisions compared to using any one method alone, but also enables intelligent drones to avoid current obstacles in ultra-low-altitude urban dense traffic flow environments, while reducing secondary conflicts, and significantly shortens the training time of real-time conflict resolution methods based on reinforcement learning.
[0083] Furthermore, the process of constructing a mixed integer nonlinear programming model based on a conflict resolution strategy includes:
[0084] (1) Based on the flight plan and the real-time flight data of intelligent drones, an initial four-dimensional route is constructed for all drones in the airspace.
[0085] (2) Check whether there are any flight conflicts on the four-dimensional route and assign a resolution strategy to each flight conflict detected.
[0086] (3) Define the optimization objective function of the i-th flight conflict as f i (ζ i ), where ζ i is a feasible solution in the search space, in are the decision variables of the conflict resolution strategy, which represent the position, take-off time and cruising speed of the intelligent drone. It should be noted that the position change is a discrete change based on the grid space, which means that the position can only jump between predefined grid nodes, but cannot be continuously displaced. In addition, hovering action is not allowed, that is, continuous decision variables cannot point to the same position point to ensure that each step of the action reflects clear mobility and directionality. The expression of the objective function is:
[0087] min:f obj =ω risk R CP +ω t_delay T delay +ω t_air T air
[0088] In the formula, ωrisk ,ω t_delay ,ω t_air are the weight factors of risk, delay and flight time respectively, and ω risk +ω t_delay +ω t_air = 1. In simple terms, the goal is to solve f1, f2, ... f F , in order to find a better strategy i .
[0089] Among them, R CP Represents the potential conflict risk, which is related to the trajectory intersection point and can be written as:
[0090]
[0091] In the formula, Indicates intelligent drone U g Distance from route intersection, W cp Indicates the distance the drone passes through all route intersections;
[0092] T delay The total delay time of all drones is expressed as:
[0093]
[0094] In the formula, Indicates intelligent drone U g The actual departure time, Indicates the scheduled departure time. air Represents the total flight time of all drones, expressed as:
[0095]
[0096] In the formula, w i and w j is the adjacent waypoint, W r represents a set of waypoints, and the neighbor relationship is defined by the absolute difference between their indexes i and j, denoted by: |ij| = 1; d ij Represents arc J ij (w i ,w j ), Indicates drone U g In Arc J ij Cruising speed on.
[0097] (4) Construct the constraints of the objective function, expressed as:
[0098]
[0099] In the formula, Indicates drone U g The actual departure time, Indicates drone U g The actual flight time, Indicates drone U g Scheduled departure time, Indicates drone U g Scheduled flight time, Indicates drone U g The delay time, N uav Represents the total number of drones; N p Represents the total number of adjacent waypoint pairs, d ij Represents arc J ij The Euclidean distance of Indicates the battery life. represents the number of drones departing from the starting point, represents the number of drones that have reached the destination, Indicates drone U g In Arc J ij Cruising speed on and Respectively represent the minimum and maximum cruising speed of the drone.
[0100] The first constraint is to limit the arrival delay time of the intelligent drone, which means that the time for the intelligent drone to arrive at the destination should not exceed the threshold. The second constraint is to ensure that the flight time of the smart drone does not exceed its battery limit; the third constraint is to ensure that the traffic flow remains balanced, and the number of smart drones departing from the starting point The number of smart drones required to reach the destination Maintain balance; the fourth constraint is to ensure that the cruising speed of the intelligent drone does not exceed the physical limit of the drone.
[0101] (5) The mixed integer nonlinear programming model is composed of the objective function and constraints.
[0102] In one example, at any time, the flight plan in an airspace is visible in real time. Based on the flight plan and the real-time flight data of the drone, a four-dimensional flight trajectory can be constructed for all drones in the airspace. Potential conflicts may occur at the waypoints where the trajectories intersect. The four-dimensional route can be defined as a directed graph G = (W, J), where W = {0, 2, ..., w + 1} is a set of waypoints, where waypoint 0 represents the starting position and waypoint w + 1 represents the destination. Set is the set of arcs of adjacent waypoints, where and is the three-dimensional grid coordinate after running the spatial domain grid, Nnode represents the total number of nodes i and j in the directed graph. The waypoint where the trajectories cross is defined as the intersection of at least two arcs, expressed as CWP k represents the waypoint where the trajectories intersect. The set of estimated arrival times of each UAV at any waypoint is expressed as The set of all drones in the airspace is represented as the set Where N uav It is represented as the total number of drones. The set of four-dimensional routes is represented as Where N UAV Represents the total number of all drones in the airspace, Indicates drone U g The four-dimensional route, is a three-dimensional vector, representing the three-dimensional grid coordinates of the four-dimensional aerial point, Indicates the time when the drone arrives at the 4D waypoint.
[0103] Flight conflicts are identified by finding intersection waypoints, which are formed by a pair of flight tracks. When the time difference between the two drones reaching the intersection waypoint is less than the conflict threshold Δt conflict When , it is considered that there is a conflict, which can be expressed as:
[0104]
[0105] In the formula, represents the time it takes for a single drone to reach a crossing waypoint; in this example, Δt conflict Set to 10 seconds.
[0106] The set of detected flight conflicts is represented as C = {c1, c2, ..., c p},c p represents the pth flight conflict, p∈N fc , N fc Indicates the total number of flight conflicts currently detected.
[0107] Furthermore, the process of optimizing the mixed integer nonlinear programming model using the improved random fractal search algorithm includes the following steps:
[0108] In step 101, assuming that a volume particle is regarded as a potential solution, the particle is initialized randomly within the problem condition constraints, which can be expressed as:
[0109] P i =B lower +λ(B upper -B lower )
[0110] Where Pi represents the i-th initial particle in the population, B lower and B upper Respectively represent the lower and upper limits of the constraint vector, λ is a random number, and λ∈[0,1];
[0111] Step 102, generating new particles in the search space through Gaussian random walk distribution, expressed as:
[0112]
[0113] Where P i η is composed of particle P i The i-th new particle generated, η is the number of new particles generated, μ P is the mean of the Gaussian distribution of the new particle position, μ P =|P i |, σ is the standard deviation, P best is the particle with the best solution in the current group, and Represents the weight factor for adjusting the degree of development and exploration, which follows a uniform distribution and ranges from [0,1];
[0114] The calculation formula of standard deviation σ is:
[0115]
[0116] Where N Iter Indicates the total number of iterations;
[0117] Step 103, all generated particles are evaluated by the fitness function, the solution particles with fitness values greater than the threshold remain unchanged, and the remaining solution particles are updated when the following conditions are met, and the conditions are expressed as:
[0118]
[0119] In the formula, Indicates the probability of a particle being updated, rank(P i ) represents the ranking of the particle after the fitness function is evaluated, which is used to determine the relative superiority or inferiority of the particle in the population. p represents the total number of particles;
[0120] Among them, the remaining solution particle update process is divided into two stages, and the first stage update is expressed as:
[0121]
[0122] Where P r and P t They are other particles randomly selected from the current population;
[0123] Then the second stage update is performed according to the following formula, expressed as:
[0124]
[0125] Where P t ' and P r ' are other particles randomly selected from the current population;
[0126] Step 104, record the position of the best particle in the population as X best , calculate the distance between the updated solution particle and the best particle All distance values are summed up into a distance vector D P =[d1d2…d n ] T ;
[0127] The calculation expression of the distance value is:
[0128]
[0129] Where v is the dimension of the particle, Represented as particle P i The value of dimension v, Denoted as the best particle P best The value in dimension v;
[0130] Step 105, calculate the fitness value of each solution particle through the fitness function, and summarize all fitness values into a fitness vector F P =[f1 f2…f n ] T ;
[0131] Step 106, normalize the distance vector and the fitness vector, and calculate the score vector of each solution particle, expressed as:
[0132] S P =ω FDB F P_norm +(1-ω FDB )D P_norm
[0133] In the formula, F P_norm is the normalized value of the fitness vector, D P_norm is the normalized value of the distance vector, ω FDB is a weight parameter, which is used to adjust the weights of fitness value and distance value during the update process. Its value range is [0,1]. The score vector of each particle is determined by the fitness function and distance to combine the needs of development and exploration;
[0134] Step 107, sort all the score vectors numerically, select the top 5% of the score vectors as the dominant population, and update the probability of each particle being updated according to the following formula:
[0135]
[0136] In the formula, l rate represents the update rate, P dominant represents the probability matrix of particles in the dominant population, N dominant represents the total number of particles in the dominant population;
[0137] Step 108, repeating steps 101 to 107 until the maximum number of iterations is reached, stopping the iteration, and selecting the solution particle with the highest fitness function value as the final solution to the optimization problem.
[0138] Furthermore, the fitness function consists of an objective function and a penalty function, expressed as:
[0139] f value =f obj (x,t,v)+ψ(x,t,v)
[0140] Where ψ(w,v,t) is the penalty function, expressed as:
[0141]
[0142] In the formula, ω fc Represents the weight factor of flight conflict, N fc Represents the total number of flight conflicts in a solution.
[0143] Furthermore, the dual deep Q network includes a policy network and a target network. The policy network and the target network have the same structure, both of which are composed of multi-layer perceptrons and attention mechanisms.
[0144] Furthermore, constructing a dual deep Q network and performing network training on it includes the following steps:
[0145] Step 201, initializing the experience pool, and initializing the parameters θ of the strategy network and the parameters θ' of the target network;
[0146] Step 202, in the simulation environment, the intelligent drone uses a policy network and combines the ε-greedy algorithm to generate actions, and uses a reward function to calculate the reward of the current action;
[0147] The reward function is defined as:
[0148]
[0149] Among them, the parameter r crossThe calculation expression is:
[0150]
[0151] In the formula, It represents the track deviation of the intelligent UAV relative to its own reference track, α and β are constant coefficients, which are 4000 and 3.5 respectively in this embodiment;
[0152] Parameter r goal The calculation expression is:
[0153]
[0154] In the formula, and represents the Euclidean distance between the position of the intelligent drone at the previous moment and the next moment and the destination, respectively, and μ is a constant coefficient, which is 6 in this embodiment;
[0155] Parameter r intru The calculation expression is:
[0156]
[0157] In the formula, rh NMAC represents the radius of the intelligent drone, which can be 2.5 meters; d near represents the Euclidean distance between the smart drone and its nearest drone; ρ is a constant coefficient, which can be 50;
[0158] Parameter r dec The calculation expression is:
[0159]
[0160] In the formula, and They represent the number of potential conflicts of the drone under the current action and the number of potential conflicts under the next given subsequent action; h is a constant coefficient, which can be 2; dec t To model a real-time quantitative definition of secondary conflict, the expression is:
[0161]
[0162] Parameter r pen The calculation expression is:
[0163]
[0164] In the above reward function, r goal and r intruThe main function of the robot is to guide the intelligent drone towards the target without conflict; dec Responsible for guiding the drone to learn to minimize secondary conflicts while performing evasive actions. When only these three parts are used as rewards, the strategy of the intelligent drone will cause it to move to a corner of the map and take small incremental actions, waiting for all other drones to leave the airspace before moving towards the target. Therefore, r is needed. cross and r pen To prevent smart drones from learning this extremely inefficient strategy.
[0165] Step 203: Collecting experience of intelligent drones Add to the experience pool. When the number of experiences in the experience pool is greater than or equal to the batch size, random batch sampling is performed from the experience pool, and the parameters θ of the policy network are updated according to the loss function; represents the complete observation vector of the intelligent UAV at time t, represents the action of the intelligent drone at time t, r t represents the reward obtained by the intelligent drone at time t;
[0166] Among them, the expression of the loss function is:
[0167]
[0168] In the formula, represents the expected value calculation of the experience data sampled in the experience replay pool D, r t represents the reward obtained by the current intelligent drone at time t, γ∈[0,1) represents the discount factor, represents the estimated Q value, which is calculated by the target network, and the parameters of the target network are represented by θ'; Indicates that the drone selects the current action The complete state after that; Q represents the current Q value, which is calculated by the policy network, and the parameters of the policy network are represented by θ.
[0169] Step 204, updating the number of iterations and periodically copying the parameters of the policy network to the parameters of the target network;
[0170] Step 205, repeating steps 202 to 204 until the maximum number of iterations is reached, and then stopping the iteration.
[0171] After the loss function is obtained, back propagation is used to calculate the gradient of the loss function to the policy network parameters, and then the parameters of the policy network are updated through the optimization algorithm Adam to minimize the loss. Finally, at regular training rounds, such as once every 50 updates, the policy network parameters are copied to the target network to maintain the stability of the target network. When the return of each round of training tends to be stable, it means that the policy network training is completed.
[0172] Furthermore, the trained dual deep Q network is used to enable the intelligent drone to output conflict-free maneuvers at each time step, including:
[0173] Step 301, construct and initialize the simulation environment;
[0174] Step 302, creating a pre-planned path for each drone;
[0175] Step 303: The intelligent drone obtains the current complete observation vector This includes the intelligent drone's own state vector The state vectors of the surrounding K neighboring drones And the state vector related to the static obstacles around the intelligent drone
[0176] Step 304, input the complete observation vector into the dual deep Q network to obtain all actions and corresponding Q values of the intelligent drone under the current observation state;
[0177] Step 305: The intelligent drone selects the action corresponding to the maximum Q value to execute.
[0178] Further, step 304 includes:
[0179] First, use a multi-layer perceptron to transform the complete observation vector of the intelligent drone Abstract a vector with a length of 256, then use the abstract vector representing the state of the intelligent drone itself as the query vector of the attention mechanism, and use the abstract vector representing the state of other drones around the intelligent drone as the key vector and value vector. The weight ω of the value vector is calculated by the key vector and the query vector. The expression is:
[0180]
[0181] In the formula, MLP represents multi-layer perceptron, dim represents the dimension of input and output;
[0182] Calculate the context vector V according to the weight c , represents the impact of the surrounding neighbor drones on the current intelligent drone, expressed as:
[0183] V c =ωV tneighbor
[0184] Where V t neighbor represents the feature vector extracted from the vector representing the surrounding neighbor drones through a multilayer perceptron neural network, and is expressed as:
[0185]
[0186] The context vector V c A link operation is performed with the abstract vector of the static obstacles around the drone, and then it passes through a multi-layer perceptron to obtain the final Q value, and the optimal action for the next step is selected from it, which is expressed as:
[0187]
[0188] In the formula, It means finding the action that can maximize the objective function in a set of possible actions a. The output of the objective function here is also the final Q value.
[0189] In one example, if Figure 2 As shown in the figure, a 210×130 meter airspace is selected from a city map as the research airspace, including the detection range of the drone, their respective starting and ending points, waypoints, and intersection waypoints. In this example, only the cruising state is considered, so the height of the airspace is not considered, and the discretization is as follows Figure 6 The 10-meter equidistant grid shown in Figure 2 In the model, the drone is modeled as a circle with a radius of 2.5 meters. 25 non-intelligent drones are generated. The non-intelligent drones have their own starting and ending points, but they cannot avoid obstacles autonomously. Intelligent drones are generated and the starting and ending points are generated. Figure 6 The environment actually used in computer simulation includes the modeling of buildings after discretization of the entire environment, as well as the layout of the initial flight routes of 25 non-intelligent drones and one intelligent drone in the airspace. After the simulation environment is built, the offline training phase of the dual deep Q network can be started, setting the batch size to 64, the learning rate to 0.00001, the total number of training rounds to 50,000, and a maximum of 45 time steps per round. When using the ε-greedy algorithm, ε starts from 1 in the first round and decreases linearly in each round, dropping to 0.05 in the 25,000th round, and then does not continue to decrease. The jumping point algorithm is used to create a pre-planned path for each drone. The turning points on the path will be used as its waypoints. At the same time, the planned path will be analyzed to find intersections, which are identified and set as intersection waypoints. The value range of the decision variable, the constraint threshold and the weight factor are then set. Potential conflict risk R in the fitness function CPBefore takeoff, each position in the airspace that can be used as a waypoint will be assigned a higher risk value. Figure 3 As shown in the figure, after the airspace is gridded using the concept of AirMatrix, the optional positions in the airspace are, grid ∈[1,21],y grid ∈[1,13]. The decision variables of each drone’s take-off time (unit: seconds) and cruising speed (unit: meters / second) are t etd ∈[1,30] and v level ∈[5,15]. In addition, the delay time threshold Set to 10 seconds, the battery life threshold Set to 900 seconds. Before takeoff, conflict resolution is defined as the minimum time difference between consecutive drones passing through the intersection waypoint is at least 3 seconds, and the weight of potential conflict before takeoff ω risk is 0.5, the flight delay weight ω t_delay and the air time weight ω t_air are all 0.25, and the conflict penalty weight ω fc Set to 0.1. Before the Z intelligent drone takes off, the mixed integer nonlinear programming model is optimized using an improved random fractal search algorithm to find the take-off time and cruising speed of each aircraft, which will be used as input parameters for subsequent deep reinforcement learning.
[0190] Load the trained dual deep Q network. The environment perceived by the intelligent drone includes its own state vector, the state vectors from neighboring drones, and the state vectors of static obstacles within the detection range. The complete state space of the intelligent drone is defined as in Represents the state vector of the intelligent drone, expressed as: where e t c Indicates the lateral deviation between the position of the intelligent drone and the reference path, and Represents the difference between the current speed and the previous speed of the smartphone on the x-axis and y-axis respectively. and Respectively represent the current speed of the intelligent drone on the x-axis and y-axis, and Represents the Euclidean distance between the intelligent drone and its next waypoint in the x-axis and y-axis directions respectively. There are a total of K neighboring drones within the detection range of the intelligent drone, represented by i = {1, 2, ..., K}. The state of a single i-th neighboring drone is It is expressed as: in, and are the Euclidean distances between the intelligent and the ith neighboring drone in the x-axis and y-axis directions, and represents the speed of the i-th neighboring drone on the x-axis and y-axis. Represents the state vector related to the static obstacles around the drone. Using the spatial data of the building, which contains the top view outline of the building (in polygonal form) and the height from the ground to the highest point, the concept of AirMatrix is used to discretize the airspace into cubes of equal size. When a cube intersects or overlaps with any polygon, the corresponding position of the cube in the matrix becomes 1. Therefore, a three-dimensional binary matrix can represent the entire airspace. Since only horizontal flight is considered, a two-dimensional matrix can be used To indicate the occupancy of the surrounding area by the building within the detection range.
[0191] The complete observation vector of the intelligent drone will be used to obtain the Q value corresponding to all actions of the intelligent drone in the current observation state through the trained dual-depth Q network. Figure 4 is the network structure diagram of the dual deep Q network used in this embodiment, which also includes the attention mechanism module, such as Figure 5 As shown in Figure 2, the three observation vectors of the intelligent drone use three MLP structures to extract features. Each MLP consists of one fully connected layer, including 256 neurons. The state vector related to the static obstacles around the intelligent drone It needs to be flattened before being fed into the MLP structure. and After extracting the features, the context vector V is extracted through the attention mechanism c , represents the influence of the surrounding neighboring drones on the current main drone. The feature and context vector V c as well as The features are spliced together, and then the Q value of all actions of the intelligent drone in the current state is output through the final MLP structure. The final MLP consists of 3 fully connected layers, with the number of neurons being 512, 512, and 9 respectively. Except for the last fully connected layer with 9 neurons, there is no activation function, and the activation functions of other fully connected layers are all rectified linear units.
[0192] The intelligent drone selects the action corresponding to the maximum Q value and executes it. The action space of the intelligent drone is discrete, with 9 options, each of which consists of a fixed value of acceleration in the x-axis and y-axis directions. The composition is as follows:
[0193]
[0194] Among them, h a is the physical limit of the acceleration of the intelligent drone in this embodiment, which is a constant real number and can be set to 4m / s 2 ,The state transition model of the intelligent drone is: It is a conditional probability that describes the likelihood of transitioning to another state given the current state and the selected action. The actual state transition model is hidden to the smart drone, but it can be derived from the dynamics given to the smart drone. The dynamics used by the smart drone in this embodiment are defined as follows:
[0195]
[0196] When the current speed of the intelligent drone exceeds the maximum speed allowed by its physical performance, The speed of the smart drone changes to:
[0197]
[0198] Among them, v max The maximum speed allowed by the physical performance of the intelligent drone in this embodiment is set to 15m / s. It is the heading direction of the current intelligent drone obtained by the velocity vector.
[0199] The intelligent drone executes the action corresponding to the maximum Q value output by the current maximum strategy network at each time step, and will steadily fly to its destination while trying not to deviate too far from the predetermined route, and reduce secondary conflicts while avoiding collisions. Figure 7 As shown in the figure, when the deep reinforcement learning algorithm proposed in the present invention is used in the intelligent drone in this embodiment, the number of secondary conflicts accumulated after it reaches the end point under different airspace congestion levels (5, 15, and 25 other drones respectively). At the same time, the number of secondary conflicts accumulated when the most advanced non-learning algorithm Optimal Reciprocal Collision Avoidance (ORCA) is used is also counted for comparison. Through the comparison, it can be seen that when the airspace is more crowded, the use of the deep reinforcement learning algorithm DDQN-attention proposed in the present invention can more effectively reduce the number of accumulated secondary conflicts, thereby taking into account the stability of the airspace.
Claims
1. A dual-layer conflict resolution method for UAVs in urban low-altitude dense traffic airspace, characterized by: include: Before the intelligent UAV takes off, a mixed integer nonlinear programming model based on conflict resolution strategy is constructed, and the improved random fractal search algorithm is used to optimize the mixed integer nonlinear programming model to obtain the initial parameters of the intelligent UAV takeoff. After the intelligent drone takes off, a dual deep Q network is constructed and trained. The introduced attention mechanism is used to simulate the influence of surrounding neighbors on the intelligent drone. The trained dual deep Q network is used to enable the intelligent drone to output conflict-free maneuvers at each time step until the intelligent drone reaches its final destination.
2. The method for resolving UAV double-layer conflicts in urban low-altitude dense traffic flow airspace according to claim 1 is characterized in that: The process of constructing a mixed integer nonlinear programming model based on conflict resolution strategy includes: Based on the flight plan and the real-time flight data of intelligent drones, an initial four-dimensional route is constructed for all drones in the airspace; Check whether there is a flight conflict on the four-dimensional route, and assign a resolution strategy to each flight conflict detected; Define the optimization objective function of the i-th flight conflict as f i (ζ i ), where ζ i is a feasible solution in the search space, in are the decision variables of the conflict resolution strategy, representing the position, take-off time and cruising speed of the intelligent drone respectively. The expression of the objective function is: min:f obj =ω risk R CP +oh t_delay T delay +oh t_air T air In the formula, ω risk ,ω t_delay ,ω t_air are the weight factors of risk, delay and flight time respectively, and ω risk +ω t_delay +ω t_air =1, R CP Indicates potential conflict risk, T delay represents the total delay time of all drones, T air Represents the total flight time of all drones; The constraints for constructing the objective function are expressed as: In the formula, Indicates drone U g The actual departure time, Indicates drone U g The actual flight time, Indicates drone U g Scheduled departure time, Indicates drone U g Scheduled flight time, Indicates drone U g The delay time, N uav Represents the total number of drones; N p Represents the total number of adjacent waypoint pairs, d ij Represents arc J ij The Euclidean distance of Indicates the battery life. represents the number of drones departing from the starting point, represents the number of drones that have reached the destination, Indicates drone U g In Arc J ij Cruising speed on and Respectively represent the minimum and maximum values of the UAV cruising speed; The mixed integer nonlinear programming model is composed of the objective function and constraints.
3. The method for resolving UAV double-layer conflicts in urban low-altitude dense traffic flow airspace according to claim 2 is characterized in that: The process of optimizing the mixed integer nonlinear programming model using the improved random fractal search algorithm includes the following steps: In step 101, assuming that a volume particle is regarded as a potential solution, the particle is initialized randomly within the constraints of the problem conditions, which can be expressed as: P i =B lower +λ(B upper -B lower ) Where P i represents the i-th initial particle in the population, B lower and B upper Respectively represent the lower and upper limits of the constraint vector, λ is a random number, and λ∈[0,1]; Step 102, generating new particles in the search space through Gaussian random walk distribution, expressed as: Where P i η is composed of particle P i The i-th new particle generated, η is the number of new particles generated, μ P is the mean of the Gaussian distribution of the new particle position, μ P =|P i |, σ is the standard deviation, P best is the particle with the best solution in the current group, and Represents the weight factor for adjusting the degree of development and exploration, which follows a uniform distribution and ranges from [0,1]; Step 103, all generated particles are evaluated by the fitness function, the solution particles with fitness values greater than the threshold remain unchanged, and the remaining solution particles are updated when the following conditions are met, and the conditions are expressed as: In the formula, Indicates the probability of a particle being updated, rank(P i ) represents the ranking of the particle after evaluation by the fitness function, n p represents the total number of particles; Step 104, record the position of the best particle in the population as X best , calculate the distance between the updated solution particle and the best particle All distance values are summed up into a distance vector D P =[d1d2…d n ] T ; Step 105, calculate the fitness value of each solution particle through the fitness function, and summarize all fitness values into a fitness vector F P =[f1 f2…f n ] T ; Step 106, normalize the distance vector and the fitness vector, and calculate the score vector of each solution particle, expressed as: S P =ω FDB F P_norm +(1-ω FDB )D P_norm In the formula, F P_norm is the normalized value of the fitness vector, D P_norm is the normalized value of the distance vector, ω FDB is the weight parameter; Step 107, sort all the score vectors numerically, select the top 5% of the score vectors as the dominant population, and update the probability of each particle being updated according to the following formula: In the formula, l rate represents the update rate, P dominant represents the probability matrix of particles in the dominant population, N dominant represents the total number of particles in the dominant population; Step 108, repeating steps 101 to 107 until the maximum number of iterations is reached, stopping the iteration, and selecting the solution particle with the highest fitness function value as the final solution to the optimization problem.
4. The method for resolving the double-layer conflict of unmanned aerial vehicles in the urban low-altitude dense traffic flow airspace according to claim 3 is characterized in that: The fitness function consists of an objective function and a penalty function, expressed as: f value =f obj (x,t,v)+ψ(x,t,v) Where ψ(w,v,t) is the penalty function, expressed as: In the formula, ω fc Represents the weight factor of flight conflict, N fc Represents the total number of flight conflicts in a solution.
5. The method for resolving the double-layer conflict of unmanned aerial vehicles in the urban low-altitude dense traffic flow airspace according to claim 4 is characterized in that: The dual deep Q network includes a policy network and a target network. The policy network and the target network have the same structure, both of which are composed of multi-layer perceptrons and attention mechanisms.
6. The method for resolving UAV double-layer conflicts in urban low-altitude dense traffic airspace according to claim 5 is characterized in that: Building a dual deep Q network and training it includes the following steps: Step 201, initializing the experience pool, and initializing the parameters θ of the strategy network and the parameters θ' of the target network; Step 202, in the simulation environment, the intelligent drone uses a policy network and combines the ε-greedy algorithm to generate actions, and uses a reward function to calculate the reward of the current action; Step 203: Collecting experience of intelligent drones Add to the experience pool. When the number of experiences in the experience pool is greater than or equal to the batch size, random batch sampling is performed from the experience pool, and the parameters θ of the policy network are updated according to the loss function; represents the complete observation vector of the intelligent UAV at time t, represents the action of the intelligent drone at time t, r t represents the reward obtained by the intelligent drone at time t; Step 204, updating the number of iterations and periodically copying the parameters of the policy network to the parameters of the target network; Step 205, repeating steps 202 to 204 until the maximum number of iterations is reached, and then stopping the iteration.
7. The method for resolving UAV double-layer conflicts in urban low-altitude dense traffic flow airspace according to claim 6 is characterized in that: The trained dual deep Q network is used to enable the intelligent drone to output conflict-free maneuvers at each time step, including: Step 301, construct and initialize the simulation environment; Step 302, creating a pre-planned path for each drone; Step 303: The intelligent drone obtains the current complete observation vector This includes the intelligent drone's own state vector The state vectors of the surrounding K neighboring drones And the state vector related to the static obstacles around the intelligent drone Step 304, input the complete observation vector into the dual deep Q network to obtain all actions and corresponding Q values of the intelligent drone under the current observation state; Step 305: The intelligent drone selects the action corresponding to the maximum Q value to execute.
8. The method for resolving UAV double-layer conflicts in urban low-altitude dense traffic flow airspace according to claim 7 is characterized in that: Step 304 includes: First, use a multi-layer perceptron to transform the complete observation vector of the intelligent drone Abstract a vector with a length of 256, then use the abstract vector representing the state of the intelligent drone itself as the query vector of the attention mechanism, and use the abstract vector representing the state of other drones around the intelligent drone as the key vector and value vector. The weight ω of the value vector is calculated by the key vector and the query vector. The expression is: In the formula, MLP represents multi-layer perceptron, dim represents the dimension of input and output; Calculate the context vector V according to the weight c , represents the influence of the surrounding neighboring drones on the current main drone, expressed as: V c =ωV t neighbor Where V t neighbor It is a feature vector extracted from the vector representing the surrounding neighboring drones by a multi-layer perceptron, and is expressed as: The context vector V c A link operation is performed with the abstract vector of the static obstacles around the drone, and then it passes through a multi-layer perceptron to obtain the final Q value, and the optimal action for the next step is selected from it, which is expressed as:
Citation Information
Patent Citations
Real-time path planning method for unmanned aerial vehicle based on deep reinforcement learning
CN110488872A
Unmanned aerial vehicle online collaborative airspace conflict resolution method based on iterative space mapping
CN112883493A
Unmanned aerial vehicle path planning method based on hierarchical reinforcement learning
CN115268494A
Conflict minimization flight path collaborative planning method considering high-altitude wind time variation
CN115938162A
Urban scene-oriented unmanned aerial vehicle adaptive conflict resolution method
CN116880541A
Cited By
Flight conflict resolution method considering uncertainty of man-machine interaction and air-ground communication
CN121053823A
Flight conflict resolution method considering human-machine interaction and air-ground communication uncertainty
CN121053823B
Intelligent management and control method and system for low-altitude corridor
CN121093479A
Unmanned aerial vehicle inspection path planning method and system based on AI
CN121115870A