Split type flying vehicle learning planning method for obstacle areas with different densities

By dividing the obstacle area and optimizing the action space in the flying vehicle path planning, the problem of low learning efficiency of flying vehicles in complex environments is solved, and efficient air-ground collaborative mission path planning is achieved.

CN120779945APending Publication Date: 2025-10-14BEIJING INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510916827.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

When a flying vehicle plans its path in a complex environment, due to the large dimension of the action space, the computational complexity of updating the Q table in Q-learning is high and the learning efficiency is low, which cannot meet the needs of efficient path planning.

Method used

An obstacle density threshold is introduced to divide the map environment into obstacle areas of different densities. Different action spaces are used for different density areas. In particular, redundant actions are eliminated in sparse obstacle areas, the action space is optimized, and a corresponding reward or penalty feedback mechanism is designed.

Benefits of technology

The computational complexity of Q-table updates in Q-learning is reduced, the learning efficiency of the flying vehicle is improved, and a shorter air-ground collaborative mission path can be planned more efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120779945A_ABST
    Figure CN120779945A_ABST
Patent Text Reader

Abstract

The invention provides a split type flying vehicle learning planning method for obstacle areas with different densities, and belongs to the technical field of aircraft planning. Comprising the following steps: step 1, dividing an obstacle area into a dense obstacle area and a sparse obstacle area; step 2, motion space optimization is carried out; different action spaces are adopted for obstacle areas with different densities; step 3, planning a task path; according to the method, a flying vehicle adopts different action spaces and environments to perform efficient interaction, an action sequence capable of accumulating rewards to the maximum, namely an optimal action sequence, is obtained, and finally, a short air-ground cooperative task path is planned with higher learning efficiency. According to the method, the concept of an obstacle density threshold value is introduced, especially for a sparse obstacle area, the flying vehicle mainly keeps the action towards the target position direction, other redundant actions are removed, and the action space dimension is effectively reduced. Therefore, the learning efficiency is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention provides a split-type flying vehicle learning planning method for areas with different density obstacles, belonging to the technical field of aircraft planning. Background Art

[0002] To address the pain points of ground traffic congestion and limited emergency rescue on inaccessible short-distance roads, flying vehicles, a fusion of aerial vehicles and ground vehicles, have emerged. They possess both aerial and ground modes of motion and can be used to perform highly maneuverable missions involving air-ground coordination. Currently, flying vehicles come in a wide variety of configurations. Among these, split flying vehicles utilize a modular and shared approach, cleverly combining ground and air functions. This effectively integrates urban spatial resources, alleviates ground traffic congestion, and is expected to become a key component of future smart city transportation networks.

[0003] Autonomous mission execution for flying vehicles requires perception, planning, and control technologies. Path planning, in particular, can provide a collision-free and optimal path for air-ground collaborative missions, helping them successfully complete their missions. Current path planning algorithms primarily include random sampling algorithms, artificial potential field algorithms, graph search algorithms, and intelligent algorithms. Q-Learning, a popular intelligent algorithm, uses a reward-and-penalty principle, interacting with the environment through a trial-and-error approach to maximize cumulative rewards. It exhibits excellent adaptability to various complex environments and can converge to the optimal path solution through multiple iterations of learning. Q-learning consists of five basic elements: an intelligent agent, an environment, a state s, an action a, and a reward R. During path planning for air-ground collaborative missions, the flying vehicle, acting as an intelligent agent, interacts with the complex environment by executing specific actions a, thereby changing its positional state s and receiving immediate feedback in the form of reinforcement signals (rewards or penalties). During this interactive process, the flying vehicle receives a penalty signal when encountering an obstacle, signaling the irrationality of its current action. When it successfully avoids an obstacle or reaches its destination, it receives a reward signal, reinforcing its effective action. Through continuous interaction with the environment, the flying vehicle gradually learns and accumulates experience, ultimately achieving the action sequence that maximizes cumulative rewards, effectively completing path planning and achieving optimal navigation from its starting position to its destination.

[0004] For ground vehicles, the above-mentioned action a includes driving on the ground in 8 directions, and its action space dimension is 8. Flying vehicles have two motion modes: driving on the ground and flying in the air. Its action a is more complex, including: driving on the ground in 8 directions, taking off, landing, and flying in 8 directions in the air. It can be seen that the action space dimension of a flying vehicle is 18, which is much larger than the action space dimension of a ground vehicle. In a complex map environment, the use of Q-learning-based methods faces significant challenges when planning air-ground collaborative mission paths for flying vehicles. Due to the large action space dimension, the computational complexity of updating the Q table in Q-learning will increase significantly, and the learning efficiency will be greatly reduced, which cannot meet the needs of efficient path planning for flying vehicles. Summary of the Invention

[0005] In response to the above technical problems, the present invention provides a split-type flying vehicle learning planning method for obstacle areas of different densities. The method introduces the concept of obstacle density threshold and divides the complex map environment into obstacle areas of different densities, specifically into dense obstacle areas and sparse obstacle areas. Different action spaces are adopted for obstacle areas of different densities. In particular, for sparse obstacle areas, the flying vehicle mainly retains the action in the direction of the target position and eliminates the remaining redundant actions, effectively reducing the dimension of the action space. This optimization operation reduces the computational complexity of the Q table update in Q-learning, thereby significantly improving the learning efficiency. The flying vehicle can more efficiently plan a shorter air-ground collaborative task path.

[0006] The specific technical solutions are:

[0007] The split-type flying vehicle learning planning method for areas with different obstacle densities includes the following steps:

[0008] Step 1: Obstacle area division. Rasterize the map, input the starting position and target position. Divide the complex map environment into n areas, and calculate the obstacle density in each area as follows:

[0009]

[0010] Where D i is the obstacle density in each area, O i N is the number of grid cells occupied by obstacles in each area. i is the total number of grids in each region, i=1,…,n.

[0011] Introducing the obstacle density threshold T, according to the above calculated value D i , divide the complex map environment into obstacle areas of different densities. If D i ≥T, the current area is a dense obstacle area; if D i<T,则当前区域为稀疏障碍物区域。

[0012] Step 2: Action space optimization: Different action spaces are adopted for areas with different obstacle densities.

[0013] Facing a dense obstacle area, the flying vehicle adopts action space 1, which includes driving in 8 directions on the ground, switching between air and ground modes, and flying in 8 directions in the air.

[0014] For areas with sparse obstacles, the flying vehicle uses Action Space 2, which is optimized to retain air-to-ground mode switching and target-direction movements, while also selecting non-target directions based on a probabilistic selection mechanism. The optimized Action Space 2 includes: ground travel toward the target, air-to-ground mode switching, air flight toward the target, probabilistically selected ground travel toward the non-target, and probabilistically selected air flight toward the non-target.

[0015] Define the 18 actions in action space 1 as a i , i=1,…,18. The action space 2 is specifically expressed as follows:

[0016]

[0017] Where: (C x ,C y ) is the current position of the flying vehicle in the x and y directions, (E x ,E y ) is the target position state in the x and y directions; the action sets [a3,a4,a5], [a1,a2,a3], [a2,a3,a4], [a5,a6,a7], [a1,a7,a8], [a6,a7,a8], [a4,a5,a6], and [a1,a2,a8] are the ground facing the target direction under different conditions; the action set [a 13 ,a 14 ,a 15 ]、[a 11 ,a 12 ,a 13 ]、[a 12 ,a 13 ,a 14 ]、[a 15 ,a 16 ,a 17 ]、[a 11 ,a 17 ,a 18 ]、[a 16 ,a 17 ,a 18 ]、[a 14 ,a15 ,a 16 ]、and [a 11 ,a 12 ,a 18 ] is flying towards the target in the air under different circumstances; Figure 2 As shown, action a9 is takeoff, action a 10 To land; action For the probability selection, the ground moves in the non-target direction. After choosing the probability, fly in the direction of non-target in the air.

[0018] In different situations, it is used to select and The probability selection mechanism is as follows:

[0019]

[0020] Where: N g It is the set of ground actions moving in non-target directions; Yes N g The probability of different actions being selected; Yes Select N g The angular deviation between the actual heading of the flying vehicle and the target heading after different actions; N f It is a set of actions that fly in the air in a non-target direction; Yes N f The probability of different actions being selected; Yes Select N f The angular deviation between the actual heading and the target heading after different actions;

[0021] Based on the above probability and Roulette algorithm is used to select ground and air actions towards non-target directions and add them to the action space 2 of the flying vehicle, and and

[0022] Step 3: Mission Path Planning. The flying vehicle efficiently interacts with the environment in different action spaces to obtain the action sequence that maximizes the cumulative reward, i.e., the optimal action sequence. Ultimately, a shorter air-ground collaborative mission path is planned with higher learning efficiency.

[0023] The specific process is as follows: Different actions taken by the flying vehicle will correspond to different rewards or penalties. When the flying vehicle encounters an obstacle, it will receive a penalty signal; when it successfully avoids the obstacle or reaches the target point, it will receive a reward signal. The reward or penalty feedback of the flying vehicle is defined as follows:

[0024] R=R a-R o

[0025] Among them, R is the reward or penalty feedback of the flying vehicle; R a is the reward feedback of the above different action spaces; R o It is the penalty feedback for obstacle collision.

[0026] R o As shown below:

[0027]

[0028] Among them, h f is the required flight altitude facing the current obstacle; h m is the maximum flight altitude of the flying vehicle.

[0029] R a As shown below:

[0030]

[0031] Among them, D g is the ground distance corresponding to the current action; D f is the air flight distance corresponding to the current action; a and b are weight factors; It is the ground in action space 1 moving in 8 directions, It is flying in 8 directions in the air in action space 1, is the space-ground mode switching action in action space 1, is the ground in action space 2 moving in different directions, It is flying in different directions in the air in action space 2, is the space-ground mode switching action in action space 2,

[0032] The flying vehicle takes different actions to continuously interact with the environment. Based on the above reward or penalty feedback, it eventually obtains a series of actions that maximize the cumulative reward, that is, the optimal action sequence, and completes the final plan.

[0033] This invention introduces the concept of obstacle density thresholds, dividing complex environment maps into zones of varying obstacle density. Based on this, the flight vehicle's action space is further optimized, designing two action spaces for zones of varying obstacle density. Furthermore, corresponding reward or penalty feedback is proposed for each action space.

[0034] In areas with sparse obstacles, the vehicle primarily retains actions toward the target location while simultaneously selecting actions toward non-target locations based on a probabilistic selection mechanism. This optimization eliminates redundant actions, reduces the dimensionality of the action space, and thus reduces the computational complexity of updating the Q-table in Q-learning. This enables the vehicle to plan a shorter path for the air-ground collaborative mission with higher learning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is the algorithm flow chart of the present invention;

[0036] Figure 2 Schematic diagram of action space 1 of the present invention;

[0037] Figure 3 Schematic diagram of action space 2 of the present invention;

[0038] Figure 4 This is the path diagram of the flying vehicle air-ground collaborative mission of the present invention. DETAILED DESCRIPTION

[0039] The specific technical solution of the present invention is described in conjunction with the accompanying drawings. Figure 1 shown.

[0040] A split-type flying vehicle learning and planning method for areas with different obstacle densities is characterized by comprising the following steps:

[0041] Step 1: Obstacle area division. Rasterize the map, input the starting position and target position. Divide the complex map environment into n areas, and calculate the obstacle density in each area as follows:

[0042]

[0043] Where D i is the obstacle density in each area, O i N is the number of grid cells occupied by obstacles in each area. i is the total number of grids in each region, i=1,…,n.

[0044] Introducing the obstacle density threshold T, according to the above calculated value D i , divide the complex map environment into obstacle areas of different densities. If D i ≥T, the current area is a dense obstacle area; if D i <T,则当前区域为稀疏障碍物区域。

[0045] Step 2, action space optimization. Different action spaces are adopted for areas with different density obstacles. Facing dense obstacle areas, the flying vehicle needs sufficient actions to ensure the success rate of its interaction with the environment, so as to effectively avoid complex obstacles and ensure the feasibility and safety of path planning. Therefore, facing dense obstacle areas, the flying vehicle adopts action space 1, which includes driving in 8 directions on the ground, air-ground mode switching actions (take-off and landing), and flying in 8 directions in the air. Figure 2 As shown, the dimension of the action space is 18.

[0046] In areas with sparse obstacles, however, the flying vehicle has fewer obstacles to avoid, and redundant actions can be appropriately eliminated to reduce the dimensionality of the action space, lower the computational complexity of Q-table updates, and improve the learning efficiency of Q-learning. Therefore, in areas with sparse obstacles, the flying vehicle uses action space 2, which is optimized to retain air-to-ground mode switching actions and actions toward the target position, while selecting actions toward non-target directions based on a probabilistic selection mechanism. The optimized action space 2 includes: driving toward the target on the ground, switching between air and ground modes, flying toward the target in the air, driving toward the non-target on the ground after probabilistic selection, and flying toward the non-target in the air after probabilistic selection.

[0047] Define the 18 actions in action space 1 as a i , i=1,…,18. The action space 2 is specifically expressed as follows:

[0048]

[0049] Where: (C x ,C y ) is the current position of the flying vehicle in the x and y directions, (E x ,E y ) is the target position state in the x and y directions; the action sets [a3,a4,a5], [a1,a2,a3], [a2,a3,a4], [a5,a6,a7], [a1,a7,a8], [a6,a7,a8], [a4,a5,a6], and [a1,a2,a8] are the ground facing the target direction under different conditions; the action set [a 13 ,a 14 ,a 15 ]、[a 11 ,a 12 ,a 13 ]、[a 12 ,a 13 ,a 14 ]、[a 15 ,a 16 ,a 17 ]、[a 11,a 17 ,a 18 ]、[a 16 ,a 17 ,a 18 ]、[a 14 ,a 15 ,a 16 ]、and [a 11 ,a 12 ,a 18 ] is flying towards the target in the air under different circumstances; Figure 2 As shown, action a9 is takeoff, action a 10 To land; action For the probability selection, the ground moves in the non-target direction. After choosing the probability, fly in the direction of non-target in the air.

[0050] It can be seen that action space 2 eliminates redundant actions, and its dimension is 10, which is smaller than that of action space 1. The optimization of action space 2 effectively reduces the computational complexity of Q-table update in Q-learning and achieves higher learning efficiency. As an example, the action space 2 is as follows Figure 3 Assume

[0051] It should be further explained that in different cases, the and The probability selection mechanism is as follows:

[0052]

[0053] Where: N g It is the set of ground actions moving in non-target directions; Yes N g The probability of different actions being selected; Yes Select N g The angular deviation between the actual heading of the flying vehicle and the target heading after different actions; N f It is a set of actions that fly in the air in a non-target direction; Yes N f The probability of different actions being selected; Yes Select N f The angular deviation between the actual heading and the target heading after different actions;

[0054] Based on the above probability and Roulette algorithm is used to select ground and air actions towards non-target directions and add them to the action space 2 of the flying vehicle, and and

[0055] Step 3: Mission Path Planning. The flying vehicle uses the aforementioned different action spaces to efficiently interact with the environment, obtaining the action sequence that maximizes the cumulative reward, i.e., the optimal action sequence. Ultimately, it plans a shorter air-ground collaborative mission path with higher learning efficiency.

[0056] The specific process is as follows: Different actions taken by the flying vehicle will correspond to different rewards or penalties. Generally, when the flying vehicle encounters an obstacle, it will receive a penalty signal; when it successfully avoids the obstacle or reaches the target point, it will receive a reward signal. The reward or penalty feedback of the flying vehicle is defined as follows:

[0057] R=R a -R o

[0058] Among them, R is the reward or penalty feedback of the flying vehicle; R a is the reward feedback of the above different action spaces; R o It is the penalty feedback for obstacle collision.

[0059] R o As shown below:

[0060]

[0061] Among them, h f is the required flight altitude facing the current obstacle; h m is the maximum flight altitude of the flying vehicle.

[0062] R a As shown below:

[0063]

[0064] Among them, D g is the ground distance corresponding to the current action; D f is the air flight distance corresponding to the current action; a and b are weight factors; It is the ground in action space 1 moving in 8 directions, It is flying in 8 directions in the air in action space 1, is the air-ground mode switching action (take-off and landing) in action space 1, is the ground in action space 2 moving in different directions, It is flying in different directions in the air in action space 2, is the air-ground mode switching action (take-off and landing) in action space 2,

[0065] The flying vehicle takes different actions to continuously interact with the environment. Based on the above reward or penalty feedback, it eventually obtains a series of actions that maximize the cumulative reward, that is, the optimal action sequence, and completes the final plan.

[0066] The planned air-ground collaborative mission path is shown in Figure 4 .

[0067] like Figure 4 As shown, in the construction of future smart city transportation networks, the points where the air-to-ground mode switching occurs in the above-mentioned mission paths can be precisely set as the ground-to-air / air-to-ground transition points of the flying vehicle. At these key transition points, the split-body flying vehicle can efficiently switch between the ground driving module and the air flying module, providing people with three-dimensional travel services and greatly enhancing the flexibility and efficiency of urban transportation.

[0068] The present invention retains the three directions of ground / air movements toward the target position, and the subsequent addition of new movements in multiple directions can be considered as the protection scope of the present invention.

[0069] This paper proposes an efficient learning and planning algorithm for a split-type flying vehicle in areas with varying density obstacles, enabling efficient path planning in complex map environments. The key technical points involved are as follows:

[0070] 1. Calculate the obstacle density in different areas of the complex environment map and introduce the concept of obstacle density threshold. If the calculated obstacle density is greater than the threshold, the current area is a dense obstacle area; otherwise, the current area is a sparse obstacle area.

[0071] 2. Different action spaces are adopted for areas with different density obstacles. For dense obstacle areas, in order to ensure the success rate of the flying vehicle's interaction with the environment, the flying vehicle's action space contains enough actions, specifically driving in 8 directions on the ground, air-ground mode switching actions (take-off and landing), and flying in 8 directions in the air. For sparse obstacle areas, in order to improve the learning efficiency of the algorithm, the flying vehicle's action space is optimized, redundant actions are eliminated, and air-ground mode switching actions and actions towards the target position are retained. At the same time, actions towards non-target directions are selected based on a probabilistic selection mechanism. In addition, corresponding reward or penalty feedback is proposed for different action spaces. The above operations effectively avoid complex obstacles and ensure the feasibility and safety of path planning, while reducing the dimension of the action space and the computational complexity of the Q table update in Q-learning. The flying vehicle can more efficiently plan a shorter air-ground collaborative task path.

Claims

1. A split-type flying vehicle learning and planning method for areas with different obstacle densities, characterized by: The following steps are involved: Step 1: Obstacle area division: rasterize the map, input the starting position and target position; divide the complex map environment into n areas, calculate the obstacle density of each area, and divide it into dense obstacle area and sparse obstacle area; Step 2: Action space optimization: adopt different action spaces for areas with different obstacle densities; Facing dense obstacle areas, the flying vehicle adopts action space 1, which includes driving in 8 directions on the ground, switching between air and ground modes, and flying in 8 directions in the air; For areas with sparse obstacles, the flying vehicle uses action space 2, which is optimized to retain the air-to-ground mode switching action and the action towards the target position, while selecting the action towards the non-target direction based on the probabilistic selection mechanism. The optimized action space 2 includes: ground driving towards the target direction, air-to-ground mode switching action, flying towards the target direction in the air, ground driving towards the non-target direction after probabilistic selection, and flying towards the non-target direction in the air after probabilistic selection. Step 3: Mission path planning: The flying vehicle adopts different action spaces to interact efficiently with the environment, obtaining the action sequence that can maximize the cumulative reward, that is, the optimal action sequence, and ultimately planning a shorter air-ground collaborative mission path with higher learning efficiency.

2. The split-type flying vehicle learning and planning method for areas with different density obstacles according to claim 1 is characterized in that: The following steps are involved: Step 1: The obstacle area is divided as follows: Where D i is the obstacle density in each area, O i N is the number of grid cells occupied by obstacles in each area. i is the total number of grids in each region, i=1,…,n; Introduce the obstacle density threshold T, and according to the above calculated value D i , divide the complex map environment into obstacle areas with different densities; if D i ≥T, then the current area is a dense obstacle area; if D i <T, then the current area is a sparse obstacle area.

3. The split-type flying vehicle learning and planning method for areas with different density obstacles according to claim 1 is characterized in that: Step 2 The specific method is: Define the 18 actions in action space 1 as a i , i=1,…,18; the action space 2 is specifically expressed as follows: Where: (C x ,C y ) is the current position of the flying vehicle in the x and y directions, (E x ,E y ) is the target position state in the x and y directions; the action sets [a3,a4,a5], [a1,a2,a3], [a2,a3,a4], [a5,a6,a7], [a1,a7,a8], [a6,a7,a8], [a4,a5,a6], and [a1,a2,a8] are the ground facing the target direction under different conditions; the action set [a 13 ,a 14 ,a 15 ]、[a 11 ,a 12 ,a 13 ]、[a 12 ,a 13 ,a 14 ]、[a 15 ,a 16 ,a 17 ]、[a 11 ,a 17 ,a 18 ]、[a 16 ,a 17 ,a 18 ]、[a 14 ,a 15 ,a 16 ]、and [a 11 ,a 12 ,a 18 ] is flying towards the target direction in the air under different conditions; as shown in Figure 2, action a9 is taking off, action a 10 To land; action For the probability selection, the ground moves in the non-target direction. After choosing the probability, fly in the direction of non-target in the air; In different situations, it is used to select and The probability selection mechanism is as follows: Where: N g It is the set of ground actions moving in non-target directions; Yes N g The probability of different actions being selected in Y i g Yes Select N g The angular deviation between the actual heading of the flying vehicle and the target heading after different actions; N f It is a set of actions that fly in the air in a non-target direction; Yes N f The probability of different actions being selected in Y i f Yes Select N f The angular deviation between the actual heading and the target heading after different actions; Based on the above probability and Roulette algorithm is used to select ground and air actions towards non-target directions and add them to the action space 2 of the flying vehicle, and and 4. The learning and planning method for a split-type flying vehicle facing obstacles of different densities according to claim 3 is characterized in that: The specific process of Step 3 is as follows: Different actions taken by the flying vehicle will correspond to different rewards or penalties. When the flying vehicle encounters an obstacle, it will receive a penalty signal; when it successfully avoids an obstacle or reaches the target point, it will receive a reward signal. The reward or penalty feedback of the flying vehicle is defined as follows: R=R a -R o Among them, R is the reward or penalty feedback of the flying vehicle; R a is the reward feedback in different action spaces; R o It is the penalty feedback for obstacle collision; R o As shown below: Among them, h f is the required flight altitude facing the current obstacle; h m is the maximum flight altitude of the flying vehicle; R a As shown below: Among them, D g is the ground distance corresponding to the current action; D f is the air flight distance corresponding to the current action; a and b are weight factors; It is the ground in action space 1 moving in 8 directions, It is flying in 8 directions in the air in action space 1. is the space-ground mode switching action in action space 1, is the ground in action space 2 moving in different directions, It is flying in different directions in the air in action space 2, is the space-ground mode switching action in action space 2, The flying vehicle takes different actions to continuously interact with the environment. Based on the above reward or penalty feedback, it eventually obtains a series of actions that maximize the cumulative reward, that is, the optimal action sequence, and completes the final plan.