A human-unmanned aerial vehicle cooperative path planning and decision method

By adjusting the parameters of the perturbed fluid field and generating multiple preference paths through reinforcement learning, and optimizing paths using the analytic hierarchy process (AHP), the multi-objective comprehensive optimization problem of path planning in dynamic threat environments is solved, improving the efficiency and safety of manned-unmanned collaborative operations.

CN121007558BActive Publication Date: 2026-03-31AIR FORCE UNIV PLA
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing path planning methods lack a multi-objective integrated optimization mechanism in dynamic threat environments, making it difficult to flexibly adapt to the needs of collaborative missions. Furthermore, the decision-making process lacks transparency, affecting the efficiency and safety of manned-unmanned collaborative operations.

Method used

By employing reinforcement learning to dynamically adjust the parameters of the disturbing fluid field, combined with multi-preference path generation and analytic hierarchy process (AHP), an environmentally adaptive path guidance is achieved. Furthermore, AHP is used for path optimization and decision-making to assist manned and machine-machine collaboration in selecting the optimal path.

Benefits of technology

It improves the environmental adaptability and decision-making transparency of path planning, enhances the decision-making efficiency and execution reliability of collaborative combat systems, and meets the needs of multi-objective missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121007558B_ABST
    Figure CN121007558B_ABST
Patent Text Reader

Abstract

The present application provides a manned-unmanned aerial vehicle cooperative path planning and decision-making method, which comprises at least one manned aerial vehicle and one unmanned aerial vehicle, and the method comprises: environment perception and obstacle modeling, construction of interference fluid field and adaptive adjustment parameter, flight state coding and reinforcement learning training, multi-preference path generation, path optimization and auxiliary decision-making based on analytic hierarchy process; the present application dynamically adjusts the interference fluid field parameter through reinforcement learning, can flexibly respond to threat changes according to real-time environment perception data, has stronger environmental adaptability and guiding flexibility, and avoids path from falling into local minimum value; through the multi-preference path generation network, multiple candidate flight paths with different optimization tendencies are output in real time, providing diversified path selection for multi-target task execution; the analytic hierarchy process is introduced for path optimization, and the safety, time efficiency and energy consumption criteria are quantitatively evaluated, so that the explainability and transparency of path decision-making are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle path planning technology, specifically relating to a manned-unmanned cooperative path planning and decision-making method. Background Technology

[0002] With the rapid development of unmanned aerial vehicle (UAV) technology, manned-UAV collaborative operations have become an important means of improving air combat effectiveness. In collaborative missions, UAVs must not only autonomously complete path planning and obstacle avoidance tasks, but also closely cooperate with manned aircraft, obey command instructions, and collaboratively complete complex combat missions. When aircraft perform missions in dynamic threat environments, they must simultaneously cope with multiple challenges such as electronic interference, sudden threats, and dynamic obstacles. Path planning must not only ensure flight safety, but also take into account mission requirements such as time efficiency and energy consumption, and possess good environmental adaptability and decision-making transparency to support the needs of high-intensity, highly dynamic collaborative operations.

[0003] Path planning methods, exemplified by the artificial potential field method, are widely used for path guidance in static environments due to their simplicity and high computational efficiency. However, this method relies on fixed parameters and lacks dynamic response capabilities to changes in environmental conditions, easily leading to paths getting trapped in local minima, affecting path safety and effectiveness, and making it unsuitable for tasks in dynamic threat environments (as described in CN112180954B). Building on this, deep reinforcement learning methods have been gradually applied to path planning in complex scenarios, exhibiting strong environmental adaptability and the ability to acquire efficient obstacle avoidance and navigation strategies through training. However, existing methods generally suffer from the following shortcomings:

[0004] On the one hand, reinforcement learning path planning methods often focus on a single optimization objective, lacking a systematic multi-objective comprehensive optimization mechanism. This makes it difficult to achieve a dynamic balance among multiple task requirements such as safety, time efficiency, and energy consumption, resulting in path selection that is difficult to flexibly adapt to the diverse needs of collaborative tasks. On the other hand, existing path decision-making processes are mostly "black box" outputs, lacking clear path optimization and decision support mechanisms. This makes it difficult to meet the pilot's needs for path understandability and decision transparency in manned-unmanned collaborative operations, affecting the reliability and execution efficiency of collaborative decision-making (as described in CN113110592B). Summary of the Invention

[0005] To overcome the shortcomings of existing technologies in terms of poor path adaptability, insufficient multi-objective decision support, and weak path selection transparency under dynamic threat environments, this invention proposes a novel path planning method capable of real-time perception of environmental changes, adaptive generation of multiple optimized paths, and human-machine collaborative decision-making mechanism to assist in trajectory selection. This aims to improve the mission completion capability and decision-making efficiency of collaborative combat systems in complex environments. The technical approach of this method is as follows: First, the UAV completes flight environment modeling and local threat field description; second, based on reinforcement learning, fluid guidance parameters are dynamically adjusted to achieve environmentally adaptive path guidance; subsequently, a multi-preference path generation mechanism is designed to form multiple candidate paths with different optimization tendencies; finally, combined with the analytic hierarchy process (AHP), each path is comprehensively evaluated based on indicators such as safety, time efficiency, and energy consumption to assist the manned aircraft in completing the optimal path selection, achieving a collaborative closed loop of "UAV detection and manned aircraft decision-making."

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A manned-unmanned cooperative path planning and decision-making method, comprising at least one manned aircraft and one unmanned aircraft, wherein the unmanned aircraft, as a forward reconnaissance unit, is deployed ahead of the manned aircraft and is responsible for situational awareness and threat modeling; the manned aircraft, as the decision-making and mission command entity, completes path optimization and trajectory decision-making, including:

[0008] Step S1: Environmental perception and obstacle modeling, that is, the UAV uses its onboard multimodal sensor system to identify ground and air obstacles in the flight area, extract their spatial distribution, and perceive electromagnetic threats, thereby obtaining real-time environmental perception data and obstacle threat potential field;

[0009] Step S2: Construct the interference fluid field and adaptive adjustment parameters, that is, based on the obstacle threat potential field and combined with the task target information, construct the interference fluid field with global path guidance capability;

[0010] Step S3: Flight state encoding and reinforcement learning training, which involves designing the flight state feature vector of the UAV and training the path control policy function to guide the UAV to select the optimal path action in the guidance field;

[0011] Step S4: Multi-preference path generation, which is to generate a multi-preference path candidate set based on the path control strategy function and the disturbance fluid field;

[0012] Step S5: Path optimization and decision support based on the analytic hierarchy process (AHP), which involves taking a multi-preference path candidate set as input and combining it with real-time environmental perception data to conduct a comprehensive evaluation based on multiple criteria using the AHP.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0014] (1) This invention dynamically adjusts the parameters of the disturbance fluid field through reinforcement learning, and can flexibly respond to threat changes based on real-time environmental perception data. Compared with the traditional fixed parameter potential field method, it has stronger environmental adaptability and guidance flexibility, and avoids the path from getting trapped in local minima.

[0015] (2) By generating a multi-preference path network, multiple candidate paths with different optimization tendencies are output in real time, which breaks the limitation of single-objective optimization in existing path planning methods and provides diversified path selection for multi-objective task execution.

[0016] (3) The analytic hierarchy process is introduced for path optimization. Based on safety, time efficiency and energy consumption criteria, quantitative evaluation is carried out, which improves the interpretability and transparency of path decision-making, makes it easier for manned aircraft pilots to make reasonable decisions within a limited time, and enhances the decision-making efficiency and execution reliability of the collaborative combat system. Attached Figure Description

[0017] Figure 1 This is a flowchart of the method of the present invention;

[0018] Figure 2 This is the three-level decision structure in step S5 of the method of the present invention;

[0019] Figure 3 These are the drone tracks generated by different path preference strategies in the embodiments of the present invention;

[0020] Figure 4 This is the initial node of the flight process in this embodiment of the invention;

[0021] Figure 5 This is the path expansion node of the flight process in this embodiment of the invention.

[0022] Figure 6 This refers to the threat interaction node in the flight process of this embodiment of the invention.

[0023] Figure 7 This refers to the mid-to-late stage adjustment node of the flight process in this embodiment of the invention;

[0024] Figure 8 This refers to the approach to the target node during the flight process in this embodiment of the invention.

[0025] Figure 9 This is a flight trajectory presented in a three-dimensional view according to an embodiment of the present invention.

[0026] Figure 10 This is a curve showing the overall score variation of different preference path strategies in the embodiments of the present invention; Detailed Implementation

[0027] To make the objectives, technical methods, and advantages of the present invention clearer, the content of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0028] This invention includes at least one manned aircraft and one unmanned aircraft. The unmanned aircraft, as a forward reconnaissance unit, is deployed ahead of the manned aircraft and is responsible for situational awareness and threat modeling. The manned aircraft, as the decision-making and mission command entity, completes path selection and trajectory decision-making. The method flow of this invention is as follows: Figure 1 As shown, it specifically includes:

[0029] Step S1: Environmental perception and obstacle modeling. The UAV, using its onboard multimodal sensor system, including an image imaging unit and an electromagnetic detection unit, identifies ground and air obstacles within its flight area, extracts their spatial distribution, and senses electromagnetic threats, providing input data for subsequent path planning and decision-making. Real-time acquired data includes the obstacle's location p. i (t) and velocity variance Var(v obs ), nearest obstacle distance d min The initial distance D from the target point in the mission goal The relative distance ΔP from the target point goal and the target point location P goal The imaging unit possesses high-resolution ground imaging capabilities within its operating range, enabling it to identify potential obstacles in all weather conditions and complex terrain, and output the spatial location and outline of the obstacles. The electromagnetic detection unit, within its operating range, detects and determines the direction of enemy radar, communication, and other electromagnetic signals, inferring the location, type, and signal strength of threat sources, providing data support for threat level assessment.

[0030] After completing its perception phase, the UAV inputs the acquired obstacle data into its onboard computing unit for real-time processing. This onboard computing unit, as the core information processing module of the UAV platform, is responsible for obstacle data analysis and local potential field modeling. First, it extracts core feature parameters from the perceived obstacle information. Then, based on the analysis results, it generates an obstacle threat potential field using artificial potential field theory, creating a repulsive field with the obstacle as the source point to simulate its repulsive effect on the aircraft's path. The obstacle threat potential field U... obs (p,t) is modeled as follows:

[0031] In the formula: U obs (p,t) represents the total strength of the obstacle threat potential field perceived by the UAV at position p and time t; N obs R represents the number of obstacles detected within the current sensing range. s p is the set sensor sensing radius; p is the current position of the drone. i(t) represents the position of the i-th obstacle, r i Let $\frac{\pi}{\pi}$ be the effective radius of the obstacle, $k$ be the potential field strength coefficient, and $Θ(·)$ be the Heaviside function used to determine whether the obstacle is within the sensing range: when |pp i (t)|<R s When it is 1, it is 0.

[0032] During the modeling process, all obstacle information sensed by the UAV is simultaneously pushed to the manned aircraft, achieving real-time sharing of situational data. In subsequent path selection decisions, the manned aircraft uses this local threat model as input to ensure close linkage between decision-making and the environmental situation. Through this modeling, local environmental threats can be dynamically characterized, providing fundamental information for path guidance. The output of this step is the obstacle threat potential field U. obs (p,t) provides a foundation for the subsequent construction of the perturbed fluid field.

[0033] Step S2: Construct the disturbance fluid field and adaptively adjust the parameters. This step is based on the obstacle threat potential field U obtained by the forward-deployed UAV in Step S1, which is sensed and modeled by the forward-deployed UAV. obs (p,t), combined with mission objective information, a disruptive fluid field with global path guidance capabilities is constructed; reinforcement learning is used to adjust the density and diffusion rate of the fluid field in real time. This disruptive fluid field simulates the mechanical behavior of fluid flowing around obstacles, providing UAVs with continuous and directional path guidance, avoiding path abrupt changes and local minima traps. Specifically, it includes:

[0034] Step S2.1: Setting initial parameters for the interference fluid field, i.e., setting the initial density and diffusivity parameters of the fluid field based on the mission range and obstacle density. The density is inversely proportional to the mission range, and the diffusivity is directly proportional to the obstacle density, to adapt to the path guidance requirements under different mission scenarios.

[0035] The initial density ρ of the disturbed fluid field base With the initial diffusion rate σ base A parameterized method is used to set the initial density ρ based on the mission range and obstacle density. base Distance D from the mission objective goal They are inversely proportional, and the expression is: Where k ρ This is the density adjustment factor, with a value ranging from [10, 50]. The initial diffusivity σ base Based on the obstacle density λ within the sensing range obs To set, the expression is σ. base =k σ ·λ obs ,in N obs Represents the sensing radius R sThe number of obstacles per unit area within the range, k σ This is the diffusion rate adjustment coefficient, with a value range of [0.1, 1.0]. It takes a higher value in scenarios with dense obstacles to enhance path avoidance capabilities.

[0036] Step S2.2: Reinforcement learning adaptive parameter adjustment. A reinforcement learning agent based on the DDPG algorithm adjusts the density and diffusion rate of the fluid field in real time, enhancing the system's adaptability to complex dynamic environments. This agent takes perceived threat information and task status as input to dynamically adjust the path guidance direction and intensity, ensuring a balance between safety and mission advancement efficiency in scenarios such as low-altitude penetration.

[0037] To enhance the adaptability of the guiding field to dynamic environments, this invention designs an adaptive parameter adjustment mechanism based on reinforcement learning, which adjusts the guiding field density ρ in real time through an agent. t With diffusion rate σ t This enables dynamic adaptive adjustment of the path guidance direction and intensity. The guidance field parameter adjustment mechanism is designed as follows:

[0038]

[0039] In the formula: ρ t This represents the current density value of the disturbing fluid field; σ t ρ represents the current diffusion rate of the disturbing fluid field. base and σ base These are the initial density and diffusivity set during system initialization, respectively; Q safe (s t This represents the safety assessment of the current environmental state, reflecting the threat intensity under the current flight conditions. The calculation formula is: Among them The distance between the drone and the nearest obstacle is represented by d0, obtained from step 1; d0 is the normalized safety distance constant, which is taken as the sensor's sensing radius R. s One-fifth. Q time (s t () represents the mission time efficiency assessment, reflecting the propulsion efficiency of the aircraft relative to the mission objective. The calculation formula is: Where ΔP goal D represents the relative distance between the drone's current position and the target point. goalThe initial target distance is obtained from step 1. α and β are adjustable gain coefficients. α ranges from [0.1, 0.5], with larger values ​​for high-threat missions to increase safety redundancy and smaller values ​​for relatively stable routes. β ranges from [0.1, 0.3], with larger values ​​for time-sensitive missions to accelerate trajectory advancement and prioritizing safety. In low-altitude penetration mission scenarios, experiments verified that α = 0.3 and β = 0.2 can ensure safety while balancing path advancement efficiency, thus ensuring the successful completion of the flight mission. tanh(·) is the hyperbolic tangent function used for smoothing; sigmoid(·) is the logistic compression function used to limit the range of parameter changes. This step employs a policy gradient-based reinforcement learning algorithm, based on the real-time perception results obtained in step 1, including the obstacle position p. i (t) and velocity variance Var(v obs ), nearest obstacle distance d min The initial distance D from the target point in the mission goal The relative distance ΔP from the target point goal and the target point location P goal Key parameters such as environmental state are used as inputs for fluid field parameter regulation. The agent is trained using the DDPG (Deep Deterministic Policy Gradient) algorithm to predict the output fluid field regulation action [Δρ]. t ,Δσ t ], where Δρ t To guide the adjustment increment of the field density, Δσ t To guide the adjustment increment of the field diffusion rate, ρ t =ρ base +Δρ t ,σ t =σ base +Δσ t The updated ρ t With σ t It is used to generate the interfering fluid field F(p), providing continuous guiding field support for subsequent path strategy training (step S3). The density ρ of the interfering fluid field can be adjusted through this process. t With σ t This enables environmental adaptation to disrupt the fluid field.

[0040] Step S2.3: Generate a disturbance fluid field, which is constructed based on the negative direction of the obstacle threat potential field gradient to achieve a spatially continuous and directionally controllable guiding force field. The disturbance fluid field will serve as the input for subsequent path strategy training, providing continuous support for path generation and optimization.

[0041] The interfering fluid field F(p) is the obstacle threat potential field U. obsThe negative direction of the gradient of (p,t) (step 1), i.e.

[0042] ρ t σ represents the current density value of the disturbing fluid field. t To dynamically adjust U, the current diffusion rate of the perturbed fluid field is used. obs The spatial distribution structure of (p,t). Therefore, ρ t With σ t The directional distribution and guiding intensity of the disturbance fluid field F(p) are effectively controlled, determining the response of the subsequent path planning system to the environmental structure. This disturbance fluid field F(p) serves as the input basis for the path strategy training and execution in step 3, providing dynamic and continuous directional guidance and spatial constraints.

[0043] Step S3: Flight State Encoding and Reinforcement Learning Training. Based on the obstacle distribution and state information provided in Step S1, and the disturbance fluid field F(p) constructed in Step S2, this step designs the UAV's flight state feature vector s. t And train the path control policy function π(a|s) t This guides the drone to select the optimal path and maneuver in the guidance field, achieving coordinated optimization of obstacle avoidance and navigation.

[0044] Step S3.1: Flight state and reward function design, i.e., designing the flight state vector, which includes features such as the distance from the current position to the target point, the distance to the nearest obstacle, the expected speed and variance of obstacles, to describe the UAV's navigation situation. The reward function is designed based on safe distance, path length, and direction change angle, guiding the agent to achieve a balance between safety and task efficiency.

[0045] In this step, the flight state vector s is defined. t This describes the UAV's flight state in the current navigation environment and serves as input to the reinforcement learning path policy network. The flight state vector is constructed based on obstacle information perceived in step 1 and the location of the mission target point.

[0046] In the formula: s t ΔP is the flight state vector at the current time step. goal This represents the relative distance from the current location to the target point. The distance to the nearest obstacle. With Var(v obs ) represent the expected value and variance of the obstacle's velocity, respectively.

[0047] The corresponding reward function R t Designed as follows:

[0048] In the formula: R t d represents the instantaneous reward value obtained by the drone at the current time step. min The minimum distance between the current obstacle and the nearest obstacle, L is the cumulative path length, and θ is the distance between the current obstacle and the nearest obstacle. t Let be the angle of change in flight direction, and d0, λ, and μ be weighting coefficients. Reward function R. t Used to evaluate in state s t The advantages and disadvantages of different path actions are closely related to the design of the agent's state characteristics. Through the design of the reward function, the agent is guided to achieve a balance between obstacle avoidance safety and flight efficiency.

[0049] The output action 'a' is the core control parameter used to dynamically adjust the disturbed fluid field, specifically a = [Δρ t ,Δσ t ], where Δρ t To guide the adjustment increment of the field density; Δσ t This is to guide the adjustment increment of the field diffusion rate. Action 'a' represents the current flight state vector 's'. t The reinforcement learning agent adaptively corrects the parameters of the disturbing fluid field to respond to changes in environmental threat situations and mission requirements.

[0050] Step S3.2: Train the policy function objective, i.e., define the path policy function, with the optimization objective being to maximize the expected reward obtained in each state. Iteratively update the network weights using the policy gradient method to train the optimal policy function, providing decision support for subsequent path generation.

[0051] Based on the above design, the path strategy function π(a|s) t ) indicates that in state s t The probability of taking action 'a' represents the decision rule of the path control policy. The training objective is to maximize the reward obtained by the agent after choosing an action in each state. Specifically, the optimization objective is:

[0052] In the formula, π * (a|s t R represents the optimal strategy. t For state s t The immediate reward obtained by taking action a (step S3.1) represents the evaluation signal of the quality of the path action.

[0053] Through the path strategy function π(a|s) t Optimize and iterate to obtain the optimal path policy function π. * (a|s t In actual deployment, the current state can be considered. tReal-time output of optimal path action a = [Δρ t ,Δσ t The probability distribution of [ ] is used to achieve comprehensive optimization of obstacle avoidance, efficient advancement, and path smoothness. To achieve the above functions, the policy network designed in this invention adopts a "state perception - feature extraction - action mapping" structure, with an overall actor-critic network architecture. The feature extraction layer adopts a 3-layer fully connected neuron structure, with the number of neurons in each layer being [128, 64, 32] respectively, and the activation function being ReLU, to complete the nonlinear mapping of state features and high-dimensional feature extraction. After training, the policy π * (a|s t This will serve as an important input to the multi-preference path generation module in step S4, supporting the subsequent trajectory generation and optimization decision-making process.

[0054] Step S4: Multi-preference path generation, i.e., the path control policy function π(a|s) obtained in step S3. t Based on the previous step, and combined with the disturbance fluid field F(p) obtained in step S2, multiple candidate flight paths with different optimization tendencies (multi-preference path candidate set) are generated to meet the comprehensive requirements of multiple objectives such as safety, time efficiency, and energy consumption during mission execution. Each path is generated through independent strategy branches, taking flight state and fluid field as inputs, forming a multi-objective path set, providing diversified scheme support for subsequent path optimization.

[0055] To address this, a multi-preference path generation network is designed, with each branch corresponding to a different optimization preference (safety first, time first, energy balance). The optimization objective of each branch is expressed through a loss function. The definition is as follows:

[0056] In the formula: L k d represents the loss function value corresponding to the optimized branch of the k-th path; safe This is the preset minimum safe distance threshold; This indicates the total flight time for the route; This is an estimated value for path energy consumption; This represents a path smoothness index (such as total steering angle); k∈safety,time,balance represent different optimization preferences, ω1, ω2, ω3 are weighting coefficients, and d safe This is a set safe distance threshold. This mechanism allows for the simultaneous output of multiple candidate tracks that meet different decision-making needs.

[0057] Multi-preference path generation network with a unified flight state vector s t Taking the disturbing fluid field F(p) as input, the system outputs several path action sequences through a branching structure. Generate the corresponding path Pk , where k represents the path preference branch number. The feature extraction layer of the network uses a 3-layer fully connected neural network (with node sizes of 128, 64, and 32 respectively), with the ReLU activation function, to extract high-dimensional feature representations from the joint input. The output layer of the network is a set composed of K different preference policies:

[0058] π multi (a|s t )={π (1) (a|s t ),π (2) (a|s t ),...,π (K) (a|s t )} (8);

[0059] Each branch shares input features but uses different target weights for policy training, generating paths with specific optimization tendencies. Each preference policy network independently generates a complete sequence of path points in the perturbed fluid field F(p), resulting in a set of candidate paths.

[0060]

[0061] Each path P k All of them have clear preference attributes and performance indicators, as well as clear safety, time efficiency and energy consumption attributes, which serve as inputs for the subsequent step 5 "path selection and decision-making", providing manned aircraft pilots with diverse and interpretable basis for path selection.

[0062] Step S5: Path selection and decision support based on the analytic hierarchy process. This step uses the multi-preference path candidate set generated in step S4. Using the real-time environmental perception data provided in step S1 as input, the Analytic Hierarchy Process (AHP) is used to conduct a multi-criteria comprehensive evaluation.

[0063] This step constructs a three-tiered decision-making structure: "Objective Layer—Criterion Layer—Alternative Layer." For example... Figure 2 As shown, the top layer is the target layer, corresponding to the core objective of this invention—path optimization—aiming to select the optimal trajectory that best meets the current operational requirements from multiple candidate paths generated by the UAV. The middle layer is the criteria layer, which, combined with the actual needs of manned-UAV cooperative combat scenarios, sets three major evaluation criteria: safety, time efficiency, and energy consumption, respectively measuring the path's performance in threat avoidance, mission completion speed, and resource consumption. The bottom layer is the solution layer, consisting of multiple candidate paths with different optimization preferences output in step S4, serving as the specific objects for the optimization decision.

[0064] During the path selection process, the manned aircraft pilot sets the weight vector W = [w] of the evaluation criteria based on the current combat situation and mission requirements.safe ,w time ,w energy This ensures that the optimal results align with mission priorities and real-time tactical requirements. Weight values ​​can be initially configured based on pre-set mission scenarios (e.g., prioritizing security during low-altitude penetration) or dynamically adjusted according to the battlefield situation (e.g., automatically weighted based on threat density, mission time urgency, etc.). For example, by automatically calculating environmental complexity C... env Assisted weighting. Environment complexity C env The environmental information sensed by the UAV in step S1 is calculated in real time, including the distance to the nearest obstacle. Variance Var(v) of perceived obstacle velocity obs The relative distance between the current spacecraft and the mission target.

[0065] |ΔP goal |and the initial target distance D of the mission goal The calculation formula is as follows:

[0066]

[0067] Where R s For the sensing radius (100 nautical miles), v max Let λ1 = 0.5, λ2 = 0.3, and λ3 = 0.2, representing the maximum speed of the obstacle. The environmental complexity is C. env The value ∈ [1, 10] indicates a higher level of environmental threat and greater difficulty in path planning. This complexity calculation result can be used to assist in dynamically adjusting the weights in W, for example, automatically increasing the security weight w in a high-threat environment. safe This enhances path avoidance capabilities. During path selection, the weight vector W is determined by the task's preset base weights and the environmental complexity C. env Linear adjustment generation. Specifically, w is set as follows: safe =0.4 + 0.4 × C env w time =0.4 - 0.3 × C env and w energy =0.2. This design ensures that the weight allocation during the path selection process can be dynamically adjusted according to the real-time environmental threat level, so that the evaluation results not only meet the operational security requirements, but also take into account path efficiency and energy consumption optimization.

[0068] After setting the weights, based on the environmental perception and threat modeling results completed by the forward-deployed UAV in step S1, and the candidate path data generated in step S4, the performance indicators of each path in terms of security, time efficiency, and energy consumption are quantitatively calculated. Specifically, this includes the nearest threat distance d. min Cumulative obstacle avoidance path length L avoid (Safety); Total flight time ttotal (Time efficiency); Path length L path Turning frequency f θ (Energy consumption). To ensure the comparability of different indicators in the comprehensive evaluation, each indicator is normalized to form a path evaluation matrix E, where each row corresponds to a candidate path and each column corresponds to an evaluation criterion.

[0069] After normalizing the path evaluation matrix E, the final path comprehensive score S is calculated using the following formula:

[0070] S = E·W T (11); where: S represents the comprehensive score vector of each candidate path; E is the normalized path evaluation matrix, where each row of the matrix corresponds to a path and each column corresponds to an evaluation criterion (such as safety, time efficiency, energy consumption); W is the criterion weight vector calculated by the analytic hierarchy process; E T W represents the standard vector multiplication operation, used to calculate the total score of each path under the weighted criteria. A higher score S indicates that the path better meets the comprehensive requirements. The candidate path set is then analyzed based on the score vector S. Sort the paths and select the one with the highest score:

[0071]

[0072] Path P * This is the final execution path, which is pushed to the manned aircraft pilot through a collaborative decision-making interface. The pilot can confirm and fine-tune the path based on the quantitative score and the battlefield situation, forming a human-machine collaborative closed-loop decision-making mechanism of "intelligent generation of UAVs - autonomous optimization of manned aircraft". This not only ensures the scientific nature and transparency of the path planning, but also ensures the flexibility and reliability of the combat decision.

[0073] In a specific embodiment of the present invention, a "low-altitude penetration mission" scenario is used as the background, and the entire process of the present invention is verified based on a simulation platform. The collaborative system consists of one manned aircraft and four unmanned aircraft, with the formation structure as follows: one manned aircraft is positioned at the rear of the formation for command and path selection, two SAR and ESM unmanned aerial vehicles (UAVs) conduct forward reconnaissance, and two escort UAVs are positioned on the flanks of the formation to perform path tracking and threat avoidance tasks. The mission space is set to 1000nm × 1000nm (nautical miles), including both fixed and dynamic threats. The simulation scenario is based on the Python simulation framework, and trajectory and situation display are visualized using Tacview. The sensor parameters are based on the performance of a certain type of UAV platform, with an effective SAR imaging range of 50 nautical miles and an ESM detection range of 100 nautical miles. The number of dynamic obstacles in the simulation is set to 10, simulating ground air defense fire zones and moving targets, with the speed of the dynamic obstacles ranging from 200 to 1000 nautical miles per hour.

[0074] Step S1: Environmental perception and threat modeling.

[0075] In this embodiment, the forward reconnaissance UAV is equipped with an Electronic Support Measures (ESM) sensor to perform environmental threat perception tasks. The ESM sensor parameters are referenced from the performance parameters of a certain type of electronic reconnaissance pod equipped on a certain type of UAV, with a detection range of 100 nautical miles and a main beam horizontal coverage angle of 45°, used for passively detecting the location and type of enemy radar emission sources. This ESM sensor acquires the electromagnetic radiation signals of enemy radar emission sources in real time through passive reception and estimates the location p of obstacles by combining information such as azimuth angle, time difference of arrival (TDOA), and frequency characteristics. i (t) and velocity variance Var(v obs ), nearest obstacle distance d min Key parameters include: [Parameters such as] the initial distance D from the target point. goal The relative distance ΔP from the target point goal and the target point location P goal Information such as these is preset by the system during the simulation initialization phase. Two UAVs advance in formation, each responsible for reconnaissance within a 30-nautical-mile radius to the left and right of the mission area, respectively, detecting obstacles and threats in real time. The raw data sensed by the UAVs is input into the embedded high-performance processor module (NVIDIA Jetson module) for analysis and processing.

[0076] Sensed threat data is synchronously transmitted to the manned aircraft cockpit via the UAV-manned aircraft data link system and aggregated to form battlefield situational information. The manned aircraft, acting as the collaborative decision-making hub, sets path evaluation criteria weights (safety, time efficiency, energy consumption) based on threat location and type information provided by the forward-deployed UAV, combined with mission objective requirements, and issues decision commands or preferred path selection commands, thus realizing the operational process of UAV sensing and manned aircraft selecting the best path. This process ensures that the passive perception capabilities of the forward-deployed UAV and the active decision-making capabilities of the manned aircraft complement each other, constructing a complete manned-UAV collaborative path planning and threat avoidance framework.

[0077] The obstacle-induced potential field function U is defined using equation (1). obs (p,t) dynamically constructs a local threat potential field for subsequent construction of the perturbed fluid field.

[0078]

[0079] Where k = 5 is the potential field strength coefficient, and Θ(·) is the Heaviside function, activated when an obstacle enters the perception range. This potential field is implemented through an onboard computing unit, updating the local threat situation in real time, serving as input for the subsequent construction of the interference fluid field. The above data is synchronized from the UAV to the manned-aircraft collaborative command system via a tactical data link. The manned aircraft receives multi-source perception data from the forward-deployed UAV, generates a dynamic threat situation, and serves as the basis for subsequent path planning and optimal decision-making.

[0080] Step S2: Construction of the disturbance fluid field and adaptive adjustment of parameters.

[0081] Based on the obstacle potential field generated in step S1, initialize the path-guiding fluid field and set the basic parameter as ρ. base =1.0, σ base =0.5.

[0082]

[0083] Among them, Q safe (s t This represents the safety assessment of the current environmental state, reflecting the threat intensity under the current flight conditions. The calculation formula is: Among them The distance between the drone and the nearest obstacle is represented by d0, obtained from step 1; d0 is the normalized safety distance constant, which is taken as the sensor's sensing radius R. s One-fifth of that, or 20 nautical miles. Q time (s t () represents the mission time efficiency assessment, reflecting the propulsion efficiency of the aircraft relative to the mission objective. The calculation formula is: Where ΔP goal D represents the relative distance between the drone's current position and the target point. goal The initial target distance is obtained from step 1. α and β are adjustable gain coefficients. α ranges from [0.1, 0.5], with larger values ​​for high-threat missions to increase safety redundancy, and smaller values ​​for relatively stable flight paths. β ranges from [0.1, 0.3], with larger values ​​for time-sensitive missions to accelerate trajectory advancement, and vice versa, prioritizing safety. In a low-altitude penetration mission scenario, experiments verified that choosing α = 0.3 and β = 0.2 ensures both safety and path advancement efficiency, guaranteeing the successful completion of the flight mission. The negative direction of the guide field gradient... As input for flight navigation.

[0084] The DDPG algorithm is used to train a reinforcement learning agent, enabling it to learn and adjust p under different environmental perceptions. t With σ tThe strategy enables the guided field to adaptively respond to dynamic threats. For example, in a training exercise, when the drone is close to an area with dense obstacles and far from the target point, the DDPG algorithm guides the agent to learn to increase Δρ. t (Increase fluid field density to enhance repulsive force), moderately reduce Δσ t The strategy (to reduce diffusion effects and improve guidance directionality) effectively avoids high-threat areas and rapidly advances towards the target. When the drone is in a relatively safe area close to the target, the strategy guides Δρ... t Reduce (reduce repulsive force to minimize unnecessary detours), Δσ t Slightly increase (smooth the path to reduce abrupt changes in heading) to achieve a balance between path smoothness and energy consumption optimization.

[0085] DDPG training employs an offline batch update and experience replay mechanism, sampling 10 samples per training round. 5 The interaction data of next state-action-reward-next state was optimized, after approximately 3×10 5 After the next iteration, the agent's policy converges and it can output Δρ in real time in diverse task scenarios. t With $Δσ t Adjustment significantly improves the environmental adaptability and robustness of path planning.

[0086] Step S3: Flight state coding and reinforcement learning training.

[0087] Design flight state vector Used to describe target distance, threat density, and uncertainty. The reward function, as shown in equation (5), guides the agent to achieve a dynamic balance between obstacle avoidance and efficient advancement by weighting three aspects: safety, time efficiency, and path smoothness. The weights are set to λ = 0.3, μ = 0.5, and d0 = 20 meters. Path policy function π * (a|s t The network is trained using a deep deterministic policy gradient algorithm. The network input is a 4-dimensional flight state vector s. t State features are extracted through three fully connected hidden layers (with the number of neurons in each layer being [128, 64, 32] respectively), and the output layer is a 2D action vector a. t =[Δρ t ,Δσ t The training set contains 50,000 state-action-reward samples. The Adam optimizer is used for policy iteration, and the initial learning rate is set to 1×10⁻⁶. -4 The batch size was 256, and the total number of training steps was 200,000. On an NVIDIA GTX 1060 GPU platform, the total training time was approximately 24 hours, and the policy network eventually converged.

[0088] The path strategy function π obtained in this step of training * (a|s t As a core control strategy, it will be invoked in the multi-preference path generation module in step S4 to support subsequent path planning and collaborative decision-making.

[0089] Step S4: Generation of multiple preference paths.

[0090] The multi-preference path generation network adopts a "shared feature encoding + multi-branch optimization" architecture, constructing a multi-preference path generation network with three branches, corresponding to three path optimization preferences: safety priority, time efficiency priority, and energy balance. The input, the flight state vector output from step 3, is first processed uniformly through a shared feature encoding layer, i.e., sequentially passing through two fully connected (Dense) layers. The first layer contains 128 neurons, and the second layer contains 64 neurons, both using the ReLU activation function to extract high-dimensional feature representations of the flight state. The final output is a shared feature vector f. t (Dimension 64) serves as a general representation for subsequent multi-branch path optimization.

[0091] Based on shared encoding, the network is designed with three independent preference branches, which are specifically modeled for three types of path optimization objectives: "safety first," "time efficiency first," and "energy balance." Each branch has the same structure, consisting of two fully connected layers (32→16 neurons, ReLU activation) and a final Tanh-activated output layer (2 neurons), used to predict the path adjustment action 'a'. t =[Δρ t ,Δσ t This allows for differentiated adjustment of the parameters of the disturbing fluid field.

[0092] Each branch is trained using the following loss function (Equation (7)):

[0093]

[0094] Where k∈safety,time,balance, the weight parameters are set as follows:

[0095] - Safety first: ω1 = 0.5, ω2 = 0.3, ω3 = 0.2;

[0096] - Prioritize time efficiency: ω1 = 0.2, ω2 = 0.6, ω3 = 0.2;

[0097] -Energy balance: ω1=0.3, ω2=0.3, ω3=0.4.

[0098] The network outputs multiple candidate sets of paths with different preference attributes, such as... Figure 3As shown, this illustrates the impact of different path preference strategies on UAV trajectories under the same mission and threat environment: Figure 3 (a) indicates that the path generated by the security priority strategy is significantly far from the threat area, with the closest distance being 69.65 nautical miles, sacrificing path length for higher security redundancy; Figure 3 (b) indicates that the speed-first strategy selects a path closer to the edge of the threat, with a minimum distance of 43.94 nautical miles, to rapidly approach the target via the shortest route. This comparison illustrates the differences in decision-making behavior of the multi-preference path generation mechanism under different task trade-offs. The path candidate set is used for subsequent decision-making.

[0099] Step S5: Path optimization and decision support based on the analytic hierarchy process.

[0100] The Analytic Hierarchy Process (AHP) was used to analyze the path candidate set. A comprehensive evaluation was conducted to construct a three-tiered decision-making structure:

[0101] -Target layer: Select the optimal path;

[0102] - Criteria layer: security, time efficiency, energy consumption;

[0103] - Schematic layer: Paths P1, P2, P3.

[0104] In practice, the manned aircraft pilot sets a criterion weight vector W = [0.5, 0.3, 0.2] based on the real-time combat situation and mission priority. T .

[0105] Based on the environmental perception and path generation in steps S1-4, the performance metrics of the three candidate paths P1, P2, and P3 are as follows:

[0106] • Security Indicators (Minimum Threat Distance d) min Cumulative obstacle avoidance length L avoid P1 = 1.62 nm, L avoid =5nm; P2=2.70nm, L avoid =8nm; P3=1.35nm, L avoid =4nm.

[0107] • Time efficiency index (total flight time t) total ): P1=1200s, P2=1100s, P3=1300s.

[0108] • Energy consumption index (path length L) path Turning frequency f θ ): P1=80nm, P2=78nm, P3=85nm.

[0109] After normalizing the above original indicators, the path evaluation matrix E is formed:

[0110]

[0111] The normalized path evaluation matrix E is weighted and calculated to obtain the overall score S = E. T ·W:

[0112]

[0113] Finally, the path with the highest score, P2, was selected and sent to the manned aircraft pilot for path confirmation.

[0114] Figures 4 to 8 This demonstrates the entire process of a path strategy performing a trajectory planning task in a dynamic adversarial environment, showcasing the path agent's ability to adaptively avoid multiple dynamic threats during flight. Figures 4 to 8 These correspond sequentially to key nodes in the flight process, including Figure 4 Task start Figure 5 Path expansion Figure 6 Threat interaction, Figure 7 Mid-to-late stage adjustments and Figure 8 Approaching the target phase. Target points are indicated by five-pointed stars in the diagram. The circular area represents the enemy radar detection range, and the fan-shaped area represents the UAV's real-time perception range.

[0115] As the mission progresses, enemy radars are continuously activated and maneuver within the scenario, dynamically constructing a threat space landscape. Based on the current perceived situation, the computing unit continuously invokes the policy network to output path adjustment actions, optimizing flight heading and obstacle avoidance strategies in real time. Results show that the flight path effectively avoids high-threat areas while maintaining a gradual approach to the target point, demonstrating strong environmental adaptability and path safety control capabilities. It achieves a good balance between safety and efficiency in multi-threat areas, with energy consumption controlled within the expected mission range.

[0116] also, Figure 9 Presenting the complete flight trajectory from a three-dimensional perspective, the study demonstrates the spatial bypass effect of the path in a three-dimensional threat field, further verifying the rationality of the path generation and dynamic obstacle avoidance performance of the proposed method under complex spatial constraints.

[0117] To verify the adaptability and effectiveness of the method of the present invention, path optimization performance tests were conducted under different environmental complexities. Figure 10This paper presents the comprehensive score changes of different preferred path strategies under the condition of gradually increasing dynamic environmental complexity. The experiment simulated the process of environmental complexity changing from 1 to 10 by gradually increasing the number of enemy radars and increasing radar movement speed. At each complexity level, the comprehensive scores of three paths (security priority, time efficiency priority, and energy balance) were calculated based on multiple indicators. In the figure, the solid line represents the security priority path, whose score continuously increases with increasing environmental complexity, reflecting the advantage of security optimization strategies in high-threat environments; the dashed line represents the time efficiency priority path, which scores the highest in low-complexity environments, but its score decreases significantly with increasing complexity; the dotted line represents the energy balance path, whose score is relatively stable overall, maintaining a moderate level in most scenarios.

[0118] Experimental results verify that the multi-preference path generation and optimization mechanism proposed in this invention can dynamically generate diverse path schemes according to mission requirements and environmental conditions, and achieve scientific optimization through hierarchical evaluation. It has good environmental adaptability, decision-making transparency and collaborative operability, and significantly improves the path planning and execution capabilities of manned-unmanned collaborative combat missions.

Claims

1. A manned-unmanned aerial vehicle cooperative path planning and decision-making method, comprising at least one manned aerial vehicle and one unmanned aerial vehicle, wherein the unmanned aerial vehicle is deployed in front of the manned aerial vehicle as a front reconnaissance unit, responsible for situation awareness and threat modeling; the manned aerial vehicle is the main body of decision-making and task command, responsible for path optimization and flight path decision-making, and the method comprises the following steps: S1: environment perception and obstacle modeling, the unmanned aerial vehicle completes the identification of ground and air obstacles in the flight area, spatial distribution extraction and electromagnetic threat perception through a multi-modal sensor system, and obtains real-time environment perception data and obstacle threat potential field; S2: interference fluid field construction and adaptive adjustment of parameters, based on the obstacle threat potential field and task target information, an interference fluid field with global path guiding ability is constructed; S3: flight state coding and reinforcement learning training, the flight state feature vector of the unmanned aerial vehicle is designed, and the path control strategy function is trained to guide the unmanned aerial vehicle to select the optimal path action in the guiding field; S4: multi-preference path generation, based on the path control strategy function and the interference fluid field, a multi-preference path candidate set is generated; S5: path optimization and auxiliary decision-making based on the analytic hierarchy process, the multi-preference path candidate set is taken as input, the real-time environment perception data is combined, and the analytic hierarchy process is used for multi-criteria comprehensive evaluation. After the unmanned aerial vehicle completes perception, the obtained obstacle data is input into the airborne computing unit for real-time processing; the airborne computing unit is responsible for obstacle data analysis and local potential field modeling; During modeling, all obstacle information perceived by the unmanned aerial vehicle is synchronously pushed to the manned aerial vehicle to realize real-time sharing of situation data; the manned aerial vehicle takes the local threat model as input basis for subsequent path optimization and decision-making. S2.1: initial parameter setting of interference fluid field, the initial density and diffusion rate parameters of the fluid field are set according to the task range and obstacle density; S2.2: adaptive adjustment of parameters by reinforcement learning, the density and diffusion rate of the fluid field are adjusted in real time by a reinforcement learning agent based on the DDPG algorithm; S2.3: generation of interference fluid field, the interference fluid field is constructed based on the negative direction of the obstacle threat potential field gradient to realize a spatially continuous and directionally controllable guiding force field.

2. The manned-unmanned teaming path planning and decision-making method of claim 1, wherein: The unmanned aerial vehicle in step S1 acquires, through a multi-modal sensor system carried by the unmanned aerial vehicle, including an image imaging unit and an electromagnetic detection unit, real-time data including position and speed variance of the obstacle , distance to the nearest obstacle , distance to the target point at the beginning of the task , relative distance to the target point , and target point position S3.1: flight state and reward function design, the flight state vector is designed, including the distance from the current position to the target point, the distance to the nearest obstacle, the obstacle speed expectation and variance; S3.2: training of strategy function target, the path strategy function is defined, and the optimization target is to maximize the expected reward obtained in each state.

3. The manned-unmanned teaming path planning and decision-making method of claim 2, wherein: In step S1, core feature parameters are extracted from the perceived obstacle information, and then based on the analysis result, an obstacle threat potential field is generated using the artificial potential field theory, a repulsive force field is generated with the obstacle as the source point, and the obstacle threat potential field Modeling as follows: ; wherein is the total intensity of the obstacle threat potential field perceived by the UAV at position and time ; represents the number of obstacles detected within the current perception range; is the set sensor perception radius; is the current position of the UAV, is the position of the th obstacle, is the effective radius of the obstacle, is the potential field intensity coefficient; is the Heaviside function, which is used to determine whether the obstacle is within the perception range: when , it is 1, otherwise it is 0.

4. The manned-unmanned teaming path planning and decision-making method of claim 1, wherein: 8.The manned-unmanned aerial vehicle cooperative path planning and decision-making method according to claim 7, wherein the specific optimization target in S3.2 is: ​ ​ ​ 5. The manned-unmanned teaming path planning and decision-making method of claim 4, wherein: Initial density in step S2.1 Distance from mission objective They are inversely proportional, and the expression is: ,in This is the density adjustment factor, and its value range is... Initial diffusion rate Based on the density of obstacles within the sensing range Configure it; the expression is: ,in , Represents the radius of perception The number of obstacles per unit area within the range. This is the diffusivity adjustment coefficient, with a value range of [value range missing]. ; The guiding field density is adjusted in real time by the agent in step S2.2 With the diffusion rate The guiding field parameter adjustment mechanism is designed as follows: ; In the formula: This indicates the current density value of the disturbing fluid field; This indicates the current diffusion rate of the interfering fluid field; and These are the initial density and diffusivity set during system initialization, respectively. The safety assessment representing the current environmental state is calculated using the following formula: , among them Indicates the distance between the drone and the nearest obstacle. The normalized safety distance constant is set to the sensor's sensing radius. One-fifth; The formula for evaluating task time efficiency is as follows: , among them The distance between the drone's current position and the target point. This is the target distance at the start of the task. and The gain coefficient is adjustable. The range of values ​​is , The range of values ​​is , This is the hyperbolic tangent function, used for smoothing adjustments; This is a logistic compression function used to limit the range of parameter variation; a policy gradient-based reinforcement learning algorithm is employed, based on the location of obstacles. and velocity variance Distance to nearest obstacle Distance from the target point at the start of the task Relative distance from the target point and target point location The environmental state is used as input for fluid field parameter regulation. The DDPG algorithm is used for agent training to predict and output fluid field regulation actions. ,in To guide the adjustment increment of the field density, To guide the adjustment increment of the field diffusion rate, ; The disturbance fluid field in step S2.3 The gradient negative direction of the obstacle threat potential field i.e. ; is the current density value of the disturbance fluid field, is the current diffusion rate of the disturbance fluid field.

6. The manned-unmanned teaming path planning and decision-making method of claim 5, wherein: In step S2.2 , .

7. The manned-unmanned teaming path planning and decision-making method of claim 1, wherein: ​ ​ ​ ​ The flight state vector is defined in step S3.1 : where: is the flight state vector at the current time step; is the relative distance to the goal point, is the distance to the nearest obstacle, and are the mean and variance of the obstacle velocity, respectively; Corresponding reward function Designed to: ; In the formula: is an instantaneous reward value obtained by the UAV at the current time step; is the minimum distance from the current and the nearest obstacle, is the cumulative path length, is the flight direction change angle, , and is a weight coefficient; Output action is for dynamically adjusting an interfering fluid field, in particular wherein is an adjustment increment for the guiding field density; is an adjustment increment for the guiding field diffusivity; ​ ; wherein represents the optimal policy, is the immediate reward obtained for taking action in state ​ By path policy function Optimization iterations, training to obtain optimal path policy function .

9. The manned-unmanned teaming path planning and decision-making method of claim 1, wherein: The branch-wise optimization objectives in step S4 are defined by loss functions are defined as follows: ; In the formula: is the loss function value corresponding to the mth path optimization branch; is a preset minimum safety distance threshold; represents the total flight time of the path; is the path energy consumption estimation value; represents the path smoothness index, i.e., the total steering angle; represents different optimization preferences, , , is a weight coefficient, is a set safety distance threshold;​ Multi-preference path generation network with unified flight state vector and a disturbance fluid field For input, output several path action sequences through branch structure , generate corresponding paths , wherein Indicates the path preference branch number, the feature extraction layer of the network adopts a 3-layer fully connected neural network, and the activation function is ReLU, which is used to extract a high-dimensional feature representation from the joint input, and the output layer of the network is a set composed of Strategies with different preferences: ; Each branch shares the input features but adopts different target weights for policy training, generating paths with specific optimization tendencies, and each preferred policy network independently generates a complete path point sequence in the interference fluid field to obtain a candidate path set: ; Each path has a clear preference attribute and performance indicator.

10. The manned-unmanned teaming path planning and decision-making method of claim 9, wherein: The environmental complexity is calculated automatically in step S5 The environmental complexity is calculated automatically in step S5 The environmental information perceived by the UAV in step S1 is calculated in real time, including the distance to the nearest obstacle , the variance of the perceived obstacle speed , the relative distance between the current UAV and the mission target , and the initial target distance of the mission , whose calculation formula is: ; wherein is a perception radius, specifically 100 nautical miles; is a maximum speed of the obstacle, specifically Mach 1, , , , environmental complexity ; weight vector by task preset base weight and environment complexity linear adjustment generation; specifically set as: , and ; path evaluation matrix final path synthesis score after normalization is calculated by the following equation: ; wherein: represents the comprehensive scoring vector of each candidate path; is the normalized path evaluation matrix, each row of the matrix corresponds to a path, and each column corresponds to an evaluation criterion; is the criterion weight vector calculated by the analytic hierarchy process method; is the standard vector multiplication operation, which is used to calculate the total score of each path under the weighted criterion; according to the scoring vector the candidate path set is sorted, and the path with the highest score is selected: ; path i.e. the final execution path.

Citation Information

Patent Citations

  • A drone obstacle avoidance method based on artificial potential field

    CN112180954B

  • A method for obstacle avoidance and path planning of unmanned aerial vehicles

    CN113110592B

  • Unexpected threat-oriented unmanned bee colony cooperative route planning method

    CN116448119A

  • Unmanned aerial vehicle adaptive path planning system and method based on information fusion and cooperation

    CN120489128A