Decision method for patrol path of unmanned aerial vehicle in complex environment

By constructing an undirected topology map and dynamic environment evaluation combined with Actor-Critic method, safe and effective drone motion commands are generated, which solves the self-loop characteristics and dynamic obstacle handling problems of drone path planning in complex environments, and realizes efficient and intelligent decision-making of drones in complex environments.

CN120521607AActive Publication Date: 2025-08-22SUN YAT SEN UNIV

Patent Information

Application Number
CN202510876863.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-08-22
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing drone path planning technology has problems such as self-ring characteristics that are easily attacked, difficult to handle dynamic obstacles, and difficult to ensure safety in dynamic environments.

Method used

The steady-state distribution is constructed using undirected topology graphs, multiple candidate state transfer matrices are generated, and the decision network weight is adjusted by combining dynamic environment evaluation and Actor-Critic methods to generate safe and effective motion commands.

Benefits of technology

Realize efficient and intelligent decision-making for drones in complex environments to ensure safety and adaptability, especially when facing flying birds, other drones or unpredictable obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120521607A_ABST
    Figure CN120521607A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle patrol path decision-making method in a complex environment, and relates to the technical field of unmanned aerial vehicle path planning, and the method comprises the steps: taking each to-be-patrolled position as a node to construct an undirected topological graph; generating steady-state distribution of each node patrolled by the unmanned aerial vehicle according to topological constraints and node importance of the undirected topological graph; generating a plurality of candidate state transition matrixes with steady state distribution; performing dynamic environment evaluation on obstacles of paths corresponding to the candidate state transition matrixes so as to determine selectable speeds and directions of the unmanned aerial vehicle on the selected paths; dynamically adjusting the weight of the decision-making network according to the dynamic environment evaluation result, and generating a motion execution command by using the decision-making network after weight adjustment; and controlling the unmanned aerial vehicle to move among the to-be-patrolled positions according to the motion command. According to the invention, dynamic environment evaluation, path planning, decision network weight adjustment and motion command generation are carried out on the obstacle, so that efficient intelligent decision making of the unmanned aerial vehicle in a complex environment can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of drone path planning, and in particular to a method for determining a drone patrol path in a complex environment. Background Art

[0002] Existing technologies use the Metropolis-Hastings algorithm to generate the state transition matrix of a Markov chain. Global path planning is generated using the A* algorithm, while local path planning is achieved using the traditional dynamic windowing approach (DWA). This generates motor control commands for translational and rotational speeds, ensuring a collision-free trajectory in a static environment. A neural network-based navigation strategy uses a combination of RGB data, depth images, and laser data as input, and uses reinforcement learning (RL) to train the network, directly outputting control commands for the drone based on current observations.

[0003] First, existing technologies are fragmented, with no connections between independent systems. Second, the state transition matrices of the Markov chains generated by current technology are highly self-looping, making them easy for potential attackers (such as attack drones) to learn and predict patrol paths. Third, the performance of traditional dynamic window methods (DWAs) relies heavily on manually adjusted weight parameters, and different environments require different weight configurations, making them unable to effectively handle dynamic obstacles and changing environments. Finally, deep learning methods make predictions based on current observations, but when faced with unforeseen inputs, they cannot guarantee the safety of the generated commands (which may lead to collisions). Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a method for making patrol path decisions for drones in complex environments, so as to enable efficient and intelligent decision-making of drones in complex environments.

[0005] To achieve the above objectives, an embodiment of the present application provides a method for determining a patrol path of a UAV in a complex environment, the method comprising the following steps:

[0006] Construct an undirected topological graph with each patrol location as a node;

[0007] Generate a steady-state distribution of each of the nodes patrolled by drones according to the topological constraints and node importance of the undirected topological graph;

[0008] Generating a plurality of candidate state transfer matrices having the steady-state distribution; wherein each of the candidate state transfer matrices includes a transition probability of the UAV between each of the nodes;

[0009] Performing a dynamic environmental assessment of obstacles on the paths corresponding to the candidate state transfer matrices, thereby determining an optional speed and direction of the UAV on the selected path;

[0010] Dynamically adjust the weights of the decision network according to the dynamic environment assessment results, and then use the weight-adjusted decision network to generate execution movement commands;

[0011] The UAV is controlled to move between each of the positions to be patrolled according to the execution motion command.

[0012] In some embodiments, generating a steady-state distribution of each node patrolled by a drone based on the topological constraints and node importance of the undirected topological graph comprises the following steps:

[0013] The steady-state distribution is generated according to the topological constraints of the undirected topological graph to increase the probability of the drone transferring to the target node; wherein, high-value target nodes and nodes with short attack time need to obtain more patrol resources, and the patrol drone has a higher probability of visiting these nodes;

[0014] The expression of the steady-state distribution is:

[0015] Among them, v i ∈V,v i represents the i-th node, V represents the set of each node, w i Represents the target node v i The value of a i represents the attack time, μ i Representative node v i The importance weight parameter, μ j Representative node v j The importance weight parameter, n is the total number of nodes, Represents the target node v i The target steady-state distribution represents the probability distribution of the patroller at each target node after the Markov chain runs for a long time.

[0016] In some embodiments, generating a plurality of candidate state transfer matrices having the steady-state distribution comprises the following steps:

[0017] Generating a plurality of candidate state transfer matrices having the steady-state distribution and the target network topology structure according to a multivariate random adaptive Markov steady-state matrix algorithm includes the following steps:

[0018] Phase 1: Stratified random matrix generation;

[0019] Generate a first-stage random matrix that satisfies the topological constraints of the undirected topological graph, wherein the self-loop probability p is assigned using four distribution intervals. ii , and then generate the remaining non-self-loop probability p ij ;

[0020] The second stage: relative error iterative adjustment algorithm;

[0021] Adjusting the first-stage random matrix according to the error between the probability distribution of the first-stage random matrix and the steady-state distribution, specifically obtaining the second-stage random matrix through relative error calculation, error ratio adjustment, row normalization and random periodic perturbation;

[0022] The third stage: similarity elimination algorithm;

[0023] Eliminate matrices whose similarity values ​​are greater than a preset similarity threshold value in the random matrix of the second stage. Specifically, similarity value detection, random adjustment, and row normalization are performed to ensure that there are no values ​​that are too close in the matrix, thereby enhancing randomness and creating a more natural random distribution to obtain the random matrix of the third stage;

[0024] Phase 4: Matrix diversity guarantee algorithm;

[0025] The random matrix of the third stage is adjusted according to the average absolute error between the probability distributions of the multiple random matrices in the third stage, specifically by matrix difference calculation, diversity determination, large-scale random perturbation and readjustment of steady-state distribution to ensure that the generated multiple Markov state transfer matrices P k With sufficient diversity rather than variations of similar matrices, a random matrix in the fourth stage is finally obtained as the candidate state transfer matrix.

[0026] In some embodiments, the step of performing a dynamic environmental assessment on obstacles along the paths corresponding to the candidate state transfer matrices comprises the following steps:

[0027] Performing a dynamic environmental assessment on obstacles in the paths corresponding to the candidate state transfer matrices according to a dynamic environmental adaptability function, summing the risk weights of each obstacle, and quantifying the environmental adaptability;

[0028] The expression of the dynamic environmental adaptability function (DEAF) is: Where, e(ν,ω): dynamic environmental adaptability function, which represents the environmental adaptability evaluation value of the UAV under a given speed command;

[0029] ν: The linear velocity command of the UAV, which controls the speed of the UAV forward or backward;

[0030] ω: angular velocity command of the UAV, which controls the speed and direction of the UAV's steering;

[0031] The current system state estimated using weighted observation data, including the environment representation of obstacle positions and velocities, where z iis the observation data of the UAV at the i-th moment, and the observation weight Calculated by the following expression:

[0032]

[0033] Observation data z i The weight of quantifies the importance of observation in state estimation;

[0034] w a , w r , w d It is the memory weight coefficient for autocorrelation, time proximity and deviation, starting from the initial empirical value, set according to the nature of the environment, and using an online adaptive method to regularly update the memory weight;

[0035] A s (z i ,Θ): the autocorrelation score of the observation, reflecting the temporal correlation of the environmental dynamics;

[0036] Θ: a set of parameters that controls the calculation of the autocorrelation score;

[0037] R s (z i ): the temporal proximity score of the observation, emphasizing the importance of recent data;

[0038] D s (z i ,ψ): Deviation score of the observation, capturing abnormal changes in the environment;

[0039] ψ: bias sensitivity parameter, which controls the degree of emphasis the algorithm places on bias;

[0040] M: the number of dynamic obstacles currently perceived, including other drones;

[0041] The probability of obstacle i existing based on the state estimate Ω is in the range [0,1], where a larger value indicates a higher collision risk. It measures the probability that obstacle i intersects the drone's path when the drone executes the speed command (ν,ω);

[0042] The relative velocity vector between the UAV and obstacle i is calculated based on the UAV velocity command (ν, ω) and the obstacle velocity information in the state estimate Ω, which measures the speed and direction of the relative motion between the UAV and the obstacle;

[0043] d i (ν,ω,Ω): The predicted minimum distance between the UAV and obstacle i when the UAV adopts the speed (ν,ω). The smaller the value, the higher the potential collision risk.

[0044] K: normalization constant, ensuring the appropriate range of values ​​for the fractional term, determined experimentally based on environmental characteristics and UAV dynamic parameters;

[0045] ε: A positive number used as a safety parameter to prevent division by zero.

[0046] In some embodiments, dynamically adjusting the weights of the decision network according to the dynamic environment assessment results includes the following steps:

[0047] Dynamically adjust the weight of the decision network according to the system state of the path corresponding to the target state transfer matrix, including the following steps:

[0048] Design and adjust the system state, output action, and reward function required by the decision-making network Actor-Critic;

[0049] By adding the dynamic environment adaptability function to the traditional DWA, a time-varying environment adaptive dynamic window method is obtained;

[0050] The Actor-Critic decision network and the time-varying environment adaptive dynamic window method are combined to obtain the AC-time-varying environment adaptive dynamic window method, which uses the state value estimate V(s) output by the Critic network in the decision network and adjusts the weight of the Actor network in the decision network through the advantage function, i.e., the weight parameter in the time-varying environment adaptive dynamic window method;

[0051] The Actor-Critic method replaces the recurrent neural network in the traditional Actor-Critic method with the MoEMamba network. MoEMamba is a network that integrates and mixes expert networks and Mamba networks. It alternates between the two networks: one layer of mixed expert network and one layer of Mamba network, and so on, with a total of k layers.

[0052] System Status t It describes all the information contained in the state of the UAV at time t, which is the basis of the AC-TEADWA method and provides the complete information required for the neural network to predict the AC-TEADWA weight. The expression of the system state is:

[0053] Among them, s t : The system status at time t, which contains the complete status information of the UAV at time t;

[0054] The distance measurement value from the drone to the obstacle at the current time t in the current drone coordinate system;

[0055] The distance measurement k control cycles ago expressed in the current UAV coordinate system;

[0056] The distance measurement value l control cycles ago expressed in the current UAV coordinate system;

[0057] The distance measurement g control cycles ago expressed in the current UAV coordinate system;

[0058] d t : The distance from the current position of the UAV to the target;

[0059] φ t : The angle from the current position of the drone to the target;

[0060] ν t : The current translation speed of the drone;

[0061] ω t : The current rotation speed of the drone;

[0062] The obstacle distance changes between consecutive time steps to capture dynamic information;

[0063] The time derivative of the distance from the drone to the target, indicating whether it is approaching the target;

[0064] The current acceleration state vector of the drone is the linear acceleration of the drone at the current time t, is the angular acceleration of the UAV at the current time t;

[0065] ξ t : Environmental complexity index, calculated by the entropy value of obstacle distribution;

[0066] The specific form of the distance measurement value is:

[0067] τ represents the past time step and refers to k, l, and g in the formula;

[0068] The superscript t-τ indicates the measurement time;

[0069] The subscript t represents the coordinate system of the current UAV position;

[0070] N represents the number of measurement points of the lidar, which is the distance between the drone's position and the nearest obstacle in the set direction;

[0071] The system state is input into the decision network to obtain the action output by the decision network. The action is expressed as:

[0072] a t =(α t ,β t ,γ t ,δ t ,λ t );

[0073] Among them, α t : heading item weight;

[0074] β t : gap term weight;

[0075] γ t : speed term weight;

[0076] δ t : curvature distance component, additional weight of the new extended cost function;

[0077] λ t : Dynamic environment adaptability function component, additional weight of the new extended cost function;

[0078] The drone receives observations or states of the environment and outputs an action a. The action is to set the cost function weights for the next control cycle. These weights will be used to generate motion commands for the drone.

[0079] Design of reward function:

[0080] If the drone successfully reaches the target location, it will be given a large positive reward R1; if the drone stays at the same location for multiple control cycles, it will be similar to being stuck and will be given a negative reward R2; if the drone collides with an obstacle / or other drones / birds during movement, it will be given a large penalty, negative R3; if the drone collides with a moving obstacle or other drones while stationary, it will be given a negative reward R4; the design of this reward function guides the reinforcement learning algorithm to learn safe and effective navigation strategies.

[0081] In some embodiments, the steps of implementing the time-varying environment adaptive dynamic window method include the following steps:

[0082] The dynamic environment adaptability function is added to the traditional DWA to obtain a time-varying environment adaptive dynamic window method;

[0083] The objective function of the time-varying environment adaptive dynamic window method is defined as:

[0084] G(ν,ω)=α·h(ν,ω)+β·c(ν,ω)+γ·v(ν,ω)+δ·d(ν,ω)+λ·e(ν,ω);

[0085] Among them, ν: the linear velocity command of the UAV, which controls the speed of the UAV forward or backward;

[0086] ω: angular velocity command of the UAV, which controls the speed and direction of the UAV's steering;

[0087] h(ν,ω): heading component, which represents the angular difference between the drone’s heading and the target, and measures the degree to which the drone is heading towards the target;

[0088] c(ν,ω): clearance component, which represents the distance from the drone to the nearest obstacle on the curvature currently considered, and measures the safety of the motion;

[0089] v(ν,ω): velocity component, representing the translation speed of the drone and measuring the efficiency of the movement;

[0090] d(ν,ω): curvature distance component, which represents the shortest distance from the target point to the curvature currently considered, and measures the relevance of the path to the target;

[0091] e(ν,ω): Dynamic environment adaptability function component, which indicates the degree of adaptability of the current speed command (ν,ω) to the predicted dynamic changes in the environment. It measures whether the UAV can effectively cope with the expected movement of dynamic obstacles in the environment when executing the current speed command;

[0092] α, β, γ, δ, λ: weights corresponding to each component;

[0093] The curvature distance component d(ν,ω) represents the shortest distance from the target point to the curvature currently considered, and measures the relevance of the path to the target. The expression is as follows:

[0094]

[0095] β: proximity gain coefficient (0-1), which increases the reward when the target is very close to the curvature;

[0096] λ: proximity sensitivity, which controls the steepness of the approach reward;

[0097] δ: Directional penalty factor, considering the deviation between the curvature line direction and the target direction;

[0098] θ curv : tangent direction of the curvature line;

[0099] θ goal : The direction of the target point relative to the drone;

[0100] robot_radius: The radius of the drone, used to normalize the distance value;

[0101] distance_to_goal: the shortest distance between the motion curvature generated by the current velocity combination (ν, ω) and the target point, which is the truncation parameter;

[0102] max_dist: Maximum distance threshold, used for normalization and limiting the range of distance influence.

[0103] In some embodiments, the steps of implementing the AC-time-varying environment adaptive dynamic window method include the following steps:

[0104] Obtain historical observation data and preprocess it, convert the historical observations to the current drone coordinate system, align historical observations at different times to the same reference coordinate system of the current drone, and define the coordinate system conversion formula as follows:

[0105]

[0106] Where, t: current time;

[0107] τ: universal time offset parameter, including k, l or g, indicating the time difference;

[0108] The converted distance measurement value represents the distance measurement of the historical observation before time τ in the coordinate system of the current time t;

[0109] At time t-τ, the distance measurement value referenced to the drone coordinate system at time t-τ;

[0110] The coordinate transformation matrix from time t-τ to time t represents the relative motion speed;

[0111] Θ t : A set of state estimates of dynamic objects detected in the environment;

[0112] η: mixing coefficient, dynamically adjusted with time and observation quality Control the weight of the two conversion methods;

[0113] The time derivative of the coordinate transformation represents the relative motion speed.

[0114] ψ(): evaluation function used to calculate the mixing coefficient η;

[0115] σ(): Swish activation function;

[0116] cart(): A function that converts distance scan data in polar coordinate form into Cartesian coordinate points;

[0117] F(): forward projection function based on velocity prediction, considering the trajectory of moving obstacle candidates including other aircraft;

[0118] pol(): A function that converts points back into the distance scan form in polar coordinates;

[0119] The neural network structure in the AC-time-varying environment adaptive dynamic window method is defined as follows:

[0120] 1D four-channel input;

[0121] The feature extraction layer processes the spatially correlated scan data;

[0122] The convolutional layer uses the Swish activation function;

[0123] Feature fusion flattens the convolution output into a vector, concatenating the target distance, angle, and current speed;

[0124] In the Actor-Critic architecture, the Actor and Critic networks share the same infrastructure, except for the last layer.

[0125] Define the training and output decisions in the AC-time-varying environment adaptive dynamic window method:

[0126] Three-stage training with different levels of difficulty;

[0127] The Actor network inputs system status information, including historical observations, target position, and current speed, and outputs actions, including the weight values ​​of the five TEADWA components: heading component weight α, gap component weight β, velocity component weight γ, curvature distance component weight δ, and dynamic environment adaptability function component weight λ;

[0128] The critic network outputs a state value estimate V(s), which guides the improvement of the actor through the advantage function.

[0129] To achieve the above objectives, another aspect of the present invention provides a device for determining a patrol path of a UAV in a complex environment, the device comprising:

[0130] A graph construction unit, used to construct an undirected topological graph using each location to be patrolled as a node;

[0131] A steady-state distribution generating unit, configured to generate a steady-state distribution of each of the nodes patrolled by the drone according to the topological constraints and node importance of the undirected topological graph;

[0132] A transfer matrix generating unit, configured to generate a plurality of candidate state transfer matrices having the steady-state distribution; wherein each of the candidate state transfer matrices includes a transfer probability of the UAV between each of the nodes;

[0133] An environmental assessment unit, configured to perform a dynamic environmental assessment of obstacles along the path corresponding to each candidate state transfer matrix, thereby determining an optional speed and direction of the UAV along the selected path;

[0134] A command generation unit is used to dynamically adjust the weights of the decision network according to the dynamic environment assessment results, and then use the weight-adjusted decision network to generate execution movement commands;

[0135] A mobile control unit is used to control the UAV to move between the various positions to be patrolled according to the execution motion command.

[0136] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.

[0137] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above-mentioned method when executed by a processor.

[0138] The embodiments of the present application include at least the following beneficial effects:

[0139] This application constructs an undirected topological graph using each patrol location as a node; generates a steady-state distribution of each of the nodes patrolled by a drone based on the topological constraints and node importance of the undirected topological graph; performs a dynamic environmental assessment of obstacles along the path corresponding to each candidate state transition matrix. This function comprehensively evaluates the safety and effectiveness of the drone's speed command in a dynamic environment by coupling autocorrelation score, temporal proximity score, and deviation score functions. By weighting the risk of each obstacle, the environmental adaptability is quantified, thereby determining the optional speed and direction of the drone on the selected path; dynamically adjusts the weights of the decision network based on the dynamic environmental assessment results, and then uses the weighted decision network to generate execution motion commands; controls the drone to move between each of the patrol locations based on the motion commands, enabling the drone system to dynamically adjust the behavior of the time-varying environment adaptive dynamic window method (TEADWA) according to the environmental state, achieving both adaptability and safety assurance in a time-varying partially observable environment, especially when the system faces birds, other drones / aircraft, or unpredictable moving obstacles, making safer and more efficient decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0140] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0141] Figure 1A flowchart of a method for determining a patrol path of a UAV in a complex environment provided by an embodiment of the present application;

[0142] Figure 2 This is an example diagram of the Markov chain topology provided in the embodiments of the present application;

[0143] Figure 3 Flowchart of the MEDRO algorithm provided in the embodiment of this application;

[0144] Figure 4 A schematic diagram of the structure of a UAV patrol path decision-making device in a complex environment provided by an embodiment of the present application;

[0145] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0146] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0147] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0148] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.

[0149] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0150] Before describing the embodiments of the present application in detail, some of the related technologies involved in the embodiments of the present application are first described as follows:

[0151] DWA: Dynamic Window Approach, a dynamic window method.

[0152] Actor-Critic: A reinforcement learning approach,Actor-Critic.

[0153] DRAMSM: Diverse Random Adaptive Markov Stationary Matrix, multivariate random adaptive Markov stationary matrix.

[0154] DEAF: Dynamic Environment Adaptability Function, dynamic environment adaptability function.

[0155] TEADWA: Time-varying Environment Adaptive Dynamic Window Approach, time-varying environment adaptive dynamic window method.

[0156] MoEMamba: Mixture of Experts + Mamba, a network customized by the inventor of this application (mixture of experts + Mamba) to replace the recurrent neural network in Actor-Critic.

[0157] AC-TEADWA: Actor-Critic Time-varying Environment Adaptive DynamicWindow Approach, AC-time-varying environment adaptive dynamic window method.

[0158] Reference Figure 1 The embodiment of the present application provides a method for determining a patrol path of a UAV in a complex environment. The method may include but is not limited to steps S100 to S130, as follows:

[0159] S100: constructing an undirected topological graph with each patrol location as a node;

[0160] S110: generating a steady-state distribution of each of the nodes patrolled by the drone according to the topological constraints and node importance of the undirected topological graph;

[0161] S120: Generate multiple candidate state transfer matrices having the steady-state distribution; wherein each candidate state transfer matrix includes a transition probability of the UAV between each of the nodes;

[0162] S130: Performing a dynamic environmental assessment on obstacles along the paths corresponding to the candidate state transfer matrices, thereby determining an optional speed and direction for the UAV along the selected path;

[0163] S140: Dynamically adjusting the weight of the decision network according to the dynamic environment assessment result, and then using the weight-adjusted decision network to generate an execution motion command;

[0164] S150: Control the UAV to move between the positions to be patrolled according to the execution motion command.

[0165] Optionally, generating a steady-state distribution of each of the nodes patrolled by a drone according to the topological constraints of the undirected topological graph comprises the following steps:

[0166] The steady-state distribution is generated according to the topological constraints of the undirected topological graph to increase the probability of the drone transferring to the target node; wherein, high-value target nodes and nodes with short attack time need to obtain more patrol resources, and the patrol drone has a higher probability of visiting these nodes;

[0167] The expression of the steady-state distribution is:

[0168]

[0169] Among them, v i ∈V,v i represents the i-th node, V represents the set of each node, w i Represents the target node v i The value of a i represents the attack time, μ i Representative node v i The importance weight parameter, μ j Representative node v j The importance weight parameter, n is the total number of nodes, Represents the target node v i The target steady-state distribution represents the probability distribution of the patroller at each target node after the Markov chain runs for a long time.

[0170] Optionally, generating a plurality of candidate state transfer matrices having the steady-state distribution comprises the following steps:

[0171] Generating a plurality of candidate state transfer matrices having the steady-state distribution and the target network topology structure according to a multivariate random adaptive Markov steady-state matrix algorithm includes the following steps:

[0172] Phase 1: Stratified random matrix generation;

[0173] Generate a first-stage random matrix that satisfies the topological constraints of the undirected topological graph, wherein the self-loop probability p is assigned using four distribution intervals. ii , and then generate the remaining non-self-loop probability p ij ;

[0174] The second stage: relative error iterative adjustment algorithm;

[0175] Adjusting the first-stage random matrix according to the error between the probability distribution of the first-stage random matrix and the steady-state distribution, specifically obtaining the second-stage random matrix through relative error calculation, error ratio adjustment, row normalization and random periodic perturbation;

[0176] The third stage: similarity elimination algorithm;

[0177] Eliminate matrices whose similarity values ​​are greater than a preset similarity threshold value in the random matrix of the second stage. Specifically, similarity value detection, random adjustment, and row normalization are performed to ensure that there are no values ​​that are too close in the matrix, thereby enhancing randomness and creating a more natural random distribution to obtain the random matrix of the third stage;

[0178] Phase 4: Matrix diversity guarantee algorithm;

[0179] The random matrix of the third stage is adjusted according to the average absolute error between the probability distributions of the multiple random matrices in the third stage, specifically by matrix difference calculation, diversity determination, large-scale random perturbation and readjustment of steady-state distribution to ensure that the generated multiple Markov state transfer matrices P k With sufficient diversity rather than variations of similar matrices, a random matrix in the fourth stage is finally obtained as the candidate state transfer matrix.

[0180] Optionally, the performing dynamic environmental assessment on obstacles in the paths corresponding to the candidate state transfer matrices comprises the following steps:

[0181] Performing a dynamic environmental assessment on obstacles in the paths corresponding to the candidate state transfer matrices according to a dynamic environmental adaptability function, summing the risk weights of each obstacle, and quantifying the environmental adaptability;

[0182] The expression of the dynamic environmental adaptability function (DEAF) is:

[0183]

[0184] Where, e(ν,ω): dynamic environmental adaptability function, which represents the environmental adaptability evaluation value of the UAV under a given speed command;

[0185] ν: The linear velocity command of the UAV, which controls the speed of the UAV forward or backward;

[0186] ω: angular velocity command of the UAV, which controls the speed and direction of the UAV's steering;

[0187] The current system state estimated using weighted observation data, including the environment representation of obstacle positions and velocities, where z i is the observation data of the UAV at the i-th moment, and the observation weight Calculated by the following expression:

[0188]

[0189] Observation data z i The weight of quantifies the importance of observation in state estimation;

[0190] w a , w r , w d It is the memory weight coefficient for autocorrelation, time proximity and deviation, starting from the initial empirical value, set according to the nature of the environment, and using an online adaptive method to regularly update the memory weight;

[0191] A s (z i ,Θ): the autocorrelation score of the observation, reflecting the temporal correlation of the environmental dynamics;

[0192] Θ: a set of parameters that controls the calculation of the autocorrelation score;

[0193] R s (z i ): the temporal proximity score of the observation, emphasizing the importance of recent data;

[0194] D s (z i ,ψ): Deviation score of the observation, capturing abnormal changes in the environment;

[0195] ψ: bias sensitivity parameter, which controls the degree of emphasis the algorithm places on bias;

[0196] M: the number of dynamic obstacles currently perceived, including other drones;

[0197] The probability of obstacle i existing based on the state estimate Ω is in the range [0,1], where a larger value indicates a higher collision risk. It measures the probability that obstacle i intersects the drone's path when the drone executes the speed command (ν,ω);

[0198] The relative velocity vector between the UAV and obstacle i is calculated based on the UAV velocity command (ν, ω) and the obstacle velocity information in the state estimate Ω, which measures the speed and direction of the relative motion between the UAV and the obstacle;

[0199] d i (ν,ω,Ω): The predicted minimum distance between the UAV and obstacle i when the UAV adopts the speed (ν,ω). The smaller the value, the higher the potential collision risk.

[0200] K: normalization constant, ensuring the appropriate range of the fractional term, determined based on environmental characteristics and UAV dynamics parameter experiments;

[0201] ε: A positive number used as a safety parameter to prevent division by zero.

[0202] Optionally, dynamically adjusting the weight of the decision network according to the dynamic environment assessment result includes the following steps:

[0203] Dynamically adjust the weight of the decision network according to the system state of the path corresponding to the target state transfer matrix, including the following steps:

[0204] Design and adjust the system state, output action, and reward function required by the decision-making network Actor-Critic;

[0205] By adding the dynamic environment adaptability function to the traditional DWA, a time-varying environment adaptive dynamic window method is obtained;

[0206] The Actor-Critic decision network and the time-varying environment adaptive dynamic window method are combined to obtain the AC-time-varying environment adaptive dynamic window method, which uses the state value estimate V(s) output by the Critic network in the decision network and adjusts the weight of the Actor network in the decision network through the advantage function, i.e., the weight parameter in the time-varying environment adaptive dynamic window method;

[0207] The Actor-Critic method replaces the recurrent neural network in the traditional Actor-Critic method with the MoEMamba network. MoEMamba is a network that integrates and mixes expert networks and Mamba networks. It alternates between the two networks: one layer of mixed expert network and one layer of Mamba network, and so on, with a total of k layers.

[0208] System Status t It describes all the information contained in the state of the UAV at time t, which is the basis of the AC-TEADWA method and provides the complete information required for the neural network to predict the AC-TEADWA weight. The expression of the system state is:

[0209]

[0210] Among them, s t : The system status at time t, which contains the complete status information of the UAV at time t;

[0211] The distance measurement value from the drone to the obstacle at the current time t in the current drone coordinate system;

[0212] The distance measurement k control cycles ago expressed in the current UAV coordinate system;

[0213] The distance measurement value l control cycles ago expressed in the current UAV coordinate system;

[0214] The distance measurement g control cycles ago expressed in the current UAV coordinate system;

[0215] d t : The distance from the current position of the UAV to the target;

[0216] φ t : The angle from the current position of the drone to the target;

[0217] ν t : The current translation speed of the drone;

[0218] ω t : The current rotation speed of the drone;

[0219] The obstacle distance changes between consecutive time steps to capture dynamic information;

[0220] The time derivative of the distance from the drone to the target, indicating whether it is approaching the target;

[0221] The current acceleration state vector of the drone is the linear acceleration of the drone at the current time t, is the angular acceleration of the UAV at the current time t;

[0222] ξ t : Environmental complexity index, calculated by the entropy value of obstacle distribution;

[0223] The specific form of the distance measurement value is:

[0224] τ represents the past time step and refers to k, l, and g in the formula;

[0225] The superscript t-τ indicates the measurement time;

[0226] The subscript t represents the coordinate system of the current UAV position;

[0227] N represents the number of measurement points of the lidar, which is the distance between the drone's position and the nearest obstacle in the set direction;

[0228] The system state is input into the decision network to obtain the action output by the decision network. The action is expressed as:

[0229] a t =(α t ,β t ,γ t ,δ t ,λ t );

[0230] Among them, α t : heading item weight;

[0231] β t : gap term weight;

[0232] γ t : speed term weight;

[0233] δ t : curvature distance component, additional weight of the new extended cost function;

[0234] λ t : Dynamic environment adaptability function component, additional weight of the new extended cost function;

[0235] The drone receives observations or states of the environment and outputs an action a. The action is to set the cost function weights for the next control cycle. These weights will be used to generate motion commands for the drone.

[0236] Design of reward function:

[0237] If the drone successfully reaches the target location, it will be given a large positive reward R1; if the drone stays at the same location for multiple control cycles, it will be similar to being stuck and will be given a negative reward R2; if the drone collides with an obstacle / or other drones / birds during movement, it will be given a large penalty, negative R3; if the drone collides with a moving obstacle or other drones while stationary, it will be given a negative reward R4; the design of this reward function guides the reinforcement learning algorithm to learn safe and effective navigation strategies.

[0238] Optionally, the steps of implementing the time-varying environment adaptive dynamic window method include the following steps:

[0239] The dynamic environment adaptability function is added to the traditional DWA to obtain a time-varying environment adaptive dynamic window method;

[0240] The objective function of the time-varying environment adaptive dynamic window method is defined as:

[0241] G(ν,ω)=α·h(ν,ω)+β·c(ν,ω)+γ·v(ν,ω)+δ·d(ν,ω)+λ·e(ν,ω);

[0242] Among them, h(ν,ω): heading component, which represents the angular difference between the drone's forward direction and the target, and measures the degree to which the drone is heading towards the target;

[0243] c(ν,ω): clearance component, which represents the distance from the drone to the nearest obstacle on the curvature currently considered, and measures the safety of the motion;

[0244] v(ν,ω): velocity component, representing the translation speed of the drone and measuring the efficiency of the movement;

[0245] d(ν,ω): curvature distance component, which represents the shortest distance from the target point to the curvature currently considered, and measures the relevance of the path to the target;

[0246] e(ν,ω): Dynamic environment adaptability function component, which indicates the degree of adaptability of the current speed command (ν,ω) to the predicted dynamic changes in the environment. It measures whether the UAV can effectively cope with the expected movement of dynamic obstacles in the environment when executing the current speed command;

[0247] α, β, γ, δ, λ: weights corresponding to each component;

[0248] The curvature distance component d(ν,ω) represents the shortest distance from the target point to the curvature currently considered, and measures the relevance of the path to the target. The expression is as follows:

[0249]

[0250] α: distance attenuation factor (≥1), which controls the nonlinearity of the distance effect;

[0251] β: proximity gain coefficient (0-1), which increases the reward when the target is very close to the curvature;

[0252] λ: proximity sensitivity, which controls the steepness of the approach reward;

[0253] δ: Directional penalty factor, considering the deviation between the curvature line direction and the target direction;

[0254] θ curv : tangent direction of the curvature line;

[0255] θgoal : The direction of the target point relative to the drone;

[0256] robot_radius: The radius of the drone, used to normalize the distance value;

[0257] distance_to_goal: the shortest distance between the motion curvature generated by the current velocity combination (ν, ω) and the target point, which is the truncation parameter;

[0258] max_dist: Maximum distance threshold, used for normalization and limiting the range of distance influence.

[0259] Optionally, the steps of implementing the AC-time-varying environment adaptive dynamic window method include the following steps:

[0260] Obtain historical observation data and preprocess it, convert the historical observations to the current drone coordinate system, align historical observations at different times to the same reference coordinate system of the current drone, and define the coordinate system conversion formula as follows:

[0261]

[0262] Where, t: current time;

[0263] τ: universal time offset parameter, including k, l or g, indicating the time difference;

[0264] The converted distance measurement value represents the distance measurement of the historical observation before time τ in the coordinate system of the current time t;

[0265] At time t-τ, the distance measurement value referenced to the drone coordinate system at time t-τ;

[0266] The coordinate transformation matrix from time t-τ to time t represents the relative motion speed;

[0267] Θ t : A set of state estimates of dynamic objects detected in the environment;

[0268] η: mixing coefficient, dynamically adjusted with time and observation quality Control the weight of the two conversion methods;

[0269] The time derivative of the coordinate transformation represents the relative motion speed.

[0270] ψ(): evaluation function used to calculate the mixing coefficient η;

[0271] σ(): Swish activation function;

[0272] cart(): A function that converts distance scan data in polar coordinate form into Cartesian coordinate points;

[0273] F(): forward projection function based on velocity prediction, considering the trajectory of moving obstacle candidates including other aircraft;

[0274] pol(): A function that converts points back into the distance scan form in polar coordinates;

[0275] The neural network structure in the AC-time-varying environment adaptive dynamic window method is defined as follows:

[0276] 1D four-channel input;

[0277] The feature extraction layer processes the spatially correlated scan data;

[0278] The convolutional layer uses the Swish activation function;

[0279] Feature fusion flattens the convolution output into a vector, concatenating the target distance, angle, and current speed;

[0280] In the Actor-Critic architecture, the Actor and Critic networks share the same infrastructure, except for the last layer.

[0281] Define the training and output decisions in the AC-time-varying environment adaptive dynamic window method:

[0282] Three-stage training with different levels of difficulty;

[0283] The Actor network inputs system status information, including historical observations, target position, and current speed, and outputs actions, including the weight values ​​of the five TEADWA components: heading component weight α, gap component weight β, velocity component weight γ, curvature distance component weight δ, and dynamic environment adaptability function component weight λ;

[0284] The critic network outputs a state value estimate V(s), which guides the improvement of the actor through the advantage function.

[0285] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.

[0286] Specifically, this embodiment includes the following technical solutions:

[0287] This embodiment proposes a novel UAV intelligent patrol path decision system and method (DRAMSM-AC-TEADWA): two algorithms are connected through a system environment module, and the system environment includes information such as network topology, steady-state conditions, and obstacle constraints.

[0288] First, a Diverse Random Adaptive Markov Stationary Matrix (DRAMSM) algorithm was developed. It is a four-stage combinatorial optimization algorithm. On the basis of satisfying the target network topology structure, through the optimization processes such as layer random matrix generation, relative error iterative adjustment, similarity value elimination and matrix diversity assurance, the generated multiple Markov state transition matrices (path selection) not only have sufficient diversity but also meet the target steady-state distribution requirements of special requirements.

[0289] Secondly, a Dynamic Environment Adaptability Function (DEAF) was developed to handle dynamic obstacles in complex environments. This function comprehensively evaluates the safety and effectiveness of UAV speed commands in dynamic environments by coupling autocorrelation score, temporal proximity score and deviation score functions, and quantifies environmental adaptability by summing the risk weighted of each obstacle.

[0290] Finally, the AC-TEADWA (Actor-Critic Time-varying Environment Adaptive Dynamic Window Approach) algorithm was developed. By incorporating the dynamic environment adaptability function (DEAF) into the traditional dynamic window approach (DWA), the time-varying environment adaptive dynamic window approach (TEADWA) was formed. This enables the system to more accurately estimate the current environmental state and use the weight adjustment of its objective function as the action space for reinforcement learning. The TEADWA parameters are adjusted using the Actor-Critic (MoEMamba network) algorithm to obtain the final AC-TEADWA algorithm. This enables the UAV system to dynamically adjust the TEADWA behavior according to the environmental state, achieving both adaptability and safety in time-varying partially observable environments. In particular, when the system faces birds, other UAVs / aircraft, or unpredictable moving obstacles, it can make safer and more efficient decisions.

[0291] More specifically, the specific implementation of this embodiment is described below.

[0292] 1. System modeling.

[0293] The physical environment is discretized according to the actual patrol situation of the UAV and modeled by an undirected topological graph, G = {V, E}, where V = {v1, v2, ..., v n} represents the set of vertices (drone patrol locations), each node v in the graph irepresents the location in the real environment where drone patrol and surveillance is required, E={(i,j)|i,j∈V,p ij ≥0}, that is, the set of paths between two patrol locations, p ij >0 represents the slave node v i Transfer to node v j The probability that there are actual flight paths of drones at the two patrol locations satisfies p ij With p ji Not necessarily equal. ij Representative node v i and v j Connected between, Representative node v i and v j The path length between i ∈R + Representative node v i The value of a i ∈R + Represents the attacker (attack drone) attacking node v i The time required. Figure 2 Represents a Markov chain with five nodes and its corresponding transition matrix P, which is irreducible and aperiodic. An example of a state transition matrix is ​​as follows:

[0294]

[0295] The transition probabilities between nodes are encoded in the transition matrix P. Each row of the matrix P represents a probability distribution (i.e., all values ​​in each row add up to 1). The patrol drone can start from any initial vertex (initial patrol position) and then sample according to the probability distribution in the row corresponding to its current vertex to determine which area it should patrol / inspect next.

[0296] 2. Multivariate Random Adaptive Markov Steady-State Matrix (DRAMSM) algorithm.

[0297] The Diverse Random Adaptive Markov Stationary Matrix (DRAMSM) algorithm is a four-stage combinatorial algorithm that efficiently solves the problem of generating a random Markov state transition matrix with specific constraints by generating a steady-state distribution that meets special requirements while satisfying the target network topology.

[0298] 2.1. Generate optimal steady-state distribution.

[0299] like Figure 2 The Markov chain shown, different nodes vi Represents different locations that require drone patrols. Different nodes (patrol locations) have different values, and attackers need different amounts of time to attack different nodes. Therefore, high-value target nodes and nodes with short attack times need to obtain more patrol resources, and patrol drones have a higher probability of visiting these nodes. π=[π1,π2,…,π n ] is a general steady-state distribution that represents the probability distribution of patrollers at various target points after the Markov chain runs for a long time. Therefore, it is necessary to design a steady-state distribution with specific optimization properties. Formula (1) shows how to calculate the steady-state distribution that takes into account the target value and attack time.

[0300]

[0301] in, v i ∈V,v i represents the i-th node, V represents the set of each node, w i Represents the target node v i The value of a i represents the attack time, μ i Representative node v i The importance weight parameter, μ j Representative node v j The importance weight parameter, n is the total number of nodes, Represents the target node v i The target steady-state distribution represents the probability distribution of the patroller at each target node after the Markov chain runs for a long time.

[0302] 2.2. DRAMSM algorithm generates transfer matrix.

[0303] In order to increase the difficulty of predicting patrol drones, the multivariate random adaptive Markov steady-state matrix (DRAMSM) algorithm is used to generate multiple * (Formula 1) but with different transfer characteristics, the transfer matrix P is used to achieve a random path selection strategy with time-space decoupling, such as Figure 3 Specifically, the DRAMSM algorithm is a multivariate random adaptive Markov steady-state matrix (DRAMSM) algorithm. The four stages are responsible for layered random matrix generation, relative error iterative adjustment, multiple similarity value elimination, and matrix diversity assurance. The matrix is ​​continuously improved in an iterative manner to meet the constraints of network topology, row sum to 1, and steady-state conditions.

[0304] Phase 1: Stratified random matrix generation.

[0305] 1. Self-loop probability generation:

[0306] For each node i, the self-loop probability p is assigned using four distribution intervals according to the self-loop type ij .

[0307] 2. Non-self-loop transition probability generation:

[0308] Determine the set of connected nodes C for each node i i (based on the adjacency matrix), calculate the residual probability, generate the weight vector w using three distributions, add random perturbations, normalize the weights and assign transition probabilities.

[0309] In this stage, the initial random matrix Q is constructed to satisfy Figure 2 The network topology constraints shown above do not usually satisfy the predefined steady-state distribution π *

[0310] The second stage: relative error iterative adjustment algorithm.

[0311] 1. Steady-state distribution verification:

[0312] Calculate the steady-state effect of the current matrix P and calculate the L1 error.

[0313] 2. Relative error calculation:

[0314] For each state j, the relative error is calculated, where a positive value indicates that the transition to the state needs to be increased, and a negative value indicates that the transition to the state needs to be reduced.

[0315] 3. Error ratio adjustment:

[0316] Set the decaying learning rate, calculate the adjustment factor, and apply a multiplicative adjustment to each valid transition probability to ensure a minimum probability threshold.

[0317] 4. Row Normalization:

[0318] The sum of each row is calculated and normalized to ensure that the Markov property is satisfied.

[0319] 5. Periodic random disturbances:

[0320] A random perturbation of ±5% was added every 10 iterations.

[0321] In this stage, the initial matrix Q is iteratively adjusted to make its steady-state distribution close to the target distribution π * , the computational process is similar to gradient descent with momentum and decaying learning rate, but the object of operation is the transfer matrix instead of the parameter vector in the standard optimization problem.

[0322] The third stage: similarity elimination algorithm.

[0323] 1. Similarity value detection:

[0324] For each valid element in row i (i.e. pij >0 and the topology allows it), check the difference between adjacent sorted value pairs, and if the difference is less than the threshold, it is determined to be similar.

[0325] 2. Random adjustment:

[0326] Randomly select one value in a similar pair to generate an adjustment, and randomly decide whether to increase or decrease the adjustment.

[0327] 3. Row Normalization:

[0328] The rows are renormalized to keep the sum of probabilities equal to 1.

[0329] This stage ensures that there are no values ​​in the matrix that are too close together, increasing randomness and creating a more natural random distribution.

[0330] Phase 4: Matrix diversity guarantee algorithm.

[0331] 1. Matrix difference calculation:

[0332] Calculate any two matrices P a and P b The mean absolute error between .

[0333] 2. Diversity determination:

[0334] If the mean absolute error is less than the error diversity threshold, the two matrices are considered too similar and need to be adjusted.

[0335] 3. Large random disturbances:

[0336] One of the matrices is randomly selected and subjected to a large random perturbation of ±10%, followed by normalization.

[0337] 4. Re-adjust the steady-state distribution:

[0338] Reapply the second-stage iterative adjustment algorithm to the perturbed matrix to ensure that the adjusted matrix still satisfies the steady-state distribution π * Require.

[0339] This stage ensures that the generated multiple Markov state transition matrices P k There is enough diversity, rather than variations of similar matrices.

[0340] 3. Dynamic environmental adaptability function.

[0341] The Dynamic Environment Adaptability Function (DEAF) comprehensively evaluates the safety and effectiveness of the UAV's speed command in a dynamic environment. It quantifies environmental adaptability by weighting the risk of each obstacle. A function value close to 1 indicates that the speed command is safe and adapts to environmental changes; a value close to 0 indicates that there may be a conflict with dynamic obstacles and that the command is not adapted to the current environmental state, as shown in Formula (2).

[0342]

[0343] Where, e(ν,ω): dynamic environmental adaptability function, which represents the environmental adaptability evaluation value of the UAV under a given speed command;

[0344] ν: The linear velocity command of the UAV, which controls the speed of the UAV forward or backward;

[0345] ω: angular velocity command of the UAV, which controls the speed and direction of the UAV's steering;

[0346] The current system state estimated using weighted observation data, including the environment representation of obstacle positions and velocities, where z i is the observation data of the UAV at the i-th moment, and the observation weight Calculated by the following expression:

[0347]

[0348] Observation data z i The weight of quantifies the importance of observation in state estimation;

[0349] w a , w r , w d It is the memory weight coefficient for autocorrelation, time proximity and deviation, starting from the initial empirical value, set according to the nature of the environment, and using an online adaptive method to regularly update the memory weight;

[0350] A s (z i ,Θ): the autocorrelation score of the observation, reflecting the temporal correlation of the environmental dynamics;

[0351] Θ: a set of parameters that controls the calculation of the autocorrelation score;

[0352] R s (z i ): the temporal proximity score of the observation, emphasizing the importance of recent data;

[0353] Ds (z i ,ψ): Deviation score of the observation, capturing abnormal changes in the environment;

[0354] ψ: bias sensitivity parameter, which controls the degree of emphasis the algorithm places on bias;

[0355] M: the number of dynamic obstacles currently perceived, including other drones;

[0356] The probability of obstacle i existing based on the state estimate Ω is in the range [0,1], where a larger value indicates a higher collision risk. It measures the probability that obstacle i intersects the drone's path when the drone executes the speed command (ν,ω);

[0357] The relative velocity vector between the UAV and obstacle i is calculated based on the UAV velocity command (ν, ω) and the obstacle velocity information in the state estimate Ω, which measures the speed and direction of the relative motion between the UAV and the obstacle;

[0358] d i (ν,ω,Ω): The predicted minimum distance between the UAV and obstacle i when the UAV adopts the speed (ν,ω). The smaller the value, the higher the potential collision risk.

[0359] K: normalization constant, ensuring the appropriate range of values ​​for the fractional term, determined experimentally based on environmental characteristics and UAV dynamic parameters;

[0360] ε: A positive number used to prevent division by zero. The value is between 10 -6 to 10 -3 between.

[0361] 3.1. Autocorrelation score.

[0362] The autocorrelation function measures the similarity between time series at different points in time. High autocorrelation values ​​indicate that environmental changes have a certain regularity or periodicity, making past observations useful for predicting future states. In a mobile obstacle environment, capturing the temporal correlation patterns of environmental dynamics allows the system to identify which historical observation patterns are similar to the current situation, thereby prioritizing those observations that are more likely to provide useful information, enabling the system to predict the future position of obstacles, as shown in Equation (4).

[0363]

[0364] Among them, A s (z i ,Θ): Observation value z i At time t i The comprehensive autocorrelation score at time , controlled by the parameter set Θ;

[0365] z i : The observation value at the jth time point in the time series, which represents the observation of the obstacle position by the UAV's lidar ranging, position estimation, or sensor;

[0366] z j : The observation value at the jth time point in the time series;

[0367] The j+|ttth time series i | observations at each time point;

[0368] The mean of all observations represents the baseline or reference value of the observations, reflecting the average level of the environmental state;

[0369] t: current time;

[0370] t i : Observation time point i;

[0371] |tt i |: the time difference between the current time t and the observation time ti, which is used as the lag of the autocorrelation;

[0372] ω j : The correlation weight at time point j, emphasizing the contribution of observations at certain time points to the correlation calculation;

[0373] and The weight coefficients of observations at different time points can be used to deal with missing data or uncertain observations;

[0374] λ: time decay parameter, j0 is the reference time point, so that the pairs farther away from the reference point contribute less to the correlation;

[0375] γ: time lag penalty coefficient, which makes the score decay faster as the time difference increases;

[0376] δ: Periodic impact factor, reflecting periodic changes in the environment (such as tides and diurnal cycles);

[0377] T cycle : The length of the cycle of environmental change;

[0378] A set of parameters that controls the calculation of autocorrelation scores;

[0379] n: The length of the observation sequence, which indicates the size of the historical data window considered.

[0380] The system maintains an observation sequence {z1,z2,…,z n}, the current time is t, the historical observation z iOccurs at time t i First, calculate the mean z of the observation sequence, and then for each historical observation z i , calculate the time difference from the observation to the current moment (tt i ), using this time difference as the lag parameter, evaluate the autocorrelation A of the entire observation sequence at this lag value s (z i ,Θ).

[0381] 3.2. Temporal proximity score.

[0382] The temporal proximity score gives higher weight to more recent observations. In a mobile obstacle environment, it reflects the temporal locality of the environmental state. The most recent observations are usually more representative of the current environmental state than distant observations, as shown in formula (5).

[0383]

[0384] Among them, R s (z i ):Observation z i The temporal proximity score quantifies how close the observation time is to the current time.

[0385] t: current time;

[0386] t i : Observation time point i;

[0387] η: normalization constant to ensure that the score is within a reasonable range;

[0388] κ: time decay exponent, which controls the decay rate of timeliness over time;

[0389] ε: a small positive number, a safety parameter to prevent division by zero, with a value between 10-6 and 10-3;

[0390] τ d : exponential decay time constant, indicating the observation "half-life";

[0391] φ: additional weight coefficient for the most recent observation;

[0392] τ c : Critical time window, observations within this window receive extra weight.

[0393] 3.3. Deviation score.

[0394] The deviation score identifies observations that deviate significantly from the norm. In a moving obstacle environment, it helps capture sudden events or key transition points, such as an obstacle suddenly changing direction or speed, and alerts the system to this potentially important state transition, as shown in formula (6).

[0395]

[0396] Among them, D s (z i ,ψ): parameterized deviation score of observation zi, quantifying the degree of deviation of the observation from its mean within the time window;

[0397] z i : The observation value at the jth time point in the time series;

[0398] ψ: bias sensitivity parameter, which controls the degree of emphasis the algorithm places on bias;

[0399] p: bias power, increasing the degree of nonlinearity;

[0400] The standard deviation of the observations, used to normalize the deviations;

[0401] ξ: stability constant, preventing the denominator from being zero;

[0402] ν: outlier gain coefficient;

[0403] 1(·): indicator function, which is 1 when the condition is met and 0 otherwise;

[0404] χ: outlier threshold coefficient, defining when an observation is considered an outlier;

[0405] Currently the average of all observations within the time window is considered.

[0406] 4. AC-time-varying environment adaptive dynamic window method (ACTEADWA).

[0407] To enable drones to better track their chosen / planned paths, an algorithm combining deep reinforcement learning (Actor-Critic) and the Time-Varying Environment Adaptive Dynamic Window Approach (TEADWA) is employed. The actor (MoEMamba) network outputs an action (five TEADWA weights) based on the input environment state (historical observations, target position, and current velocity), while the critic (MoEMamba) network outputs a state value estimate V(s). This is guided by an advantage function to improve the actor (policy), resulting in a more stable training process. The AC-TEADWA objective function uses artificial intelligence to adjust the TEADWA parameters rather than directly generating motion commands. This achieves both the adaptability of deep learning and the safety guarantees of traditional DWA.

[0408] 4.1. Setting of system related status.

[0409] Design of system state: system state s t It describes all the information contained in the state of the UAV at time t and is the basis of the AC-TEADWA method. It provides the neural network with the complete information required to predict the AC-TEADWA weight, as shown in formula (7).

[0410]

[0411] Among them, s t : The system status at time t, which contains the complete status information of the UAV at time t;

[0412] The distance measurement value from the drone to the obstacle at the current time t in the current drone coordinate system;

[0413] The distance measurement k control cycles ago expressed in the current UAV coordinate system;

[0414] The distance measurement value l control cycles ago expressed in the current UAV coordinate system;

[0415] The distance measurement g control cycles ago expressed in the current UAV coordinate system;

[0416] dt : The distance from the current position of the UAV to the target;

[0417] φ t : The angle from the current position of the drone to the target;

[0418] ν t : The current translation speed of the drone;

[0419] ω t : The current rotation speed of the drone;

[0420] The obstacle distance changes between consecutive time steps to capture dynamic information;

[0421] The time derivative of the distance from the drone to the target, indicating whether it is approaching the target;

[0422] The current acceleration state vector of the drone is the linear acceleration of the drone at the current time t, is the angular acceleration of the UAV at the current time t;

[0423] ξ t : Environmental complexity index, calculated by the entropy value of obstacle distribution;

[0424] The specific form of the distance measurement value is:

[0425] τ represents the past time step and refers to k, l, and g in the formula;

[0426] The superscript t-τ indicates the measurement time;

[0427] The subscript t represents the coordinate system of the current UAV position;

[0428] N represents the number of measurement points of the lidar, which is the distance between the drone's position and the nearest obstacle in the set direction.

[0429] Using observations at multiple time points can help predict the motion of dynamic obstacles, including target information can help with goal-directed navigation, and including the current motion state can help generate smooth control commands.

[0430] Action design: The drone receives observations (or states) of the environment and outputs an action a. The action is to set the cost function weights for the next control cycle. These weights will be used to generate motion commands for the drone.

[0431] a t =(α t ,β t ,γ t,δ t ,λ t ).

[0432] α t : Heading item weight.

[0433] β t : Gap term weight.

[0434] γ t : Speed ​​term weight.

[0435] δ t : Curvature distance component, additional weight of the new extended cost function.

[0436] λ t : Dynamic environment adaptability function component, additional weight of the new extended cost function.

[0437] This design uses the weight adjustment of ACTEADWA as the action space for reinforcement learning, enabling the system to dynamically adjust the behavior of TEADWA according to the environment state.

[0438] Reward function design: If a drone successfully reaches the target location, it receives a large positive reward, R1. If the drone remains in the same location for multiple control cycles, similar to being stuck, it receives a negative reward, R2. If the drone collides with an obstacle, another drone, or a bird while moving, it receives a large penalty, a negative reward, R3. If the drone collides with a moving obstacle or another drone while stationary, it receives a negative reward, R4. This reward function design guides the reinforcement learning algorithm to learn safe and effective navigation strategies.

[0439] 4.2. AC-Time-Varying Environment Adaptive Dynamic Window Method (AC-TEADWA) algorithm design.

[0440] The Actor-Critic Time-varying Environment Adaptive Dynamic Window Approach (AC-TEADWA) makes TEADWA more robust in time-varying partially observable environments, especially when the system faces flying birds, other drones, or unpredictable moving obstacles. This weighted calculation based on time, correlation, and anomaly enables the system to more accurately estimate the current environmental state, thereby making safer and more efficient decisions. Its objective function is shown in Equation (8).

[0441] G(ν,ω)=α·h(ν,ω)+β·c(ν,ω)+γ·v(ν,ω)+δ·d(ν,ω)+λ·e(ν,ω)(8)

[0442] h(ν,ω): Heading component, which represents the angular difference between the drone’s heading and the target, and measures the degree to which the drone is heading towards the target.

[0443] c(ν,ω): Clearance component, which represents the distance from the drone to the nearest obstacle on the curvature currently considered, and measures the safety of the motion.

[0444] v(ν,ω): Velocity component, representing the translational speed of the drone and measuring the efficiency of the movement.

[0445] d(ν,ω): The curvature distance component represents the shortest distance from the target point to the currently considered curvature (i.e., the trajectory that the drone will form under a specific speed combination), which measures the relevance of the path to the target.

[0446] e(ν,ω): Dynamic environment adaptability function component, which indicates the degree of adaptability of the current speed command (ν,ω) to the predicted dynamic changes in the environment. It measures whether the UAV can effectively cope with the expected movement of dynamic obstacles in the environment when executing the current speed command.

[0447] α, β, γ, δ, λ: weights corresponding to each component.

[0448] The curvature distance component d(ν,ω) represents the shortest distance from the target point to the currently considered curvature (i.e., the motion trajectory that the UAV will form under a specific speed combination), and measures the relevance of the path to the target, as shown in formula (9).

[0449]

[0450] α: distance attenuation factor (≥1), which controls the nonlinearity of the distance effect;

[0451] β: proximity gain coefficient (0-1), which increases the reward when the target is very close to the curvature;

[0452] λ: proximity sensitivity, which controls the steepness of the approach reward;

[0453] δ: Directional penalty factor, considering the deviation between the curvature line direction and the target direction;

[0454] θ curv : tangent direction of the curvature line;

[0455] θ goal : The direction of the target point relative to the drone;

[0456] robot_radius: The radius of the drone, used to normalize the distance value;

[0457] max_dist: Maximum distance threshold, used for normalization and limiting the range of distance influence.

[0458] Here, distance_to_goal is the shortest distance between the curvature of the motion generated by the current velocity combination (ν, ω) and the target point, and serves as a cutoff parameter. This formula cleverly converts physical distance into a score in the [0, 1] interval, ensuring that curvatures closer to the target receive higher scores, meeting the goal-oriented requirements of navigation. This provides an "obstacle bypass" perspective, helping drones make better decisions when facing unfavorable heading components and addressing the issue of drone oscillation in front of obstacles.

[0459] Each speed combination (ν, ω) in the ACTEADWA search space corresponds to a possible curvature. A cost function evaluates the "goodness" of each curvature (safety, proximity to the target, etc.). Reinforcement learning is used to optimize the corresponding weights in the cost function to select the optimal curvature (i.e., the optimal speed combination).

[0460] 4.3. Implementation of AC-Time-Varying Environment Adaptive Dynamic Window Method (AC-TEADWA) algorithm.

[0461] The AC-Time-Varying Environment Adaptive Dynamic Window Method (AC-TEADWA) algorithm uses artificial intelligence to adjust the parameters of the traditional DWA method. The Actor (MoEMamba) network outputs actions (five TEADWA weight values) based on input state information (historical observations, target position, current speed), and the Critic (MoEMamba) network outputs the state value estimate V(s). The advantage function is used to guide the improvement of the Actor (strategy), thereby achieving a more stable training process.

[0462] 4.3.1. Acquisition and preprocessing of historical observations.

[0463] 4.3.1.1 Multi-time observation data collection.

[0464] Collect laser scanning data from the drone at four time points, and the data at the current time (t) Data from k control cycles ago (tk) Data from l control cycle ago (tl) Data from g control cycles ago (tg) Get N distance values ​​at each time point.

[0465] 4.3.1.2 Coordinate system conversion.

[0466] Use formula (10) to convert historical observations to the current drone coordinate system. Align historical observations at different times to the same reference coordinate system of the current drone, so that historical observations can be directly compared with current observations, which facilitates the neural network to learn the obstacle movement pattern.

[0467]

[0468] Where, t: current time;

[0469] τ: universal time offset parameter (k, l or g), indicating the time difference;

[0470] The converted distance measurement value represents the distance measurement of the historical observation before time τ in the coordinate system of the current time t;

[0471] At time t-τ, the distance measurement value referenced to the drone coordinate system at time t-τ;

[0472] The coordinate transformation matrix from time t-τ to time t represents the relative motion speed;

[0473] Θ t : A set of state estimates of dynamic objects detected in the environment;

[0474] η: mixing coefficient, dynamically adjusted with time and observation quality Control the weight of the two conversion methods;

[0475] The time derivative of the coordinate transformation represents the relative motion speed.

[0476] ψ(): evaluation function used to calculate the mixing coefficient η;

[0477] σ(): Swish activation function;

[0478] cart(): function that converts distance scan data (in polar coordinate form) into Cartesian coordinate points;

[0479] F(): forward projection function based on velocity prediction, considering candidate trajectories of moving obstacles (including other aircraft);

[0480] pol(): function that converts the point back to distance scan form (polar coordinate form);

[0481] first step:

[0482] enter: is the original distance measurement before time τ.

[0483] Function: Convert distance scan data (polar coordinates) into Cartesian coordinate points.

[0484] Output: A series of Cartesian coordinate points.

[0485] Step 2:

[0486] Input: The Cartesian coordinate point obtained in the previous step.

[0487] The coordinate transformation matrix from time t-τ to time t.

[0488] Function: Transform the point at time t-τ to the drone coordinate system at time t.

[0489] Output: The point in the current drone coordinate system (Cartesian coordinate form).

[0490] Step 3: pol(...).

[0491] Input: Cartesian coordinate point in the current coordinate system.

[0492] Function: Convert the point back to distance scan format (polar coordinate format).

[0493] Processing: Use binning and interpolation to handle unknown values.

[0494] Output: Distance measurement in the current coordinate system (polar coordinates).

[0495] 4.3.1.3 Data processing and standardization.

[0496] Handling "unknown values": Use binning techniques to discretize the angular space and use interpolation to estimate missing values.

[0497] Distance clipping: Limit all distance values ​​to within dclip to ensure data range consistency.

[0498] 4.3.2. Neural network input design.

[0499] 4.3.2.1 Input structure.

[0500] Form a 1D four-channel input:

[0501] First channel: current observation

[0502] Second channel: observations from k control periods ago (already transformed to the current coordinate system).

[0503] The third channel: observation before 1 control cycle (already transformed to the current coordinate system).

[0504] Channel 4: Observation g control cycles ago (already transformed to the current coordinate system).

[0505] 4.3.2.2 Additional status information.

[0506] Target related information: distance d t and angle φ t .

[0507] Drone motion state: current translation speed ν t and the rotational speed ω t .

[0508] 4.3.3. Network architecture design.

[0509] 4.3.3.1 Feature extraction layer.

[0510] Three layers of one-dimensional convolution: Processing spatially correlated scan data.

[0511] All convolutional layers use the Swish activation function.

[0512] 4.3.3.2 Feature fusion.

[0513] Flatten the convolution output into a vector, concatenating information such as target distance, angle, current speed, etc.

[0514] 4.3.3.3 Decision-making level.

[0515] There are m fully connected layers with n four linear neurons, and the outputs correspond to the DWA weights α, β, γ, δ, and λ.

[0516] 4.3.4. Implicit prediction step.

[0517] 4.3.4.1 Time structure learning.

[0518] Instead of explicitly outputting the future positions of obstacles, the network learns the temporal pattern of observations at three time points through convolutional layers.

[0519] The convolution operation creates connections between different time points at the same angle, enabling the network to detect patterns of distance changes (i.e., movement).

[0520] 4.3.4.2 Motion pattern recognition.

[0521] The position of the static obstacle is basically fixed during the observations at three time points.

[0522] The position of a dynamic obstacle changes during observation at different time points.

[0523] The network learns to identify these changing patterns and predict their trends.

[0524] 4.3.4.3 Weight adaptation mechanism.

[0525] When the movement of an obstacle is detected, α (heading weight), β (gap weight), γ (speed weight), δ (distance to curvature weight), and λ (dynamic environment adaptability function weight) are adjusted. Different motion modes correspond to different weight combinations.

[0526] 4.3.5. Training and optimization.

[0527] 4.3.5.1 Reinforcement Learning Framework

[0528] Using the Actor-Critic (MoEMamba network) method, the MoEMamba network replaces the recurrent neural network in the traditional Actor-Critic method. MoEMamba is a new type of network designed in this paper, which integrates and mixes expert networks (Mixture of Experts, MoE) and Mamba networks, and alternates between the two networks: one layer of mixed expert network and one layer of Mamba network, and so on, with a total of k layers. The critic network is responsible for evaluating the value of the state and guiding the improvement of the Actor (strategy) through the advantage function, thereby achieving a more stable training process.

[0529] Actor's objective function:

[0530]

[0531] Among them, s: system state, environmental state information input to the Actor network;

[0532] a: Control action output by the Actor network;

[0533] w k : Current Actor network parameters (parameters of the kth iteration);

[0534] w: Actor network parameter variable to be optimized;

[0535] π w : The current updated strategy, a strategy function based on the parameter w;

[0536] strategies for generating experiences;

[0537] Advantage function, which measures the advantage of taking action a relative to the average level in state s;

[0538] η: Clipping constant, which limits the range of policy changes and controls the update step size;

[0539] clip(): Clipping function, limiting the value to the range [1-η,1+η];

[0540] Actor updates network parameters:

[0541]

[0542] Actor updates the parameter w by maximizing the objective function L k is the current Actor parameter, w k+1 are the updated Actor parameters.

[0543] Critic's advantage function:

[0544]

[0545] Among them, s t : The system state at time t, the environmental state information input to the Actor network;

[0546] a t : The control action output by the Actor network at time t;

[0547] r t : The instant reward value obtained at time t;

[0548] r t+1 : The reward value obtained at time t+1;

[0549] r t+n : The reward value obtained at time t+n;

[0550] γ: Discount factor, used to attenuate the weight of future rewards, with a value between 0 and 1;

[0551] γ t+n : The (t+n) power of the discount factor, which indicates the degree of discount on the reward at time t+n.

[0552] Critic's value function:

[0553]

[0554] Among them, π w : policy function based on parameter w;

[0555] n: time step, indicating the upper limit of the number of time steps considered;

[0556] s t : The system state at time t, the environmental state information input to the Actor network;

[0557] r0: reward value at the initial moment (step 0);

[0558] r1: reward value of step 1;

[0559] rn : reward value of step n;

[0560] γ: discount factor, used to decay the weight of future rewards;

[0561] γ n : The nth power of the discount factor, indicating the degree of discount on the reward of the nth step;

[0562] In strategy π w Next state s t The value function represents the value of the state s t Start by following the strategy π w The expected cumulative reward for execution;

[0563] In the Actor-Critic (MoEMamba network) architecture, the Actor network outputs an action (five DWA weight values) based on input state information (historical observations, target position, current speed), and the Critic network outputs the state value estimate V(s). The Actor and Critic networks share the same basic structure, except for the last layer.

[0564] 4.3.5.2 Three-stage training.

[0565] Phase 1: There are no moving obstacles, and the drone first learns basic navigation.

[0566] Phase 2: There are obstacles with the same moving speed, and the drone learns to handle simple dynamic environments.

[0567] Phase 3: There are obstacles with different moving speeds, and the drone learns to handle complex dynamic environments.

[0568] In summary, this embodiment includes the following technical solutions:

[0569] 1. A UAV intelligent patrol path decision system and method in complex environments is proposed: Multivariate Random Adaptive Markov Steady-State Matrix-AC-Time-Varying Environment Adaptive Dynamic Window Method (DRAMSM-AC-TEADWA), which takes into account random path selection, path tracking and dynamic obstacle avoidance.

[0570] 2. A Diverse Random Adaptive Markov Stationary Matrix (DRAMSM) algorithm was developed. Through optimization processes such as layer random matrix generation, iterative adjustment of relative errors, similarity value elimination, and matrix diversity assurance, the generated multiple Markov state transition matrices (path selection) have sufficient diversity while satisfying constraints such as network topology, probability conservation, and target steady-state distribution.

[0571] 3. A Dynamic Environment Adaptability Function (DEAF) was developed to handle dynamic obstacles in complex environments. This function comprehensively evaluates the safety and effectiveness of UAV speed commands in dynamic environments by coupling autocorrelation score, temporal proximity score, and deviation score functions, and quantifies environmental adaptability by summing the risk weights of each obstacle.

[0572] 4. We developed the AC-TEADWA (Actor-Critic Time-varying Environment Adaptive Dynamic Window Approach) algorithm. By incorporating the dynamic environment adaptability function (DEAF) into the traditional dynamic window approach (DWA), we formed the time-varying environment adaptive dynamic window approach (TEADWA). This enables the system to more accurately estimate the current environmental state and use the weight adjustment of its objective function as the action space for reinforcement learning. The actor-critic (MoEMamba network) algorithm is used to adjust the TEADWA parameters, resulting in the final AC-TEADWA algorithm. This enables the UAV system to dynamically adjust the TEADWA behavior according to the environmental state, achieving both adaptability and safety in time-varying partially observable environments. In particular, when the system faces birds, other drones / aircraft, or unpredictable moving obstacles, it can make safer and more efficient decisions.

[0573] The beneficial effects of this embodiment include:

[0574] This embodiment integrates an intelligent drone patrol decision-making system and method. First, a Diverse Random Adaptive Markov Stationary Matrix (DRAMSM) algorithm is developed. It is particularly suitable for optimizing random patrol strategies in time-varying Markov decision processes and generates highly randomized drone patrol paths that are more difficult to predict than those generated by existing algorithms.

[0575] Secondly, a dynamic environment adaptability function (DEAF) is added on the basis of traditional DWA. This function comprehensively evaluates the safety and effectiveness of the UAV speed command in a dynamic environment by coupling the autocorrelation score, temporal proximity score and deviation score functions, and quantifies the environmental adaptability by summing the risk weighted of each obstacle.

[0576] Finally, combining the Actor-Critic (MoEMamba network) algorithm, we proposed the AC-Time-Varying Environment Adaptive Dynamic Window Approach (AC-TEADWA). A neural network is designed to predict the weights of TEADWA, and the weight adjustment of TEADWA is used as the action space for reinforcement learning. This enables the UAV system to dynamically adjust the behavior of TEADWA according to the environmental state, rather than directly generating motion commands. This achieves both the adaptability of deep learning and the safety guarantee of TEADWA. It can better handle dynamic obstacles / other aircraft than either method alone, and its performance is superior to traditional DWA and pure deep learning methods. It increases the system's adaptability to environmental changes, especially when the system faces birds, other drones / aircraft, or unpredictable moving obstacles, making safer and more efficient decisions.

[0577] Reference Figure 4 The embodiment of the present application further provides a device for determining a patrol path of a UAV in a complex environment, which can implement the above-mentioned method for determining a patrol path of a UAV in a complex environment. The device includes:

[0578] A graph construction unit, used to construct an undirected topological graph using each location to be patrolled as a node;

[0579] A steady-state distribution generating unit, configured to generate a steady-state distribution of each of the nodes patrolled by the drone according to the topological constraints and node importance of the undirected topological graph;

[0580] A transfer matrix generating unit, configured to generate a plurality of candidate state transfer matrices having the steady-state distribution; wherein each of the candidate state transfer matrices includes a transfer probability of the UAV between each of the nodes;

[0581] An environmental assessment unit, configured to perform a dynamic environmental assessment of obstacles along the path corresponding to each candidate state transfer matrix, thereby determining an optional speed and direction of the UAV along the selected path;

[0582] A command generation unit is used to dynamically adjust the weights of the decision network according to the dynamic environment assessment results, and then use the weight-adjusted decision network to generate execution movement commands;

[0583] A mobile control unit is used to control the UAV to move between the various positions to be patrolled according to the execution motion command.

[0584] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0585] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method of the present application. The electronic device can be any smart terminal, such as a tablet computer or an in-vehicle computer.

[0586] It can be understood that the contents of the above method embodiments are all applicable to the embodiments of the present device, the functions specifically implemented by the embodiments of the present device are the same as those of the method of the present application, and the beneficial effects achieved are also the same as those achieved by the method of the present application.

[0587] See also Figure 5 , Figure 5 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0588] The processor 501 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0589] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 502 and is called by the processor 501 to execute the methods of the embodiments of this application.

[0590] Input / output interface 503, used to implement information input and output;

[0591] Communication interface 504, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0592] Bus 505 , which transmits information between various components of the device (e.g., processor 501 , memory 502 , input / output interface 503 , and communication interface 504 );

[0593] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via a bus 505 .

[0594] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method of the present application is implemented.

[0595] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0596] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0597] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0598] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0599] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0600] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0601] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0602] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0603] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0604] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0605] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0606] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0607] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A method for determining the patrol path of a UAV in a complex environment, characterized by: The method comprises the following steps: Construct an undirected topological graph with each patrol location as a node; Generate a steady-state distribution of each of the nodes patrolled by drones according to the topological constraints and node importance of the undirected topological graph; Generating a plurality of candidate state transfer matrices having the steady-state distribution; wherein each of the candidate state transfer matrices includes a transition probability of the UAV between each of the nodes; Performing a dynamic environmental assessment of obstacles on the paths corresponding to the candidate state transfer matrices, thereby determining an optional speed and direction of the UAV on the selected path; Dynamically adjust the weights of the decision network according to the dynamic environment assessment results, and then use the weight-adjusted decision network to generate execution movement commands; The UAV is controlled to move between each of the positions to be patrolled according to the execution motion command.

2. The method for determining a patrol path of a UAV in a complex environment according to claim 1 is characterized in that: Generating the steady-state distribution of each node patrolled by a drone according to the topological constraints and node importance of the undirected topological graph comprises the following steps: The steady-state distribution is generated according to the topological constraints of the undirected topological graph to increase the probability of the drone transferring to the target node; wherein, high-value target nodes and nodes with short attack time need to obtain more patrol resources, and the patrol drone has a higher probability of visiting these nodes; The expression of the steady-state distribution is: Among them, v i ∈V,v i represents the i-th node, V represents the set of each node, w i Represents the target node v i The value of a i represents the attack time, μ i Representative node v i The importance weight parameter, μ j Representative node v j The importance weight parameter, n is the total number of nodes, Represents the target node v i The target steady-state distribution represents the probability distribution of the patroller at each target node after the Markov chain runs for a long time.

3. The method for determining a patrol path of a UAV in a complex environment according to claim 1 is characterized in that: The generating of a plurality of candidate state transfer matrices having the steady-state distribution comprises the following steps: Generating a plurality of candidate state transfer matrices having the steady-state distribution and the target network topology structure according to a multivariate random adaptive Markov steady-state matrix algorithm includes the following steps: Phase 1: Stratified random matrix generation; Generate a first-stage random matrix that satisfies the topological constraints of the undirected topological graph, wherein the self-loop probability p is assigned using four distribution intervals. ii , and then generate the remaining non-self-loop probability p ij ; The second stage: relative error iterative adjustment algorithm; Adjusting the first-stage random matrix according to the error between the probability distribution of the first-stage random matrix and the steady-state distribution, specifically obtaining the second-stage random matrix through relative error calculation, error ratio adjustment, row normalization and random periodic perturbation; The third stage: similarity elimination algorithm; Eliminate matrices whose similarity values ​​are greater than a preset similarity threshold value in the random matrix of the second stage. Specifically, similarity value detection, random adjustment, and row normalization are performed to ensure that there are no values ​​that are too close in the matrix, thereby enhancing randomness and creating a more natural random distribution to obtain the random matrix of the third stage; Phase 4: Matrix diversity guarantee algorithm; The random matrix of the third stage is adjusted according to the average absolute error between the probability distributions of the multiple random matrices in the third stage, specifically by matrix difference calculation, diversity determination, large-scale random perturbation and readjustment of steady-state distribution to ensure that the generated multiple Markov state transfer matrices P k With sufficient diversity rather than variations of similar matrices, a random matrix in the fourth stage is finally obtained as the candidate state transfer matrix.

4. The method for determining a patrol path of a UAV in a complex environment according to claim 1, wherein: The step of dynamically evaluating the obstacles on the paths corresponding to the candidate state transfer matrices comprises the following steps: Performing a dynamic environmental assessment on obstacles in the paths corresponding to the candidate state transfer matrices according to a dynamic environmental adaptability function, summing the risk weights of each obstacle, and quantifying the environmental adaptability; The expression of the dynamic environmental adaptability function (DEAF) is: Where, e(ν,ω): dynamic environmental adaptability function, which represents the environmental adaptability evaluation value of the UAV under a given speed command; ν: The linear velocity command of the UAV, which controls the speed of the UAV forward or backward; ω: angular velocity command of the UAV, which controls the speed and direction of the UAV's steering; The current system state estimated using weighted observation data, including the environment representation of obstacle positions and velocities, where z i is the observation data of the UAV at the i-th moment, and the observation weight Calculated by the following expression: Observation data z i The weight of quantifies the importance of observation in state estimation; w a , w r , w d It is the memory weight coefficient for autocorrelation, time proximity and deviation, starting from the initial empirical value, set according to the nature of the environment, and using an online adaptive method to regularly update the memory weight; A s (z i ,Θ): the autocorrelation score of the observation, reflecting the temporal correlation of the environmental dynamics; Θ: a set of parameters that controls the calculation of the autocorrelation score; R s (z i ): the temporal proximity score of the observation, emphasizing the importance of recent data; D s (z i ,ψ): Deviation score of the observation, capturing abnormal changes in the environment; ψ: bias sensitivity parameter, which controls the degree of emphasis the algorithm places on bias; M: the number of dynamic obstacles currently perceived, including other drones; The probability of obstacle i existing based on the state estimate Ω is in the range [0,1], where a larger value indicates a higher collision risk. It measures the probability that obstacle i intersects the drone's path when the drone executes the speed command (ν,ω); The relative velocity vector between the UAV and obstacle i is calculated based on the UAV velocity command (ν, ω) and the obstacle velocity information in the state estimate Ω, which measures the speed and direction of the relative motion between the UAV and the obstacle; d i (ν,ω,Ω): The predicted minimum distance between the UAV and obstacle i when the UAV adopts the speed (ν,ω). The smaller the value, the higher the potential collision risk. K: normalization constant, ensuring the appropriate range of values ​​for the fractional term, determined experimentally based on environmental characteristics and UAV dynamic parameters; ε: A positive number used as a safety parameter to prevent division by zero.

5. The method for determining a patrol path of a UAV in a complex environment according to claim 1 is characterized in that: The method of dynamically adjusting the weight of the decision network according to the dynamic environment assessment result includes the following steps: Dynamically adjust the weight of the decision network according to the system state of the path corresponding to the target state transfer matrix, including the following steps: Design and adjust the system state, output action, and reward function required by the decision-making network Actor-Critic; By adding the dynamic environment adaptability function to the traditional DWA, a time-varying environment adaptive dynamic window method is obtained; The Actor-Critic decision network and the time-varying environment adaptive dynamic window method are combined to obtain the AC-time-varying environment adaptive dynamic window method, which uses the state value estimate V(s) output by the Critic network in the decision network and adjusts the weight of the Actor network in the decision network through the advantage function, i.e., the weight parameter in the time-varying environment adaptive dynamic window method; The Actor-Critic method replaces the recurrent neural network in the traditional Actor-Critic method with the MoEMamba network. MoEMamba is a network that integrates and mixes expert networks and Mamba networks. It alternates between the two networks: one layer of mixed expert network and one layer of Mamba network, and so on, with a total of k layers. System Status t It describes all the information contained in the state of the UAV at time t, which is the basis of the AC-TEADWA method and provides the complete information required for the neural network to predict the AC-TEADWA weight. The expression of the system state is: Among them, s t : The system status at time t, which contains the complete status information of the UAV at time t; The distance measurement value from the drone to the obstacle at the current time t in the current drone coordinate system; The distance measurement k control cycles ago expressed in the current UAV coordinate system; The distance measurement value l control cycles ago expressed in the current UAV coordinate system; The distance measurement g control cycles ago expressed in the current UAV coordinate system; d t : The distance from the current position of the UAV to the target; φ t : The angle from the current position of the drone to the target; ν t : The current translation speed of the drone; ω t : The current rotation speed of the drone; The obstacle distance changes between consecutive time steps to capture dynamic information; The time derivative of the distance from the drone to the target, indicating whether it is approaching the target; The current acceleration state vector of the drone is the linear acceleration of the drone at the current time t, is the angular acceleration of the UAV at the current time t; ξ t : Environmental complexity index, calculated by the entropy value of obstacle distribution; The specific form of the distance measurement value is: τ represents the past time step and refers to k, l, and g in the formula; The superscript t-τ indicates the measurement time; The subscript t represents the coordinate system of the current UAV position; N represents the number of measurement points of the lidar, which is the distance between the drone's position and the nearest obstacle in the set direction; The system state is input into the decision network to obtain the action output by the decision network. The action is expressed as: a t =(a t ,b t ,c t ,d t ,l t ); Among them, α t : heading item weight; β t : gap term weight; γ t : speed term weight; δ t : curvature distance component, additional weight of the new extended cost function; λ t : Dynamic environment adaptability function component, additional weight of the new extended cost function; The drone receives observations or states of the environment and outputs an action a. The action is to set the cost function weights for the next control cycle. These weights will be used to generate motion commands for the drone. Design of reward function: If the drone successfully reaches the target location, it will be given a large positive reward R1; if the drone stays at the same location for multiple control cycles, it will be similar to being stuck and will be given a negative reward R2; if the drone collides with an obstacle / or other drones / birds during movement, it will be given a large penalty, negative R3; if the drone collides with a moving obstacle or other drones while stationary, it will be given a negative reward R4; the design of this reward function guides the reinforcement learning algorithm to learn safe and effective navigation strategies.

6. The method for determining a patrol path of a UAV in a complex environment according to claim 5, characterized in that: The steps for implementing the time-varying environment adaptive dynamic window method include the following steps: The dynamic environment adaptability function is added to the traditional DWA to obtain a time-varying environment adaptive dynamic window method; The objective function of the time-varying environment adaptive dynamic window method is defined as: G(ν,ω)=α·h(ν,ω)+β·c(ν,ω)+γ·v(ν,ω)+δ·d(ν,ω)+λ·e(ν,ω); Among them, ν: the linear velocity command of the UAV, which controls the speed of the UAV forward or backward; ω: angular velocity command of the UAV, which controls the speed and direction of the UAV's steering; h(ν,ω): heading component, which represents the angular difference between the drone’s heading and the target, and measures the degree to which the drone is heading towards the target; c(ν,ω): clearance component, which represents the distance from the drone to the nearest obstacle on the curvature currently considered, and measures the safety of the motion; v(ν,ω): velocity component, representing the translation speed of the drone and measuring the efficiency of the movement; d(ν,ω): curvature distance component, which represents the shortest distance from the target point to the curvature currently considered, and measures the relevance of the path to the target; e(ν,ω): Dynamic environment adaptability function component, which indicates the degree of adaptability of the current speed command (ν,ω) to the predicted dynamic changes in the environment. It measures whether the UAV can effectively cope with the expected movement of dynamic obstacles in the environment when executing the current speed command; α, β, γ, δ, λ: weights corresponding to each component; The curvature distance component d(ν,ω) represents the shortest distance from the target point to the curvature currently considered, and measures the relevance of the path to the target. The expression is as follows: α: distance attenuation factor (≥1), which controls the nonlinearity of the distance effect; β: proximity gain coefficient (0-1), which increases the reward when the target is very close to the curvature; λ: proximity sensitivity, which controls the steepness of the approach reward; δ: Directional penalty factor, considering the deviation between the curvature line direction and the target direction; θ curv : tangent direction of the curvature line; θ goal : The direction of the target point relative to the drone; robot_radius: The radius of the drone, used to normalize the distance value; distance_to_goal: the shortest distance between the motion curvature generated by the current velocity combination (ν, ω) and the target point, which is the truncation parameter; max_dist: Maximum distance threshold, used for normalization and limiting the range of distance influence.

7. The method for determining a patrol path of a UAV in a complex environment according to claim 5, characterized in that: The steps for implementing the AC-time-varying environment adaptive dynamic window method include the following steps: Obtain historical observation data and preprocess it, convert the historical observations to the current drone coordinate system, align historical observations at different times to the same reference coordinate system of the current drone, and define the coordinate system conversion formula as follows: Where, t: current time; τ: universal time offset parameter, including k, l or g, indicating the time difference; The converted distance measurement value represents the distance measurement of the historical observation before time τ in the coordinate system of the current time t; At time t-τ, the distance measurement value referenced to the drone coordinate system at time t-τ; The coordinate transformation matrix from time t-τ to time t represents the relative motion speed; Θ t : A set of state estimates of dynamic objects detected in the environment; η: mixing coefficient, dynamically adjusted with time and observation quality Control the weight of the two conversion methods; The time derivative of the coordinate transformation represents the relative motion speed. ψ(): evaluation function used to calculate the mixing coefficient η; σ(): Swish activation function; cart(): A function that converts distance scan data in polar coordinate form into Cartesian coordinate points; F(): forward projection function based on velocity prediction, considering the trajectory of moving obstacle candidates including other aircraft; pol(): A function that converts points back into the distance scan form in polar coordinates; The neural network structure in the AC-time-varying environment adaptive dynamic window method is defined as follows: 1D four-channel input; The feature extraction layer processes the spatially correlated scan data; The convolutional layer uses the Swish activation function; Feature fusion flattens the convolution output into a vector, concatenating the target distance, angle, and current speed; In the Actor-Critic architecture, the Actor and Critic networks share the same infrastructure, except for the last layer. Define the training and output decisions in the AC-time-varying environment adaptive dynamic window method: Three-stage training with different levels of difficulty; The Actor network inputs system status information, including historical observations, target position, and current speed, and outputs actions, including the weight values ​​of the five TEADWA components: heading component weight α, gap component weight β, velocity component weight γ, curvature distance component weight δ, and dynamic environment adaptability function component weight λ; The critic network outputs a state value estimate V(s), which guides the improvement of the actor through the advantage function.

8. A UAV patrol path decision-making device in a complex environment, characterized by: The device comprises: A graph construction unit, used to construct an undirected topological graph using each location to be patrolled as a node; A steady-state distribution generating unit, configured to generate a steady-state distribution of each of the nodes patrolled by the drone according to the topological constraints and node importance of the undirected topological graph; A transfer matrix generating unit, configured to generate a plurality of candidate state transfer matrices having the steady-state distribution; wherein each of the candidate state transfer matrices includes a transfer probability of the UAV between each of the nodes; An environmental assessment unit, configured to perform a dynamic environmental assessment of obstacles along the path corresponding to each candidate state transfer matrix, thereby determining an optional speed and direction of the UAV along the selected path; A command generation unit is used to dynamically adjust the weights of the decision network according to the dynamic environment assessment results, and then use the weight-adjusted decision network to generate execution movement commands; A mobile control unit is used to control the UAV to move between the various positions to be patrolled according to the execution motion command.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Unmanned aerial vehicle routing inspection real-time video transmission system based on network slices

    CN112672110A

  • Mobile robot local path planning method based on parameter adaptive dynamic window method

    CN117991786A

  • Unmanned aerial vehicle path planning method and device based on maximum entropy safety reinforcement learning

    CN118192668A

  • Multi-unmanned aerial vehicle hunting method and system

    CN120044983A

  • Heterogeneous multi-unmanned aerial vehicle cooperative path planning method based on multi-agent deep reinforcement learning

    CN120103855A

Cited By

  • Path planning method and robot

    CN120820167A

  • Target driving advancing and online avoiding method for ocean underwater robot

    CN121165757A