UAV (Unmanned Aerial Vehicle) turbulence avoidance method and system for spatio-temporal sequence prediction based on physical guidance
By constructing a physically guided spatiotemporal sequence prediction system, and combining multi-source radar data and atmospheric dynamic constraints, a future multi-step turbulence prediction field is generated, which solves the coupling problem between meteorological forecasting and trajectory decision-making in UAV autonomous navigation and enables UAVs to fly safely and efficiently in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-03-11
- Publication Date
- 2026-05-15
AI Technical Summary
In existing UAV autonomous navigation and turbulence avoidance systems, weather forecasting and flight path decision-making lack close coupling, making it difficult for the decision-making module to obtain a complete risk characterization and to effectively handle the inherent coupling relationship between physical constraints and decision uncertainty. This makes it difficult to balance flight efficiency and safety in complex environments.
A physics-guided spatiotemporal sequence prediction system is constructed. By combining three-dimensional wind field inversion, physics-guided spatiotemporal convolutional networks, and deep deterministic policy gradient algorithms with multi-source radar data and atmospheric dynamic constraints, a future multi-step turbulence prediction field is generated. This field is then embedded into a composite observation space with the dynamic state of low-altitude vehicles, a flight state transition model is defined, and evasion strategies are dynamically adjusted.
It improves the flight path safety and decision-making reliability of UAVs in complex turbulent environments, achieves risk avoidance and flight efficiency optimization while satisfying atmospheric dynamic constraints, and enhances the generalization ability and online adaptive performance of UAVs in extreme scenarios under time-varying three-dimensional meteorological fields.
Smart Images

Figure CN122043923A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of UAV trajectory planning technology, and in particular to a UAV turbulence avoidance method and system based on physical guidance and spatiotemporal sequence prediction. Background Technology
[0002] In reinforcement learning applications for autonomous navigation and turbulence avoidance of low-altitude UAVs, existing technologies typically treat weather forecasting and trajectory decision-making as independent sub-tasks. Specifically, the environmental forecasting module and the trajectory decision-making module often lack a tightly coupled collaborative optimization mechanism. The forecasting module optimizes solely to improve overall statistical accuracy, and the uncertainty information it outputs is often oversimplified or ignored when passed to the decision-making module, making it difficult for the decision-making end to obtain a complete risk characterization. Furthermore, because the decision-making module lacks explicit modeling of underlying atmospheric physical processes such as mass conservation and stability, it is prone to making infeasible decisions that violate physical laws in actual flight environments with forecast errors. In addition, the existing fragmented architecture cannot effectively handle the inherent coupling between physical constraints and decision uncertainty, making it difficult for the system to achieve a globally optimal trade-off between risk avoidance and flight efficiency, significantly increasing the risk of cascading failures when encountering sudden turbulence threats. Summary of the Invention
[0003] To address the aforementioned shortcomings, the present invention aims to propose a UAV turbulence avoidance method and system based on physics-guided spatiotemporal sequence prediction. This method constructs a physics-guided spatiotemporal convolutional network to predict three-dimensional turbulent fields and deeply couples it with reinforcement learning decision-making. This allows for dynamic adjustment of avoidance strategies based on the physical consistency confidence of the predicted field, while satisfying atmospheric dynamic constraints, thereby improving the trajectory safety and decision reliability of UAVs in complex turbulent environments.
[0004] To achieve this objective, the present invention adopts the following technical solution: UAV turbulence avoidance methods based on physics-guided spatiotemporal sequence prediction include: A three-dimensional wind field inversion system was constructed, and the acquired multi-source radar observation data was subjected to variational assimilation processing to generate a comprehensive turbulence intensity field that reflects the real-time meteorological environment. A physics-guided spatiotemporal convolutional sequence network is constructed to encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation. Physical process constraints based on atmospheric dynamics equations are introduced to pre-train the network to generate a future multi-step turbulence prediction field that conforms to physical consistency. A reinforcement learning environment based on a deep deterministic policy gradient algorithm is constructed, and the future multi-step turbulence prediction field and its corresponding physical residual confidence are embedded together with the real-time dynamic state of the low-altitude vehicle into a composite observation space. A flight state transition model including takeoff, cruise and landing phases is defined. Initialize the Actor network and Critic network in the deep deterministic policy gradient algorithm, and construct a composite reward function. The composite reward function includes a staged flight constraint term corresponding to the flight state transition model, and a physical violation penalty term whose weights are dynamically adjusted according to the physical residual confidence. The deep deterministic policy gradient algorithm is used for iterative exploration. In each iteration, the Actor network outputs a three-dimensional acceleration control action based on the current composite observation space, interacts with the reinforcement learning environment, obtains a reward value according to the composite reward function, and stores the transfer sample. The parameters of the Critic network and the Actor network are updated using the transferred samples, and the target network is synchronized using a soft update mechanism. The iterative steps are repeated until the path planning result meets the preset convergence condition, and the optimal turbulence avoidance trajectory is output.
[0005] Preferably, a three-dimensional wind field inversion system is constructed to perform variational assimilation processing on the acquired multi-source radar observation data, generating a comprehensive turbulence intensity field reflecting the real-time meteorological environment, including: By loading radial data from wind profiler radar and volumetric scan data from Doppler radar, a background field error covariance matrix is constructed using a 3D-Var variational assimilation algorithm. Covariance matrix of observation error The three-dimensional wind field state is updated by minimizing the cost function to obtain the analysis field. : ; in, Indicates the background field. Represents the observation operator. Represents the observation vector; Based on the analysis field Calculation of turbulent kinetic energy using three-dimensional wind speed components Vertical wind shear energy The turbulent kinetic energy Satisfying the relation: ; The vertical wind shear energy Satisfying the relation: ; in, , , They represent the actual wind speed at... , , Instantaneous components of three-dimensional wind speed in three directions. , , They represent , , The average wind speed in three directions under a preset spatiotemporal dimension. and These represent the gradients of the two horizontal wind speed components along the vertical direction; The turbulent kinetic energy obtained by calculation With the vertical wind shear energy Synthetic comprehensive turbulence intensity index The comprehensive turbulence intensity field is generated, and the comprehensive turbulence intensity index is... Satisfying the relation: ; in, This represents the comprehensive turbulence intensity index. and These represent the preset weighting coefficients. Represents the squared terms of the mixed length. Indicates the first Average wind speed at altitude.
[0006] Preferably, constructing a physically guided spatiotemporal convolutional sequence network to encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation includes: A physically guided network architecture is constructed, comprising vertical dimensionality reduction blocks, temporal convolutional blocks, spatial convolutional blocks, and physical constraint blocks. Factorized 3D convolutional kernels are used to extract and separate spatiotemporal features into blocks of size [size missing]. Spatial features and dimensions are Temporal characteristics; By imposing causal constraints in the time dimension and ensuring that predictions do not violate temporal causality through time reversal functions or causal convolutions, the reconstructed spatiotemporal feature map is obtained. Satisfying the relation: ; in, This represents the reconstructed spatiotemporal feature map. This represents the time reversal function. The intermediate temporary feature map representing time reversal. Indicates from the first layer to the first Layer index range marker, Represents a non-linear activation function. This represents the weight matrix of the convolutional layer. Indicates the input feature map, Represents the bias vector; Construct a physical loss function to encode the atmospheric dynamics equations into differentiable residual terms, where the mass-conserved residuals... Satisfying the relation: ; Stability residual Satisfying the relation: ; Vertical gradient residual Satisfying the relation: ; in, This represents the partial derivative of turbulence intensity with respect to time. Represents a three-dimensional velocity vector. Represents the spatial gradient operator, Represents the space Laplace operator, Represents the turbulent diffusion coefficient. This represents the reference turbulence intensity baseline value. This represents a piecewise adjustment function based on the Richardson number. Indicates atmospheric characteristic scale altitude, Indicates vertical height.
[0007] Preferably, generating a physically consistent future multi-step turbulence prediction field includes: The physically guided spatiotemporal convolutional sequence network is pre-trained using a defined overall objective function, wherein the overall loss function is defined as follows: Satisfying the relation: ; in, Represents the total loss function. This represents the data loss weighting coefficient. This represents the turbulence intensity field predicted by the network. Indicates the observed turbulence intensity field. This represents the residual weighting coefficient for mass conservation. This represents the residual due to mass conservation. This represents the stability residual weighting coefficient. Indicates stability residuals, This represents the weighting coefficient of the vertical gradient residual. Represents the vertical gradient residual. Indicates the first The weighting coefficients of the hard constraints, Indicates the first digit of the ReLU penalty form. Hard constraint loss function; Will include Historical timeline window at each time step As network input, the trained physical-guided spatiotemporal convolutional sequence network is used for inference to output the future. Turbulence intensity prediction field of step ; Freeze the physically guided spatiotemporal convolutional sequence network weights after training is complete. A low learning rate fine-tuning mode is set to ensure that the future multi-step turbulence prediction field satisfies the physical consistency constraint, which satisfies the following relationship: ; in, This represents the sum of the mass conservation residual, the stability residual, and the vertical gradient residual. This indicates the preset allowable threshold for physical consistency deviation.
[0008] Preferably, a reinforcement learning environment based on a deep deterministic policy gradient algorithm is constructed, in which the future multi-step turbulence prediction field and its corresponding physical residual confidence are embedded together with the real-time dynamic state of the low-altitude vehicle into a composite observation space, and a flight state transition model including takeoff, cruise, and landing phases is defined as follows: A dual-cascaded architecture is constructed, using the physically guided spatiotemporal convolutional sequence network as an environment cognition module, receiving the real-world field at each step. And update the history buffer. Generate a local prediction window ; Extending the observation space of the deep deterministic policy gradient algorithm, where the composite observation space Satisfying the relation: ; in, Represents the dynamic state. Indicates the spatiotemporal characteristics of the prediction field. Indicates the confidence level of the physical residual; Among them, dynamic state Satisfying the relation: ; in, This represents the three-dimensional position vector of the vehicle. This represents the three-dimensional velocity vector of the vehicle. This represents the three-dimensional acceleration vector of the vehicle. Indicates the vehicle's yaw angle. Indicates the vehicle's pitch angle; Among them, the spatiotemporal features of the prediction field Satisfying the relation: ; in, This represents the feature extraction mapping of a three-dimensional convolutional neural network; Among them, the physical residual confidence level Satisfying the relation: ; in, This represents the residual due to mass conservation. Indicates stability residuals, Represents the vertical gradient residual; The three-stage flight states in the flight state transition model are defined based on vehicle altitude, speed, and target distance: During the takeoff phase, when the conditions are met and At that time, a speed constraint prioritizing vertical climb is applied, where Indicates the vehicle's current altitude. This indicates the preset cruising altitude. Indicates the vehicle's current horizontal speed. This indicates the preset cruising speed; During the cruise phase, when the horizontal distance is met... At that time, the height hold and speed hold constraints are activated, where This represents the horizontal Euclidean distance between the vehicle and the target point. This indicates the preset stage switching distance threshold; During the descent phase, when the horizontal distance is met... At that time, the forced descent and deceleration constraints are activated.
[0009] Preferably, constructing the composite reward function includes: Calculate phased rewards The staged rewards Satisfying the relation: ; in, This indicates phased rewards. This indicates a reward during the takeoff phase. This indicates a reward during the cruise phase. Indicates rewards during the landing phase; The takeoff phase reward Satisfying the relation: ; in, and These represent the height reward weight and the speed reward weight, respectively. and These represent the preset normalization coefficients. Indicates the current altitude. Indicates the altitude of the cruise target. Represents velocity in the vertical direction. This represents a penalty function for excessive horizontal velocity. Indicates the magnitude of horizontal velocity. Indicates an indicator function; Cruise phase rewards Satisfying the relation: ; in, This indicates that the reward is maintained at a high level. This indicates the deviation between the current altitude and the planned cruising altitude. This indicates a vertical velocity stability bonus. This indicates a smooth posture reward. Indicates the pitch angle; The landing phase reward Satisfying the relation: ; in, and These represent the accuracy reward weight and the deceleration reward weight, respectively. This represents the horizontal Euclidean distance from the target landing point. Indicates the accuracy attenuation coefficient. Indicates the current airspeed. Indicates the safe landing speed threshold; Calculate security constraint rewards The security constraint reward Satisfying the relation: ; in, This indicates a three-dimensional turning reward. Based on the change in heading angle With pitch angle change Calculations show that This indicates that based on the comprehensive turbulence intensity index The calculated turbulence avoidance reward, Indicates distance from the nearest boundary Calculated boundary security reward; Introducing physical residual adaptive weights , Satisfying the relation: ; in, Indicates the adaptive weights of the physical residuals. This represents the preset sensitivity coefficient. This represents the sum of the mass conservation residual, the stability residual, and the vertical gradient residual.
[0010] Preferably, utilizing the Actor network to output three-dimensional acceleration actions based on the composite observation space, performing three-dimensional dynamic model-environment interaction, and updating network parameters includes: Get the current status of the vehicle The Actor network outputs three-dimensional acceleration commands. The three-dimensional acceleration command Satisfying the relation: ; in, This indicates the output three-dimensional acceleration command. Indicates that the Actor network parameter set is Deterministic policy mapping at time, This represents the current state of the composite observation space. This represents the sampled values from the Ornstein-Uhlenbeck exploration of the noise process; The vehicle state is updated using a three-dimensional dynamic model, where the updated velocity vector... With position vector They respectively satisfy the following relations: ; ; in, and These represent the three-dimensional velocity vectors before and after the update, respectively. and These represent the 3D position vectors before and after the update, respectively. Indicates the discrete time step. Indicates the damping coefficient; Apply acceleration and attitude angle constraints to ensure the vehicle satisfies the following relationship: ,in Indicates the magnitude of horizontal velocity. Indicates the upper limit of horizontal speed. Indicates pitch angle, This indicates the upper limit of the absolute value of the pitch angle; Calculate the total reward value And store the transferred samples to the experience replay buffer, the total reward value Satisfying the relation: ; in, Indicates the total compound reward. This indicates phased rewards. Indicates safety constraint rewards. This represents the final reward for successfully reaching the target point. This indicates the direction of arrival at the target point. The network is updated by sampling batch data from the experience replay buffer, wherein the Critic network minimizes the TD error. Update the TD error. Satisfying the relation: ; in, This represents the loss function of the Critic network. Indicates the number of samples in the batch. Represents the action value function. The target value is represented; the Actor network parameters are updated by maximizing the action value function, and a soft update mechanism is used to synchronize the target network parameters, which satisfies the following relationship: ; in, Indicates the target network parameters. Indicates the current network parameters. This represents the preset soft update coefficient.
[0011] Preferably, the output of the optimal turbulence avoidance trajectory includes: Step S61: Initialize the policy gradient network structure of the deep deterministic policy gradient algorithm, copy the current network parameters to the target network to generate the target policy, initialize the experience replay buffer, and construct Ornstein-Uhlenbeck stochastic process noise for action exploration; Step S62: At the beginning of each training round, initialize the environment round, reset the three-dimensional turbulent environment state to obtain the initial observations, and reset the noise state of the Ornstein-Uhlenbeck stochastic process; Step S63: Obtain and process the composite observation space state at the current moment, and input it into the Actor network. Generate a deterministic action vector based on the input observation, and after superimposing the sampled value of the Ornstein-Uhlenbeck random process noise, trim the action to the legal action space range to obtain the final action to be executed. Step S64: Input the final action into the three-dimensional dynamic model environment to perform the interaction, update the agent state, and obtain the instant reward value and round termination flag; Step S65: Store the current transfer sample in the experience replay buffer. When the number of samples stored in the buffer reaches a preset batch size, randomly sample batch data and update the Critic network by minimizing the mean square error. The loss function of the Critic network is... Satisfying the relation: ; in, This represents the loss function of the Critic network. Indicates the number of samples in the batch. This represents the estimated value of the action value function at the current moment. Indicates the current state in the sampled data. Indicates the action performed in the sampled data. Indicates the target Q value; Step S66: Update the Actor network using a gradient ascent strategy and perform norm clipping on the network gradient, wherein the update strategy of the Actor network aims to maximize the action value function output by the Critic network; Step S67: Gradually synchronize the target network parameters using a Polyak averaging method, perform a soft update of the target network, and then synchronize the target network parameters. Satisfying the relation: ; in, Indicates the target network parameters. Indicates the current network parameters. This represents the preset soft update coefficient; Repeat steps S61-S67 until the preset maximum number of training rounds is met or the path planning result meets the preset convergence condition, and output the optimal turbulence avoidance path.
[0012] A UAV turbulence avoidance system based on physics-guided spatiotemporal sequence prediction includes: The 3D field construction module is used to build a 3D wind field inversion system, which performs variational assimilation processing on the acquired multi-source radar observation data to generate a comprehensive turbulence intensity field that reflects the real-time meteorological environment. The physics-guided spatiotemporal sequence prediction module is used to construct a physics-guided spatiotemporal convolutional sequence network, encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation, and introduce physical process constraints based on atmospheric dynamics equations to pre-train the network to generate a future multi-step turbulence prediction field that conforms to physical consistency. The fusion sequence observation module is used to construct a reinforcement learning environment based on a deep deterministic policy gradient algorithm. It embeds the future multi-step turbulence prediction field and its corresponding physical residual confidence with the real-time dynamic state of the low-altitude vehicle into a composite observation space, and defines a flight state transition model including takeoff, cruise and landing phases. The network initialization and reward shaping module is used to initialize the Actor network and Critic network in the deep deterministic policy gradient algorithm and construct a composite reward function. The composite reward function includes a staged flight constraint term corresponding to the flight state transition model and a physical violation penalty term whose weights are dynamically adjusted according to the physical residual confidence. The iterative exploration module is used to perform iterative exploration through the deep deterministic policy gradient algorithm. In each iteration, the Actor network outputs a three-dimensional acceleration control action based on the current composite observation space and interacts with the reinforcement learning environment. It obtains a reward value according to the composite reward function and stores the transfer sample. The parameter update and trajectory decision module is used to update the parameters of the Critic network and the Actor network using the transfer samples, and to synchronize the target network using a soft update mechanism. The iterative steps are repeated until the path planning result meets the preset convergence condition, and the optimal turbulence avoidance trajectory is output.
[0013] One of the above technical solutions has the following advantages or beneficial effects: This invention constructs a three-dimensional wind field inversion system to perform variational assimilation processing on multi-source radar observation data, generating a comprehensive turbulence intensity field that reflects the real-time meteorological environment, thus providing high-quality initial field data for subsequent predictions. It constructs a physics-guided spatiotemporal convolutional sequence network to encode the historical time-series comprehensive turbulence intensity field into a spatiotemporal feature representation, and introduces physical process constraints based on atmospheric dynamics equations to pre-train the network, ensuring that the generated future multi-step turbulence prediction field meets physical consistency requirements such as mass conservation, stability, and vertical gradient, effectively improving the physical reliability of future turbulence evolution predictions. Based on a deep deterministic policy gradient algorithm-based reinforcement learning environment, the future multi-step turbulence prediction field and its corresponding physical residual confidence are embedded together with the real-time dynamic state of low-altitude vehicles into a composite observation space, defining a space encompassing takeoff, cruise, and landing. The phased flight state transition model enables the decision-making module to fully perceive prediction uncertainties and adjust its strategy accordingly. By initializing the Actor and Critic networks and constructing a composite reward function that includes phased flight constraints and a physical violation penalty term with dynamically adjusted weights based on physical residual confidence, the agent is guided to take optimal actions that conform to physical laws and adapt to prediction confidence at different flight phases. Finally, through a closed-loop training mechanism of iterative exploration, experience storage, and soft updating of network parameters, the three-dimensional acceleration control actions output by the Actor network can effectively avoid turbulent regions while ensuring flight safety. This achieves end-to-end collaborative optimization from physical constraint perception, spatiotemporal prediction to risk adaptive decision-making, significantly enhancing the UAV's generalization ability and online adaptive performance in extreme scenarios under time-varying three-dimensional meteorological environments. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0015] Figure 1 This is a flowchart of a UAV turbulence avoidance method based on physical guidance and spatiotemporal sequence prediction provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a UAV turbulence avoidance system based on physical guidance and spatiotemporal sequence prediction provided in an embodiment of the present invention. Detailed Implementation
[0016] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0017] In this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0018] A preferred embodiment of the UAV turbulence avoidance method based on physics-guided spatiotemporal sequence prediction includes the following steps: S1: Construct a three-dimensional wind field inversion system, perform variational assimilation processing on the acquired multi-source radar observation data, and generate a comprehensive turbulence intensity field that reflects the real-time meteorological environment; It should be noted that the three-dimensional wind field inversion system refers to a computational system that reconstructs the three-dimensional atmospheric wind field structure by fusing multi-source observation data. Its core lies in using variational assimilation techniques to fuse observation data from different sources and at different resolutions into a physically consistent analytical field. Multi-source radar observation data mainly includes radial data from wind profiler radar and volumetric scanning data from Doppler radar. Wind profiler radar detects the backscattered signals of precipitation particles or aerosols in the atmosphere by emitting electromagnetic waves, acquiring radial velocity information along the radar beam direction. This data has high temporal resolution but limited spatial coverage. Doppler radar volumetric scanning data, on the other hand, acquires the radial velocity distribution in three-dimensional space by performing azimuth scanning at different elevation angles. It has a wider spatial coverage but relatively lower temporal resolution. Variational assimilation is a data fusion method based on optimal estimation theory. By constructing a cost function and minimizing this function, it solves for the optimal analytical field that is consistent with both the observation data and the background field. The background field is usually provided by a priori estimates from numerical weather prediction models, the observation field is the actual radar detection data, and the analytical field is the fused optimal estimate. The comprehensive turbulence intensity field is a scalar field that characterizes the intensity of turbulence at each point in three-dimensional space. Its magnitude reflects the potential threat level of the location to UAV flight, providing an environmental understanding basis for subsequent prediction and decision-making.
[0019] Understandably, the purpose of step S1 is to establish a transformation link from raw radar observations to usable environmental cognitive information. Since single radar sources have blind spots or limited accuracy, variational assimilation and integration of multi-source data can compensate for their respective deficiencies, generating a spatially complete and physically consistent three-dimensional wind field. Furthermore, through turbulence intensity index calculation, the vector wind field is transformed into a scalar threat field, enabling subsequent prediction modules to directly process environmental features related to flight safety. This invention, on the one hand, ensures the mathematical rigor of data fusion through optimal estimation theory, making the analysis field statistically closest to the real atmospheric state; on the other hand, through turbulence intensity quantification, it transforms complex meteorological information into a unified representation usable for decision-making, providing a high-quality initial field for physically guided prediction.
[0020] S2: Construct a physics-guided spatiotemporal convolutional sequence network to encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation, and introduce physical process constraints based on atmospheric dynamics equations to pre-train the network to generate a future multi-step turbulence prediction field that conforms to physical consistency. It should be noted that the Physics-Guided Spatiotemporal Convolutional Sequence Network is a deep learning architecture that integrates data-driven learning capabilities with physical constraints. Its core components include vertical dimensionality reduction blocks, temporal convolutional blocks, spatial convolutional blocks, and physical constraint blocks. The vertical dimensionality reduction block compresses the vertical information of the three-dimensional field through convolutional operations along the height dimension, reducing the dimensionality complexity of subsequent calculations. The temporal convolutional block uses causal convolution or time reversal mechanisms to extract the temporal dependencies of turbulent evolution, ensuring that predictions do not violate temporal causality. The spatial convolutional block extracts features on the horizontal plane, capturing the spatial correlation of turbulent structures. The physical constraint block encodes atmospheric dynamics equations into differentiable loss functions, guiding the network output to satisfy physical laws. Spatiotemporal feature representation refers to the high-dimensional abstract features obtained by transforming the original three-dimensional temporal data through a neural network. These features simultaneously contain spatial structure information and temporal evolution patterns. Physical process constraints refer to transforming fundamental principles of atmospheric dynamics such as mass conservation, energy conservation, and stability conditions into mathematical residual forms, which are added as regularization terms to the network training process, ensuring that the predicted field not only fits historical observations but also inherently satisfies the structural constraints of the physical equations. Pre-training refers to the process of initially optimizing network parameters using historical datasets before formal deployment. This involves learning the relationship between the statistical laws governing turbulence evolution and physical constraints through a large number of samples. The future multi-step turbulence prediction field refers to extrapolating the spatial distribution of turbulence intensity across multiple time steps from the current moment, providing the decision-making module with forward-looking environmental information. Physical consistency refers to the degree to which the prediction field satisfies the fundamental equations of atmospheric dynamics, quantified and evaluated using indicators such as mass conservation residuals and stability residuals.
[0021] Understandably, the purpose of step S2 is to construct a prediction model with physical credibility, replacing the potentially absurd physical outputs of a purely data-driven model. Traditional neural networks learn mapping relationships solely through data fitting, which may produce predictions that violate mass or energy conservation in areas with insufficient training data coverage, such as the emergence of turbulent energy out of thin air or the appearance of negative concentrations—non-physical phenomena. By embedding physical constraint blocks into the spatiotemporal convolutional architecture and introducing atmospheric dynamics equations as soft constraints into the loss function, the network must simultaneously minimize data fitting errors and physical residuals during optimization, thereby learning a prediction pattern that conforms to both observational statistics and physical laws. This invention, on the one hand, separates spatiotemporal feature extraction through factorized convolutional design, improving the model's ability to model long-term time-series dependencies; on the other hand, it ensures the physical rationality of extrapolated predictions through physical residual constraints, maintaining atmospheric dynamic consistency in the predicted field at distant future times and avoiding physical collapse due to increased prediction step size.
[0022] S3: Construct a reinforcement learning environment based on a deep deterministic policy gradient algorithm, embed the future multi-step turbulence prediction field and its corresponding physical residual confidence into a composite observation space together with the real-time dynamic state of the low-altitude vehicle, and define a flight state transition model including takeoff, cruise and landing phases. It should be noted that the Deep Deterministic Policy Gradient (DDPG) algorithm is a deep reinforcement learning method applicable to continuous action spaces, consisting of an Actor network and a Critic network: the Actor network acts as the policy network, directly mapping states to deterministic actions; the Critic network acts as the value network, evaluating the long-term expected reward of the current state-action pair. The reinforcement learning environment refers to the external system with which the agent (here, a drone) interacts, including state transition dynamics, reward functions, and termination conditions. Physical residual confidence is a quantitative indicator characterizing the physical consistency of the prediction field, typically calculated from a combination of physical residuals such as mass conservation, stability, and vertical gradient; higher confidence indicates more reliable predictions. The composite observation space refers to extending the state representation in traditional reinforcement learning by fusing environmental prediction information, physical confidence indicators, and the vehicle's own state into a high-dimensional observation vector, enabling the decision network to simultaneously perceive external threats, prediction uncertainties, and its own motion state. Real-time dynamic states include physical quantities describing kinematic characteristics, such as the vehicle's three-dimensional position, velocity, acceleration, and attitude angles (yaw and pitch). The flight state transition model refers to the logical rules that automatically determine the flight stage (takeoff, cruise, landing) based on the vehicle's current motion parameters and trigger corresponding constraints, thereby achieving segmented management of the entire flight profile.
[0023] Understandably, the purpose of step S3 is to establish a tight coupling mechanism between the prediction module and the decision-making module, addressing the problems of distorted prediction information transmission and insufficient utilization of uncertainty in traditional fragmented architectures. By encoding and compressing the future multi-step prediction field into a high-dimensional feature vector, and cascading it with physical residual confidence and real-time dynamic state to form a composite observation, the Actor network can proactively consider future turbulence evolution trends when outputting actions, while adjusting the aggressiveness of decision-making based on prediction confidence. By defining a three-stage flight state model, the system can apply differentiated motion constraint priorities according to different stages of the flight mission, such as emphasizing vertical climb efficiency during takeoff and emphasizing position accuracy and speed control during landing, thereby achieving optimization throughout the entire mission cycle. The dual-cascade architecture of this invention forms a closed loop between environmental cognition and trajectory decision-making, with prediction information directly embedded in the input space of the policy network, avoiding information loss in intermediate links. The explicit introduction of physical residual confidence enables decision-making to have uncertainty perception capabilities, actively utilizing forward-looking information to avoid turbulence when predictions are reliable, and automatically switching to a conservative strategy when predictions are unreliable. Staged state management ensures that decisions meet the basic requirements of flight mechanics and mission logic.
[0024] S4: Initialize the Actor network and Critic network in the deep deterministic policy gradient algorithm, and construct a composite reward function. The composite reward function includes a staged flight constraint term corresponding to the flight state transition model, and a physical violation penalty term whose weights are dynamically adjusted according to the physical residual confidence. It should be noted that the Actor network is the policy approximation network in the DDPG algorithm. Its input is the current observed state, and its output is a deterministic action (in this case, a 3D acceleration command). The network structure typically employs a multilayer perceptron or a convolutional neural network. The Critic network is a value approximation network. Its input is a state-action pair, and its output is a Q-value estimate (long-term expected reward). It is used to evaluate the quality of the Actor's output action and guide its optimization direction. The composite reward function is a scalar signal formed by a weighted combination of multiple sub-reward items. It is used to quantify the agent's behavior at each interaction step, guiding the policy to evolve in the desired direction. The staged flight constraint term is a function component that applies differentiated rewards or penalties based on the current flight stage (takeoff, cruise, landing), reflecting the different requirements for state variables such as altitude, speed, and position at different mission stages. The physical violation penalty term is a negative reward applied when the vehicle's actions violate dynamic constraints or the physical consistency of the predicted field deviates significantly. It is used to suppress non-physical or dangerous decision-making behaviors. Dynamic weight adjustment refers to changing the contribution ratio of a certain reward in the total reward through function mapping based on the real-time changes in the confidence level of the physical residual, thereby achieving adaptive reward shaping for uncertainty.
[0025] Understandably, the purpose of step S4 is to establish a refined behavior evaluation mechanism, guiding the agent to learn the optimal strategy under complex constraints through multi-objective reward design. A single global reward is insufficient to characterize the phased features of a flight mission and the hierarchical nature of safety constraints. By decomposing it into phased constraint terms, the reward function can accurately express the differentiated pursuit of takeoff efficiency, cruise stability, and landing accuracy. By introducing a physical violation penalty term, the system can suppress dangerous behaviors early in the training process, accelerating convergence. Through dynamic weight adjustment, the reward function can automatically balance the trade-off between forward-looking avoidance and conservative safety based on prediction reliability, avoiding excessive reliance on erroneous predictions that could lead to risks when predictions are unreliable. The initialization of the Actor-Critic network in this invention ensures a reasonable starting point for policy optimization; the multi-component design of the composite reward makes credit allocation more explicit, reducing the learning difficulty in sparse reward environments; and the physical residual adaptive mechanism makes the reward signal environmentally adaptable, improving the generalization ability of the strategy in out-of-distribution scenarios.
[0026] S5: Iterative exploration is performed through the deep deterministic policy gradient algorithm. In each iteration, the Actor network outputs a three-dimensional acceleration control action based on the current composite observation space and interacts with the reinforcement learning environment. The reward value is obtained according to the composite reward function and the transfer sample is stored. It should be noted that iterative exploration refers to the training process in which an agent interacts with the environment in multiple rounds, gradually collecting experience and optimizing its strategy. Each round is called an episode, and each episode contains multiple time steps. Three-dimensional acceleration control actions refer to the continuous vectors output by the Actor network. This represents the acceleration command in three directions within the body coordinate system or inertial coordinate system. The velocity change can be obtained through integration, and the position can then be updated. A transition sample is a tuple that records complete information about a single-step interaction, including the current state, the action performed, the reward received, the next state, and whether the interaction has terminated. It is the basic data unit for training neural networks. Experience storage refers to the process of saving transition samples to an experience replay buffer. This buffer typically uses a fixed-capacity queue structure, where new samples overwrite old samples, to break data correlations and support random sampling training.
[0027] Understandably, the purpose of step S5 is to collect training data through online interaction to achieve incremental policy optimization. The Actor network generates actions based on the current composite observations. These actions, superimposed with exploration noise, are executed in the environment, ensuring both deterministic policy output and maintaining necessary exploratory activity. The environment updates its state and calculates rewards based on the dynamic model, forming a complete interactive closed loop. The storage of transferred samples allows for the reuse of historical experience, supporting off-policy learning and improving data efficiency. This invention maintains temporal exploration continuity through the Ornstein-Uhlenbeck process, suitable for continuous control problems of inertial systems; it achieves efficient use of samples through an experience replay mechanism, stabilizing the neural network training process; and through iterative accumulation, it gradually converges the policy from random behavior to the optimal avoidance mode.
[0028] S6: Update the parameters of the Critic network and the Actor network using the transferred samples, and synchronize the target network using a soft update mechanism. Repeat the iterative steps until the path planning result meets the preset convergence condition, and output the optimal turbulence avoidance trajectory.
[0029] It's important to note that network parameter updates refer to the process of adjusting neural network weights using gradient descent (or ascent) algorithms to optimize the objective function. Critic network updates aim to minimize temporal difference (TD) error, i.e., the mean squared error between the predicted Q-value and the target Q-value, making value estimation more accurate. Actor network updates aim to maximize the Q-value estimated by the Critic network, improving action quality through policy gradient methods. The target network refers to a lag network (Target Network) that is periodically synchronized to calculate the target Q-value, including Target-Actor and Target-Critic networks. A soft update mechanism is used to slowly track the current network, improving training stability. The soft update mechanism refers to updating the target network parameters using Polyak averaging, achieving a smooth transition of parameters. Convergence conditions are the criteria for determining training termination, typically including: reaching the maximum number of training epochs, the average reward no longer significantly increasing over the last N epochs, and the success rate stabilizing above a threshold. The optimal turbulence avoidance trajectory refers to the flight path of the Actor network from the starting point to the ending point in deterministic mode (no exploration noise) after training. This trajectory statistically achieves the optimal trade-off between risk avoidance and flight efficiency.
[0030] Understandably, the purpose of step S6 is to converge the policy to its optimal value through iterative optimization of the neural network, and to ensure the stability of the convergence process through a stable target network and a soft update mechanism. Accurate value estimation by the Critic network is the foundation for Actor network optimization; their alternating updates create synergistic improvement. The lag characteristic of the target network avoids target value chasing during bootstrapping, alleviating the instability of neural network function approximation. Soft updates further smooth target value changes compared to hard copying, making training more stable. This invention achieves efficient sample utilization through policy learning; alleviates the non-stationary target problem through a dual-network architecture and soft update mechanism; and avoids overfitting or ineffective training through convergence judgment, ultimately outputting a physically feasible and risk-controllable evasion trajectory.
[0031] Preferably, a three-dimensional wind field inversion system is constructed to perform variational assimilation processing on the acquired multi-source radar observation data, generating a comprehensive turbulence intensity field reflecting the real-time meteorological environment, including: By loading radial data from wind profiler radar and volumetric scan data from Doppler radar, a background field error covariance matrix is constructed using a 3D-Var variational assimilation algorithm. Covariance matrix of observation error The three-dimensional wind field state is updated by minimizing the cost function to obtain the analysis field. : ; in, Indicates the background field. Represents the observation operator. Represents the observation vector; Based on the analysis field Calculation of turbulent kinetic energy using three-dimensional wind speed components Vertical wind shear energy The turbulent kinetic energy Satisfying the relation: ; The vertical wind shear energy Satisfying the relation: ; in, , , They represent the actual wind speed at... , , Instantaneous components of three-dimensional wind speed in three directions. , , They represent , , The average wind speed in three directions under a preset spatiotemporal dimension. and These represent the gradients of the two horizontal wind speed components along the vertical direction; The turbulent kinetic energy obtained by calculation With the vertical wind shear energy Synthetic comprehensive turbulence intensity index The comprehensive turbulence intensity field is generated, and the comprehensive turbulence intensity index is... Satisfying the relation: ; in, This represents the comprehensive turbulence intensity index. and These represent the preset weighting coefficients. Represents the squared terms of the mixed length. .
[0032] It should be noted that wind profiler radar radial data refers to the radial velocity measurements along the radar beam direction obtained by wind profiler radar through emitting electromagnetic waves into the atmosphere and receiving the backscattered signals. This data has high temporal resolution (usually every few minutes) but only provides single-point vertical profile information, and serves as the primary source of observation for the vertical wind field structure in this step. Doppler radar volumetric scan data refers to the three-dimensional spatial radial velocity distribution data obtained by Doppler weather radar through 360-degree azimuth scanning at different elevation angles. This data has a wide horizontal coverage (up to hundreds of kilometers) but relatively low temporal resolution (usually every 6-10 minutes), and serves as the primary source of observation for the horizontal wind field structure in this step. The 3D-Var variational assimilation algorithm is a three-dimensional variational data assimilation method. By constructing and minimizing a cost function that includes background field bias and observation bias, it solves for the optimal analysis field that is consistent with both the prior background field and the actual observation. Its core lies in using the background field error covariance matrix. The spatial correlation structure of uncertainty in the background field is characterized by the observation error covariance matrix. Characterizing the random error properties of the observation instruments, both the background field and the observed field together determine the weighting of the background field and the observed field during assimilation. Analysis field This is the output of the assimilation algorithm, representing the three-dimensional wind field estimate that most closely approximates the true atmospheric state in a statistically optimal sense. The instantaneous components of the three-dimensional wind speed represent the instantaneous velocity components of atmospheric motion in the east-west, north-south, and vertical directions, respectively, and are fundamental physical quantities describing the wind field state. The average wind speed is given under a preset spatiotemporal dimension. , , This refers to the climatological or background wind speed obtained by averaging instantaneous wind speeds within a specific time window (e.g., 10 minutes) and spatial neighborhood (e.g., 5km × 5km horizontally), used to separate turbulent fluctuation components from the instantaneous field. Vertical wind shear energy. The rate of change of horizontal wind speed with height is a crucial power source for mechanical turbulence, and its square form ensures that the energy remains constantly positive. (Comprehensive turbulence intensity index) It is a dimensionless parameter that, by integrating the weighted contributions of turbulent kinetic energy (reflecting turbulence intensity) and vertical wind shear (reflecting shear generation), and normalizing it to the average wind speed, achieves a unified quantification of the degree of turbulence threat under different altitudes and climatic conditions. Weighting coefficient and Used for adjustment and The relative importance of a component in a comprehensive index is usually determined through observational data calibration or theoretical analysis. Mixed length It is a characteristic scale parameter in turbulence theory, representing the average size of turbulent eddies, and is usually related to height in the boundary layer. make It provides comparability and avoids false high turbulence indications caused solely by increased wind speed.
[0033] Understandably, wind profiler radar and Doppler radar are complementary in their observation characteristics: the former excels at high-frequency vertical monitoring but has limited horizontal coverage, while the latter excels at wide-area horizontal scanning but has limited vertical resolution. By fusing the two through 3D-Var variational assimilation, their complementary advantages can be achieved within the optimal estimation framework, generating an analysis field with complete three-dimensional spatial coverage and physical consistency. Furthermore, through... The turbulent fluctuation energy is calculated and separated by... Calculate and extract shear instability information, and finally synthesize The index transforms the vector wind field into a scalar threat field, converting meteorological information from different sources and scales into a unified format directly usable by the decision-making system. The mathematical rigor of variational assimilation ensures the statistical optimality of the analysis field; through... and The physical combination ensures the completeness of turbulence characterization (covering both thermodynamic and mechanical turbulence); through make It has cross-scenario comparability, providing a stable input distribution for subsequent generalization learning of neural networks.
[0034] Preferably, constructing a physically guided spatiotemporal convolutional sequence network to encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation includes: A physically guided network architecture is constructed, comprising vertical dimensionality reduction blocks, temporal convolutional blocks, spatial convolutional blocks, and physical constraint blocks. Factorized 3D convolutional kernels are used to extract and separate spatiotemporal features into blocks of size [size missing]. Spatial features and dimensions are Temporal characteristics; By imposing causal constraints in the time dimension and ensuring that predictions do not violate temporal causality through time reversal functions or causal convolutions, the reconstructed spatiotemporal feature map is obtained. Satisfying the relation: ; in, This represents the reconstructed spatiotemporal feature map. This represents the time reversal function. The intermediate temporary feature map representing time reversal. Indicates from the first layer to the first Layer index range marker, Represents a non-linear activation function. This represents the weight matrix of the convolutional layer. Indicates the input feature map, Represents the bias vector; Construct a physical loss function to encode the atmospheric dynamics equations into differentiable residual terms, where the mass-conserved residuals... Satisfying the relation: ; Stability residual Satisfying the relation: ; Vertical gradient residual Satisfying the relation: ; in, This represents the partial derivative of turbulence intensity with respect to time. Represents a three-dimensional velocity vector. Represents the spatial gradient operator, Represents the space Laplace operator, Represents the turbulent diffusion coefficient. This represents the reference turbulence intensity baseline value. This represents a piecewise adjustment function based on the Richardson number. Indicates atmospheric characteristic scale altitude, Indicates vertical height.
[0035] It should be noted that vertical dimensionality reduction blocks are network components that reduce the dimensionality complexity of data by performing convolution operations along the vertical dimension (height). They are typically implemented using... The convolutional kernels slide vertically, compressing multi-layer height information into fewer feature representations. In the current step, this is used to process the vertical structure of the 3D turbulent field, reducing the number of parameters required for subsequent computations. Temporal convolutional blocks are network components that extract sequence dependencies along the time dimension. The convolutional kernels perform one-dimensional convolutions along the time axis, used in the current step to capture the evolution of turbulence intensity over time. Spatial convolutional blocks are network components that extract features on the horizontal plane. The convolutional kernel slides in the xy-plane and is used in the current step to identify the spatial morphology and horizontally related features of turbulent structures. The physical constraint block is a network component that transforms atmospheric dynamics equations into a computable loss function. It calculates the physical residuals using automatic differentiation techniques and guides the network output to satisfy physical laws in the current step. The factorized 3D convolutional kernel refers to an architecture design that decomposes the traditional 3D convolutional kernel (size t×d×d) into independent convolutions in the temporal and spatial dimensions. It replaces the joint spatiotemporal convolution by cascading 1×d×d spatial convolutions and t×1×1 temporal convolutions, significantly reducing the number of parameters and improving feature extraction efficiency in the current step. Causal constraints refer to temporal constraints that ensure predictions rely only on historical and current information and do not use future data. In the current step, this is achieved through a time reversal function. This can be achieved by either reversing the sequence, performing convolution, and then reversing it again, or by causal convolution (where the kernel only covers the current and past positions), ensuring the physical feasibility of the prediction process. Mass conservation residuals are also considered. This encodes the local rate of change of turbulence intensity, the balance between advection transport and turbulent diffusion, into mathematical residuals. Its physical basis is the continuity equation of the turbulent scalar, which is used in the current step to constrain the mass conservation properties of the predicted field. Stability residuals It is a constraint form used to determine the stability of stratification and adjust the intensity of turbulence. For piecewise functions, in unstable strata ( Sufficient turbulence is allowed in stable stratification. This suppresses turbulence and is used in the current step to constrain the thermodynamic stability of the predicted field. Vertical gradient residuals. It is a mathematical form that constrains the exponential decay of turbulence intensity with height. Its theoretical basis is the boundary layer similarity theory, which is used in the current step to constrain the rationality of the vertical structure of the prediction field.
[0036] Understandably, by using factorized convolution design, the 3D convolution is decomposed into spatially and temporally independent processes, significantly reducing the number of parameters while maintaining the receptive field, enabling the network to handle high-resolution 3D turbulent field sequences. Through causal constraint mechanisms, each prediction step is ensured to strictly follow the arrow of time, avoiding false accuracy caused by information leakage and making the prediction physically achievable. By using joint constraints of three types of physical residuals, fundamental principles of atmospheric dynamics such as mass conservation, thermodynamic stability, and vertical structure are embedded into the network training process, allowing the network to not only learn statistical correlations in the data but also intrinsically grasp the structural constraints of physical equations, thereby improving the physical consistency of far-future predictions. Factorization design improves computational efficiency and scalability, causal constraints guarantee the temporal logical legitimacy of the predictions, and physical residual constraints suppress the generation of non-physical predictions, enhancing the model's generalization ability and credibility in out-of-distribution scenarios.
[0037] Preferably, generating a physically consistent future multi-step turbulence prediction field includes: The physically guided spatiotemporal convolutional sequence network is pre-trained using a defined overall objective function, wherein the overall loss function is defined as follows: Satisfying the relation: ; in, Represents the total loss function. This represents the data loss weighting coefficient. This represents the turbulence intensity field predicted by the network. Indicates the observed turbulence intensity field. This represents the residual weighting coefficient for mass conservation. This represents the residual due to mass conservation. This represents the stability residual weighting coefficient. Indicates stability residuals, This represents the weighting coefficient of the vertical gradient residual. Represents the vertical gradient residual. Indicates the first The weighting coefficients of the hard constraints, Indicates the first digit of the ReLU penalty form. Hard constraint loss function; Will include Historical timeline window at each time step As network input, the trained physical-guided spatiotemporal convolutional sequence network is used for inference to output the future. Turbulence intensity prediction field of step ; Freeze the physically guided spatiotemporal convolutional sequence network weights after training is complete. A low learning rate fine-tuning mode is set to ensure that the future multi-step turbulence prediction field satisfies the physical consistency constraint, which satisfies the following relationship: ; in, This represents the sum of the mass conservation residual, the stability residual, and the vertical gradient residual. This indicates the preset allowable threshold for physical consistency deviation.
[0038] It should be noted that the overall objective function is a scalar function guiding the optimization of neural network parameters. It consists of a weighted average of data fitting terms and physical constraint terms, balancing prediction accuracy and physical reliability in the current step. (Weight coefficients) , , , The relative importance of data loss and each physical residual in the overall objective is controlled separately. Their values reflect the trade-off between prediction accuracy and physical consistency, and are determined in this step through hyperparameter tuning or uncertainty quantification methods. Hard constraint loss function. It is a penalty term designed for inviolable physical conditions (such as nonnegativity and boundedness), and adopts the ReLU form. or In the current step, ensure that the predicted field satisfies rigid boundary conditions. Historical time series window. It is the observation sequence of the past T time steps input to the network, serving as the initial condition for predicting the future state, and providing contextual information for the temporal evolution in the current step. Future Turbulence intensity prediction field of step It is the core output of the network, representing the future from the current moment. The spatiotemporal distribution of three-dimensional turbulence intensity at each time step provides the decision-making module with forward-looking environmental awareness in the current step. Network weights This is a fixed set of parameters after training. The freeze operation prevents gradient updates, ensuring the stability and reproducibility of the prediction module in the current step. The low learning rate fine-tuning mode is a mechanism that, after freezing the backbone network, fine-tunes only specific layers (such as the output layer or batch normalization layer) with an extremely small learning rate (e.g., 1e-6). In the current step, it is used to fine-tune the prediction field to meet a strict physical consistency threshold. Physical consistency deviation tolerance threshold. It is the critical value for determining whether the predicted field meets the physical requirements. When the residual sum is lower than this threshold, the prediction is considered to be physically consistent and is used as a quality control standard in the current step.
[0039] Understandably, through the design of a multi-objective loss function, the network must simultaneously approximate the observed data and reduce the physical residuals during the optimization process, learning a consistent representation of data distribution and physical laws. The input mechanism of historical time-series windows ensures that predictions are based on sufficient past context, improving the accuracy of time-series extrapolation. The weight freezing and fine-tuning mechanism allows for fine-tuning of physical constraints while maintaining pre-trained knowledge, ensuring that the output strictly meets the preset physical consistency standard. Multi-objective training gives the network dual optimization objectives, avoiding physical absurdities caused by pure data fitting; the time-series window input provides evolutionary context, supporting multi-step prediction; the freeze-fine-tuning mechanism balances stability and fine-tuning requirements, ensuring that the physical credibility of the predicted field meets decision-making requirements.
[0040] Preferably, a reinforcement learning environment based on a deep deterministic policy gradient algorithm is constructed, in which the future multi-step turbulence prediction field and its corresponding physical residual confidence are embedded together with the real-time dynamic state of the low-altitude vehicle into a composite observation space, and a flight state transition model including takeoff, cruise, and landing phases is defined as follows: A dual-cascaded architecture is constructed, using the physically guided spatiotemporal convolutional sequence network as an environment cognition module, receiving the real-world field at each step. And update the history buffer. Generate a local prediction window ; Extending the observation space of the deep deterministic policy gradient algorithm, where the composite observation space Satisfying the relation: ; in, Represents the dynamic state. Indicates the spatiotemporal characteristics of the prediction field. Indicates the confidence level of the physical residual; Among them, dynamic state Satisfying the relation: ; in, This represents the three-dimensional position vector of the vehicle. This represents the three-dimensional velocity vector of the vehicle. This represents the three-dimensional acceleration vector of the vehicle. Indicates the vehicle's yaw angle. Indicates the vehicle's pitch angle; Among them, the spatiotemporal features of the prediction field Satisfying the relation: ; in, This represents the feature extraction mapping of a three-dimensional convolutional neural network; Among them, the physical residual confidence level Satisfying the relation: ; in, This represents the residual due to mass conservation. Indicates stability residuals, Represents the vertical gradient residual; The three-stage flight states in the flight state transition model are defined based on vehicle altitude, speed, and target distance: During the takeoff phase, when the conditions are met and At that time, a speed constraint prioritizing vertical climb is applied, where Indicates the vehicle's current altitude. This indicates the preset cruising altitude. Indicates the vehicle's current horizontal speed. This indicates the preset cruising speed; During the cruise phase, when the horizontal distance is met... At that time, the height hold and speed hold constraints are activated, where This represents the horizontal Euclidean distance between the vehicle and the target point. This indicates the preset stage switching distance threshold; During the descent phase, when the horizontal distance is met... At that time, the forced descent and deceleration constraints are activated.
[0041] It should be noted that the dual-cascade architecture refers to a system structure that strings together the environmental cognition module (prediction network) and the decision-making module (reinforcement learning). The output of the upstream module serves as the input of the downstream module, forming an end-to-end information flow, achieving tight integration of prediction and decision-making in the current step. Historical buffer. It is a circular queue storing the real turbulent field from the most recent few time steps, used to construct the input time window of the prediction network, maintaining temporal continuity in the current step. Local prediction window. It is a predicted field clipped to the area surrounding the vehicle's current location, rather than a global field, reducing computational burden and focusing on local threats in the current step. Composite observation space. It is a high-dimensional vector that extends the traditional reinforcement learning state representation, integrating ontology state, environmental prediction, and uncertainty information to provide the policy network with comprehensive decision-making support in the current step. (3D position vector) Describes the spatial coordinates of the vehicle in the inertial coordinate system, and its three-dimensional velocity vector. Describes the speed and direction of motion; three-dimensional acceleration vector. Describes the rate of change of velocity, yaw angle Describes the direction of the aircraft's nose in the horizontal plane (usually measured clockwise with north as the reference), and the vehicle's pitch angle. Describing the angle between the body's longitudinal axis and the horizontal plane, these dynamic states collectively describe the vehicle's complete kinematic configuration, supporting accurate dynamic modeling and constraint application in the current step. 3D convolutional neural network feature extraction mapping. This is a function that compresses the 3D spatiotemporal prediction field into a low-dimensional feature vector. It typically contains several 3D convolutional layers, pooling layers, and fully connected layers, transforming high-dimensional prediction information into a form that the policy network can process in the current step. (Stage switching distance threshold) It is the horizontal distance threshold for determining whether to enter the landing phase. It is usually set based on the airport's airspace clearance or the drone's deceleration capability, and enables automatic switching of flight phases in the current step.
[0042] Understandably, the dual-cascade architecture allows predictive information to be directly embedded into the policy network's observation space, avoiding information transmission losses inherent in traditional fragmented architectures. Through the design of a composite observation space, the decision network can simultaneously perceive the current motion state, future environmental threats, and prediction reliability, enabling decision-making based on uncertainty awareness. The three-stage flight state model allows the system to automatically adjust constraint priorities according to mission progress, ensuring that behavior at each stage conforms to flight mechanics and mission requirements. The dual-cascade architecture eliminates the information bottleneck between prediction and decision-making; the composite observation space supports policy learning with uncertainty awareness; and staged state management ensures the rationality and safety of behavior throughout the entire mission lifecycle.
[0043] Preferably, constructing the composite reward function includes: Calculate phased rewards The staged rewards Satisfying the relation: ; in, This indicates phased rewards. This indicates a reward during the takeoff phase. This indicates a reward during the cruise phase. Indicates rewards during the landing phase; The takeoff phase reward Satisfying the relation: ; in, and These represent the height reward weight and the speed reward weight, respectively. and These represent the preset normalization coefficients. Indicates the current altitude. Indicates the altitude of the cruise target. Represents velocity in the vertical direction. This represents a penalty function for excessive horizontal velocity. Indicates the magnitude of horizontal velocity. Indicates an indicator function; Cruise phase rewards Satisfying the relation: ; in, This indicates that the reward is maintained at a high level. This indicates the deviation between the current altitude and the planned cruising altitude. This indicates a vertical velocity stability bonus. This indicates a smooth posture reward. Indicates the pitch angle; The landing phase reward Satisfying the relation: ; in, and These represent the accuracy reward weight and the deceleration reward weight, respectively. This represents the horizontal Euclidean distance from the target landing point. Indicates the accuracy attenuation coefficient. Indicates the current airspeed. Indicates the safe landing speed threshold; Calculate security constraint rewards The security constraint reward Satisfying the relation: ; in, This indicates a three-dimensional turning reward. Based on the change in heading angle With pitch angle change Calculations show that This indicates that based on the comprehensive turbulence intensity index The calculated turbulence avoidance reward, Indicates distance from the nearest boundary Calculated boundary security reward; Introducing physical residual adaptive weights , Satisfying the relation: ; in, Indicates the adaptive weights of the physical residuals. This represents the preset sensitivity coefficient. This represents the sum of the mass conservation residual, the stability residual, and the vertical gradient residual.
[0044] It should be noted that phased rewards are reward components that apply differentiated incentives based on the current flight phase. They reflect the different priority requirements for state variables such as altitude, speed, and position at different mission phases, guiding the strategy to learn behavioral patterns consistent with the flight profile in the current step. Takeoff phase rewards. The aim is to incentivize rapid ascent to a safe altitude while suppressing high-speed low-altitude flight, with altitude bonuses included. Encourage increased altitude and speed bonus items Encourage vertical ascent speed, penalty items When the altitude is below 30% of the cruising altitude, horizontal speed is suppressed to prevent the risk of collisions caused by high speeds at low altitudes. Cruise phase bonus Aimed at maintaining a stable cruise status, altitude retention bonus Penalty for height deviation, reward for vertical speed stability Suppressing vertical velocity fluctuations and smoothing attitude rewards Encourages low pitch angle flight. The landing phase reward (Reward_landing) aims to achieve precise landing, with accuracy-based reward items. Gaussian form is used to excite the approach to the target point, with a deceleration bonus. Positive incentives are provided when the speed falls below a threshold to jointly guide a safe approach. Safety constraint reward. It is a reward component that ensures flight safety from three dimensions: handling quality, environmental threats, and spatial constraints; a three-dimensional steering reward. Punishing drastic changes in posture to improve ride quality and energy efficiency, turbulence avoidance rewards To balance safety and efficiency, a strong penalty is imposed for entering high-risk turbulence zones, while a weak reward is given for remaining in low-turbulence zones, with a boundary safety reward. Prevent vehicles from approaching geographical boundaries or obstacles. Physical residual adaptive weights. It is a weighting factor that dynamically adjusts the aggressiveness of decision-making based on the physical consistency of the predicted field, and adopts an exponential decay form, when the residuals and Small (reliable prediction) hours When the residual sum is large (prediction is unreliable) In the current step, strategy adjustments are made based on uncertainty awareness.
[0045] Understandably, the phased reward component ensures that the strategy adopts actions consistent with mission logic at different flight phases, avoiding the problem that a single global reward is insufficient to characterize phased features; the safety constraint reward component suppresses dangerous behaviors from multiple dimensions, establishing safety boundary awareness early in training and accelerating convergence; the physical residual adaptive weights achieve uncertainty adaptation in reward shaping, enabling the strategy to dynamically balance forward avoidance and conservative safety based on prediction reliability, avoiding excessive reliance on erroneous information leading to risks when predictions are unreliable. The multi-component reward structure makes credit allocation more explicit, reducing the learning difficulty in sparse reward environments; the phased design ensures that behavior conforms to flight profile requirements; safety constraints establish a multi-layered protection mechanism; and the physical adaptive mechanism enhances the robustness of the strategy under prediction uncertainty conditions.
[0046] Preferably, utilizing the Actor network to output three-dimensional acceleration actions based on the composite observation space, performing three-dimensional dynamic model-environment interaction, and updating network parameters includes: Get the current status of the vehicle The Actor network outputs three-dimensional acceleration commands. The three-dimensional acceleration command Satisfying the relation: ; in, This indicates the output three-dimensional acceleration command. Indicates that the Actor network parameter set is Deterministic policy mapping at time, This represents the current state of the composite observation space. This represents the sampled values from the Ornstein-Uhlenbeck exploration of the noise process; The vehicle state is updated using a three-dimensional dynamic model, where the updated velocity vector... With position vector They respectively satisfy the following relations: ; ; in, and These represent the three-dimensional velocity vectors before and after the update, respectively. and These represent the 3D position vectors before and after the update, respectively. Indicates the discrete time step. Indicates the damping coefficient; Apply acceleration and attitude angle constraints to ensure the vehicle satisfies the following relationship: ,in Indicates the magnitude of horizontal velocity. Indicates the upper limit of horizontal speed. Indicates pitch angle, This indicates the upper limit of the absolute value of the pitch angle; Calculate the total reward value And store the transferred samples to the experience replay buffer, the total reward value Satisfying the relation: ; in, Indicates the total compound reward. This indicates phased rewards. Indicates safety constraint rewards. This represents the final reward for successfully reaching the target point. This indicates the direction of arrival at the target point. The network is updated by sampling batch data from the experience replay buffer, wherein the Critic network minimizes the TD error. Update the TD error. Satisfying the relation: ; in, This represents the loss function of the Critic network. Indicates the number of samples in the batch. Represents the action value function. The target value is represented; the Actor network parameters are updated by maximizing the action value function, and a soft update mechanism is used to synchronize the target network parameters, which satisfies the following relationship: ; in, Indicates the target network parameters. Indicates the current network parameters. This represents the preset soft update coefficient.
[0047] It should be noted that deterministic policy mapping An Actor network is a function approximation from an observed state to a deterministic action. It typically employs a multilayer perceptron structure, where the output layer is mapped to the action space via tanh activation, providing the optimal action estimate for a given state in the current step. Ornstein-Uhlenbeck exploration noise is also relevant. It is a time-dependent stochastic process that adds a smooth exploratory perturbation to the action in the current step, making it suitable for continuous control of inertial systems. Discrete time step. The sampling period of the simulation or control system determines the time granularity of state updates, affecting the accuracy and computational efficiency of the dynamic model in the current step. Damping coefficient. Characterizing energy dissipation due to air resistance or system friction, this makes velocity updates more consistent with real physics in the current step, avoiding unbounded acceleration. Horizontal velocity upper limit. The maximum permissible speed and pitch angle limit are set based on the vehicle's aerodynamic performance or regulatory restrictions. Based on the maximum permissible pitch angle set during the flight phase (e.g., a larger climb angle is allowed during takeoff, and a smaller angle is limited during cruise), the physical feasibility and flight safety of the action are ensured in the current step. The experience replay buffer is a finite-capacity queue storing historical transfer samples, supporting random sampling to break data correlations and enabling efficient reuse of samples and offline policy learning in the current step. Action value function. It is the state-action pair of the Critic network that estimates the long-term reward and target value. Composed of the current reward and the discounted target network estimate, it provides the objective for temporal difference learning in the current step. Soft update coefficients. It is a small amount (usually 0.001-0.01) that controls the rate at which the target network parameters are mixed with the current network parameters, ensuring that the target value changes slowly in the current step and improving training stability.
[0048] Understandably, the deterministic output of the Actor network, superimposed with OU noise, balances the utilization and exploration of strategies; updating the state through a 3D dynamic model ensures consistency between the simulation environment and real physics; imposing constraints guarantees the legality and safety of actions; experience replay and TD learning achieve efficient sample utilization and accurate estimation of the value function; and a soft update mechanism avoids training instability caused by drastic fluctuations in the target value. OU noise provides smooth exploration suitable for continuous control; the dynamic model ensures the physical credibility of state transitions; constraint handling guarantees the feasibility of actions; experience replay and TD learning achieve stable value estimation; and the soft update mechanism improves the stability of training convergence.
[0049] Preferably, the output of the optimal turbulence avoidance trajectory includes: Step S61: Initialize the policy gradient network structure of the deep deterministic policy gradient algorithm, copy the current network parameters to the target network to generate the target policy, initialize the experience replay buffer, and construct Ornstein-Uhlenbeck stochastic process noise for action exploration; Step S62: At the beginning of each training round, initialize the environment round, reset the three-dimensional turbulent environment state to obtain the initial observations, and reset the noise state of the Ornstein-Uhlenbeck stochastic process; Step S63: Obtain and process the composite observation space state at the current moment, and input it into the Actor network. Generate a deterministic action vector based on the input observation, and after superimposing the sampled value of the Ornstein-Uhlenbeck random process noise, trim the action to the legal action space range to obtain the final action to be executed. Step S64: Input the final action into the three-dimensional dynamic model environment to perform the interaction, update the agent state, and obtain the instant reward value and round termination flag; Step S65: Store the current transfer sample in the experience replay buffer. When the number of samples stored in the buffer reaches a preset batch size, randomly sample batch data and update the Critic network by minimizing the mean square error. The loss function of the Critic network is... Satisfying the relation: ; in, This represents the loss function of the Critic network. Indicates the number of samples in the batch. This represents the estimated value of the action value function at the current moment. Indicates the current state in the sampled data. Indicates the action performed in the sampled data. Indicates the target Q value; Step S66: Update the Actor network using a gradient ascent strategy and perform norm clipping on the network gradient, wherein the update strategy of the Actor network aims to maximize the action value function output by the Critic network; Step S67: Gradually synchronize the target network parameters using a Polyak averaging method, perform a soft update of the target network, and then synchronize the target network parameters. Satisfying the relation: ; in, Indicates the target network parameters. Indicates the current network parameters. This represents the preset soft update coefficient; Repeat steps S61-S67 until the preset maximum number of training rounds is met or the path planning result meets the preset convergence condition, and output the optimal turbulence avoidance path.
[0050] It should be noted that the policy gradient network structure refers to the topological configuration of the Actor and Critic networks in the DDPG algorithm, including hyperparameters such as the number of layers, units, and activation functions. In this step, it needs to be designed appropriately based on task complexity to ensure expressive power. The target policy is a lag policy parameterized by the target network, used to calculate the target Q-value, providing a stable learning objective in this step. Round initialization is the operation of resetting the environment and noise state at the beginning of each training round, ensuring the independence and comparability of each round, and maintaining the statistical validity of training in this step. The round termination flag is the signal that determines the end of the current interaction sequence, including conditions such as successfully reaching the target, a collision, exceeding the boundary, and timeout, defining the scope of a single learning iteration in this step. The batch size is the number of samples sampled from the experience buffer each time the network is updated, affecting the variance and computational efficiency of gradient estimation in this step. The target Q-value is the regression objective used to supervise the Critic network's learning, composed of the current reward and the discounted target network estimate, implementing temporal difference learning in this step. Gradient norm clipping is a regularization technique that imposes an upper bound on the gradient of network parameters. Scaling is performed when the gradient norm exceeds a threshold, preventing training instability caused by gradient explosion in the current step. Polyak averaging is the mathematical expression of a soft update mechanism. It mixes current and target parameters through exponential moving averages, achieving a smooth parameter transition in the current step. Convergence criteria are comprehensive criteria for determining training termination, which may include maximum number of epochs, average reward, success rate threshold, etc., preventing overfitting and ineffective training in the current step. The optimal turbulence avoidance path is the three-dimensional trajectory from the starting point to the ending point generated by the Actor network in deterministic mode (no exploration noise) after training, representing the final output of policy learning.
[0051] Understandably, a stable training pipeline is built through a complete initialization-interaction-storage-update-synchronization loop. Round initialization and termination flags define independent learning units, supporting task learning. Batch sampling and gradient updates enable incremental optimization of neural network parameters. Gradient pruning and soft updates ensure numerical stability during training. Multi-dimensional convergence judgment terminates training and outputs the optimal policy at appropriate times. The standardized training process ensures systematic learning; convergence judgment avoids resource waste and overfitting; and the final optimal trajectory represents the best performance of the policy on the training distribution, which can be directly used for deployment or as a reference trajectory.
[0052] UAV turbulence avoidance system based on physics-guided spatiotemporal sequence prediction, such as Figure 2 As shown, it includes: The 3D field construction module is used to build a 3D wind field inversion system, which performs variational assimilation processing on the acquired multi-source radar observation data to generate a comprehensive turbulence intensity field that reflects the real-time meteorological environment. The physics-guided spatiotemporal sequence prediction module is used to construct a physics-guided spatiotemporal convolutional sequence network, encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation, and introduce physical process constraints based on atmospheric dynamics equations to pre-train the network to generate a future multi-step turbulence prediction field that conforms to physical consistency. The fusion sequence observation module is used to construct a reinforcement learning environment based on a deep deterministic policy gradient algorithm. It embeds the future multi-step turbulence prediction field and its corresponding physical residual confidence with the real-time dynamic state of the low-altitude vehicle into a composite observation space, and defines a flight state transition model including takeoff, cruise and landing phases. The network initialization and reward shaping module is used to initialize the Actor network and Critic network in the deep deterministic policy gradient algorithm and construct a composite reward function. The composite reward function includes a staged flight constraint term corresponding to the flight state transition model and a physical violation penalty term whose weights are dynamically adjusted according to the physical residual confidence. The iterative exploration module is used to perform iterative exploration through the deep deterministic policy gradient algorithm. In each iteration, the Actor network outputs a three-dimensional acceleration control action based on the current composite observation space and interacts with the reinforcement learning environment. It obtains a reward value according to the composite reward function and stores the transfer sample. The parameter update and trajectory decision module is used to update the parameters of the Critic network and the Actor network using the transfer samples, and to synchronize the target network using a soft update mechanism. The iterative steps are repeated until the path planning result meets the preset convergence condition, and the optimal turbulence avoidance trajectory is output.
[0053] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0054] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A UAV turbulence avoidance method based on physically-guided spatiotemporal sequence prediction, characterized in that, include: A three-dimensional wind field inversion system was constructed, and the acquired multi-source radar observation data was subjected to variational assimilation processing to generate a comprehensive turbulence intensity field that reflects the real-time meteorological environment. A physics-guided spatiotemporal convolutional sequence network is constructed to encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation. Physical process constraints based on atmospheric dynamics equations are introduced to pre-train the network to generate a future multi-step turbulence prediction field that conforms to physical consistency. A reinforcement learning environment based on a deep deterministic policy gradient algorithm is constructed, and the future multi-step turbulence prediction field and its corresponding physical residual confidence are embedded together with the real-time dynamic state of the low-altitude vehicle into a composite observation space. A flight state transition model including takeoff, cruise and landing phases is defined. Initialize the Actor network and Critic network in the deep deterministic policy gradient algorithm, and construct a composite reward function. The composite reward function includes a staged flight constraint term corresponding to the flight state transition model, and a physical violation penalty term whose weights are dynamically adjusted according to the physical residual confidence. The deep deterministic policy gradient algorithm is used for iterative exploration. In each iteration, the Actor network outputs a three-dimensional acceleration control action based on the current composite observation space, interacts with the reinforcement learning environment, obtains a reward value according to the composite reward function, and stores the transfer sample. The parameters of the Critic network and the Actor network are updated using the transferred samples, and the target network is synchronized using a soft update mechanism. The iterative steps are repeated until the path planning result meets the preset convergence condition, and the optimal turbulence avoidance trajectory is output.
2. The physically-guided spatio-temporal sequence prediction based UAV turbulence avoidance method of claim 1, wherein, A three-dimensional wind field inversion system is constructed, and the acquired multi-source radar observation data is subjected to variational assimilation processing to generate a comprehensive turbulence intensity field reflecting the real-time meteorological environment, including: The radial data of wind profile radar and the volume scanning data of Doppler radar are loaded, a background field error covariance matrix is constructed by a 3D-Var variation assimilation algorithm and an observation error covariance matrix , and a three-dimensional wind field state is updated by minimizing a cost function, so that an analysis field is obtained : ; in, Indicates the background field. Represents the observation operator. Represents the observation vector; Based on the analysis field Calculation of turbulent kinetic energy using three-dimensional wind speed components Vertical wind shear energy The turbulent kinetic energy Satisfying the relation: ; The vertical wind shear energy Satisfying the relation: ; in, , , They represent the actual wind speed at... , , Instantaneous components of three-dimensional wind speed in three directions. , , They represent , , The average wind speed in three directions under a preset spatiotemporal dimension. and These represent the gradients of the two horizontal wind speed components along the vertical direction; The turbulent kinetic energy obtained by calculation With the vertical wind shear energy Synthetic comprehensive turbulence intensity index The comprehensive turbulence intensity field is generated, and the comprehensive turbulence intensity index is... Satisfying the relation: ; in, This represents the comprehensive turbulence intensity index. and These represent the preset weighting coefficients. Represents the squared terms of the mixed length. Indicates the first Average wind speed at altitude.
3. The UAV turbulence avoidance method based on physically guided spatiotemporal sequence prediction according to claim 1, characterized in that, Constructing a physically guided spatiotemporal convolutional sequence network to encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation includes: A physically guided network architecture is constructed, comprising vertical dimensionality reduction blocks, temporal convolutional blocks, spatial convolutional blocks, and physical constraint blocks. Factorized 3D convolutional kernels are used to extract and separate spatiotemporal features into blocks of size [size missing]. Spatial features and dimensions are Temporal characteristics; By imposing causal constraints in the time dimension and ensuring that predictions do not violate temporal causality through time reversal functions or causal convolutions, the reconstructed spatiotemporal feature map is obtained. Satisfying the relation: ; in, This represents the reconstructed spatiotemporal feature map. This represents the time reversal function. The intermediate temporary feature map representing time reversal. Indicates from the first layer to the first Layer index range marker, Represents a non-linear activation function. This represents the weight matrix of the convolutional layer. Indicates the input feature map, Represents the bias vector; Construct a physical loss function to encode the atmospheric dynamics equations into differentiable residual terms, where the mass-conserved residuals... Satisfying the relation: ; Stability residual Satisfying the relation: ; Vertical gradient residual Satisfying the relation: ; in, This represents the partial derivative of turbulence intensity with respect to time. Represents a three-dimensional velocity vector. Represents the spatial gradient operator, Represents the space Laplace operator, Represents the turbulent diffusion coefficient. This represents the reference turbulence intensity baseline value. This represents a piecewise adjustment function based on the Richardson number. Indicates atmospheric characteristic scale altitude, Indicates vertical height.
4. The UAV turbulence avoidance method based on physical-guided spatiotemporal sequence prediction according to claim 1, characterized in that, Generating a physically consistent future multi-step turbulence prediction field includes: The physically guided spatiotemporal convolutional sequence network is pre-trained using a defined overall objective function, wherein the overall loss function is defined as follows: Satisfying the relation: ; in, Represents the total loss function. This represents the data loss weighting coefficient. This represents the turbulence intensity field predicted by the network. Indicates the observed turbulence intensity field. This represents the residual weighting coefficient for mass conservation. This represents the residual due to mass conservation. This represents the stability residual weighting coefficient. Indicates stability residuals, This represents the weighting coefficient of the vertical gradient residual. Represents the vertical gradient residual. Indicates the first The weighting coefficients of the hard constraints, Indicates the first digit of the ReLU penalty form. Hard constraint loss function; Will include Historical timeline window at each time step As network input, the trained physical-guided spatiotemporal convolutional sequence network is used for inference to output the future. Turbulence intensity prediction field of step ; Freeze the physically guided spatiotemporal convolutional sequence network weights after training is complete. A low learning rate fine-tuning mode is set to ensure that the future multi-step turbulence prediction field satisfies the physical consistency constraint, which satisfies the following relationship: ; in, This represents the sum of the mass conservation residual, the stability residual, and the vertical gradient residual. This indicates the preset allowable threshold for physical consistency deviation.
5. The UAV turbulence avoidance method based on physically guided spatiotemporal sequence prediction according to claim 1, characterized in that, A reinforcement learning environment based on a deep deterministic policy gradient algorithm is constructed, embedding the future multi-step turbulence prediction field and its corresponding physical residual confidence level together with the real-time dynamic state of the low-altitude vehicle into a composite observation space. A flight state transition model including takeoff, cruise, and landing phases is defined, comprising: A dual-cascaded architecture is constructed, using the physically guided spatiotemporal convolutional sequence network as an environment cognition module, receiving the real-world field at each step. And update the history buffer. Generate a local prediction window ; Extending the observation space of the deep deterministic policy gradient algorithm, where the composite observation space Satisfying the relation: ; in, Represents the dynamic state. Indicates the spatiotemporal characteristics of the prediction field. Indicates the confidence level of the physical residual; Among them, dynamic state Satisfying the relation: ; in, This represents the three-dimensional position vector of the vehicle. This represents the three-dimensional velocity vector of the vehicle. This represents the three-dimensional acceleration vector of the vehicle. Indicates the vehicle's yaw angle. Indicates the vehicle's pitch angle; Among them, the spatiotemporal features of the prediction field Satisfying the relation: ; in, This represents the feature extraction mapping of a three-dimensional convolutional neural network; Among them, the physical residual confidence level Satisfying the relation: ; in, This represents the residual due to mass conservation. Indicates stability residuals, Represents the vertical gradient residual; The three-stage flight states in the flight state transition model are defined based on vehicle altitude, speed, and target distance: During the takeoff phase, when the conditions are met and At that time, a speed constraint prioritizing vertical climb is applied, where Indicates the vehicle's current altitude. This indicates the preset cruising altitude. Indicates the vehicle's current horizontal speed. This indicates the preset cruising speed; During the cruise phase, when the horizontal distance is met... At that time, the height hold and speed hold constraints are activated, where This represents the horizontal Euclidean distance between the vehicle and the target point. This indicates the preset stage switching distance threshold; During the descent phase, when the horizontal distance is met... At that time, the forced descent and deceleration constraints are activated.
6. The UAV turbulence avoidance method based on physical-guided spatiotemporal sequence prediction according to claim 1, characterized in that, Constructing a compound reward function includes: Calculate phased rewards The staged rewards Satisfying the relation: ; in, This indicates phased rewards. This indicates a reward during the takeoff phase. This indicates a reward during the cruise phase. Indicates rewards during the landing phase; The takeoff phase reward Satisfying the relation: ; in, and These represent the height reward weight and the speed reward weight, respectively. and These represent the preset normalization coefficients. Indicates the current altitude. Indicates the altitude of the cruise target. Represents velocity in the vertical direction. This represents a penalty function for excessive horizontal velocity. Indicates the magnitude of horizontal velocity. Indicates an indicator function; Cruise phase rewards Satisfying the relation: ; in, This indicates that the reward is maintained at a high level. This indicates the deviation between the current altitude and the planned cruising altitude. This indicates a vertical velocity stability bonus. This indicates a smooth posture reward. Indicates the pitch angle; The landing phase reward Satisfying the relation: ; in, and These represent the accuracy reward weight and the deceleration reward weight, respectively. This represents the horizontal Euclidean distance from the target landing point. Indicates the accuracy attenuation coefficient. Indicates the current airspeed. Indicates the safe landing speed threshold; Calculate security constraint rewards The security constraint reward Satisfying the relation: ; in, This indicates a three-dimensional turning reward. Based on the change in heading angle With pitch angle change Calculations show that This indicates that based on the comprehensive turbulence intensity index The calculated turbulence avoidance reward, Indicates distance from the nearest boundary Calculated boundary security reward; Introducing physical residual adaptive weights , Satisfying the relation: ; in, Indicates the adaptive weights of the physical residuals. This represents the preset sensitivity coefficient. This represents the sum of the mass conservation residual, the stability residual, and the vertical gradient residual.
7. The UAV turbulence avoidance method based on physically guided spatiotemporal sequence prediction according to claim 1, characterized in that, The Actor network outputs three-dimensional acceleration actions based on the composite observation space, performs three-dimensional dynamic model-environment interaction, and updates network parameters, including: Get the current status of the vehicle The Actor network outputs three-dimensional acceleration commands. The three-dimensional acceleration command Satisfying the relation: ; in, This indicates the output three-dimensional acceleration command. Indicates that the Actor network parameter set is Deterministic policy mapping at time, This represents the current state of the composite observation space. This represents the sampled values from the Ornstein-Uhlenbeck exploration of the noise process; The vehicle state is updated using a three-dimensional dynamic model, where the updated velocity vector... With position vector They respectively satisfy the following relations: ; ; in, and These represent the three-dimensional velocity vectors before and after the update, respectively. and These represent the 3D position vectors before and after the update, respectively. Indicates the discrete time step. Indicates the damping coefficient; Apply acceleration and attitude angle constraints to ensure the vehicle satisfies the following relationship: ,in Indicates the magnitude of horizontal velocity. Indicates the upper limit of horizontal speed. Indicates pitch angle, This indicates the upper limit of the absolute value of the pitch angle; Calculate the total reward value And store the transferred samples to the experience replay buffer, the total reward value Satisfying the relation: ; in, Indicates the total compound reward. This indicates phased rewards. Indicates safety constraint rewards. This represents the final reward for successfully reaching the target point. This indicates the direction of arrival at the target point. The network is updated by sampling batch data from the experience replay buffer, wherein the Critic network minimizes the TD error. Update the TD error. Satisfying the relation: ; in, This represents the loss function of the Critic network. Indicates the number of samples in the batch. Represents the action value function. The target value is represented; the Actor network parameters are updated by maximizing the action value function, and a soft update mechanism is used to synchronize the target network parameters, which satisfies the following relationship: ; in, Indicates the target network parameters. Indicates the current network parameters. This represents the preset soft update coefficient.
8. The UAV turbulence avoidance method based on physical-guided spatiotemporal sequence prediction according to claim 1, characterized in that, The optimal turbulence avoidance trajectory includes: Step S61: Initialize the policy gradient network structure of the deep deterministic policy gradient algorithm, copy the current network parameters to the target network to generate the target policy, initialize the experience replay buffer, and construct Ornstein-Uhlenbeck stochastic process noise for action exploration; Step S62: At the beginning of each training round, initialize the environment round, reset the three-dimensional turbulent environment state to obtain the initial observations, and reset the noise state of the Ornstein-Uhlenbeck stochastic process; Step S63: Obtain and process the composite observation space state at the current moment, and input it into the Actor network. Generate a deterministic action vector based on the input observation, and after superimposing the sampled value of the Ornstein-Uhlenbeck random process noise, trim the action to the legal action space range to obtain the final action to be executed. Step S64: Input the final action into the three-dimensional dynamic model environment to perform the interaction, update the agent state, and obtain the instant reward value and round termination flag; Step S65: Store the current transfer sample in the experience replay buffer. When the number of samples stored in the buffer reaches a preset batch size, randomly sample batch data and update the Critic network by minimizing the mean square error. The loss function of the Critic network is... Satisfying the relation: ; in, This represents the loss function of the Critic network. Indicates the number of samples in the batch. This represents the estimated value of the action value function at the current moment. Indicates the current state in the sampled data. Indicates the action performed in the sampled data. Indicates the target Q value; Step S66: Update the Actor network using a gradient ascent strategy and perform norm clipping on the network gradient, wherein the update strategy of the Actor network aims to maximize the action value function output by the Critic network; Step S67: Gradually synchronize the target network parameters using a Polyak averaging method, perform a soft update of the target network, and then synchronize the target network parameters. Satisfying the relation: ; in, Indicates the target network parameters. Indicates the current network parameters. This represents the preset soft update coefficient; Repeat steps S61-S67 until the preset maximum number of training rounds is met or the path planning result meets the preset convergence condition, and output the optimal turbulence avoidance path.
9. A UAV turbulence avoidance system based on physically guided spatiotemporal sequence prediction, characterized in that, include: The 3D field construction module is used to build a 3D wind field inversion system, which performs variational assimilation processing on the acquired multi-source radar observation data to generate a comprehensive turbulence intensity field that reflects the real-time meteorological environment. The physics-guided spatiotemporal sequence prediction module is used to construct a physics-guided spatiotemporal convolutional sequence network, encode the comprehensive turbulence intensity field of historical time series into a spatiotemporal feature representation, and introduce physical process constraints based on atmospheric dynamics equations to pre-train the network to generate a future multi-step turbulence prediction field that conforms to physical consistency. The fusion sequence observation module is used to construct a reinforcement learning environment based on a deep deterministic policy gradient algorithm. It embeds the future multi-step turbulence prediction field and its corresponding physical residual confidence with the real-time dynamic state of the low-altitude vehicle into a composite observation space, and defines a flight state transition model including takeoff, cruise and landing phases. The network initialization and reward shaping module is used to initialize the Actor network and Critic network in the deep deterministic policy gradient algorithm and construct a composite reward function. The composite reward function includes a staged flight constraint term corresponding to the flight state transition model and a physical violation penalty term whose weights are dynamically adjusted according to the physical residual confidence. The iterative exploration module is used to perform iterative exploration through the deep deterministic policy gradient algorithm. In each iteration, the Actor network outputs a three-dimensional acceleration control action based on the current composite observation space and interacts with the reinforcement learning environment. It obtains a reward value according to the composite reward function and stores the transfer sample. The parameter update and trajectory decision module is used to update the parameters of the Critic network and the Actor network using the transfer samples, and to synchronize the target network using a soft update mechanism. The iterative steps are repeated until the path planning result meets the preset convergence condition, and the optimal turbulence avoidance trajectory is output.