Unmanned aerial vehicle simulation adaptive management system based on multi-modal neural network

By fusing drone and ocean wave data using a multimodal neural network, a sea surface deformation map is generated and a hovering strategy is optimized, solving the dynamic adaptability problem of drone hovering management in maritime rescue and improving hovering stability and rescue efficiency.

CN121477938BActive Publication Date: 2026-04-17XIAMEN OCEAN VOCATIONAL & TECH COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN OCEAN VOCATIONAL & TECH COLLEGE
Filing Date
2026-01-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In maritime rescue operations, the hovering management strategy for drones is affected by the downdraft of the rotor and the wake of the rescue vessel, causing the hovering method in the simulation to be unable to effectively adapt to dynamic changes in sea waves, resulting in distortion.

Method used

An adaptive management system based on multimodal neural networks is adopted. The data fusion unit acquires UAV flight, rotor and sea wave data to generate sea surface deformation map. Combined with the hovering analysis unit and pose planning unit, the system uses deep reinforcement learning and temporal attention neural network to generate adaptive management commands and optimize hovering strategy.

Benefits of technology

It accurately captures the effects of the rotor's downdraft and the rescue vessel's wake, reducing simulation distortion under dynamic wave changes and improving hovering stability and rescue efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121477938B_ABST
    Figure CN121477938B_ABST
Patent Text Reader

Abstract

This invention relates to the field of hovering management technology and discloses a UAV simulation adaptive management system based on a multimodal neural network, comprising: a data fusion unit, a hovering analysis unit, a pose planning unit, and a pose management unit. This solution generates an accurate sea surface deformation map by fusing flight, power, and wave modal data, and understands in real time the dynamic changes in sea surface morphology caused by rotor downdraft and rescue vessel wake. The generated dynamic hovering tolerance domain can effectively offset wave-induced displacement, generate multi-step optimal pose trajectories and adaptive management commands, reduce simulation distortion, improve the stability and targeting of UAV hovering, ensure the effective implementation of management strategies, and enhance the reliability of maritime rescue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hovering management technology, specifically to an adaptive management system for unmanned aerial vehicle (UAV) simulation based on a multimodal neural network. Background Technology

[0002] Currently, in drone simulation, flight data of the drone and environmental data in the simulation are often collected, and then feature analysis is performed using CNN / LSTM neural networks to achieve adaptive optimization and management of the drone's hovering state during the simulation.

[0003] However, the above management method still has the following defects in maritime rescue: In actual maritime rescue, the downdraft of the drone rotor will affect the wave pattern of the sea surface, and the wake of the rescue ship will also produce similar interference. These mutually affecting air and fluid will cause the drift trajectory of the person who fell into the water to change. This will cause the hovering method in the simulation to be distorted in the face of the dynamic changes of the actual sea waves, and cause the actual adaptive hovering management strategy of the drone to fail. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an adaptive management system for UAV simulation based on multimodal neural networks, which solves the aforementioned problems.

[0005] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0006] The UAV simulation adaptive management system based on multimodal neural networks includes:

[0007] The data fusion unit is used to acquire flight mode data of UAV, dynamic mode data of UAV rotor, and wave mode data of ocean in real time in a simulation environment. Based on a multimodal fusion neural network, the three preprocessed mode data are fused to obtain a sea surface deformation map.

[0008] The hovering analysis unit is used to obtain the coordinates of the target hovering point, combine the coordinates with the sea surface deformation map, analyze the wave-induced displacement that the UAV needs to offset, and obtain the dynamic hovering tolerance domain.

[0009] The pose planning unit uses the sea surface deformation map as a disturbance source and the dynamic hovering tolerance domain as a constraint condition. It performs multi-step planning through a deep reinforcement learning neural network to generate multi-step optimal pose trajectories.

[0010] The pose management unit is used to perform collaborative analysis of multi-step optimal pose trajectories and sea surface deformation maps based on a temporal attention neural network to obtain adaptive management instructions.

[0011] Furthermore, the preprocessed three modal data are fused using a multimodal fusion neural network to obtain a sea surface deformation map, including:

[0012] The flight mode data, dynamic mode data, and ocean wave mode data were respectively processed by three independent deep encoder networks to extract features, generating flight status codes, dynamic response codes, and ocean wave codes, respectively.

[0013] The interaction relationship between flight situation coding, dynamic response coding and ocean wave coding is analyzed by using a bidirectional attention mechanism to generate a cross-modal interaction factor matrix.

[0014] Furthermore, the preprocessed three modal data are fused using a multimodal fusion neural network to obtain a sea surface deformation map, which also includes:

[0015] Based on the cross-modal interaction factor matrix, the spatiotemporal propagation law of wave-induced displacement is analyzed, and a spatiotemporal state vector is generated.

[0016] A preliminary sea surface deformation map is obtained by decoding the spatiotemporal state vector using a deconvolutional generative network.

[0017] After calibrating the preliminary sea surface deformation map, the sea surface deformation map is obtained.

[0018] Furthermore, by combining this coordinate system with a sea surface deformation map, the wave-induced displacement that the UAV needs to offset is analyzed, resulting in a dynamic hovering tolerance domain, including:

[0019] Using the horizontal coordinates of the target hovering point as the center point, analyze the main direction and intensity changes of the deformation gradient within the window in the sea surface deformation map to generate hovering disturbance baseline values.

[0020] The dynamic response index of the UAV is obtained by calculating the flight mode data and dynamic mode data;

[0021] Based on the hovering disturbance baseline and the UAV dynamic response index, we analyze the UAV pose shift prediction sequence under disturbance.

[0022] Furthermore, by combining this coordinate system with a sea surface deformation map, the wave-induced displacement that the UAV needs to offset is analyzed, resulting in a dynamic hovering tolerance domain, which also includes:

[0023] By analyzing the pose offset prediction sequence, the probabilistic safety boundary of the UAV pose at each time step is obtained.

[0024] Within the probabilistic safety boundary at each moment, an efficient hovering subspace is defined based on the hovering energy consumption of the drone.

[0025] Furthermore, by combining this coordinate system with a sea surface deformation map, the wave-induced displacement that the UAV needs to offset is analyzed, resulting in a dynamic hovering tolerance domain, which also includes:

[0026] Based on the probabilistic safety boundary and efficient hovering subspace, and based on task priority, an adaptive tolerance rule is generated;

[0027] By integrating probabilistic safety boundaries, efficient hovering subspaces, and adaptive tolerance rules, a dynamic hovering tolerance domain is generated.

[0028] Furthermore, using the sea surface deformation map as the disturbance source and the dynamic hovering tolerance domain as the constraint, a deep reinforcement learning neural network is used for multi-step planning to generate multi-step optimal pose trajectories, including:

[0029] Based on a deep reinforcement learning neural network, the sea surface deformation map and the dynamic hovering tolerance domain are fused and encoded to obtain the global planning state vector.

[0030] Based on the adaptive tolerance rule, the pose adjustment of the UAV is transformed into multiple basic actions with different amplitudes and directions, resulting in a hierarchical control action set.

[0031] For each basic action, the drone's pose at the next moment is analyzed based on the selected basic action, and a composite reward value is generated.

[0032] Furthermore, using the sea surface deformation map as a disturbance source and the dynamic hovering tolerance domain as a constraint, a deep reinforcement learning neural network is used for multi-step planning to generate multi-step optimal pose trajectories. This also includes:

[0033] Based on the global planning state vector, a basic action is selected from the hierarchical planning action set through a deep reinforcement learning neural network. The pose of the UAV after executing the basic action is analyzed, and this process is repeated to generate a multi-step trajectory state sequence.

[0034] Based on the multi-step trajectory state sequence, the pose is adjusted according to the adaptive tolerance rule to generate a set of corrected trajectory sequences;

[0035] From the set of corrected trajectory sequences, the corresponding trajectory is selected based on the composite reward value to generate a multi-step optimal pose trajectory.

[0036] Furthermore, for each basic action, the drone's pose at the next moment is analyzed based on the selected basic action to generate a composite reward value, including:

[0037] The current UAV pose is compared with the efficient hovering subspace and the probabilistic safety boundary to calculate the multidimensional offset entropy of the UAV pose and the trajectory energy density required to reach the pose.

[0038] The spatial coordinates of the UAV's attitude are mapped to the sea surface deformation map. The influence of deformation gradient and phase on the UAV's attitude is analyzed, and a disturbance interlock factor is generated.

[0039] Based on the pose offset prediction sequence, the pose change trend caused by the execution of the current basic action is analyzed to obtain the control manifold curvature;

[0040] Based on the adaptive tolerance rule, the weight coefficients of the four parameters—multidimensional offset entropy, trajectory energy consumption density, disturbance interlock factor, and control manifold curvature—are adjusted and fused to generate a composite reward value.

[0041] Furthermore, by performing collaborative analysis of the multi-step optimal pose trajectory and sea surface deformation map using a temporal attention neural network, adaptive management instructions are obtained, including:

[0042] The multi-step optimal pose trajectory and the sea surface deformation map are input into the temporal attention neural network to calculate the coupling degree between the pose point at each time step on the trajectory and the local features of the deformation field at that time step, and generate the attitude adjustment matrix.

[0043] Based on the attitude adjustment matrix, the attention weights for the pose state at different times in the multi-step optimal pose trajectory are adjusted to generate adaptive management instructions.

[0044] In summary, the present invention has the following main beneficial effects:

[0045] The data fusion unit integrates flight modal data, dynamic modal data, and wave modal data. An independent depth encoder extracts flight status codes, dynamic response codes, and ocean wave codes. A cross-modal interaction factor matrix is ​​then generated through a bidirectional attention mechanism. Combined with a deconvolutional generator network and calibration process, a sea surface deformation map is accurately output, effectively capturing sea surface morphology changes caused by rotor downdraft and rescue vessel wake. The hovering analysis unit combines the target hovering point with the sea surface deformation map to generate hovering disturbance baseline values, UAV dynamic response index, and pose offset prediction sequences. This further constructs a probabilistic safety boundary and an efficient hovering subspace, and generates adaptive tolerance rules to accurately quantify wave-induced displacement.

[0046] Using the sea surface deformation map as the disturbance source and the dynamic hovering tolerance domain as the constraint, the pose planning unit obtains the global planning state vector through fusion encoding, constructs a hierarchical control action set, and generates a composite reward value through multi-dimensional offset entropy, trajectory energy consumption density, disturbance interlock factor, and control manifold curvature. The optimal pose trajectory in multiple steps is selected. The pose management unit generates the attitude adjustment matrix and adaptive management instructions with the help of a temporal attention neural network, and matches the pose with the sea wave deformation in real time. This reduces simulation distortion under dynamic sea wave changes, improves hovering stability and targeting, ensures the effective implementation of the adaptive hovering management strategy, and helps the UAV accurately capture the drift trajectory of a person who has fallen into the water after hovering, thus optimizing the efficiency of maritime rescue. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the UAV simulation adaptive management system based on multimodal neural networks of the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] refer to Figure 1 An adaptive management system for UAV simulation based on multimodal neural networks includes:

[0050] The data fusion unit is used to acquire flight mode data of UAV, dynamic mode data of UAV rotor, and wave mode data of ocean in real time in a simulation environment. Based on a multimodal fusion neural network, the three preprocessed mode data are fused to obtain a sea surface deformation map.

[0051] Flight modal data includes: three-dimensional position, attitude angles, linear velocity, and velocities at various angles; among which, attitude angles include: roll angle, pitch angle, and yaw angle;

[0052] Dynamic modal data include: rotor speed, rotor collective pitch, rotor thrust, rotor torque, etc.

[0053] Ocean wave modal data includes: wave height, wave direction, wave period, wave frequency, wave speed, etc.

[0054] The hovering analysis unit is used to obtain the coordinates of the target hovering point, combine the coordinates with the sea surface deformation map, analyze the wave-induced displacement that the UAV needs to offset, and obtain the dynamic hovering tolerance domain.

[0055] The pose planning unit uses the sea surface deformation map as a disturbance source and the dynamic hovering tolerance domain as a constraint condition. It performs multi-step planning through a deep reinforcement learning neural network to generate multi-step optimal pose trajectories.

[0056] The pose management unit is used to perform collaborative analysis of multi-step optimal pose trajectories and sea surface deformation maps based on a temporal attention neural network to obtain adaptive management instructions.

[0057] In one embodiment, the preprocessed three modal data are fused using a multimodal fusion neural network to obtain a sea surface deformation map, including:

[0058] The flight modal data, dynamic modal data, and ocean wave modal data were extracted using three independent deep encoder networks to generate flight status codes, dynamic response codes, and ocean wave codes, respectively. Specifically, a 4-layer temporal convolutional encoder (TCN) was used. The input was 12-dimensional flight modal data, with the 12 dimensions being 3D position + 3D attitude angle + 3D linear velocity + 3D angular velocity. The temporal window was divided with a 10ms step size. The data within the temporal window was normalized. The dynamic temporal correlation between attitude and velocity was captured by a 3×1 convolutional kernel. The coupling variance of roll angle and pitch angle was used as the channel attention weight. After weighting the convolutional features, global average pooling was performed, with a pooling kernel size of 5, resulting in a 128-dimensional code, which is the flight status code.

[0059] A 5-layer residual encoder (ResNet) is used to input 4-dimensional dynamic modal data. The 4-dimensional data is normalized to the 0-1 interval. Through 16, 32, 64, 64, and 64 channel residual blocks, the nonlinear mapping of speed-thrust is mined. After 1×1 convolution dimensionality reduction, the 64-dimensional code is output, which is the dynamic response code.

[0060] An 8-head self-attention encoder is used to input 5-dimensional ocean wave modal data. First, a fast Fourier transform is performed on the ocean wave modal data to extract 5 key frequency domain feature values ​​in the 0-5Hz frequency band and the original 5-dimensional time domain data, which are then concatenated into a 10-dimensional feature vector. The 8-head self-attention mechanism focuses on two sets of core related feature segments: wave height-wave period and wave direction-wave speed. After passing through a fully connected layer, a 96-dimensional code is output to obtain the ocean wave code.

[0061] The interaction relationship between flight status coding, dynamic response coding, and ocean wave coding is analyzed through a bidirectional attention mechanism to generate a cross-modal interaction factor matrix. Specifically, the flight status coding, dynamic response coding, and ocean wave coding are uniformly mapped to a 256-dimensional common feature space through a linear projection layer, and then input into the core bidirectional cross-attention module. The bidirectional cross-attention module calculates the pairwise interactions between the three modalities in parallel. For example, using the flight coding as the query vector and the dynamic coding as the key and value vector, an interaction feature is obtained by scaling dot product attention, and vice versa, thereby generating six bidirectional interaction feature vectors.

[0062] These bidirectional interactive feature vectors, along with the mapped flight status code, dynamic response code, and ocean wave code, are input into the dynamic gated fusion layer. The dynamic gated fusion layer uses independent gated neural networks to generate a weight vector for each modality. By weighted summing of all input features, the fused feature sequence is reshaped into a 16-row, 16-column 256-dimensional real matrix, which is the cross-modal interaction factor matrix.

[0063] In one embodiment, the three preprocessed modal data are fused based on a multimodal fusion neural network to obtain a sea surface deformation map, and the method further includes:

[0064] Based on the cross-modal interaction factor matrix, the spatiotemporal propagation law of wave-induced displacement is analyzed to generate a spatiotemporal state vector. Specifically, this involves performing dynamic non-negative matrix decomposition on the 16×16 cross-modal interaction factor matrix. The rank k of this decomposition is not fixed but is determined by analyzing the singular values ​​of the matrix. The cumulative energy contribution rate of the first 8 singular values ​​is calculated. When the cumulative energy contribution rate exceeds 85% for the first time, the corresponding index is determined as a candidate k value. If the candidate k value is < 4, then k = 4; if the candidate k value is > 8, then k = 8; otherwise, the candidate k value is directly adopted, with a value range between 4 and 8. Then, decomposition is performed to obtain a 16×k non-negative basis matrix and a k×16 non-negative coefficient matrix. Each column of the non-negative basis matrix represents a potential sea surface spatial displacement mode; each row of the non-negative coefficient matrix corresponds to the intensity change of a mode in 16 temporal dimensions.

[0065] Empirical mode decomposition is performed on each row of the non-negative coefficient matrix to extract the average period and energy ratio of the first intrinsic mode function, generating a 2D time series descriptor for each sea surface spatial displacement mode. The time series descriptors of multiple sea surface spatial displacement modes are concatenated, and the sparse spatial feature vectors obtained by flattening after threshold filtering in the non-negative basis matrix are concatenated to form a 54-dimensional joint feature. The threshold filtering is to retain elements in the non-negative basis matrix that are greater than 0.7 times the mean of all elements.

[0066] This joint feature is input into a Gaussian process coding layer, which contains nine parallel Gaussian process sub-layers. Each sub-layer uses an RBF kernel with different length scale parameters to map the 54-dimensional input to a latent space of different dimensions. Finally, the outputs of all sub-layers are concatenated into a 256-dimensional spatiotemporal state vector.

[0067] A preliminary sea surface deformation map is obtained by decoding the spatiotemporal state vector using a deconvolutional generator network. Specifically, the 256-dimensional spatiotemporal state vector is input into the deconvolutional generator network, first passing through a feature adaptation layer to map it into a 64-dimensional feature map. Then, each layer determines 4-8 variable-size convolutional kernels based on the entropy value of the preceding features. After three layers of upsampling with sampling factors of 4, 2, and 2, the output channels are gradually reduced to 1, and the output is the deformation distribution map. The value of each pixel in the deformation distribution map is the vertical displacement of the sea surface, i.e., the deformation value. The deformation values ​​in the deformation distribution map are weighted by spatial attention to obtain a preliminary sea surface deformation map with a high resolution of 256×256 pixels.

[0068] After calibrating the preliminary sea surface deformation map, a sea surface deformation map is obtained. Specifically, the deviations of the flight status code, dynamic response code, and ocean wave code from the corresponding areas of the preliminary deformation map are taken as real-time deviations. The real-time deviation of the flight status code is the average of the position error and attitude error, the real-time deviation of the dynamic response code is the average of the thrust error and rotational speed error, and the real-time deviation of the ocean wave code is the difference in wave height. Then, a 16×16×3 three-dimensional calibration factor matrix is ​​constructed. The elements of the three-dimensional calibration factor matrix are the weighted sum of the real-time deviations of the three codes, and the total weight of the three codes is 1.

[0069] The three-dimensional calibration factor matrix has a dimension of 16×16×3, where each 16×16 grid point corresponds to a 16×16 pixel grid area in the preliminary deformation map; the preliminary deformation map is divided into 8×8 sub-blocks of 32×32 pixels, and each 32×32 pixel sub-block corresponds to a 2×2 grid area in the three-dimensional calibration factor matrix.

[0070] The error variance of each sub-block is calculated. For sub-blocks with an error variance greater than 0.012, dynamic iterative correction is adopted, with 3 iterations. The correction amount in each iteration is the error variance of the sub-block multiplied by the weight mean of the corresponding three-dimensional calibration factor matrix region, thereby adjusting the deformation value of each pixel in the preliminary deformation map. The standard deviation of the wave height of the first 10 frames is calculated. 1.2 times the standard deviation is used as the threshold to filter out abnormal pixels in the preliminary sea surface deformation map whose wave height change (wave height change is the instantaneous change of deformation value) exceeds the threshold. Finally, a 256×256 sea surface deformation map is output, and the pixel value of the sea surface deformation map is the adjusted deformation value.

[0071] By deeply fusing flight modal data, dynamic modal data, and wave modal data through a multimodal fusion neural network, high-order features of each modality are extracted with the help of an encoder, and cross-modal interaction relationships are mined by combining a bidirectional attention mechanism. This generates a cross-modal interaction factor matrix, accurately capturing the spatiotemporal propagation law of wave-induced displacement. A high-precision sea surface deformation map is then obtained through a deconvolutional generator network and multi-dimensional calibration. This can effectively adapt to the interaction between air and fluid brought about by rotor downdraft and rescue vessel wake, reduce simulation distortion under dynamic wave changes, ensure the effectiveness of adaptive hovering management strategies, and improve the stability of UAV hovering during maritime rescue.

[0072] In one embodiment, the coordinates are combined with a sea surface deformation map to analyze the wave-induced displacement that the UAV needs to offset, thus obtaining the dynamic hovering tolerance domain, including:

[0073] Using the horizontal coordinates of the target hovering point as the center point, the main direction and intensity changes of the deformation gradient within the sea surface deformation map are analyzed to generate hovering disturbance base values. Specifically, this includes: using the horizontal coordinates of the target hovering point as the center point, extracting a 5×5 analysis window in the sea surface deformation map; within this analysis window, calculating the difference between the deformation values ​​of each pixel and its eight neighboring pixels to obtain the local gradients in eight directions at that point, thus forming a local deformation gradient vector set.

[0074] The relative deformation field is obtained by subtracting the deformation value of the center point from the deformation values ​​of all points within the analysis window. For each point in the relative deformation field except the center point, its two-dimensional direction vector relative to the center of the analysis window is calculated, and this direction vector is multiplied by its corresponding deformation value to obtain the weighted direction vector. The weighted direction vectors of all points within the analysis window are summed to obtain the dominant direction vector. The magnitude of the dominant direction vector is the consistency intensity of deformation in the dominant direction within the analysis window, and its direction angle is calculated using the arctangent function. The direction angle ranges from 0° to 360°, and the angle of the direction angle is taken as the dominant perturbation direction.

[0075] Calculate the direction of the vector with the largest magnitude in the local gradient vector set of each pixel within the window, calculate the cosine value of the angle between this direction and the dominant perturbation direction, and calculate the variance of the cosine values ​​of all points to obtain the direction consistency index.

[0076] Divide the magnitude of the dominant direction vector by the current window size (5), and then multiply it by the inverse of the variance of the direction consistency index to generate the hovering disturbance base value.

[0077] The dynamic response index of the UAV is obtained by calculating the flight modal data and dynamic modal data. Specifically, this includes extracting the roll angle, pitch angle, roll angular velocity, and pitch angular velocity at the current moment from the flight modal data.

[0078] Extract the current rotor total thrust and total torque from the dynamic modal data, and obtain the reference power for standard drone hovering; multiply the total thrust and total torque and divide by the reference power to obtain the power saturation.

[0079] Calculate the product of roll angle and roll velocity, and the product of pitch angle and pitch velocity separately. Add the two products together to obtain the overall recovery trend.

[0080] The power margin coefficient is obtained by taking the reciprocal of the power saturation. The dynamic response index of the UAV is obtained by multiplying the absolute value of the comprehensive recovery trend by the dynamic weight and then by the power margin coefficient.

[0081] Specifically, when the overall recovery trend is negative, the dynamic weight is 0.8; when the overall recovery trend is positive or zero, the dynamic weight is 0.3.

[0082] Based on the hovering disturbance baseline and the UAV dynamic response index, the pose deviation prediction sequence of the UAV under disturbance is analyzed. Specifically, the magnitude of the dominant direction vector is multiplied by 0.02 as the initial prediction amplitude of the prediction sequence.

[0083] The total prediction time span is set to 720 milliseconds, and it is divided into 6 prediction steps of equal length, each prediction step lasting 120 milliseconds; if the UAV dynamic response index is ≥0.75, the attenuation factor is 0.35; if the UAV dynamic response index is <0.75, the attenuation factor is 0.65.

[0084] Starting from the initial prediction amplitude, subtract the attenuation factor from 1 and multiply by the initial prediction amplitude to obtain the amplitude of the second prediction step; the amplitude of the remaining prediction steps is obtained by multiplying the amplitude of the previous prediction step by the corresponding power of (1 minus the attenuation factor). For example, to calculate the amplitude of the third prediction step, subtract 1 from 3 and the result is 2, which is the corresponding power.

[0085] The magnitude of each prediction step is multiplied by the dominant direction vector to generate an ordered sequence containing six two-dimensional horizontal offset vectors, which is the pose offset prediction sequence.

[0086] In one embodiment, the coordinates are combined with a sea surface deformation map to analyze the wave-induced displacement that the UAV needs to offset, thus obtaining the dynamic hovering tolerance domain. This also includes:

[0087] Analyzing the pose offset prediction sequence yields the probabilistic safety boundary of the UAV's pose at each moment. Specifically, this includes: constructing an initial circular region centered on the endpoint of each two-dimensional horizontal offset vector in the pose offset prediction sequence; and adding the amplitude of the current prediction step to the UAV's dynamic response index to obtain the radius of the circular region.

[0088] Secondly, based on the directional consistency index, the circular regions are subjected to directional stretching deformation: the reciprocal of the directional consistency index is used as the stretching factor, and the radius of each circular region is increased by 1 times the stretching factor along the dominant disturbance direction; in the direction perpendicular to the dominant disturbance direction, the radius is reduced to 10% of the original stretching factor; and then the circular regions are deformed into ellipses with their major axis pointing to the dominant disturbance direction.

[0089] For each deformed elliptical region, all points inside it are defined as safe points at that moment. Connecting all the safe points yields the probabilistic safety boundary.

[0090] Within the probabilistic safety boundary at each moment, an efficient hovering subspace is defined based on the hovering energy consumption of the UAV. Specifically, this includes: determining the geometric center, major axis direction, and minor semi-axis length of the probabilistic safety boundary, where the major axis direction is the dominant disturbance direction; and obtaining the minor semi-axis length by multiplying the radius of the circular region by a stretching factor of 10% in the direction perpendicular to the dominant disturbance direction.

[0091] Along the major axis, starting from the geometric center, candidate points are selected in both directions with a step size of 0.15 times the length of the minor semi-axis, until the probabilistic safety boundary is reached, resulting in multiple candidate points.

[0092] Calculate the Euclidean distance between the candidate point and the endpoint of the pose offset prediction vector at that moment to obtain the position compensation energy consumption;

[0093] Calculate the sum of the absolute values ​​of the current roll angular velocity and pitch angular velocity to obtain the attitude maintenance energy consumption;

[0094] The comprehensive energy consumption index is obtained by adding the position compensation energy consumption and the attitude maintenance energy consumption, multiplying it by the power margin coefficient, and then subtracting the directional consistency index.

[0095] Select the three points with the smallest comprehensive energy consumption index from all candidate points, and connect the three points to form a triangular region; move each vertex of the triangular region along the direction pointing to its geometric centroid, and move the distance is 0.2 times the distance from the vertex to the centroid. The triangular region formed after shrinking is the efficient hovering subspace at this moment.

[0096] In one embodiment, the coordinates are combined with a sea surface deformation map to analyze the wave-induced displacement that the UAV needs to offset, thus obtaining the dynamic hovering tolerance domain. This also includes:

[0097] Based on the probabilistic safety boundary and efficient hovering subspace, and based on task priority, an adaptive tolerance rule is generated, which specifically includes: setting a task priority value, the value of which ranges from 0 to 1. The larger the task priority value, the more emphasis is placed on position maintenance, and the smaller the task priority value, the more emphasis is placed on energy saving.

[0098] Calculate the Euclidean distance from the current UAV pose to the centroid of the efficient hovering subspace, denoted as the first distance; and calculate the Euclidean distance from the current UAV pose to the geometric center of the probabilistic safety boundary, denoted as the second distance;

[0099] The first threshold is obtained by multiplying the average of the Euclidean distances from the centroid of the efficient hovering subspace to its three vertices by 0.4; the second threshold is obtained by multiplying the length of the short semi-axis of the probabilistic safety boundary by 0.6.

[0100] The value obtained by multiplying the task priority value by 0.2 is added to the first threshold and the second threshold respectively to obtain the first adjustment threshold and the second adjustment threshold;

[0101] Multiply the length of the minor semi-axis of the ellipse in the probabilistic safety boundary by 0.8 to obtain the boundary's proximity to the threshold.

[0102] If the first distance is less than the first adjustment threshold, it is determined to be at the green level;

[0103] If the first adjustment threshold ≤ the first distance ≤ the second adjustment threshold, and at the same time the second distance ≤ the boundary proximity threshold, then it is judged as yellow level;

[0104] If the first distance is greater than the second adjustment threshold, and at the same time the second distance is greater than the boundary approach threshold, then it is judged as red level;

[0105] For the green level, control the drone to maintain its current efficient hovering subspace; for the yellow level, control the drone to return to the efficient hovering subspace; for the red level, control the drone to quickly return to the probabilistic safety boundary.

[0106] The three control rules mentioned above are the adaptive tolerance rules.

[0107] The probabilistic safety boundary, efficient hovering subspace, and adaptive tolerance rules are integrated to generate a dynamic hovering tolerance domain. Specifically, this includes: constructing a hierarchical spatiotemporal tolerance structure, and for each moment in the pose offset prediction sequence, using the probabilistic safety boundary at that moment as the maximum allowable space and the efficient hovering subspace as the core preservation space.

[0108] Adjust the maximum allowable space and core holding space: calculate the time step percentage from the current time to each time step, subtract the time step percentage from 1 to obtain the tolerance decay coefficient at that time step; multiply the short semi-axis length of the maximum allowable space by the tolerance decay coefficient; move each vertex of the core holding space towards its geometric centroid by a distance equal to the distance from the original vertex to the centroid multiplied by (1 minus the tolerance decay coefficient).

[0109] Based on the adaptive tolerance rule level at the current moment, a strategy fusion weight is assigned to the maximum allowable space after shrinkage at each moment. Under the green level, a weight of 0.8 is used for the core preservation space and a weight of 0.2 is used for the maximum allowable space; under the yellow level, the weights are all 0.5; under the red level, a weight of 0.3 is used for the core preservation space and a weight of 0.7 is used for the maximum allowable space.

[0110] Based on the strategy fusion weights, the points within the adjusted maximum allowable space and the core retention space are weighted and fused to generate a continuous resident probability distribution map.

[0111] The adjusted high-efficiency hovering subspace, as the core holding space, together with the dwell probability distribution map, constitutes the dynamic hovering tolerance domain.

[0112] By fusing the target hovering point coordinates with the sea surface deformation map, deformation features are extracted through a 5×5 analysis window to generate hovering disturbance baseline values ​​and directional consistency indices. Combined with flight and dynamic modal data, the UAV dynamic response index and power margin coefficient are calculated, and a pose offset prediction sequence is constructed. Furthermore, probabilistic safety boundaries and efficient hovering subspaces are generated. Adaptive tolerance rules are formulated and integrated into a hierarchical spatiotemporal dynamic hovering tolerance domain, which can accurately adapt to the dynamic interference of rotor downdraft and rescue vessel wake, and offset wave-induced displacement in real time, reducing simulation distortion, improving hovering stability and accuracy, ensuring the effective implementation of adaptive hovering management strategies, and taking into account both position maintenance and energy saving.

[0113] In one embodiment, the sea surface deformation map is used as the disturbance source, and the dynamic hovering tolerance domain is used as the constraint. Multi-step planning is performed using a deep reinforcement learning neural network to generate a multi-step optimal pose trajectory, including:

[0114] The sea surface deformation map and the dynamic hovering tolerance domain are fused and encoded based on a deep reinforcement learning neural network to obtain a global planning state vector. Specifically, the sea surface deformation map is input into a 5-layer convolutional encoder, and downsampling is performed using three 3×3 convolutional kernels with a stride of 2 to extract the gradient of deformation value changes, outputting a 128-dimensional deformation feature vector. From the dynamic hovering tolerance domain, the short semi-axis length of the probabilistic safety boundary, the coordinates of the three vertices of the efficient hovering subspace, and the current adaptive tolerance rule level are extracted, totaling 12 scalars.

[0115] The 12 scalars are encoded into a 64-dimensional constraint feature vector through a 3-layer fully connected network. A cross-attention mechanism is used, with the 128-dimensional deformation feature as the query and the 64-dimensional constraint feature as the key and value, to calculate an attention-weighted 96-dimensional fusion feature. This fusion feature is then concatenated with the UAV’s current 3-dimensional position and 3-dimensional angular velocity to form a 256-dimensional global planning state vector.

[0116] Based on the adaptive tolerance rule, the pose adjustment of the UAV is transformed into multiple basic actions of different amplitudes and directions, resulting in a hierarchical control action set. Specifically, this includes three control levels—green, yellow, and red—defined by the adaptive tolerance rule. This constructs a hierarchical, discrete action space for the deep reinforcement learning neural network. The actions in this action space constitute the hierarchical control action set.

[0117] For the green level, five basic actions are defined, including translation in four horizontal directions (east, south, west, and north) and height maintenance in one vertical direction. The amplitude of the actions is 0.15 times the hovering disturbance baseline value.

[0118] For the yellow level, nine basic movements are defined. Based on the green level movements, four diagonal translations (northeast, southeast, southwest, and northwest) are added. The amplitude of the movements is 0.35 times the hovering disturbance baseline value.

[0119] For the red level, 14 basic movements are defined. Based on the yellow level movements, a rapid horizontal translation in four directions (up, down, left, and right) and an emergency vertical climb are added. The amplitude benchmark of the movements is 0.7 times the hovering disturbance base value.

[0120] The final execution range of all actions is the action range benchmark multiplied by (2 minus the UAV dynamic response index), thus forming a hierarchical control action set closely related to the real-time tolerance state.

[0121] For each basic action, the drone's pose at the next moment is analyzed based on the selected basic action, and a composite reward value is generated.

[0122] In one embodiment, the sea surface deformation map is used as the disturbance source, and the dynamic hovering tolerance domain is used as the constraint condition. Multi-step planning is performed using a deep reinforcement learning neural network to generate a multi-step optimal pose trajectory. The method also includes:

[0123] Based on the global planning state vector, a basic action is selected from the hierarchical control action set through a deep reinforcement learning neural network. The UAV pose after executing the basic action is analyzed, and this process is repeated to generate a multi-step trajectory state sequence. Specifically, the global planning state vector is calculated through the policy network of the deep reinforcement learning neural network, and the action selection probability on the hierarchical control action set is output. Specifically, the deep reinforcement learning neural network contains a policy network, which takes the 256-dimensional global planning state vector as its unique input. The policy network has three fully connected layers, containing 128, 64, and 28 neurons respectively. The ReLU activation function is used for nonlinear feature transformation, and the policy network finally outputs a 28-dimensional original logical value vector, whose dimension is consistent with the total number of basic actions contained in the hierarchical control action set. According to the adaptive tolerance rule level at the current time, the dimensions corresponding to actions that do not match the current level are dynamically filtered out from these 28 logical values. For example, if the current level is green, only the first 5 dimensions are retained as valid. The remaining valid logical values ​​of the original logical value vector are normalized by the Softmax function and converted into a probability distribution. This probability distribution is the action selection probability on the hierarchical control action set.

[0124] A basic action is selected based on the action selection probability sampling. The final execution amplitude of the basic action is multiplied by the response delay, and the three-dimensional linear displacement increment caused by the basic action within the current planning step size can be obtained. The response delay is 120 milliseconds.

[0125] Simultaneously, the roll angular velocity and pitch angular velocity in the current flight mode data are multiplied by 0.25 to obtain the attitude angle increment; the three-dimensional linear displacement increment and attitude increment are superimposed on the current pose of the UAV to obtain the predicted UAV pose at the next planning moment after the execution of this action (the next planning moment is 120 milliseconds after the current moment);

[0126] Based on the predicted UAV pose, a new global planning state vector is recalculated, and the process of obtaining the predicted UAV pose is repeated. This process is repeated 5 times to generate multiple predicted pose sequences that are arranged in chronological order and contain 6 time points, i.e., multi-step trajectory state sequences.

[0127] Based on the multi-step trajectory state sequence, the pose is adjusted according to the adaptive tolerance rule to generate a corrected trajectory sequence set. Specifically, for each predicted pose point in the multi-step trajectory state sequence, the corrected reference point is determined based on the adaptive tolerance rule level at the current time: under the green level, the reference point is the geometric centroid of the triangle; under the yellow level, the reference point is the geometric center of the ellipse; under the red level, the reference point is the point on the boundary of the ellipse that is closest to the predicted pose point.

[0128] Subtract the UAV dynamic response index from 1.5 and multiply it by the directional consistency index to obtain the correction range;

[0129] Move the current predicted pose point along the direction pointing to the reference point by the correction amount to complete the single-point correction; traverse all 6 predicted pose points in the sequence, correct them in turn, and then connect all the corrected pose points in the original time order to form the corrected trajectory sequence set.

[0130] From the set of corrected trajectory sequences, the corresponding trajectory is selected based on the composite reward value to generate the multi-step optimal pose trajectory. Specifically, for each trajectory in the set of corrected trajectory sequences, the composite reward values ​​of all basic actions performed on the trajectory are added together to obtain the total composite reward value of the trajectory. The total composite reward values ​​of all trajectories are compared, and the trajectory with the largest total composite reward value is selected as the multi-step optimal pose trajectory.

[0131] By using sea surface deformation maps as disturbance sources and dynamic hovering tolerance domains as constraints, a 256-dimensional global planning state vector is generated through a 5-layer convolutional encoder and cross-attention mechanism. A hierarchical control action set is constructed in conjunction with adaptive tolerance rules to adapt to the pose adjustment requirements of different control levels (green, yellow, and red). Multi-step optimal pose trajectories are generated, which can accurately adapt to the dynamic interference of rotor sinking airflow and rescue vessel wake, reduce simulation distortion caused by wave changes, improve the real-time performance and accuracy of pose planning, and enhance the pose stability of UAVs during maritime rescue.

[0132] In one embodiment, for each basic action, the UAV pose at the next moment caused by the selected basic action is analyzed to generate a composite reward value, including:

[0133] The current UAV pose is compared with the efficient hovering subspace and the probabilistic safety boundary to calculate the multidimensional offset entropy of the UAV pose and the trajectory energy density required to reach the pose. Specifically, this includes:

[0134] Extract the horizontal coordinates from the current 3D position of the UAV, calculate the 2D Euclidean distance to the centroid of the efficient hovering subspace, and then divide it by the length of the minor semi-axis of the probabilistic safety boundary to obtain the horizontal offset; obtain the target hovering height from the target hovering point, divide the absolute value of the difference between the current height of the UAV and the target hovering height by 5 meters to obtain the altitude offset; sum the absolute values ​​of the roll angle and pitch angle and divide by 25 degrees to obtain the attitude offset.

[0135] Normalize these three offsets to the 0-1 interval to obtain three probability values. Multiply each of the three probability values ​​by the natural logarithm of its probability value and then add them together. Multiply the sum by negative one to obtain the multidimensional offset entropy of the UAV pose.

[0136] The total displacement length is taken as the straight-line distance in three-dimensional space from the previous pose to the current pose; at the same time, the sum of the absolute values ​​of the roll angle and pitch angle changes during the displacement process is calculated to obtain the angle change sum.

[0137] Multiply the total displacement length by the power margin factor, add it to the angle change, and then divide by the total displacement length to obtain the trajectory energy density required to reach the pose.

[0138] The spatial coordinates of the UAV's pose are mapped to the sea surface deformation map. The influence of deformation gradient and phase on the UAV's pose is analyzed, and a disturbance interlock factor is generated. Specifically, the following steps are taken: based on the horizontal coordinates of the UAV's current three-dimensional position, the corresponding pixel point is determined in the sea surface deformation map, and a 3×3 pixel local area is set with the pixel point as the center.

[0139] Calculate the deformation gradient of the local region in the east-west and north-south directions to obtain the gradient vector; calculate the angle between the direction of this gradient vector and the direction of the dominant perturbation, and multiply the absolute value of the cosine of the angle by the magnitude of the gradient vector to obtain the gradient interlocking component.

[0140] The phase interlock component is obtained by calculating the sine value of the phase of the current deformation value of the pixel in the most recent complete wave cycle, and then multiplying it by the difference between the current deformation value and the deformation value of the corresponding point in the previous frame.

[0141] Multiply the directional consistency index by the UAV dynamic response index to obtain the coupling interlock component; multiply the three components and normalize the calculation result to the 0-1 interval to obtain the disturbance interlock factor.

[0142] Based on the pose offset prediction sequence, the pose change trend caused by the execution of the current basic action is analyzed to obtain the control manifold curvature. Specifically, this includes: for the second vector and the first vector in the pose offset prediction sequence, the absolute value of the angle between them is calculated; the ratio of the magnitude of the second vector to the magnitude of the first vector is calculated; the absolute value is multiplied by (1 minus the absolute value of the ratio) to obtain the preliminary curvature measure for this step; this process is repeated for all 5 adjacent steps to obtain 5 preliminary curvature measures.

[0143] After this step, the square of the cosine of the angle between the direction of the last vector in the pose offset prediction sequence and the direction of the dominant disturbance is used to obtain the directional modulation coefficient. The five preliminary curvature measures are multiplied by the corresponding directional modulation coefficients and summed, then divided by the directional consistency index, and then divided by the UAV dynamic response index to obtain the control manifold curvature representing the pose change trend.

[0144] Based on the adaptive tolerance rule, the weight coefficients of four parameters—multidimensional offset entropy, trajectory energy consumption density, disturbance interlock factor, and control manifold curvature—are adjusted and fused to generate a composite reward value. Specifically, this includes assigning basic weight coefficients to the four parameters based on green, yellow, or red levels according to the adaptive tolerance rule.

[0145] Under the green level, the weighting coefficients for multidimensional offset entropy, trajectory energy consumption density, disturbance interlock factor, and control manifold curvature are 0.1, 0.5, 0.2, and 0.2, respectively; under the yellow level, the weighting coefficients for all four are 0.25; under the red level, the weighting coefficients are adjusted to 0.4, 0.2, 0.3, and 0.1.

[0146] Multiply the task priority value by 0.05 and add the calculation result to the weight coefficients of the multidimensional offset entropy and the disturbance interlock factor respectively. At the same time, subtract twice the calculation result from the weight coefficient of the trajectory energy consumption density, while the weight coefficient of the control manifold curvature remains unchanged.

[0147] The adjusted weight coefficients are multiplied by their corresponding parameters and summed. The result is then multiplied by negative one to generate a composite reward value used to evaluate the quality of the action.

[0148] By accurately calculating multidimensional offset entropy, trajectory energy consumption density, disturbance interlock factor, and control manifold curvature, and combining the green, yellow, and red levels of adaptive tolerance rules to dynamically allocate weights, a composite reward value is generated. This fully adapts to the dynamic interference caused by rotor sinking airflow and rescue vessel wake during maritime rescue, accurately captures the impact of wave deformation on UAV attitude, effectively quantifies attitude offset, energy consumption, disturbance, and changing trends, improves the accuracy of deep reinforcement learning action selection, and reduces the distortion between simulation and actual sea conditions.

[0149] In one embodiment, adaptive management instructions are obtained by co-analyzing the multi-step optimal pose trajectory and sea surface deformation map using a temporal attention neural network, including:

[0150] The multi-step optimal pose trajectory and the sea surface deformation map are input into the temporal attention neural network to calculate the coupling degree between the pose point at each moment on the trajectory and the local features of the deformation field at that moment, and generate the attitude adjustment matrix. Specifically, the temporal attention neural network consists of a bidirectional LSTM layer and a 4-head attention layer. The input is the three-dimensional position and roll angle and pitch angle sequence of the 6 prediction steps in the multi-step optimal pose trajectory.

[0151] For each time step, a 9×9 pixel region centered on the horizontal coordinates of the attitude at that time is extracted from the sea surface deformation map. The mean, variance, and deformation gradient along the dominant perturbation direction of the deformation value within this region are calculated to obtain a 3D deformation feature. The 3D deformation feature is concatenated with the attitude angle at that time step to form a 6D vector, forming a temporal input sequence. The temporal input sequence is encoded into 128-dimensional features through bidirectional LSTM. These features are used as the query and key, and the fusion vector of the power margin coefficient and the orientation consistency index at the current time step is used as the value for attention calculation. The attention weights at each time step are then output.

[0152] The attention weights are weighted with the attitude angles at the corresponding times and mapped to a 6×4 matrix, where the row index corresponds to the 6 prediction times and the column index corresponds to the combined adjustment intensity of the roll and pitch angles. This matrix is ​​the attitude adjustment matrix.

[0153] Based on the attitude adjustment matrix, the attention weights for the pose state at different times in the multi-step optimal pose trajectory are adjusted, and adaptive management instructions are generated. Specifically, the row index i (1 to 6) of the attitude adjustment matrix corresponds to the i-th prediction time in the multi-step optimal pose trajectory, and the odd and even positions of the column index j (1 to 6) are associated with the roll angle and pitch angle, respectively.

[0154] For each time i, sum the absolute values ​​of all elements in the i-th row of the attitude adjustment matrix, and then multiply by the tolerance decay coefficient corresponding to the current time to obtain the instantaneous adjustment weight; multiply the maximum value of all elements in the first and fourth rows of the attitude adjustment matrix by the reciprocal of the UAV dynamic response index to obtain the global adjustment benchmark.

[0155] The adaptive hovering management command consists of a four-dimensional vector. The first two elements are the combined adjustment amounts of the roll angle and pitch angle: the combined adjustment amount is obtained by weighting and summing all odd and even columns of the attitude adjustment matrix according to the instantaneous adjustment weights of the corresponding rows, and then multiplying by the global adjustment reference.

[0156] The next two elements are horizontal thrust: the instantaneous adjustment weight at that moment is multiplied by the magnitude of the dominant direction vector to obtain a thrust, and the thrust is uniformly decomposed into two parts, one part along the dominant disturbance direction and the other part along the direction perpendicular to the dominant disturbance direction.

[0157] Among them, the projected thrust along the dominant disturbance direction serves as the thrust required to counteract the main wave disturbance; the projected thrust along the vertical direction serves as the thrust to maintain lateral stability or counteract secondary disturbances.

[0158] These two thrusts are the last two elements of the four-dimensional vector; the vector formed by the four elements constitutes the adaptive management command, which is mainly aimed at the hovering of the drone.

[0159] By using a temporal attention neural network to coordinate multi-step optimal pose trajectory and sea surface deformation map, deformation features of a 9×9 pixel region are extracted. Attention weights are calculated by combining power margin coefficient and directional consistency index to generate a 6×4 attitude adjustment matrix. Adaptive management commands are generated to accurately allocate horizontal thrust along the dominant disturbance direction and vertical direction. This can effectively adapt to the dynamic interference of rotor sinking airflow and rescue vessel wake, and match wave deformation and pose changes in real time, reducing simulation distortion, ensuring the pertinence and effectiveness of hovering management commands, and improving the hovering stability of UAVs.

[0160] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An adaptive management system for unmanned aerial vehicle (UAV) simulation based on a multimodal neural network, characterized in that, include: The data fusion unit is used to acquire UAV flight mode data, UAV rotor dynamic mode data, and ocean wave mode data in real time within a simulation environment. Based on a multimodal fusion neural network, it fuses the preprocessed data from the three modes to obtain a sea surface deformation map, including: The flight mode data, dynamic mode data, and ocean wave mode data were respectively processed by three independent deep encoder networks to extract features, generating flight status codes, dynamic response codes, and ocean wave codes, respectively. The interaction relationship between flight situation coding, dynamic response coding and ocean wave coding is analyzed by using a bidirectional attention mechanism to generate a cross-modal interaction factor matrix; Based on the cross-modal interaction factor matrix, the spatiotemporal propagation law of wave-induced displacement is analyzed, and a spatiotemporal state vector is generated. A preliminary sea surface deformation map is obtained by decoding the spatiotemporal state vector using a deconvolutional generative network. After calibrating the preliminary sea surface deformation map, the sea surface deformation map is obtained; The hovering analysis unit is used to acquire the coordinates of the target hovering point, combine these coordinates with the sea surface deformation map, analyze the wave-induced displacement that the UAV needs to offset, and obtain the dynamic hovering tolerance domain, including: Using the horizontal coordinates of the target hovering point as the center point, analyze the main direction and intensity changes of the deformation gradient within the window in the sea surface deformation map to generate hovering disturbance baseline values. The dynamic response index of the UAV is obtained by calculating the flight mode data and dynamic mode data; Based on hovering disturbance baseline and UAV dynamic response index, we analyze the UAV pose shift prediction sequence under disturbance. By analyzing the pose offset prediction sequence, the probabilistic safety boundary of the UAV pose at each time step is obtained. Within the probabilistic safety boundary at each moment, an efficient hovering subspace is defined based on the hovering energy consumption of the UAV. Based on the probabilistic safety boundary and efficient hovering subspace, and based on task priority, an adaptive tolerance rule is generated; By integrating probabilistic safety boundaries, efficient hovering subspaces, and adaptive tolerance rules, a dynamic hovering tolerance domain is generated. The pose planning unit uses the sea surface deformation map as a disturbance source and the dynamic hovering tolerance domain as a constraint. It performs multi-step planning through a deep reinforcement learning neural network to generate multi-step optimal pose trajectories, including: Based on a deep reinforcement learning neural network, the sea surface deformation map and the dynamic hovering tolerance domain are fused and encoded to obtain the global planning state vector. Based on the adaptive tolerance rule, the pose adjustment of the UAV is transformed into multiple basic actions with different amplitudes and directions, resulting in a hierarchical control action set. For each basic action, the drone's pose at the next moment is analyzed based on the selected basic action, and a composite reward value is generated. Based on the global planning state vector, a basic action is selected from the hierarchical planning action set through a deep reinforcement learning neural network. The pose of the UAV after executing the basic action is analyzed, and this process is repeated to generate a multi-step trajectory state sequence. Based on the multi-step trajectory state sequence, the pose is adjusted according to the adaptive tolerance rule to generate a set of corrected trajectory sequences; From the set of corrected trajectory sequences, select the corresponding trajectory based on the composite reward value to generate a multi-step optimal pose trajectory; The pose management unit is used to perform collaborative analysis of multi-step optimal pose trajectories and sea surface deformation maps based on a temporal attention neural network to obtain adaptive management instructions.

2. The UAV simulation adaptive management system based on multimodal neural networks according to claim 1, characterized in that, For each basic action, the drone's pose at the next moment is analyzed based on the selected basic action to generate a composite reward value, including: The current UAV pose is compared with the efficient hovering subspace and the probabilistic safety boundary to calculate the multidimensional offset entropy of the UAV pose and the trajectory energy density required to reach the pose. The spatial coordinates of the UAV's attitude are mapped to the sea surface deformation map. The influence of deformation gradient and phase on the UAV's attitude is analyzed, and a disturbance interlock factor is generated. Based on the pose offset prediction sequence, the pose change trend caused by the execution of the current basic action is analyzed to obtain the control manifold curvature; Based on the adaptive tolerance rule, the weight coefficients of the four parameters—multidimensional offset entropy, trajectory energy consumption density, disturbance interlock factor, and control manifold curvature—are adjusted and fused to generate a composite reward value.

3. The UAV simulation adaptive management system based on multimodal neural networks according to claim 1, characterized in that, Based on the collaborative analysis of multi-step optimal pose trajectory and sea surface deformation map using a temporal attention neural network, adaptive management instructions are obtained, including: The multi-step optimal pose trajectory and the sea surface deformation map are input into the temporal attention neural network to calculate the coupling degree between the pose point at each time step on the trajectory and the local features of the deformation field at that time step, and generate the attitude adjustment matrix. Based on the attitude adjustment matrix, the attention weights for the pose state at different times in the multi-step optimal pose trajectory are adjusted to generate adaptive management instructions.

Citation Information

Patent Citations

  • Unmanned aerial vehicle vision-assisted hovering method based on multi-view-angle vision synchronization tight coupling

    CN114355961A

  • Sea rescue helicopter rotor aerodynamic force simulation method and system

    CN121118261A