A multimodal fusion real-time re-landing prediction method, medium and device
Patent Information
- Application Number
- CN202610804383.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-05
AI Technical Summary
一是基于传统机器学习方法的事后识别方案,虽然此类方案能有效识别已发生的重着陆事件,但本质上仍是一种事后分析手段,无法实现触地前的实时预警,且预测过程过于依赖人工设计的特征,对飞行过程中的动态时序特性捕捉不足
[0014] The present invention has at least the following beneficial effects: decomposing the first subsequence corresponding to the vertical acceleration parameter type into multiple intrinsic mode subsequences effectively reduces the complexity of the first subsequence, thereby improving the reliability of prediction using the prediction sub-model; by processing multiple prediction sub-models in parallel and then performing aggregate calculation, the prediction accuracy is guaranteed and the processing efficiency is improved, thereby meeting the real-time requirements; by jointly inputting the auxiliary feature vector and the intrinsic mode subsequence into the prediction sub-model, multi-modal fusion is achieved, making full use of multi-dimensional flight parameter information, improving the robustness and generalization ability of the prediction process, thereby improving the real-time performance and reliability of hard landing prediction.
Smart Images

Figure CN122347252B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of flight safety monitoring technology, and in particular to a multimodal fusion real-time hard landing prediction method, medium, and equipment. Background Technology
[0002] A hard landing is an unsafe event in which the vertical acceleration of an aircraft exceeds design limits during landing, potentially leading to structural damage and safety hazards. Traditional methods typically rely on post-flight analysis of Quick Access Recorder (QAR) data, which is insufficient for providing real-time warnings before landing. In civil aviation, QARs continuously record hundreds or thousands of flight parameters throughout the entire flight phase, covering dimensions such as aircraft attitude, engine status, control surface deflection, navigation information, and environmental data.
[0003] In existing technologies, hard landing prediction is mainly achieved through two technical approaches. One is a post-event identification scheme based on traditional machine learning methods. Although such schemes can effectively identify hard landing events that have already occurred, they are essentially still a post-event analysis method and cannot achieve real-time early warning before touchdown. Moreover, the prediction process relies too much on manually designed features and is insufficient in capturing the dynamic temporal characteristics during flight.
[0004] The second approach is a pre-landing identification scheme based on time-series prediction methods using deep neural networks. For example, multidimensional QAR time-series data can be directly input into a long short-term memory network model. The memory units capture time dependencies and output the predicted value of future vertical acceleration. However, long short-term memory network models have insufficient long-term dependency modeling capabilities, low computational efficiency, and difficulty in meeting the real-time requirements of real-time prediction. Moreover, such schemes do not adequately handle the non-stationary characteristics of vertical acceleration signals and fail to effectively fuse the complex coupling relationships between multidimensional flight parameters, resulting in low reliability of hard landing predictions.
[0005] It is evident that existing technologies exhibit poor real-time performance and reliability in predicting hard landings in advance. Furthermore, existing models often heavily rely on specific flight paths or small datasets, resulting in weak generalization capabilities. When flight environment, aircraft weight, weather conditions, or pilot operations change, model performance deteriorates further.
[0006] Therefore, improving the real-time performance and reliability of hard landing prediction has become an urgent problem to be solved. Summary of the Invention
[0007] To address the aforementioned technical problems, the present invention employs a multimodal fusion real-time relanding prediction method, which includes the following steps:
[0008] S1 decomposes the first subsequence in the real-time transmitted flight parameter sequence into N intrinsic mode subsequences, where N is a positive integer.
[0009] S2, based on the auxiliary feature encoder, extract features from the second subsequences corresponding to the M auxiliary parameter types in the flight parameter sequence to obtain auxiliary feature vectors, where M is a positive integer.
[0010] S3. For any intrinsic mode subsequence, predict the predicted subsequence corresponding to the intrinsic mode subsequence based on the prediction sub-model corresponding to the intrinsic mode subsequence, the intrinsic mode subsequence, and the auxiliary feature vector.
[0011] S4 aggregates and calculates the predicted subsequences corresponding to the N intrinsic mode subsequences to obtain the target prediction sequence, and provides real-time risk warning for aircraft hard landing by comparing the target prediction sequence with a preset threshold.
[0012] The present invention also provides a non-transitory computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to realize the above-described multimodal fusion real-time relanding prediction method.
[0013] The present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0014] The present invention has at least the following beneficial effects: decomposing the first subsequence corresponding to the vertical acceleration parameter type into multiple intrinsic mode subsequences effectively reduces the complexity of the first subsequence, thereby improving the reliability of prediction using the prediction sub-model; by processing multiple prediction sub-models in parallel and then performing aggregate calculation, the prediction accuracy is guaranteed and the processing efficiency is improved, thereby meeting the real-time requirements; by jointly inputting the auxiliary feature vector and the intrinsic mode subsequence into the prediction sub-model, multi-modal fusion is achieved, making full use of multi-dimensional flight parameter information, improving the robustness and generalization ability of the prediction process, thereby improving the real-time performance and reliability of hard landing prediction. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 The flowchart shows a real-time relanding prediction method based on multimodal fusion provided in Embodiment 1 of the present invention.
[0017] Figure 2 This is a schematic diagram of the model structure in a real-time relanding prediction method based on multimodal fusion provided in Embodiment 1;
[0018] Figure 3 This is a schematic diagram of another model structure in a real-time relanding prediction method based on multimodal fusion provided in Embodiment 1;
[0019] Figure 4 The figure shows the experimental verification results of a multimodal fusion real-time relanding prediction method provided in Example 1;
[0020] Where, x VRTG,t-L+1:t Let IMF1, IMF2, ..., IMF be the first subsequence corresponding to the time point from the (t-L+1)th time point to the current tth time point in the flight parameter sequence. N These are the 1st, 2nd, ..., Nth intrinsic mode subsequences, X t-L+1:t For the second subsequence, F z Let Informer1, Informer2, ..., InformerN be the 1st, 2nd, ..., Nth prediction sub-models, and let IMF be the auxiliary feature vector. 1,pre IMF 2,pre ..., IMF N,pre These are the 1st, 2nd, ..., Nth predicted subsequences, y(t+1:t+H) is the target predicted sequence, Embedding3 is the third embedding layer, and Q... i3 The third embedding feature vector, TCN1 is the first temporal convolutional layer, Embedding1 is the first embedding layer, TCN2 is the second temporal convolutional layer, Embedding2 is the second embedding layer, and IMF is the third embedding feature vector. i,feed_en Let J be the input sequence of the i-th encoder. i1 Let Q be the i-th first time-series feature matrix. i1 For the i-th first embedded feature vector, V i1 For the i-th first modal eigenvector, IMF i,feed_de For the i-th decoder input sequence, J i2 Let Q be the i-th second time-series feature matrix. i2 For the i-th second embedding feature vector, E i For the i-th second modal eigenvector, V i3 For the i-th first cross feature vector, V i2 Let be the i-th second cross feature vector. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the terms used to distinguish similar objects can be interchanged so that the invention can also be implemented in other embodiments besides the illustrated or described embodiments. Furthermore, the terms "including," "having," and any variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0023] Example 1
[0024] This first embodiment provides a real-time relanding prediction method based on multimodal fusion, such as... Figure 1 As shown, the real-time relanding prediction method based on multimodal fusion includes the following steps:
[0025] S1 decomposes the first subsequence in the real-time transmitted flight parameter sequence into N intrinsic mode subsequences, where N is a positive integer.
[0026] S2, based on the auxiliary feature encoder, extract features from the second subsequences corresponding to the M auxiliary parameter types in the flight parameter sequence to obtain auxiliary feature vectors, where M is a positive integer.
[0027] S3. For any intrinsic mode subsequence, predict the predicted subsequence corresponding to the intrinsic mode subsequence based on the prediction sub-model corresponding to the intrinsic mode subsequence, the intrinsic mode subsequence, and the auxiliary feature vector.
[0028] S4 aggregates and calculates the predicted subsequences corresponding to the N intrinsic mode subsequences to obtain the target prediction sequence, and provides real-time risk warning for aircraft hard landing by comparing the target prediction sequence with a preset threshold.
[0029] The Quick Access Recorder (QAR) is an airborne real-time data recorder used to store multi-dimensional parameters throughout the entire flight phase, i.e., a multi-dimensional sequence of flight parameters, enabling rapid download and analysis. In this embodiment, real-time risk warnings for aircraft hard landings are based on the real-time transmitted flight parameter sequences.
[0030] Vertical acceleration parameters refer to the core indicators of landing impact, which are physically represented as the rate of change of the aircraft's velocity along the vertical direction. In this embodiment, the first subsequence corresponding to the vertical acceleration parameter type is used as the main sequence. Auxiliary parameters refer to non-main sequence parameters that are strongly related to landing safety and are used to describe boundary conditions such as aircraft attitude, weight, speed, and altitude. The second subsequence corresponding to the auxiliary parameter type is used in this embodiment to supplement additional information.
[0031] The intrinsic mode subsequence refers to the stationary components with different frequencies and scales obtained by decomposing the first subsequence. The intrinsic mode subsequence can characterize the local fluctuation characteristics of the vertical acceleration parameter over time.
[0032] A prediction sub-model refers to a dedicated prediction unit designed for a corresponding intrinsic mode sub-sequence, which can adapt to the frequency characteristics of that intrinsic mode sub-sequence.
[0033] Predicted subsequences refer to time series obtained after time series prediction processing of corresponding intrinsic mode subsequences.
[0034] The predicted subsequences corresponding to the N intrinsic mode subsequences are aggregated and calculated to obtain the final reliable target prediction sequence. Essentially, this is the inverse operation of the decomposition algorithm used in the first subsequence decomposition. In one embodiment, the first subsequence is decomposed into multiple intrinsic mode subsequences using a variational mode decomposition algorithm. Correspondingly, the inverse operation of variational mode decomposition, i.e., linear superposition, is used to aggregate and calculate the N predicted subsequences to obtain the target prediction sequence.
[0035] In this embodiment, the aircraft hard landing risk warning is determined based on the peak value of the vertical acceleration parameter in the target prediction sequence. As an example, the warning is implemented by comparing the peak value of the vertical acceleration parameter in the target prediction sequence with a preset threshold. When the peak value of the vertical acceleration parameter in the target prediction sequence is greater than the preset threshold, an aircraft hard landing risk warning is generated. Furthermore, the implementer can set multiple threshold intervals according to the actual situation, and realize different levels of warnings according to the threshold interval to which the peak value of the vertical acceleration parameter in the target prediction sequence belongs.
[0036] Specifically, after obtaining multi-dimensional parameters for the entire flight phase, in order to improve the accuracy of subsequent predictions, the multi-dimensional parameters for the entire flight phase are preprocessed and standardized to obtain a flight parameter sequence, ensuring that all dimensions in the flight parameter sequence are processed under the same standard. Those skilled in the art will recognize that any preprocessing and standardization method in the prior art falls within the protection scope of this invention, and will not be elaborated upon here.
[0037] In one specific embodiment, the auxiliary parameter type includes at least one of yaw angle, ground speed, gross weight, altitude, indicated vertical speed, and lateral acceleration. In this embodiment, the auxiliary parameter type may simultaneously include yaw angle, ground speed, gross weight, altitude, indicated vertical speed, lateral acceleration, longitude, pitch angle, pitch rate, roll angle, roll rate, vacuum speed, wind speed, and yaw rate.
[0038] It should be noted that parameter types can be screened based on flight dynamics principles, correlation analysis, and engineering knowledge to achieve precise selection of auxiliary parameter types. This ensures the interpretability of parameters for landing loads and vertical acceleration while avoiding redundant parameters that increase model computational overhead. Specifically, an initial screening is performed based on the flight dynamics laws of the aircraft landing phase, selecting core parameters strongly correlated with vertical acceleration, including ground speed, gross weight, altitude, indicated vertical velocity, lateral acceleration, yaw angle, pitch angle, and other basic parameters. Parameters such as avionics and navigation aids that are not directly related to landing impact are eliminated. Then, the Pearson correlation coefficient is calculated between the initially screened parameters and the first subsequence corresponding to vertical acceleration. A correlation coefficient threshold is set (0.3 in this embodiment), retaining parameters with correlation coefficients exceeding the threshold and eliminating low-correlation redundant parameters. Simultaneously, multicollinearity among parameters is tested using the variance inflation factor, and parameters with a variance inflation factor less than 10 are included in the final candidate parameter set to avoid model overfitting caused by parameter coupling. Further combining QAR data from actual operations, the quantitatively screened parameters were validated in practice. Parameters that could stably characterize landing status under different aircraft types, routes, and weather conditions were retained, ultimately determining several auxiliary parameter types. Those skilled in the art will recognize that the Pearson correlation coefficient calculation method in the prior art falls within the scope of this invention, and will not be elaborated upon here.
[0039] It should be noted that in this embodiment, the auxiliary feature encoder, each prediction sub-model, and the aggregation calculation model can be trained in an end-to-end manner. The training loss function can be expressed as a weighted sum of the first sub-loss and the second sub-loss. The first sub-loss is the weighted sum of the mean squared errors of the actual outputs and target outputs of each prediction sub-model, and the second sub-loss is the mean squared error of the aggregation calculation result and the sample sequence. This ensures the overall prediction accuracy and provides independent supervision for the prediction of each intrinsic mode sequence, preserving the unique time-frequency characteristics of the intrinsic mode sequences.
[0040] The variational mode decomposition algorithm decomposes the non-stationary first subsequence into intrinsic mode subsequences with different center frequencies. Each frequency mode corresponds to different physical processes during the landing phase. The high-frequency intrinsic mode subsequence (center frequency > 10Hz) corresponds to the instantaneous tire-runway impact at the moment of aircraft touchdown and the high-frequency valve action of the landing gear dampers, directly reflecting the core dynamic characteristics of the landing impact and serving as a key indicator of hard landing risk. The mid-frequency intrinsic mode subsequence (5Hz ≤ center frequency ≤ 10Hz) corresponds to the damped oscillations of the landing gear struts and aircraft structural vibrations, representing the secondary response of the landing impact. The low-frequency intrinsic mode subsequence (center frequency < 5Hz) corresponds to the trend terms of basic operating conditions such as aircraft heave rate and total weight, exhibiting lower discriminative power for hard landing risk. Correspondingly, the high-center-frequency intrinsic mode sequences typically contain more valuable temporal variation information. Therefore, in the training loss function, a higher weight is assigned to the mean squared error corresponding to the high-center-frequency intrinsic mode sequences, enabling the model to focus on fine-grained details and sudden changes during landing, thereby improving the prediction accuracy of key events. For example, a weighting coefficient of 2.0 is assigned to intrinsic modal subsequences with high center frequencies, and a weighting coefficient of 1.0 is assigned to intrinsic modal subsequences with medium and low frequencies. The total loss function is a weighted combination of the weighted sum of the losses of each modal sub-model and the weighted sum of the final aggregated prediction loss. This allows the model to focus on fine-grained high-frequency temporal details and sudden shock changes during training, thereby improving the prediction accuracy of critical events for re-landing.
[0041] This embodiment ensures feature validity and reduces computational redundancy by providing core auxiliary parameter types in the flight parameter sequence, thereby reducing model complexity and computational overhead.
[0042] In one specific embodiment, S1 includes the following steps:
[0043] The variational mode decomposition algorithm is used to decompose the first subsequence into N intrinsic mode subsequences.
[0044] Variational mode decomposition (VMD) is an adaptive signal decomposition algorithm that finds the center frequency and bandwidth of each mode through iterative optimization and outputs stable intrinsic mode function (IMF) components, i.e., intrinsic mode subsequences.
[0045] The intrinsic mode subsequence is used to characterize the fluctuation of the corresponding vertical acceleration in a specific frequency range. For example, the high-frequency component corresponds to the instantaneous impact of tire-runway contact, the high-frequency valve action of the shock absorber, etc., the mid-frequency component reflects the damped oscillation of the landing gear strut and structural vibration, and the residual component represents the trend terms related to the sinking rate and mass configuration, etc.
[0046] Specifically, the first subsequence includes vertical acceleration parameters corresponding to several preset time points in a temporal order. Typically, the time interval between adjacent preset time points is fixed at W, where W is a positive integer; for example, the corresponding sampling frequency is 8Hz. Using the first subsequence as the initial signal, the variational mode decomposition algorithm adaptively decomposes the initial signal into multiple mode signals using a variational method without relying on recursive operations. Each mode signal is a signal with a finite frequency band, which can better adapt to the complex non-stationary characteristics of the initial signal.
[0047] In one implementation, the implementer can use decomposition algorithms such as empirical mode decomposition, ensemble empirical mode decomposition, and wavelet decomposition to decompose the first subsequence. The number N of intrinsic mode subsequences can be adaptively selected by indicators such as envelope entropy and cross-correlation coefficient.
[0048] This embodiment uses a variational mode decomposition algorithm to decompose the first subsequence into multiple intrinsic mode subsequences. The decomposition accuracy is high, and it can avoid the problem of mode aliasing. The decomposed intrinsic mode subsequences are stable and easy to model, which can effectively reduce the learning difficulty of subsequent prediction sub-models, thereby improving prediction accuracy.
[0049] In one specific implementation, the prediction sub-model includes a prediction encoder and a prediction decoder, and the intrinsic mode sub-sequence includes the intrinsic mode parameters corresponding to the (t-L+1)th time point to the current tth time point. S3 includes the following steps:
[0050] S31, For any intrinsic mode subsequence, input the intrinsic mode subsequence into the predictive encoder corresponding to the intrinsic mode subsequence to obtain the first mode feature vector.
[0051] S32, extract the intrinsic mode parameters corresponding to the (t-i+1)th time point to the current tth time point in the intrinsic mode subsequence to obtain a temporary subsequence, where i is an integer less than t.
[0052] S33, concatenate the temporary subsequence and the preset alignment sequence to obtain the decoder input sequence, wherein the preset alignment sequence includes Li zero parameters.
[0053] S34, input the first modality feature vector, the auxiliary feature vector and the decoder input sequence into the prediction decoder corresponding to the intrinsic modality subsequence to obtain the prediction subsequence corresponding to the intrinsic modality subsequence.
[0054] The predictive encoder is used to extract the temporal features of the corresponding intrinsic modal subsequence and output the first modal feature vector. The temporary subsequence refers to the data of the most recent i time points of the extracted intrinsic modal subsequence. The preset alignment sequence is used to pad the temporary subsequence to a fixed length L to meet the input format of the predictive decoder.
[0055] The predictive decoder is used to fuse the auxiliary feature vector with the output of the predictive encoder to generate a predictive subsequence for future time steps.
[0056] Specifically, temporary subsequences are used to capture short-term fluctuation information of the corresponding intrinsic mode subsequences, auxiliary feature vectors are used to provide flight status and environmental information, and alignment sequences are used to ensure temporal consistency.
[0057] As an example, temporary subsequences can be adaptively truncated using a sliding window approach. The sliding window can be implemented using a weighted sliding window to enhance the weight of recent data.
[0058] In this embodiment, i = rounddown(t / 2) + 1, where rounddown() is a round-down function, meaning the temporary subsequence is the last 50% of the most recent data of the intrinsic mode subsequence. This ensures the capture of short-term features while preserving the length format through the alignment sequence. For example, when t = 71 (the intrinsic mode subsequence length is 72, corresponding to 9s of flight data at 8Hz sampling), i = 36, the temporary subsequence length is 36, the preset alignment sequence contains 36 zero parameters, and the length of the prediction decoder input sequence is still 72.
[0059] In one embodiment, the preset alignment sequence may also use a method such as mean filling instead of zero-value filling.
[0060] In this embodiment, a predictive encoder and predictive decoder architecture is used to achieve temporal prediction for each intrinsic modal subsequence, and auxiliary feature vectors are fused with modal feature vectors to enhance the sub-model's adaptability to boundary conditions, thereby improving the stability and reliability of the sub-model's prediction.
[0061] In one specific embodiment, the predictive encoder includes a first temporal convolutional layer, a first embedding layer, and a first attention layer, and S31 includes the following steps:
[0062] S311, For any intrinsic mode subsequence, input the intrinsic mode subsequence into the first temporal convolutional layer corresponding to the intrinsic mode subsequence to obtain the first temporal feature matrix.
[0063] S312, the first temporal feature matrix is input into the first embedding layer corresponding to the intrinsic mode subsequence to obtain the first embedding feature vector.
[0064] S313, the first embedded feature vector is input into the first attention layer corresponding to the intrinsic modality subsequence to obtain the first modality feature vector corresponding to the intrinsic modality subsequence.
[0065] The first temporal convolutional layer is used to capture short-term temporal dependencies and enhance local feature fusion, outputting a first temporal feature matrix. In this embodiment, the temporal convolutional layer is implemented using a temporal convolutional network (TCN) model.
[0066] The first embedding layer is used to project scalar time series features into a continuous vector space and supplement the missing time sequence information in the architecture.
[0067] The first attention layer includes an attention module and an attention distillation module. The attention distillation module uses the attention distillation mechanism to achieve efficient representation learning while reducing computational complexity.
[0068] Specifically, the temporal convolutional network model in this embodiment consists of causal convolution and dilated convolution. Causal convolution ensures that the output at the next preset time point depends only on the input at the current and previous preset time points, thereby preventing information leakage from the future. Dilated convolution introduces a dilation factor into the convolution kernel, exponentially expanding the receptive field while maintaining computational efficiency. Furthermore, the temporal convolutional network model in this embodiment introduces residual connections, which helps gradients propagate through the temporal convolutional network model, mitigating model performance degradation and improving model stability and convergence.
[0069] The first embedding layer includes token embedding and location embedding. Token embedding is used to project scalar time-series features into a continuous vector space, thereby enabling joint modeling of different IMF components. Location embedding is used to supplement the missing temporal order information in the architecture.
[0070] The first embedding feature vector can be obtained by adding the embedding representations of the Token embedding and the position embedding element by element.
[0071] It should be noted that the first modality feature vector output by the first attention layer has undergone intermediate processing, which refers to residual connection and layer normalization to stabilize gradient propagation and enhance feature fusion.
[0072] To further compress redundant information and enhance inter-layer feature representation, an attention distillation module is introduced after the attention module. The attention distillation mechanism first applies a one-dimensional convolution to the output of the attention module, followed by max pooling, to downsample and refine significant temporal features, focusing on the temporal structure with the most information, reducing redundancy, and improving the compactness and stability of the learned representation.
[0073] It should be noted that the first attention layer can include multiple attention modules and an attention distillation module to mine deep features. The attention modules employ an improved sparse attention mechanism, retaining only the top u queries with the most information for attention calculation; the attention distillation module performs one-dimensional convolution and max pooling operations on the output of the attention modules sequentially, downsampling and refining significant temporal features, focusing on the temporal structure with the most information. The output of the first attention layer, after residual connection and layer normalization, yields the first modality feature vector.
[0074] This embodiment uses a combination of temporal convolution, information embedding, and attention mechanisms to efficiently extract temporal feature information. It can take into account both local and global information, improve the utilization rate of key information, reduce noise interference, and thus improve the reliability of the subsequent decoding process.
[0075] In one specific embodiment, the prediction decoder includes a second temporal convolutional layer, a second embedding layer, a second attention layer, a first cross-attention layer, a second cross-attention layer, and a fully connected layer. S34 includes the following steps:
[0076] S341, the decoder input sequence is input into the second temporal convolutional layer corresponding to the intrinsic modality subsequence to obtain the second temporal feature matrix.
[0077] S342, input the second temporal feature matrix into the second embedding layer corresponding to the intrinsic mode subsequence to obtain the second embedding feature vector.
[0078] S343, the second embedded feature vector is input into the second attention layer corresponding to the intrinsic modality subsequence to obtain the second modality feature vector.
[0079] S344, the auxiliary feature vector and the second modality feature vector are input into the first cross-attention layer corresponding to the intrinsic modality subsequence to obtain the first cross-feature vector.
[0080] S345, input the first cross feature vector and the first modality feature vector corresponding to the intrinsic modality subsequence into the second cross attention layer corresponding to the intrinsic modality subsequence to obtain the second cross feature vector.
[0081] S346, the second cross feature vector is input into the fully connected layer to obtain the predicted subsequence corresponding to the intrinsic mode subsequence.
[0082] The second temporal convolutional layer has the same structure as the first temporal convolutional layer. It is used to capture short-term temporal dependencies and enhance local feature fusion, and outputs a second temporal feature vector.
[0083] The second embedding layer has the same structure as the first embedding layer. It is used to project scalar time series features into a continuous vector space and supplement the missing time sequence information in the architecture.
[0084] The second attention layer includes an attention module. The output of the second attention layer is processed by intermediate processing to obtain the second modality feature vector. The intermediate processing refers to residual connection and layer normalization to stabilize gradient propagation and enhance feature fusion.
[0085] The first cross-attention layer is used to fuse the auxiliary feature vector and the second modality feature vector to output the first cross-feature vector. The second cross-attention layer is used to fuse the historical information in the first modality feature vector into the first cross-feature vector to output the second cross-feature vector.
[0086] The fully connected layer is used to map the second cross feature vector into a predicted subsequence, thus completing the dimensionality transformation.
[0087] Specifically, the first cross-attention layer transposes the input auxiliary feature vector and the second modality feature vector, then performs cross-attention operation along the feature dimension, and then transposes the calculation result to obtain the first cross feature vector. This enables the model to automatically identify and integrate the complex dynamic relationships between multi-source features, thereby improving prediction accuracy and generalization ability.
[0088] The second cross-attention layer fuses the first cross-feature vector as the query vector and the first modality feature vector as the key vector, enabling the predictive decoder to selectively focus on relevant temporal features in the encoded sequence. This effectively captures historical information.
[0089] In one embodiment, the predictive decoder further includes a feedforward network layer. After the second cross-feature vector is output from the second cross-attention layer, the second cross-feature vector is input into the feedforward network layer to employ a feedforward network (FFN) to enhance feature abstraction and improve nonlinear representation capabilities. The output of the feedforward network layer is then input into a fully connected layer.
[0090] This embodiment fully integrates auxiliary features and encoder features through a dual cross-attention layer, thereby improving prediction reliability. Moreover, multi-level feature processing can adapt to complex flight scenarios and enhance the model's generalization ability.
[0091] In one specific embodiment, the auxiliary feature encoder includes a third embedding layer, a third attention layer, and a feedforward neural network layer, and S2 includes the following steps:
[0092] S21, input the M second subsequences into the third embedding layer to obtain the third embedding feature vector.
[0093] S22, the third embedded feature vector is input into the third attention layer to obtain the third modality feature vector.
[0094] S23, input the third modality feature vector into the feedforward neural network layer to obtain the auxiliary feature vector.
[0095] The third embedding layer projects the M second subsequences of the input onto a low-dimensional latent space, while incorporating positional encoding to inject temporal order information. The third embedding layer consists of two parts: token embedding and positional embedding. Token embedding projects the scalar time series of each auxiliary parameter type onto a continuous d-dimensional latent space. model In a dimensional vector space, we obtain a dimension of L×d. model The embedding matrix is used to inject temporal positional information into the embedding matrix to preserve the temporal order. The third embedding feature vector is obtained by element-wise addition of the token embedding and the positional embedding.
[0096] The third attention layer also includes an attention module. Specifically, the auxiliary feature encoder may include multiple sets of third attention layers and feedforward neural network layers. The third attention layer adopts a multi-head improved sparse attention mechanism, retaining only the top u queries with the most information in the attention distribution for attention calculation, thus reducing computational complexity. The output of the third attention layer also needs to undergo residual connection and layer normalization processing to obtain the third modality feature vector, thereby enhancing model stability and feature fusion capability.
[0097] The third modality feature vector is input into a feedforward neural network layer, which consists of two linear transformations and a nonlinear activation function to enhance the nonlinear expressive power of the features. The output of the feedforward neural network layer is processed through residual connections and layer normalization to obtain an auxiliary feature vector.
[0098] Before passing the auxiliary feature vector output by the auxiliary feature encoder to the decoder, it is transposed to achieve better fusion between the auxiliary feature information and the feature representation corresponding to the vertical acceleration parameter.
[0099] In this embodiment, all prediction sub-models share the same auxiliary feature encoder, which enhances feature generalization ability while maintaining computational efficiency. This not only reduces the complexity of model parameters but also ensures targeted modeling of each intrinsic modality sequence, balancing model simplicity and prediction accuracy.
[0100] In one specific implementation, the first attention layer, the second attention layer, and the third attention layer employ an improved sparse attention mechanism.
[0101] The first attention layer, the second attention layer, and the third attention layer all contain attention modules. In this embodiment, the attention module adopts an improved sparse attention mechanism. The improved sparse attention mechanism only retains the top u most informative queries that play a dominant role in the attention distribution for attention calculation, where u is a positive integer. This reduces computational complexity by selectively calculating while preserving the ability to model key long-term dependencies as much as possible.
[0102] Specifically, the amount of information corresponding to each query can be represented by a sparsity metric.
[0103] In this embodiment, the first subsequence corresponding to the vertical acceleration parameter type is decomposed into multiple intrinsic mode subsequences, which effectively reduces the complexity of the first subsequence and improves the reliability of prediction using the prediction sub-model. By processing multiple prediction sub-models in parallel and then performing aggregate calculations, the prediction accuracy is guaranteed and the processing efficiency is improved, thereby meeting the real-time requirements. The auxiliary feature vector and the intrinsic mode subsequence are jointly input into the prediction model to achieve multi-modal fusion, making full use of multi-dimensional flight parameter information, improving the robustness and generalization ability of the prediction process, and thus improving the real-time performance and reliability of hard landing prediction.
[0104] Example 2
[0105] Figure 2 This is a schematic diagram of the model structure in the real-time relanding prediction method based on multimodal fusion provided in Example 1, including an auxiliary feature encoder and N prediction sub-models, namely Informer1, Informer2, ..., InformerN, specifically:
[0106] First, the variational mode decomposition algorithm is used to extract the first subsequence x corresponding to the time point from the (t-L+1)th time point to the current tth time point in the flight parameter sequence. VRTG,t-L+1:t Decomposed into N intrinsic mode subsequences, namely IMF1, IMF2, ..., IMF N .
[0107] Based on the auxiliary feature encoder, the second subsequence X corresponding to the M auxiliary parameter types in the flight parameter sequence is respectively... t-L+1:t Feature extraction is performed to obtain the auxiliary feature vector F. z Among them, X L-t+1:L ∈R d-1 d is the total data dimension of the flight parameter sequence, d=M+1, which includes the sequence dimension corresponding to M auxiliary parameter types and the sequence dimension corresponding to a first subsequence.
[0108] Each intrinsic mode subsequence corresponds to a prediction submodel. Based on the prediction submodel Informer1 corresponding to the first intrinsic mode subsequence, the intrinsic mode subsequence IMF1, and the auxiliary feature vector F... z Prediction is performed to obtain the predicted subsequence IMF corresponding to the intrinsic mode subsequence. 1,pre And IMF 1,pre ∈R (H×1) Accordingly, N predicted subsequences IMF are obtained. 1,pre IMF 2,pre ..., IMF N,pre .
[0109] Furthermore, for each of the N intrinsic mode subsequences, the predicted subsequence IMF is calculated. 1,pre IMF 2,pre ..., IMF N,pre Aggregation calculations are performed to obtain the final reliable target prediction sequence y(t+1:t+H), where H is the length of the target prediction sequence.
[0110] Example 3
[0111] Figure 3 This is a schematic diagram of another model structure in the real-time relanding prediction method based on multimodal fusion provided in Example 1, specifically:
[0112] The auxiliary feature encoder includes a third embedding layer (Embedding3), a third attention layer, and a feedforward neural network layer, which converts M second subsequences X... t-L+1:t The input is fed into the third embedding layer, Embedding3, to obtain the third embedding feature vector Q. i3 The third embedded feature vector Q i3 The input is fed into the third attention layer to obtain the third modality feature vector. This third modality feature vector is then fed into the feedforward neural network layer to obtain the auxiliary feature vector F. z In this embodiment, the third attention layer includes a multi-head ProSparse self-attention mechanism, residual connections, and normalization, while the feedforward neural network layer includes a feedforward neural network, residual connections, and a normalization layer.
[0113] Each intrinsic modality subsequence corresponds to a prediction sub-model. Each prediction sub-model includes a prediction encoder and a prediction decoder. Each prediction encoder includes a first temporal convolutional layer TCN1, a first embedding layer Embedding1, and a first attention layer. Each prediction decoder includes a second temporal convolutional layer TCN2, a second embedding layer Embedding2, a second attention layer, a first cross-attention layer, a second cross-attention layer, and a fully connected layer. In this embodiment, the temporal convolutional layer includes a causal convolutional layer, a dilated convolutional layer, and a residual connection layer; the embedding layer includes token embedding and position embedding; the first attention layer includes several sets of multi-head ProSparse self-attention mechanisms, one-dimensional convolutions, and max pooling layers; the second attention layer includes a multi-head ProSparse self-attention mechanism, residual connections, and a normalization layer; the first cross-attention layer includes a feature multi-head cross-attention mechanism, residual connections, and a normalization layer; the second cross-attention layer includes a multi-head cross-attention mechanism, residual connections, a normalization layer, a feedforward neural network, and residual connections and a normalization layer.
[0114] This embodiment uses the i-th intrinsic mode subsequence as an example for illustration. The i-th intrinsic mode subsequence IMF is input to the predictive encoder. i,feed_en Including the intrinsic mode parameters corresponding to the (t-L+1)th time point to the current tth time point, the intrinsic mode subsequence IMF is calculated. i,feed_en The input is fed into the first temporal convolutional layer TCN1 corresponding to the intrinsic modality subsequence to obtain the first temporal feature matrix J. i1 The first time-series feature matrix J i1 The input is fed into the first embedding layer, Embedding1, to obtain the first embedding feature vector Q. i1 The first embedded feature vector Q i1 The input is fed into the first attention layer to obtain the first modality feature vector V corresponding to the intrinsic modality subsequence. i1 .
[0115] Extract the intrinsic mode parameters from the (t-i+1)th time point to the current tth time point in the intrinsic mode subsequence to obtain the temporary subsequence IMF. i,token The temporary subsequence IMF i,token Alignment sequence IMF with preset i,0 By concatenating the sequences, the decoder input sequence IMF is obtained. i,feed_de .
[0116] The first mode feature vector V i1 Auxiliary feature vector F z and decoder input sequence IMF i,feed_de The input is fed into the predictor decoder corresponding to the intrinsic mode subsequence to obtain the predictor subsequence IMF corresponding to the intrinsic mode subsequence.i,pre The specific process includes: inputting the decoder into the IMF sequence. i,feed_de The input is fed into the second temporal convolutional layer TCN2 corresponding to the intrinsic modality subsequence to obtain the second temporal feature matrix J. i2 The second time-series feature matrix J i2 The input is fed into the second embedding layer Embedding2 corresponding to the intrinsic modality subsequence to obtain the second embedding feature vector Q. i2 The second embedded feature vector Q i2 The input is fed into the second attention layer corresponding to the intrinsic modality subsequence to obtain the second modality feature vector E. i The auxiliary feature vector F z Second mode eigenvector E i The input is fed into the first cross-attention layer corresponding to the intrinsic modality subsequence to obtain the first cross-feature vector V. i3 The first cross feature vector V i3 The first modal eigenvector V corresponding to this intrinsic modal subsequence i1 The input is fed into the second cross-attention layer corresponding to the intrinsic modality subsequence to obtain the second cross-feature vector V. i2 The second cross feature vector V i2 The input is fed into a fully connected layer to obtain the predicted subsequence IMF corresponding to the intrinsic mode subsequence. i,pre .
[0117] Example 4
[0118] Figure 4 The figure shows the experimental verification results of a multimodal fusion real-time relanding prediction method provided in Example 1. Specifically:
[0119] To quantitatively evaluate the performance of the multimodal fusion-based real-time hard landing prediction method in real-world flight scenarios, experimental validation was conducted using a real QAR dataset provided by an airline. The dataset contains landing phase data from 5000 B737-800 flights, covering landing scenarios under different routes, weather conditions, and pilot operating modes. The sampling frequency is 8Hz, and each flight records all flight parameters from radio altitude 50ft to 2 seconds after touchdown. The dataset is divided into training, validation, and test sets in a 7:2:1 ratio. It should be noted that all data underwent preprocessing: outliers were removed using the 3σ criterion, missing values were filled using linear interpolation, Z-score standardization was used to eliminate dimensional differences, and finally, the data was uniformly resampled to an 8Hz sampling frequency to ensure data consistency.
[0120] The core parameters of this embodiment are set as follows: the number of intrinsic mode subsequences obtained after variational mode decomposition of the vertical acceleration sequence (first subsequence) is N=6; the input time window length is t=72 (corresponding to 9s of historical data); the prediction step size is 8 (corresponding to the vertical acceleration sequence in the next 1s); the auxiliary parameter type M selects 14 auxiliary parameters that are strongly related to landing impact, including yaw angle, ground speed, total weight, altitude, indicated vertical velocity, lateral acceleration, longitude, pitch angle, pitch rate, roll angle, roll rate, vacuum speed, wind speed, and yaw rate. All comparison models use the same training data and input / output format. Specifically, the number of channels in the temporal convolutional network of the predictive encoder and predictive decoder is [128, 256, 512], and the kernel size is 3; the attention layer adopts an improved sparse attention mechanism with 8 attention heads; the model dimension d_model=512, and the feedforward network dimension dff=2048. The batch size in the training parameters is 32, and the initial learning rate is 1×10. -4 The training rounds were 100, the optimizer was Adam, and the loss function used was weighted multimodal loss (final prediction loss weight α=2.0, high center frequency mode loss weight β=2.0, and mid-to-low frequency mode weight β=1.0).
[0121] The evaluation indicators used are root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 The accuracy of vertical acceleration sequence prediction is measured; for the ability to identify hard landing events, accuracy, recall, precision, and F1 score are used for comprehensive evaluation.
[0122] exist Figure 4 In the experiment, multiple trigger altitudes were set. When the aircraft descended below the trigger altitude, flight data from the previous 9 seconds (72 sampling points) was extracted as input to predict the vertical acceleration sequence during the subsequent landing impact phase. Experiment 1 demonstrates the evolution of the prediction accuracy of the vertical acceleration sequence at different trigger altitudes (Height (ft)), specifically MAE, RMSE (left axis), and R... 2 The evolution of (right axis) is shown. Results indicate that optimal performance is achieved at a trigger height of 4ft, with RMSE=0.0305, MAE=0.024, and R... 2=0.917. Prediction accuracy improves significantly as the trigger altitude approaches ground contact. At higher altitudes (10-30 feet), prediction errors increase because the critical risk of a hard landing accumulates primarily in the final moments before touchdown, while the aircraft behaves relatively normally in the earlier stages. Once descending below 10 feet, the model achieves stable high accuracy, especially in the 3-7 foot range. These experimental results validate the feasibility of real-time warning, identifying 3-7 feet as the optimal trigger altitude range, meeting the timing requirements of airborne real-time monitoring.
[0123] Experiment 2 demonstrates the comparison between predicted and actual values of the maximum vertical load (peak vertical acceleration) at different trigger altitudes. The experiment shows that in the higher trigger altitude range (50ft–60ft), due to the aircraft still being in the early approach phase and the landing dynamics not yet fully manifested, there is a certain deviation between the predicted and actual values, with an error range of approximately 0.05–0.10g, but the basic trend direction can still be captured. In the medium trigger altitude range (20ft–40ft), as the aircraft enters the leveling phase, key dynamic characteristics gradually emerge, the prediction error narrows significantly, the RMSE drops to within 0.03, and the overlap between the predicted and actual curves significantly improves. In the low trigger altitude range (10ft–15ft), within the critical warning window before touchdown, the prediction results converge rapidly, with the deviation from the actual value remaining within 0.02g, accurately estimating the impending maximum impact load. These experimental results verify that the multimodal fusion real-time hard landing prediction method has the ability to estimate the maximum impact load in real time before touchdown, meeting the early warning requirements for high-risk events during the landing phase.
[0124] Experiment 3 illustrates the impact of different modality numbers (N value) on prediction error, specifically the prediction errors when N=4, 6, 8, and 10. When N=4, multi-scale dynamic feature separation is insufficient, failing to capture the multi-scale features of the vertical acceleration signal; RMSE=0.038, MAE=0.030, R... 2 =0.891; When N=8 or 10, the introduction of redundant components leads to over-decomposition, resulting in a decrease in prediction performance, with RMSE increasing to 0.037 and 0.047 respectively. 2 As the R² values decrease to 0.834 and 0.801, the generalization ability declines. The model achieves the best balance between accuracy and stability when N=6, with RMSE=0.025, MAE=0.020, and R²=0.025. 2 =0.953, which effectively separates high-frequency impact characteristics from low-frequency trend characteristics while avoiding noise introduced by redundant modes. This experimental result verifies the rationality of the variational mode decomposition in the specification, determining N=6 as the optimal number of modes, which can effectively reduce the complexity of the first subsequence and improve prediction reliability.
[0125] Experiment 4 demonstrates how the input sequence length varies from 2s to 9s. The results show that a 9-second input window (72 sampling points) provides the most discriminative temporal information, effectively characterizing the key dynamic features of the landing phase, with the lowest RMSE (0.025). A 6-second window, due to insufficient data, cannot capture long-term dependencies, resulting in an RMSE of 0.036. A 12-second window introduces redundant temporal information, increasing computational overhead, and while the RMSE does not decrease significantly (0.024), inference time increases by 40%. These experimental results validate the rationality of the sequence length design, determining 9 seconds as the optimal input window, balancing prediction accuracy and computational efficiency.
[0126] Furthermore, under optimal parameter settings (N=6, trigger height 3ft, input window 9 seconds), the prediction performance of the model in this embodiment on the test set and its comparison with the prediction performance of existing mainstream models (including Informer, Transformer, and Long Short-Term Memory (LSTM)) are shown in Table 1 below. The results show that this embodiment has the strongest ability to discriminate the risk of re-landing in actual operation and the lowest false alarm rate, verifying the effectiveness of multimodal fusion and temporal modeling in this embodiment.
[0127] Table 1. Prediction performance of different models for vertical acceleration prediction tasks.
[0128]
[0129] Furthermore, the recognition performance of the model in this embodiment on the test set and its recognition performance compared with existing machine learning and deep learning models (including Support Vector Machine (SVM), Random Forest, BP Neural Network, and Long Short-Term Memory (LSTM)) are shown in Table 2 below. The results show that this embodiment can effectively reduce the false alarm rate, meet the stringent reliability requirements of aviation safety, and verify the effectiveness of multimodal fusion and temporal modeling in this embodiment.
[0130] Table 2. Recognition performance of different models for the heavy landing recognition task.
[0131]
[0132] Furthermore, to verify the necessity of the key modules in this embodiment, three sets of ablation experiments were designed based on model variants: removing variational mode decomposition and directly inputting the original vertical acceleration sequence; removing the temporal convolutional layer; and removing the auxiliary feature encoder and using only the vertical acceleration sequence for prediction. The corresponding prediction performance is shown in Table 3.
[0133] Table 3. Prediction performance of different variant models for the vertical acceleration prediction task.
[0134]
[0135] The results show that removing variational mode decomposition significantly reduces model performance, indicating that multi-scale signal decomposition is key to handling non-stationary characteristics of vertical acceleration. Removing the temporal convolutional layer weakens the ability to model short-term time dependencies, verifying the role of the temporal convolutional layer in local feature fusion. Removing the auxiliary feature encoder prevents the model from utilizing the coupling information of multi-dimensional flight parameters, significantly reducing generalization ability and verifying the necessity of multi-modal fusion.
[0136] In summary, the variational mode decomposition, improved sparse attention mechanism, and dual-cross attention fusion techniques employed in this embodiment effectively address the non-stationary characteristics of vertical acceleration signals and fully exploit the coupling relationships among multi-dimensional flight parameters. The prediction accuracy and hard landing recognition accuracy are significantly superior to existing models. Furthermore, the optimal parameter settings balance prediction accuracy and computational efficiency, meeting the requirements of real-time airborne monitoring. Ablation experiments verify the necessity of core modules such as variational mode decomposition, temporal convolutional layers, and auxiliary feature encoders; their synergistic effect ensures the model's robustness and generalization ability. Experimental results fully support the technical advantages of this embodiment, demonstrating that it effectively solves the problems of poor real-time performance, low reliability, and weak generalization ability, and possesses engineering application value.
[0137] Example 5
[0138] Embodiment 5 of the present invention provides a non-transient computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the multimodal fusion real-time relanding prediction method provided in the above embodiment.
[0139] The type of non-transient computer-readable storage medium is not specifically limited, including but not limited to read-only memory (ROM), random access memory (RAM), solid-state drive (SSD), hard disk drive (HDD), USB flash drive, portable hard drive, magnetic disk, optical disk, flash memory card, cloud storage medium, and any carrier capable of persistently storing computer program instructions; the compatible electronic devices cover various hardware devices with data processing capabilities, such as airborne real-time monitoring terminals, ground flight safety monitoring platforms, aircraft health management systems, flight simulation equipment, and general-purpose computer servers. As long as the electronic device is equipped with a general-purpose processor, embedded processor, or dedicated chip (such as FPGA, DSP) capable of loading and executing the instructions / programs, the hard landing prediction method in Embodiment 1 can be implemented.
[0140] Furthermore, this non-transient computer-readable storage medium can be flexibly deployed according to actual application scenarios: In airborne real-time application scenarios, the storage medium can be built into the core processing unit of the aircraft avionics system, directly connecting to the real-time data interface of the aircraft's fast access recorder. After the processor loads and executes instructions / programs, it can realize real-time reception and processing of flight parameters and millisecond-level early warning of hard landing risks, meeting the real-time and reliability requirements of airborne equipment; In ground monitoring scenarios, the storage medium can be deployed in the server of the ground flight safety management platform, connecting to the QAR historical data or real-time streaming data transmitted by the aircraft, realizing hard landing risk monitoring and retrospective analysis of single / multiple aircraft landing processes; In simulation testing scenarios, the storage medium can be integrated into the aircraft flight simulation equipment, combined with the flight parameter sequence generated by simulation, to complete the model debugging, parameter optimization and performance verification of the prediction method of this invention, achieving iterative upgrades of the method without relying on real flight tests.
[0141] Furthermore, the instructions / programs stored in this non-transitory computer-readable storage medium possess modular, portable, and scalable characteristics: the instructions of each functional module (such as the VMD decomposition module, feature encoding module, prediction sub-model module, and risk warning module) are independent of each other and have a unified interface. Model parameters (such as the number of intrinsic modal subsequences N, the number of auxiliary parameter types M, attention mechanism hyperparameters, and warning thresholds) can be flexibly adjusted according to the operational needs of different aircraft types (such as B737 and A320) and different airlines. At the same time, the instructions / programs can be ported across operating systems and hardware platforms and can be executed in different system environments such as Windows, Linux, and embedded real-time operating systems (RTOS) without significant modifications, effectively reducing engineering deployment costs.
[0142] This embodiment solves the problem of traditional hard landing prediction methods being difficult to deploy by solidifying the multimodal fusion real-time hard landing prediction method into computer-executable instructions / programs and storing them in a non-transitory computer-readable storage medium. It achieves the reproducibility, portability, and scalability of the method. At the same time, by encapsulating the complex model calculation and feature fusion logic into standardized instructions, it significantly reduces the deployment difficulty and usage threshold in various electronic devices. High-precision real-time hard landing prediction can be achieved with simple loading and execution. It can be widely used in multiple fields such as flight safety monitoring, aircraft health management, and flight simulation testing in civil aviation, providing reliable technical support for aviation flight safety.
[0143] Example 6
[0144] Embodiment 6 of the present invention provides an electronic device, which includes a processor and a non-transitory computer-readable storage medium as described in Embodiment 5 of the present invention.
[0145] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A real-time relanding prediction method based on multimodal fusion, characterized in that, The multimodal fusion real-time relanding prediction method includes the following steps: S1, decompose the first subsequence in the real-time transmitted flight parameter sequence into N intrinsic mode subsequences, where N is a positive integer, and the first subsequence includes several vertical acceleration parameters that conform to the time order; S2, according to the auxiliary feature encoder, feature extraction is performed on the second subsequences corresponding to the M auxiliary parameter types in the flight parameter sequence to obtain auxiliary feature vectors, where M is a positive integer, and the auxiliary parameter types include at least one of yaw angle, ground speed, total weight, altitude, indicated vertical speed and lateral acceleration; S3, for any intrinsic mode subsequence, predict the predicted subsequence corresponding to the intrinsic mode subsequence based on the prediction sub-model corresponding to the intrinsic mode subsequence, the intrinsic mode subsequence, and the auxiliary feature vector, to obtain the predicted subsequence corresponding to the intrinsic mode subsequence. The prediction sub-model includes a prediction encoder and a prediction decoder. The intrinsic mode subsequence includes intrinsic mode parameters corresponding to time points t-L+1 to the current t-th time point. S3 includes the following steps: S31, For any intrinsic mode subsequence, input the intrinsic mode subsequence into the predictive encoder corresponding to the intrinsic mode subsequence to obtain the first mode feature vector; S32, extract the intrinsic mode parameters corresponding to the (t-i+1)th time point to the current tth time point in the intrinsic mode subsequence to obtain a temporary subsequence, where i is an integer less than t; S33, the temporary subsequence and the preset alignment sequence are concatenated to obtain the decoder input sequence, wherein the preset alignment sequence includes Li zero parameters; S34, input the first modality feature vector, the auxiliary feature vector and the decoder input sequence into the prediction decoder corresponding to the intrinsic modality subsequence to obtain the prediction subsequence corresponding to the intrinsic modality subsequence; S4, aggregate and calculate the prediction subsequences corresponding to the N intrinsic mode subsequences respectively to obtain the target prediction sequence, and provide real-time risk warning for aircraft hard landing by comparing the target prediction sequence with a preset threshold.
2. The real-time relanding prediction method based on multimodal fusion according to claim 1, characterized in that, S1 includes the following steps: The first subsequence is decomposed into the N intrinsic mode subsequences using a variational mode decomposition algorithm.
3. The real-time relanding prediction method based on multimodal fusion according to claim 1, characterized in that, The predictive encoder includes a first temporal convolutional layer, a first embedding layer, and a first attention layer. S31 includes the following steps: S311, For any intrinsic mode subsequence, input the intrinsic mode subsequence into the first temporal convolutional layer corresponding to the intrinsic mode subsequence to obtain the first temporal feature matrix; S312, the first temporal feature matrix is input into the first embedding layer corresponding to the intrinsic mode subsequence to obtain the first embedding feature vector; S313, the first embedded feature vector is input into the first attention layer corresponding to the intrinsic modality subsequence to obtain the first modality feature vector corresponding to the intrinsic modality subsequence.
4. The real-time relanding prediction method based on multimodal fusion according to claim 3, characterized in that, The predictive decoder includes a second temporal convolutional layer, a second embedding layer, a second attention layer, a first cross-attention layer, a second cross-attention layer, and a fully connected layer. S34 includes the following steps: S341, the decoder input sequence is input into the second temporal convolutional layer corresponding to the intrinsic modality subsequence to obtain the second temporal feature matrix; S342, the second temporal feature matrix is input into the second embedding layer corresponding to the intrinsic mode subsequence to obtain the second embedding feature vector; S343, the second embedded feature vector is input into the second attention layer corresponding to the intrinsic modality subsequence to obtain the second modality feature vector; S344, the auxiliary feature vector and the second modality feature vector are input into the first cross-attention layer corresponding to the intrinsic modality subsequence to obtain the first cross-feature vector; S345, input the first cross feature vector and the first modality feature vector corresponding to the intrinsic modality subsequence into the second cross attention layer corresponding to the intrinsic modality subsequence to obtain the second cross feature vector; S346, the second cross feature vector is input into the fully connected layer to obtain the predicted subsequence corresponding to the intrinsic mode subsequence.
5. The real-time relanding prediction method based on multimodal fusion according to claim 4, characterized in that, The auxiliary feature encoder includes a third embedding layer, a third attention layer, and a feedforward neural network layer. S2 includes the following steps: S21, input the M second sub-sequences into the third embedding layer to obtain the third embedding feature vector; S22, the third embedded feature vector is input into the third attention layer to obtain the third modality feature vector; S23, the third modality feature vector is input into the feedforward neural network layer to obtain the auxiliary feature vector.
6. The real-time relanding prediction method based on multimodal fusion according to claim 5, characterized in that, The first attention layer, the second attention layer, and the third attention layer employ an improved sparse attention mechanism.
7. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the real-time relanding prediction method of multimodal fusion as described in any one of claims 1-6.
8. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 7.
Citation Information
Patent Citations
Short-term ship attitude prediction method based on empirical mode decomposition and support vector regression
CN112052623A