Double-flow time series data abnormity early warning method based on cross attention

By using a cross-attention dual-stream time-series data anomaly early warning method, the problem of asynchronous fault identification in complex systems is solved, achieving efficient and reliable fault early warning, and applicable to complex systems with multi-source asynchronous time-series characteristics.

CN121723364APending Publication Date: 2026-03-24NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture asynchronous fault symptoms in complex systems, resulting in high false alarm rates. Furthermore, traditional methods cannot effectively distinguish between operating condition fluctuations and mechanical faults, affecting the reliability of fault warnings.

Method used

A cross-attention-based dual-stream time-series data anomaly early warning method is adopted. The operating condition and status monitoring data are processed separately through a dual-stream feature extraction network. The cross-attention mechanism is used to fuse features, and the method is combined with a deep ellipsoidal support vector data description model and a multi-criteria anomaly score scoring method to achieve early identification of faults.

Benefits of technology

It significantly reduces the false alarm rate, improves the accuracy and reliability of fault detection, and can identify early faults in complex systems in advance. It is suitable for complex systems with multi-source asynchronous timing characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121723364A_ABST
    Figure CN121723364A_ABST
Patent Text Reader

Abstract

The invention discloses a double-flow time series data abnormity early warning method based on cross attention, and aims to solve the problem of fault early warning caused by asynchronous operation conditions and state response time series in a complex system. The method comprises the following steps: firstly, independently processing operation condition and time sequence state data through a double-flow feature extraction network; then, adaptive fusion is carried out on the extracted features by using a cross attention mechanism to capture asynchronous associated information; and finally, making a decision by fusing the comprehensive score of the multi-dimensional criterion. In order to verify the feasibility of the method, the method is evaluated on an aero-engine outfield flight data set, and the result shows that the method is high in average test precision on the data set and low in false alarm rate and missing report rate, can perform sortie early warning in advance, and is superior to other comparison methods. Therefore, the method has good early warning performance, and vibration abnormity early warning of the complex system with excellent performance can be effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a feature fusion and fault diagnosis method for complex systems, and particularly to a method for early warning of anomalies in dual-stream time-series data based on cross-attention, belonging to the field of intelligent fault diagnosis technology. Background Technology

[0002] The safe and reliable operation of many critical pieces of equipment is essential to the systems they operate in. However, these devices often operate in harsh or variable environments for extended periods, making their critical components susceptible to early damage due to fatigue, wear, or external impacts. The evolution of these failures is often highly insidious and sudden, posing a significant challenge to early warning systems.

[0003] Traditional anomaly detection methods typically rely on a single condition monitoring signal (such as vibration or temperature). However, equipment condition responses are not isolated but closely related to their operating conditions. Different operating commands, load changes, or environmental parameters can significantly affect the performance of condition signals. Based solely on single condition data, models struggle to effectively distinguish signal changes caused by fluctuations in normal operating conditions from anomalies resulting from actual mechanical failures, leading to a high risk of false alarms.

[0004] Secondly, a deeper technical challenge lies in the significant temporal asynchrony between the impact of operating conditions on structural health and the response of state signals. For example, a drastic load change may induce minor mechanical damage, but the corresponding abnormal characteristics may not appear immediately, but rather be observed only in a subsequent, relatively stable operating phase. Traditional single-stream data processing or simple data stitching methods, unable to capture this delayed, dynamic causal relationship, often suffer from interference due to premature fusion of information between modalities, resulting in poor feature extraction and fusion performance.

[0005] After acquiring multi-source features such as operating conditions and status, achieving efficient fusion is the core technological bottleneck. Commonly used feature splicing or weighted summation strategies in existing technologies are essentially still just real-time or aligned information overlays. These methods cannot solve the aforementioned problem of delayed fault response.

[0006] In terms of decision-making mechanisms, most data-driven models rely on a single instantaneous anomaly score for judgment, which is severely unreliable. Normal operational commands may cause drastic fluctuations in signals within a short period of time, resulting in a high instantaneous anomaly score. If alarms are triggered solely based on this, it will lead to a large number of false alarms, severely undermining the credibility of the early warning system.

[0007] In summary, in order to address a series of technical challenges in the early warning of anomalies in complex systems, such as insufficient robustness to disturbances under changing operating conditions, inability to effectively capture asynchronous fault symptoms, and high false alarm rate of decision-making mechanisms, a novel technical solution is urgently needed. Summary of the Invention

[0008] Technical issues:

[0009] To address the aforementioned problems in existing anomaly warning technologies for complex systems, this invention aims to propose a time-series data anomaly warning method that combines a cross-attention feature fusion strategy with a dual-stream deep learning model.

[0010] Technical solution:

[0011] This invention proposes a cross-attention-based anomaly warning method for dual-stream time-series data (Cross-TEESVDD). The method includes a model training phase and a monitoring and warning phase.

[0012] The model training phase includes the following steps:

[0013] Step 1a: Obtain the operating time-series data of the monitoring system under normal conditions and the status monitoring time-series data as training data respectively;

[0014] Step 2a: A dual-stream feature extraction network is used to independently extract deep features from the operating condition time series data and the status monitoring time series data to obtain operating condition features and status features, respectively; wherein, the dual-stream feature extraction network adopts a parallel architecture, which includes an operating condition stream for processing operating condition time series data and a status stream for processing status monitoring time series data;

[0015] Step 3a: Use a cross-attention mechanism to fuse the working condition features and the state features to generate a fused feature containing asynchronous correlation information;

[0016] Step 4a: Based on the fusion features, construct and train a Deep Ellipsoid Support Vector Data Description (DeepESVDD) model to learn the compact distribution pattern of normal data;

[0017] The monitoring and early warning phase includes the following steps:

[0018] Step 1b: Acquire the real-time operating condition time-series data and status monitoring time-series data of the monitoring system respectively;

[0019] Step 2b: Using the trained dual-stream feature extraction network, feature extraction is performed on the data from Step 1b to obtain working condition features and state features;

[0020] Step 3b: Using the trained cross-attention mechanism, fuse the working condition features and state features obtained in Step 2b to generate fused features;

[0021] Step 4b: Input the fused features obtained in step 3b into the trained DeepESVDD model, calculate the instantaneous anomaly score, and calculate at least two dimensions of criteria based on the instantaneous anomaly score to obtain a comprehensive score. Compare the comprehensive score with a preset decision threshold to generate an early warning signal.

[0022] First, a brief introduction to the preparatory knowledge involved in this invention, such as Deep Ellipsoid Support Vector Data Description (DeepESVDD), Two-Stream Feature Extraction Network, Cross Attention Fusion Module, and Multi-Criterion Anomaly Score Fusion, will be given. Then, the process of using the Cross-TEESVDD method proposed in this invention for vibration anomaly early warning will be described in detail.

[0023] 1) Deep Ellipsoid Support Vector Data Description (DeepESVDD)

[0024] Deep Ellipsoidal Support Vector Data Description (DeepESVDD) is a deep learning method for anomaly detection. Its core idea is to learn a vector representation of support vectors from the input space. to feature space nonlinear mapping And in this feature space, construct a minimum volume ellipsoid that can enclose the vast majority of normal data points.

[0025] Unlike traditional Support Vector Data Description (SVDD), which uses Euclidean distance to define a hypersphere, such as As shown, DeepESVDD uses Mahalanobis distance to construct an ellipsoidal boundary. Mahalanobis distance can effectively take into account the correlation and scale differences of data across different dimensions, which allows the ellipsoid to more closely fit the actual distribution of normal data, thereby improving the sensitivity of anomaly detection.

[0026] Given a set of inputs Its distance to the center of the ellipsoid The Mahalanobis distance is defined as:

[0027]

[0028] in It is a diagonal matrix, where the diagonal elements are the standard deviations of each feature dimension. (This is achieved through a deep neural network.) After feature mapping, this distance can be transformed into a distance in the feature space. The expression form in Chinese.

[0029] Therefore, the training objective of DeepESVDD is to minimize the distance from the center of all normal samples in the feature space. The average Mahalanobis distance is given by the following objective function:

[0030]

[0031] The first term of the objective function aims to obtain a minimized ellipsoidal volume that includes all normal samples, while the second term is a regularization term for the network weights to prevent overfitting. In this way, DeepESVDD can learn a compact description of normal data patterns, making it easier to identify outliers that deviate from these patterns.

[0032] 2) Two-stream feature extraction network

[0033] This invention provides a dual-stream feature extraction network for monitoring the status of multi-source time-series data. The network employs a dual-branch parallel processing architecture to process two types of time-series data from different sources with varying characteristics, thereby accurately capturing the complex correlation between operating conditions and status responses.

[0034] The network comprises two independent data processing branches: Condition Stream, which inputs key parameters reflecting external loads or operating commands of the system; and State Stream, which inputs timing signals from key sensors.

[0035] In each of the aforementioned data streams, a Transformer encoder is deployed as the core feature extraction unit. This encoder, based on a self-attention mechanism, effectively captures long-range dependencies in time-series data. Its core calculation formula is as follows:

[0036]

[0037] The core advantage of the dual-stream design lies in its respect for and utilization of the complex temporal correlation between external loads or operational commands and internal state responses. Crucially, mechanical responses directly elicited by current operating conditions may exhibit time delays or may only be significantly expressed during specific load transients. This means that early mechanical failure symptoms caused by specific external loads may not be apparent in the current state signal but may manifest later in system operation. The independent channel architecture allows each Transformer encoder to focus on learning the most discriminative normal pattern in its own modality, thus avoiding interference from premature multimodal fusion. For the state stream, even with only one sensor, the Transformer encoder's ability to perceive the global context of long sequences allows it to delve into the dynamics of the entire temporal context. This allows for the effective capture of failure modes that are asynchronous to but closely related to attitude conditions. This design lays the foundation for a subsequent cross-attention fusion module, enabling it to search for and verify intrinsic causal relationships between the two modes on a global timescale, rather than simply splicing instantaneous features.

[0038] In the technical solution of this invention, the Transformer encoder is used as a feature extractor. For example... As shown, each branch (taking the running condition branch as an example) receives its corresponding timing input data. ( For sequence length, (for feature dimensions), and through a Transformer encoder Extracting high-level feature representations This method effectively captures long-term dependencies in the sequence, providing high-quality input for subsequent anomaly detection.

[0039] 3) Cross-attention fusion module

[0040] Cross-attention is a technique used in deep learning models. Unlike Self-Attention, which allows elements within a sequence to interact with each other (e.g., words in a sentence calculate relationships between each other), Cross-Attention establishes attention relationships between two different sequences (or data sources). The core of Cross-Attention is that it allows one sequence (called the Query) to focus on another sequence (called the Key and Value), thus achieving information fusion. The Cross-Attention mechanism aims to fuse feature representations from different runtime branches. Feature representation of state response branches The core advantage of this method lies in its efficient fusion of features from operational conditions and state responses to accurately capture the asynchronous temporal correlation between the two. This module can solve the technical problems of complex systems where early fault symptoms induced by specific flight conditions are difficult to identify, but whose delays are reflected in vibration signals.

[0041] like As shown, this method utilizes the dynamic interaction of queries, keys, and values ​​to adaptively model the correlation between two feature modes. Taking aero-engine vibration monitoring as an example, it uses attitude features that characterize the engine's real-time operating state... As a query, the vibration characteristics that contain the vibration history information of the entire flight segment Simultaneously serving as both key and value, this design ensures that the fusion process is not a simple instantaneous feature superposition, but rather a cross-temporal pattern retrieval: using the current flight status as the query condition, it actively filters and amplifies those fault symptoms most relevant to the current operating condition but potentially delayed in appearance from the entire history of the vibration sequence. This weighted fusion strategy not only overcomes the limitations of traditional concatenation or addition methods in lacking interaction between features, but also fully explores the complex nonlinear relationships between features through a multi-head attention mechanism, significantly improving the representational ability of the fused features and the early warning accuracy of the model.

[0042]

[0043] in These represent the query, key, and value matrices, respectively. This weight matrix reflects the relative importance of features at each moment in the vibration history to the current decision, within the context of the current attitude condition. The scaling dot product attention weights are calculated as follows:

[0044]

[0045] in is the dimension of the key vector, used to stabilize the gradient. Essentially, this process is a cross-time pattern retrieval: using the current flight state as an index, it actively filters and amplifies fault symptoms that are most relevant to the current condition but may be delayed from the entire history of the vibration sequence.

[0046] Residual connections are a crucial structure in deep learning networks, their core idea being to directly add the network's input to its output. The output of the attention mechanism is passed through a linear projection layer and then added to the original query (i.e., ...). Perform residual connections to obtain the fused features:

[0047]

[0048] This process allows the attitude branch to selectively incorporate highly correlated information from the vibration branch. Residual connections ensure the integrity of the original attitude information while injecting vibration context asynchronously related to the current state. The core advantage of the cross-attention mechanism lies in its dynamic weight allocation capability, which adaptively adjusts the contribution ratio of the two modalities in the final decision based on the specific flight state. This is crucial for effectively extracting temporally asynchronous but logically strongly correlated fault information from single-sensor data, significantly improving the ability to identify early, subtle faults and the model's early warning accuracy.

[0049] After the fusion calculation is completed, the classification token [cls] at the beginning of the sequence will be combined with all the fused temporal features. Interacting with each other. The [cls] token acts as a global information aggregator, ultimately generating a fixed-dimensional, generalized feature vector containing the failure modes of the entire flight segment. This vector will be directly used for subsequent early warning decisions. The ESVDD core is applied to... This step projects the final features onto a graph. The hypersphere centered on In the middle. This center This is determined during training. The goal is to tightly aggregate all normal fusion features. Surroundings. This design ensures that the model learns to be compact and uniform.

[0050] 4) Multi-criteria anomaly score scoring module

[0051] To improve the robustness and accuracy of fault detection, this invention introduces three criteria: Area Under the Curve (AUC), Trend, and Density. AUC characterizes the overall deviation by integrating the cumulative value of the anomaly score within a time window; Trend identifies fault development patterns by using slope fitting and abrupt change detection; and Density counts the percentage of points exceeding the threshold per unit time to capture local anomaly clustering features.

[0052] To optimize weight configuration, such as As shown, this invention designs an adaptive weight allocation mechanism based on a single typical fault case. The optimization objective of this mechanism is to find an optimal set of weight combinations ( This allows the resulting composite score (CS) to distinguish this type of failure case from all normal cases to the greatest extent possible. The specific process is as follows: Weight search is constructed as an optimization problem, by traversing a predefined parameter space, for each set of candidate weights ( Calculate the combined score for all normal cases and the failure case:

[0053]

[0054] The evaluation function considers the false positive rate, false negative rate, and the classification margin between normal and fault scores to find the weight combination that produces the best classification performance (i.e., maximizing the margin and minimizing false positives / false negatives). This process is based on only one of the most representative fault cases, aiming to learn a general weight allocation strategy that can effectively amplify fault signals rather than overfitting specific fault types, thus enhancing the model's generalization ability.

[0055] Ultimately, the calibration of the composite threshold (FCS) is a fully automated process. This involves determining the optimal weights. Next, the composite score (CS) for all normal training flights is calculated, and the maximum value is taken. The final composite score (FCS) is automatically set as the minimum score greater than this maximum normal score:

[0056]

[0057] This setup ensures a tight decision boundary with zero false alarms on the training set. During the monitoring phase, if the overall score (CS) of a case exceeds this threshold (FCS), an alarm signal is triggered. This approach significantly improves the reliability and adaptability of fault warnings through multi-dimensional evidence fusion and data-driven decision boundary calibration.

[0058] Based on the proposed dual-stream feature extraction network, cross-attention fusion, DeepESVDD model, and multi-criteria anomaly score fusion module, this invention constructs a vibration anomaly early warning method for complex systems, named Cross-TEESVDD. As shown, during the training phase, the normal-state attitude vibration signals, after time window segmentation and temporal feature extraction, are input into their respective branches for training. During the monitoring phase, real-time data undergoes the same process, passing through the trained Cross-TEESVDD model, and finally outputting a decision based on a multi-criteria anomaly score.

[0059] The Cross-TEESVDD method constructed in this invention utilizes a large amount of aero-engine field flight data to train network parameters and establish an aero-engine vibration anomaly early warning model. The system mainly consists of three modules: a training module, a testing module, and a monitoring and execution module, including the following steps:

[0060] Step 1: Collect flight attitude data and casing vibration signal data of the aero-engine. Preprocess the data, including selecting the effective operating segment based on engine speed, wavelet denoising, and sliding window feature extraction.

[0061] Step 2: (Corresponding to the training phase) Train the network parameters of the Cross-TEESVDD model using a large amount of normal historical flight data, including the parameters of the dual-stream Transformer encoder, cross-attention module, and DeepESVDD model, and complete the weight and threshold calibration of the multi-criteria decision mechanism.

[0062] Step 3: Validate the trained model using a test dataset containing known faults, evaluating whether it meets the performance metrics (such as AUC, false positive rate, and false negative rate). If it does not meet the metrics, return to Step 2 to readjust the training.

[0063] Step 4: (Corresponding to the monitoring and early warning stage) Deploy the trained and validated model to the monitoring system. The system receives engine attitude and vibration data in real time and inputs it into the model for calculation. Once the calculated comprehensive score exceeds the preset FCS threshold, the system immediately outputs an early warning signal;

[0064] Beneficial effects:

[0065] 1) The dual-stream independent modeling architecture proposed in this invention performs deep feature extraction on the operating conditions and state responses respectively, effectively avoiding modal interference caused by early feature fusion; on this basis, a cross-attention mechanism is introduced, which can adaptively retrieve fault features related to the current operating conditions but delayed from the state history, solving the key problem that traditional methods cannot cope with the asynchronous timing of "operating conditions-response".

[0066] 2) By integrating multi-dimensional criteria (area under the curve, curve trend, and anomaly density) to construct a comprehensive anomaly score, the method effectively distinguishes between real faults and normal operating condition fluctuations, significantly reducing the false alarm rate. Experimental results show that, on aero-engine field data, the method of this invention maintains a high detection rate (AUC of 95.9%) while achieving an extremely low false alarm rate (0.52%) and a false negative rate (0%).

[0067] 3) The Cross-TEESVDD method proposed in this invention does not rely on fault samples and can complete model training and threshold calibration using only normal data, making it suitable for situations where fault samples are scarce in real-world industrial scenarios. Furthermore, the method has a modular structure and can be extended to other complex system fault early warning tasks with multi-source asynchronous temporal characteristics.

[0068] 4) Comparison and ablation experiments on aero-engine field flight datasets show that the method of the present invention can achieve early warning of typical faults such as compressor disc cracking and turbine shaft fracture, and its performance is better than many mainstream baseline methods, demonstrating good potential for engineering applications. Attached Figure Description

[0069] This is a schematic diagram of the DeepESVDD structure.

[0070] This is a schematic diagram of the two-stream feature extraction structure.

[0071] This is a schematic diagram of the Cross-Attention structure.

[0072] This is a schematic diagram of the multi-criteria abnormal score scoring structure.

[0073] This is a schematic diagram of the Cross-TEESVDD structure.

[0074] This is a schematic diagram of a turbofan engine.

[0075] This is a schematic diagram of three vibration states.

[0076] This is a schematic diagram of speed signal filtering.

[0077] This is a schematic diagram of the signal reconstruction effect after wavelet decomposition.

[0078] This is a schematic diagram of sliding window segmentation and feature extraction.

[0079] This is a schematic diagram of the early warning results of the Cross-TEESVDD algorithm on the compressor disc cracking fault module.

[0080] This is a schematic diagram of the early warning results of the Cross-TEESVDD algorithm on the turbine shaft fracture fault module. Detailed Implementation

[0081] The present invention will be further described below with reference to the accompanying drawings and tables.

[0082] This invention utilizes aero-engine field flight datasets to verify the reliability and superiority of the proposed Cross-TEESVDD method. All experiments were conducted on a computer configured with an Intel® Core™ i7-12650H CPU 4.70 GHz, an NVIDIA GTX4060 GPU, 16GB of RAM, and Windows 11, and were run using Python.

[0083] To evaluate the effectiveness of the proposed Cross-TEESVDD model, comparative and ablation experiments were designed and validated on an aero-engine field flight dataset. The experiments primarily examined the overall performance (detection rate, false alarm rate) of the Cross-TEESVDD model compared to mainstream baseline models in engine anomaly detection tasks. The contributions of the dual-stream design, cross-attention fusion module, and multi-criteria decision mechanism in the model to the final performance were also verified. The parameters of the six comparative algorithms are as follows: As shown.

[0084] The specific implementation of the method of the present invention mainly includes two core stages: model training and model application.

[0085] (I) Model Training Phase

[0086] Step 1: Data Acquisition and Preprocessing. Collect flight data including those of abnormal vibration faults such as compressor disc cracks and turbine shaft fractures, as well as their preceding flights, to construct an aero-engine field flight dataset.

[0087] This embodiment selects a turbofan engine field flight dataset. This is a cross-sectional structural diagram of the turbofan engine, including core components such as the air intake, fan, low-pressure compressor (LPC), high-pressure compressor (HPC), combustion chamber, high-pressure turbine (HPT), low-pressure turbine (LPT), bypass duct, and nozzle. The arrows in the diagram indicate the airflow path of the turbofan engine. The intake air is first pressurized by the fan and then split into two paths. One path enters the inner duct and flows sequentially through the LPC, HPC, combustion chamber, HPT, and LPT; the other path enters the bypass duct. In the inner duct, the high-temperature, high-pressure gas drives the high- and low-pressure turbines to rotate. Finally, the two airflows mix in front of the nozzle and are discharged. If necessary, the thrust can be further increased through the afterburner. The installation location of the casing vibration sensor is clearly marked in the diagram (point "B").

[0088] The key monitoring point in this embodiment is the vibration sensor (marked "B" in the figure) installed on the central casing. This measurement point was chosen because it is located on the engine's main load-bearing structure, effectively sensing the combined vibration response of the high- and low-pressure rotor systems. This location is particularly sensitive to faults in the high-pressure system (such as HPC disc cracks) because vibration energy can be transmitted to the casing through the rotor-support structure. Simultaneously, since the dynamic characteristics of both the high- and low-pressure rotors are coupled and reflected in the casing vibration signal, this measurement point also has the potential to monitor systemic faults involving power transmission, such as turbine shaft fracture.

[0089] In this embodiment, the timing signal collected by vibration sensor B is used as the input to the Vibration Stream; the flight parameters (rotation speed N2, pitch angle, roll angle, and normal overload value) are used as the input to the Attitude Stream, which together are used to verify the ability of the proposed Cross-TEESVDD model to distinguish between faults and disturbances under various operating conditions.

[0090] This embodiment uses field flight data from turbofan aero-engines, including 10,000 normal flight segments, 4 fault segments, and the preceding 50 segments. The fault segments are marked from actual field maintenance records, covering two typical faults: compressor disk cracks and turbine shaft fatigue cracks. The sampling frequency is 10 Hz. All data were preprocessed into continuous 50-second segments and normalized. Training data consists of 50 flights in typical flight attitudes, while the rest are test data.

[0091] Step 2: Perform data preprocessing operations such as sliding window segmentation and wavelet denoising on the field flight dataset.

[0092] The purpose of an aircraft engine vibration anomaly early warning mission is to identify the engine's operating condition by analyzing vibration signals. For example... As shown, this step divides the engine's operating state into a normal state and a fault state containing early signs of vibration anomalies. Therefore, the task can be simplified to an anomaly detection task, providing a warning signal when the vibration signal indicates a fault state. During takeoff and landing, the operating conditions (such as thrust, speed, and airflow) of an aero-engine undergo drastic and rapid transient changes. This non-stationarity causes significant fluctuations in the vibration signal, primarily stemming from normal flight operations rather than mechanical failures. Directly introducing such data into model training introduces significant noise, interfering with the model's extraction and learning of mechanical fault features. To eliminate the interference from transient operating conditions and ensure the model focuses on the mechanical health state under steady-state or quasi-steady-state operation, such as... As shown, this step sets a speed-based filtering threshold based on the engine's duty cycle characteristics: only data segments with a high-pressure speed (N2) greater than 60% are retained for analysis. This threshold corresponds to the starting point when the engine enters a stable high-power state (such as climb or cruise), effectively filtering out low-power and unstable data segments such as takeoff and approach.

[0093] To accurately capture the early warning features of aero-engine vibration anomalies and achieve fault early warning at least one flight ahead, this step first preprocesses the original vibration signal using wavelet decomposition. Aero-engine vibration signals face significant noise interference due to the extremely complex environment in which they operate. This noise mainly originates from aerodynamic disturbances of high-speed rotating components, combustion chamber pressure pulsations, multi-source mechanical vibration coupling, and sensor measurement errors, resulting in a low signal-to-noise ratio in the original signal and easily obscuring early fault features. The core principle of wavelet decomposition is to use a scalable and translational basis function (wavelet) to decompose the signal at multiple scales (frequency), thereby simultaneously obtaining the localized features of the signal in both the time and frequency domains. This step uses the dmey (Discrete Meyer) wavelet, setting the decomposition level to 10 levels, and obtaining the approximate coefficients a10 of the 10th level as the denoised signal. The a10 signal is the low-frequency backbone component after deep denoising, effectively filtering out high-frequency noise and transient interference from the original signal, thus more clearly highlighting the physically meaningful vibration trend dominated by the engine's mechanical state itself. The signal reconstruction effect after wavelet decomposition is intuitively demonstrated. The a10 component retains the main trend features of the original signal while significantly suppressing high-frequency noise. This verifies the effectiveness of wavelet decomposition in improving signal quality and enhancing fault feature identification, providing a reliable data foundation for subsequent feature extraction and fault early warning.

[0094] Based on the denoised a10 signal, this step further extracts a feature set that can characterize its waveform morphology and statistical regularity. The vibration signal of the aero-engine casing is essentially a time-domain signal, while the input of the deep learning model is discrete sample points. Therefore, it is necessary to extract time-domain features by using a sliding time window to convert the continuous time-domain vibration signal into the corresponding feature matrix at each time point.

[0095] like As shown, this embodiment uses a sliding window to segment the denoised signal from the a10 signal and extracts temporal and inter-column relationship features, including mean, variance, median, minimum, maximum, peak-to-peak value, average inter-column correlation coefficient, average inter-column difference, maximum and minimum inter-column difference, and inter-column time synchronization difference. The inter-column relationship features effectively capture the spatial patterns and cooperative response characteristics of fault propagation by calculating the statistical correlation between multi-sensor signals. This feature extraction process uses a 50-second temporal window with a 10-second step for sliding sampling (sampling frequency of 10 Hz), thereby constructing a feature sequence with high temporal resolution while fully preserving the transient and steady-state information of the signal. Merging the feature sequences from the five sensors yields a feature matrix with a dimension of (15×5) at the current time point.

[0096] Step 3: Model Building and Training. The Cross-TEESVDD model is trained using a large amount of historical flight data under normal conditions.

[0097] First, construct the model architecture, the core of which includes:

[0098] Dual-stream feature extraction network: Processes flight attitude data (operational condition stream) and casing vibration signal (state stream) separately. Each stream uses a Transformer encoder as a feature extraction unit.

[0099] Cross-attention fusion module: This module uses the working condition features as the query and the state features as the key and value, performs adaptive weighted fusion, and adds the fused features to the original query features through residual connections to obtain the fused features. After fusion, a [cls] token located at the beginning of the sequence is used to interact with all the fused temporal features to aggregate and generate a fixed-dimensional global feature vector.

[0100] DeepESVDD anomaly detection module: Based on the global feature vector, a Deep Ellipsoid Support Vector Data Description (DeepESVDD) model is constructed and trained to learn the compact distribution pattern of normal data.

[0101] Multi-criteria decision module: Determines the formula for calculating the composite score (CS), which is obtained by weighted summation of three criteria: area under the curve (AUC), curve trend, and outlier density. .

[0102] Subsequently, model training and parameter calibration are performed. The model parameters are trained using preprocessed normal data. Simultaneously, based on a historical failure case (which is not involved in the model training process), an adaptive weight allocation mechanism is used to determine the weight coefficients a, b, and c. This mechanism searches the parameter space and optimizes the evaluation function to find the weight combination that best distinguishes the historical failure case from all normal cases. Finally, the overall score (CS) of all normal training flights is calculated, and its maximum value is automatically set as the decision threshold.

[0103] (II) Monitoring and Early Warning Stage

[0104] Step 4: Model validation and deployment early warning.

[0105] First, the trained model is validated and its performance evaluated using a test dataset containing known faults. In this step, to comprehensively and fairly evaluate the performance of the proposed Cross-TEESVDD model, six representative state-of-the-art anomaly detection algorithms covering different technical approaches are selected as baselines. These baseline methods can be broadly categorized into two types: traditional machine learning methods and deep learning methods. This step selects six commonly used unsupervised algorithms in anomaly detection tasks, where Isolation Forest and SVDD are shallow baselines, and DCAE, DSVDD, DAGMM, and LSTM-VAE are deep baselines. For shallow baselines, the original 15 * 5 matrix input is flattened into a one-dimensional vector of length 75. For deep baselines, no processing is performed on the input data. Simultaneously, to further verify the effectiveness of each core component in the proposed Cross-TEESVDD model and its contribution to the final performance, this step designs a system ablation study. The experiment quantitatively evaluated the impact of different modules on anomaly detection performance by progressively removing or replacing key modules in the model, including removing the pose flow, removing the cross attention module, replacing the Transformer encoder with an LSTM with equal parameters, and removing the multi-criteria fusion decision mechanism, under the same experimental settings and dataset.

[0106] Evaluate whether it meets the performance targets (such as AUC, false positive rate, and false negative rate). If it does not meet the targets, return to step 3 of the model training phase (I) to readjust the training.

[0107] Next, model deployment is performed. The trained and validated model is deployed to the monitoring system. The system receives engine attitude and vibration data in real time and inputs it into the deployed model for calculation. After performing the same preprocessing and feature extraction on the real-time data as during the training phase, the model calculates a comprehensive score (CS). Once this score exceeds the preset decision threshold during the training phase, the system immediately outputs an early warning signal.

[0108] Seven comparative algorithms, including the Cross-TEESVDD proposed in this invention, and four variant models of ablation experimental design were used to conduct experiments on the aero-engine field flight dataset constructed in step 1 to evaluate the performance of Cross-TEESVDD. The results will evaluate the effectiveness of the proposed method from two perspectives: test accuracy AUC value and false alarm / false negative rate.

[0109] The experiment was conducted on a Windows 11 operating system, using the PyTorch deep learning development framework and Python as the programming language. The CPU used in the experiment was an Intel Core i7-12650H 2.30GHz, and the GPU was an NVIDIA GeForce RTX 4060 8G.

[0110] To optimize the training process and ensure the stability and efficiency of model convergence, this step systematically sets the key hyperparameters of the Cross-TEESVDD model, as detailed below. As shown in the diagram. In terms of model structure, a 3-layer Transformer encoder is used, with 8 attention heads per layer. The feedforward network dimension is set to 2048 to balance model expressiveness and computational complexity. During training, the total number of training epochs is set to 150, the batch size to 256, and the Adam optimizer is used for parameter updates to adaptively adjust the learning rate. For the learning rate scheduling strategy, a step decay mechanism with an initial value of 0.0001 is adopted. Specifically, after every 25 training epochs, the learning rate is gradually reduced by a decay rate of 0.2. This strategy allows the model to converge quickly with a higher learning rate in the early stages of training, while finely adjusting the parameters by reducing the learning rate in the later stages, thus more stably approaching the global optimum and avoiding oscillations or overfitting.

[0111] First, the theoretical vibration anomaly early warning performance of seven comparative methods was evaluated using the AUC value as the evaluation metric. The average AUC values ​​of the seven methods on four datasets containing faults and 50 preceding faults are shown below. As shown, Cross-TEESVDD's AUC value is higher than other competing methods in each class, indicating that Cross-TEESVDD has better performance and detection capability than similar methods in the task of detecting vibration anomalies in aero-engines. This is due to Cross-TEESVDD's keen perception mechanism for temporal anomaly patterns: the dual-stream Transformer encoder can extract long-range temporal features from attitude condition and vibration response data respectively, while the cross-attention fusion module adaptively enhances key time segments in the vibration signal that are strongly correlated with the fault through dynamic weight allocation. This structure enables the model to effectively identify weak temporal signs of early faults and accurately distinguish between real anomalies and vibration fluctuations caused by operational disturbances, thereby achieving more accurate anomaly localization and detection in complex flight scenarios.

[0112] We selected one failure case from two fault modules (compressor disc crack, turbine shaft fracture) and several normal sortie cases with typical flight operation characteristics for weight optimization. The weight optimization results and the corresponding FCS thresholds are as follows: As shown.

[0113] To quantitatively evaluate the generalization ability and reliability of the Cross-TEESVDD model in practical applications, this embodiment statistically analyzes its false positive rate (FPR) and false negative rate (FNR). The false positive rate measures the probability that the model misclassifies normal samples as anomalies; an excessively high FPR value weakens the reliability of the early warning system. The false negative rate measures the probability that the model fails to identify genuine anomalies; an excessively high FNR value directly relates to flight safety. An ideal, highly reliable anomaly detection model should possess both low false positive and low false negative rates.

[0114] This embodiment constructs a test set covering two typical engine failure types (compressor disc crack and turbine shaft fracture), with each failure type treated as two independent cases. Each case's test data includes 50 normal flight sorties and 1 abnormal flight sortie injected with that type of failure, simulating a real-world scenario of occasional specific failures occurring amidst a large number of normal operations. Simultaneously, a separate test case containing 10,000 purely normal flight sorties is included to rigorously evaluate the model's false alarm level under fault-free conditions. The model's anomaly detection results (based on the CS score exceeding the FCS threshold) are statistically presented as follows: As shown.

[0115] In the fault detection of Cases 1 to 4, the model issued early warnings for all abnormal flights (≥1 early warning flight) and generated no false alarms for any preceding normal flights. Therefore, in these four fault cases, the model's false alarm rate (FPR) and false negative rate (FNR) were both 0%. This result fully demonstrates that the Cross-TEESVDD model possesses extremely high anomaly identification sensitivity and accuracy when facing various typical faults, effectively detecting fault signs lurking in normal operations without triggering false alarms.

[0116] In 10,000 flight tests with purely normal cases, the model's false alarm rate was 0.52%. This means that the model can correctly identify normal conditions in the vast majority of cases (99.48%), demonstrating good stability. The extremely low false alarm rate indicates that the model's boundary learning of normal operating conditions is tight, has strong anti-interference ability, and avoids warning fatigue caused by frequent false alarms.

[0117] Based on all the cases, the model achieved zero false alarms and zero false negatives in the fault cases, and only had an extremely low false alarm rate in the tens of thousands of normal data. This proves that the method proposed in this embodiment can maintain an extremely low false alarm level (low false alarm rate) while ensuring an extremely high fault detection rate (low false alarm rate), effectively solving the problem of balancing sensitivity and reliability in anomaly detection, and meeting the stringent reliability requirements of aero-engine safety monitoring for models.

[0118] Combination - The time-series monitoring curves show that the red area represents the number of failed flights, and the orange area represents the number of flights before the failure. The model's warnings (orange area) precisely correspond to the number of flights where the CS score exceeds the threshold (FCS), and all of them occur within the actual abnormal range. This further verifies the correctness of the above quantitative statistical results and the effectiveness of the model from a visualization perspective.

[0119] Ablation test results as follows As shown, the performance of the complete Cross-TEESVDD model is listed as a benchmark.

[0120] The variant model without attitude flow exhibits a significantly higher false alarm rate, demonstrating that without the aid of flight attitude condition information, the model struggles to distinguish between vibration anomalies caused by normal operating condition disturbances such as severe maneuvering and mechanical failures. The model without the cross-attention fusion mechanism performs worse than the complete model, proving that simple feature stitching cannot achieve deep interaction and adaptive filtering of intermodal information. The cross-attention mechanism proposed in this invention can dynamically retrieve and enhance the most relevant fault symptoms from the vibration history based on the current flight state, which is crucial for accurately capturing asynchronous temporal correlations. The performance loss of replacing the Transformer encoder with an LSTM encoder verifies the superiority of the Transformer encoder in aero-engine anomaly detection tasks. Compared to LSTM, the Transformer encoder, with its self-attention mechanism, can better model the long-range temporal dependencies between vibration signals and operating parameters, thus more effectively capturing the weak and slowly evolving features of early faults. The false alarm rate of the single-criteria model is much higher than that of the complete model. This indicates that the multi-criteria fusion strategy of AUC, Trend, and Density introduced in this invention can effectively filter out transient interference by comprehensively evaluating the overall energy, development trend, and local clustering of anomalies. Thus, it significantly reduces the false alarm rate with almost no loss of detection rate (FNR=0) and significantly improves the reliability of the early warning system.

[0121] In summary, the proposed Cross-TEESVDD outperforms all comparative methods in early warning performance of aero-engine vibration anomalies, and can serve as a novel method for daily flight parameter interpretation and fault early warning. Furthermore, the work presented in this invention can be considered a general workflow for early fault diagnosis in the aerospace field. In the future, it can be considered for application to other intelligent fault diagnosis tasks with asynchronous timing characteristics and multi-sensor fusion requirements.

[0122] surface 7 Comparison Model Structures

[0123]

[0124] surface Overview of Hyperparameter Settings

[0125]

[0126] surface AUC values ​​of 7 comparison methods in 4 failure cases

[0127]

[0128] surface Weight optimization results and corresponding FCS thresholds

[0129]

[0130] surface Anomaly detection results of the Cross-TEESVDD model

[0131]

[0132] surface Ablation test results

[0133]

Claims

1. A method for early warning of anomalies in dual-stream time-series data based on cross-attention, characterized in that, This includes the model training phase and the monitoring and early warning phase; The model training phase includes the following steps: Step 1a: Obtain the operating time-series data of the monitoring system under normal conditions and the status monitoring time-series data as training data respectively; Step 2a: A dual-stream feature extraction network is used to independently extract deep features from the operating condition time series data and the status monitoring time series data to obtain operating condition features and status features, respectively; wherein, the dual-stream feature extraction network adopts a parallel architecture, which includes an operating condition stream for processing operating condition time series data and a status stream for processing status monitoring time series data; Step 3a: Use a cross-attention mechanism to fuse the working condition features and the state features to generate a fused feature containing asynchronous correlation information; Step 4a: Based on the fusion features, construct and train a Deep Ellipsoid Support Vector Data Description (DeepESVDD) model to learn the compact distribution pattern of normal data; The monitoring and early warning phase includes the following steps: Step 1b: Acquire the real-time operating condition time-series data and status monitoring time-series data of the monitoring system respectively; Step 2b: Using the trained dual-stream feature extraction network, feature extraction is performed on the data from Step 1b to obtain working condition features and state features; Step 3b: Using the trained cross-attention mechanism, fuse the working condition features and state features obtained in Step 2b to generate fused features; Step 4b: Input the fused features obtained in step 3b into the trained DeepESVDD model, calculate the instantaneous anomaly score, and calculate at least two dimensions of criteria based on the instantaneous anomaly score to obtain a comprehensive score. Compare the comprehensive score with a preset decision threshold to generate an early warning signal.

2. The method according to claim 1, characterized in that, After step 3a in the model training phase and step 3b in the monitoring and early warning phase, the following step is also included: using a classification token ([cls]) located at the beginning of the sequence to interact with all the fused temporal features to aggregate and generate a global feature vector of fixed dimension; the training and early warning are both based on this global feature vector.

3. The method according to claim 1, characterized in that, In steps 2a and 2b, both the operating condition flow and the state flow use a Transformer encoder as the feature extraction unit.

4. The method according to claim 1, characterized in that, The cross-attention mechanism in steps 3a and 3b uses the working condition features as queries and the state features as keys and values, and calculates attention weights to achieve adaptive weighted fusion of the two types of features.

5. The method according to claim 4, characterized in that, In steps 3a and 3b, the fused features are obtained by performing a residual connection between the output of the cross-attention mechanism and the working condition features used as a query.

6. The method according to claim 1, characterized in that, The multi-criteria decision in step 4b includes the following criteria: the area under the curve of the instantaneous abnormal score over a period of time, the trend of the instantaneous abnormal score curve, and the density of abnormal points exceeding the threshold score over a period of time.

7. The method according to claim 6, characterized in that, The composite score (CS) is obtained by weighted summation of the three criteria: area under the curve, curve trend, and outlier density. The calculation is as follows: , in, These are the weighting coefficients for each criterion. , , These represent the area under the curve, the curve trend, and the density of outliers, respectively.

8. The method according to claim 7, characterized in that, The decision threshold is automatically set to the maximum value among the comprehensive scores of a set of normal training data.

9. The method according to claim 7, characterized in that, The weight coefficients a, b, c are determined through an adaptive weight allocation mechanism. This mechanism is based on a historical failure case and finds the weight combination that can best distinguish the historical failure case from all normal cases by traversing the parameter space and optimizing the evaluation function. The historical failure case does not participate in the model training process.

10. The method according to claim 1, characterized in that, Prior to step 1a, a data preprocessing step is also included: Based on the threshold that the high-pressure speed N2 is greater than 60%, time series data in the stable working stage are screened. The time series data of the state monitoring is decomposed by wavelet decomposition to denoise, and the approximation coefficient a10 of the 10th layer is extracted as the denoised signal. A sliding window is used to segment the denoised signal, and time-domain and inter-column relationship features, including mean, variance, median, minimum, maximum, peak-to-peak value, average inter-column correlation coefficient, average inter-column difference, maximum and minimum inter-column difference, and inter-column time synchronization difference, are extracted from it.