Aero-engine fault diagnosis method based on multi-view signals and physical prior perception
By using multi-view signals and physical prior perception, noise components are suppressed and multi-scale features are reconstructed, solving the problem that traditional methods are difficult to identify fault features under extreme conditions, and achieving high accuracy and robustness in the diagnosis of aero-engine bearing faults.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV OF SCI & TECH
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional aero-engine bearing fault diagnosis methods are difficult to adaptively capture the dynamic evolution of signals under extreme operating conditions, are prone to losing key information, and suffer from time-frequency image distortion in strong noise environments, making it difficult to identify fault features and resulting in insufficient model robustness and generalization.
A fault diagnosis method for aero-engines based on multi-view signals and physical prior perception is constructed. By using a wavelet domain physical prior perception module and a threshold-guided spatiotemporal cross-domain attention mechanism, noise components are suppressed, sensitive features are refined, and a parallel multi-scale bit-domain modulation architecture is constructed to reconstruct multi-scale features, thereby achieving multi-view semantic adaptive complementarity and progressive optimization.
It significantly improves the accuracy and robustness of fault diagnosis, enabling accurate identification of fault characteristics in high-noise environments and enhancing the model's noise resistance and generalization ability.
Smart Images

Figure CN121834699A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a kind of based on multi-view signal and physical priori perception aero-engine fault diagnosis method, it belongs to the field of fault diagnosis. BACKGROUND
[0002] As the core load-bearing component of aero-engine, the running state of bearing directly determines the overall performance and operation reliability of aero-engine. However, the bearings between the shafts of aero-engine are prone to wear, cracks and aging of parts under long-term extreme working conditions of alternating load and impact stress. Once bearing failure occurs, it is likely to cause serious accidents, and even cause disastrous air crash. Therefore, effective fault diagnosis of active aero-engine bearings has important engineering value, but the traditional method has significant limitations: based on signal processing method, highly dependent on artificial experience in feature extraction, unable to adaptively capture the dynamic evolution law of the diagnosis signal under extreme conditions of aero-engine, easy to lose key information; At the same time, the attention range of traditional deep learning network is limited, it is difficult to filter noise channels and amplify weak local energy of bearing at the same time, which will lead to distortion of time-frequency image and difficulty in identifying fault features, increasing the risk of missed detection and false alarm; In addition, the existing method faces complex scenes such as aero-grade strong noise, non-stationary load and multi-fault coupling, and the model robustness and generalization are insufficient, which cannot meet the high reliability diagnosis demand of aero-engine bearings. In view of the above problems, the present application proposes aero-engine fault diagnosis method based on multi-view signal and physical priori perception (MP-PFDNet), which injects the model with multi-scale wavelet domain physical priori perception information as priori knowledge, and guides the visual selective state space parameter update to suppress high-frequency noise components. On this basis, the threshold-guided spatio-temporal cross-domain attention mechanism is used to focus on the effective information in the domain and selectively suppress the noise channels through the channel difference mask of the learnable threshold and the spatio-temporal difference probabilistic gain, so as to fine-granularize the real features. The parallel multi-scale bit domain modulation architecture is constructed in time domain view, the signal is processed in parallel in multiple scales in time domain, and the wide kernel convolution is coupled with the bit domain re-calibration module to reconstruct the multi-scale features, providing cross-cycle representation for time-frequency domain view, and significantly improving the fault diagnosis capability. SUMMARY
[0003] The present application is aimed at the problems that traditional single-view method is difficult to capture the complementary information between time domain and time-frequency domain features, and time-frequency image is easy to lead to signal distortion, spectrum blur and other problems under strong noise interference. The aero-engine fault diagnosis method based on multi-view signal and physical priori perception is proposed. The specific implementation steps of the present application are as follows: 1. Constructing time-frequency domain view network, suppressing noise components and fine-granularizing sensitive features through wavelet domain physical priori perception module and threshold-guided spatio-temporal cross-domain attention gating modulation. The steps are as follows: (1a) Wavelet domain physical prior information perception Let the features of the input wavelet time-frequency image be... ,in For batch quantity, These represent the height of the image's time dimension and the width of its frequency dimension, respectively. The number of channels after image decomposition corresponds to time-frequency components at different scales; the feature dimension is expanded to [number] through linear projection. ,in , This is the channel expansion coefficient, set to 2 to balance feature capacity and computational efficiency; and it is split into feature branches. With gated branches : Next Convert to channel priority format Furthermore, depthwise convolution is used to enhance local feature edges, avoiding confusion of faulty texture features caused by cross-channel interference. The expression is shown below: Where, groups= , The kernel size is [size]. Zero-padding values are used to ensure that the feature scale after convolution is consistent with the input. Activation functions enhance the nonlinear representation of fault characteristics through nonlinear variations; Then, to Perform wavelet decomposition to obtain frequency scale information, assuming For two-dimensional wavelet decomposition functions, the following is adopted: wavelet basis Layer decomposition yields low-frequency approximation coefficients. With high-frequency detail coefficient ,in These are the high-frequency coefficients in the horizontal, vertical, and diagonal directions, respectively. This decomposition result will serve as the physical constraint basis for the selective state space, directly guiding the processing strategy of state parameters for different frequency components; further... Combined with the frequency scaling information obtained from wavelet decomposition, it is reconstructed into a feature flow in four directions: in Indicates the dimension of spatial flattening. This is the low-frequency enhancement factor, whose value is dynamically adjusted according to the intensity of high-frequency noise. When the noise increases, Increase the weights of low- and mid-frequency features to enhance them, and conversely, maintain feature balance. Based on the above multi-directional-frequency joint feature flow , the core of the selective state space model is to embed the wavelet domain high-frequency noise suppression mechanism into the state transition matrix and the calculation of dynamic time step, so that the state learning process actively follows the physical law of suppressing high-frequency noise and preserving low-frequency fault characteristics. The continuous domain state equation of the selective state space is first modeled based on the dynamic characteristics of the vibration signal, and the state variable is , the continuous domain state update formula is: Among them, the state transition matrix integrates the wavelet physical parameter prior, and the matrix of the traditional selective state space only ensures stability through negative index, while this module dynamically adjusts the decay strength of based on the sparsity of , and the mathematical expression is: Among them is the learnable basic state matrix parameter, is the high-frequency decay coefficient, is the Frobenius norm; in order to adapt to discrete adaptive image data, the above continuous domain model needs to be discretized, and by introducing the dynamic time step further integrates the wavelet domain physical information constraint; this module adjusts based on the sparsity of , and the specific calculation is as follows: Among them is the time step offset, is the noise step adjustment coefficient, is the flattened vector. The discrete state transition matrix and the input projection matrix are: (1b) Dynamic threshold channel allocation The input wavelet time-frequency feature is , where represents the batch size, represents the number of channels, and represent the height and width of the time-frequency graph respectively; first, the input is globally spatially averaged along the channel dimension to characterize the overall response strength of each channel, and the specific operation is as follows: Among them is the channel-level response vector, denotes the th sample in the channel and spatial location ; then through Depth-Wise convolution to capture the local time-frequency texture within the channel, followed by point convolution to complete the channel projection and linear combination, to get the projection feature : Then repeat the global mean operation on the projection feature to get the projected channel statistics , and define the channel response difference: Where denotes the channel linear mapping of the original statistics at , the difference directly reflects the increase and decrease trend of each channel response after the interaction of the noise channel, indicates that the channel is amplified in the aggregate representation, and has more discriminative potential, and then uses a learnable threshold parameter , and constructs a dynamic binary mask: Where 1 indicates that the channel with significant difference is selected, and 0 indicates that it is suppressed; then the difference value filtered by the mask is processed along the channel dimension: Where and the sum of each sample in the channel dimension is 1; Then, the original projection feature, the probabilistic gain and the nonlinear mapping of the difference are fused to construct the final channel attention score: Where is the attention weight in the channel dimension, is the , used to capture the nonlinear relationship of the difference; the channel attention is obtained by scaling the main elements of the projection feature after broadcasting in the spatial dimension, thereby obtaining the channel weighted feature: After obtaining the channel enhanced feature, further parallel calculation is performed on the channel response based on the channel mean and the corresponding spatial response after enhancement: Where are the channel average spatial mappings before and after fusion, respectively; construct the spatial difference and apply the flattened spatial position vector with to get the spatial position weight: where is the probability distribution of each sample in spatial position, reshaped as shape and broadcast along the channel to get the spatial weighted feature ; the channel and spatial branch are fused by learnable gating coefficients in proportion: where is the fused weighted feature; to maintain the details and numerical stability of the original wavelet image, the final result is connected back to the input with a residual connection: where is the learnable scaling factor, used to control the compensation strength of the fused feature to the original input; (1c) spatio-temporal cross-domain attention calculation calculate the channel mean of each spatial position as the global spatial statistic, for the input feature map , the global spatial statistic is defined as: where represents the global spatial statistic map, representing the average channel response intensity of each spatial position; then use point convolution to convert the input feature map; at the same time, calculate the global spatial statistics of the converted feature map as the reference benchmark for spatial enhancement; in order to highlight the spatial area of the fault, the spatial difference signal is defined as: where reflects the degree of change of the spatial position after point convolution processing, through function to normalize the difference in spatial dimension, get the spatial enhancement weight: where the enhanced weight along the spatial dimension satisfies the probability distribution in space; respectively on the converted feature map spatial max-pooling and average-pooling, then concatenate the pooling results in the channel dimension, and then generate the spatial attention map through convolution and function: wherein Capture spatial context semantic information through joint local receptive field; Then the global average pooling compresses the spatial dimension, and then Convolution and After activation, the channel attention vector is generated by function, the specific calculation process is as follows: wherein is the feature map after compressing the spatial dimension, is the channel attention vector; Finally, the feature map is weighted by the assigned weight to get the output, and the specific calculation process is as follows: wherein is the spatial difference enhancement guided attention, is the addition fusion of channel-spatial attention, denotes the normalization of the spatial enhancement part, denotes the final weighted output of the feature map.
[0004] 2. Construct the time domain perspective network, design a parallel multi-scale bit field modulation architecture, process the signal in the time domain in parallel with multiple scales, and use a wide kernel convolution coupled bit field rescaling module to reconstruct the multi-scale features, providing cross-cycle representation for the time-frequency domain perspective. The steps are as follows: (2a) Parallel multi-scale frequency band processing The original vibration sequence is , wherein the batch , channel , and sampling length ; First, construct a multi-scale segmentation operator on the time scale set , which will parallelly cut the signal sequence into subsequences with a length of , the specific operation is as follows: wherein denotes the interception operator of the scale and the segment; Then, for each scale branch , first use the wide kernel convolution operator to capture the long-time correlation and wide-band coupling effect, and then perform point-by-point projection and nonlinear activation to obtain the segmented features, the calculation process is as follows: wherein is the branch convolution kernel length, is the number of channels output by the branch, is the point-wise projection, is the activation function; Then, the outputs of each segment are fused by gating modulation, and the calculation process is as follows: wherein is the feature after gating modulation; (2b) Bit field gating re-scaling The fused features are maximum-pooled in the channel dimension and average-pooled , and the two are fused into a sequence before bit field interaction. The specific calculation process is as follows: wherein is the feature vector after maximum pooling, is the feature vector after average pooling, is the sequence before bit field interaction; Then, bit field re-scaling is performed to obtain the reset three-dimensional features, and the specific calculation formula is as follows: wherein is the reset three-dimensional feature; Subsequently, the bit field re-scaling features are maximum-pooled in the spatial dimension and average-pooled , and the two are connected to form the final feature vector: 3. Constructing a multi-level self-decoupling feature fusion mechanism, the feature vectors generated in the time-frequency domain and the time domain are spliced in the channel dimension to construct a multi-view joint representation; on this basis, a channel-level self-decoupling strategy is used to reconstruct and reverse map the features in the high-level semantic space, generate spatial features that match the semantics and pass them into the next level of encoder-decoder, so as to achieve adaptive complementary and progressive optimization of multi-view semantics, and the steps are as follows: (3a) Channel-level splicing and joint representation construction First, the initial feature vectors generated in the time-frequency domain and the time domain are fused by splicing the features in the channel dimension; specifically, for the layer, wherein = , is the total number of layers; the features extracted from the signal branch are cascaded along the channel dimension to generate multi-view joint representation : Ensure the preliminary integration of complementary information in the time-frequency domain and time domain perspective in the high-level semantic space, form a unified multi-view representation, lay the foundation for subsequent optimization link; (3b) feature reconstruction and reverse mapping under channel level self-decoupling Based on the above joint representation , through the parameterized fusion operator Reconstruct the high-level semantic space to generate decoupled fusion features : Subsequently, split along the channel dimension into two subsets and , respectively through reverse mapping transformer and Realize the generation of spatial features between the perspectives of semantic matching: (3c) progressive optimization The generated semantic space features and are transmitted to the next level of codec as the input of the first layer, ensuring progressive link optimization, this mechanism realizes the adaptive complementarity and progressive optimization of multi-view semantics, significantly improving the robustness and generalization ability of the model.
[0005] The method has the following advantages: (1) Efficient noise filtering capability: through the wavelet domain physical information prior perception to guide the visual state space parameter update, suppress the noise dominant high frequency component, so as to eliminate the noise sensitive component in the image, solve the problem of image distortion caused by background noise interference. Secondly, the threshold guided spatio-temporal cross-domain attention mechanism can adaptively select noise images at the channel level, focus on the effective information in the domain and selectively suppress noise channels, so as to granulate the real features.
[0006] (2) Cross-cycle feature perception capability: construct a parallel multi-scale bit domain modulation architecture, implement multi-scale frequency band parallel decomposition on the original sequence to capture fault information on different frequency bands; then use wide convolution kernel and bit domain module for synchronous collaborative optimization, reconstruct the spatio-temporal features of each scale, generate multi-scale joint representation, and make up for the deficiency of time-frequency domain in cross-cycle perception capability.
[0007] (3) Multi-view depth collaborative fusion capability: After the channel-level self-decoupling of each view feature, the matching semantic spatial features are input into the next layer of the encoder-decoder, thereby realizing adaptive complementation and progressive optimization of multi-view semantics, and significantly enhancing the fault classification robustness and precision under complex working conditions. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 is the overall framework diagram of the present application Figure 2 is a confusion matrix classification effect diagram of a random experiment on the HIT dataset DETAILED DESCRIPTION
[0009] 1. Data preprocessing Data acquisition and standardization: drive the aero-engine test bench, measure and record the vibration signals of the inter-shaft bearing using eddy current sensors and acceleration sensors under normal and fault working conditions of the bearing, calculate the mean and standard deviation of the training set signals, and perform min-max normalization on the original signals to eliminate dimensional differences and improve model convergence efficiency; Time domain signal acquisition: a fixed length sliding window is used to sample the time series data, the window length is set to 1024 sampling points, and the step length is 512 points, to obtain a one-dimensional time domain signal, ensuring that the sample covers the complete vibration period; Time-frequency domain image acquisition: the normalized vibration signal is converted into a two-dimensional time-frequency image through continuous wavelet transform, the cmor50-1 wavelet basis function is selected, the fixed scale number is set to 256, and the time-frequency representation is generated under the condition of a sampling frequency of 25 kHz. The absolute value of the wavelet coefficient represents the signal energy, and the contour map is used to generate a time-frequency image with a size of 224x224 pixels, in which the horizontal axis represents the time sequence and the vertical axis represents the frequency distribution, and the image color depth corresponds to the signal energy intensity, thereby forming a two-dimensional time-frequency image with time domain transient characteristics and frequency domain resonance characteristics.
[0010] 2. Constructing a time-frequency domain view network The two-dimensional time-frequency image obtained after processing is input into the time-frequency view network, and the following operations are performed in turn: First, the input feature map is divided into two groups in the channel dimension, and then respectively and in parallel input into the branch dominated by the wavelet domain physical information prior perception module and the branch dominated by the threshold-guided space-time cross-domain attention. Each branch independently processes C / 2 channels (C is the total number of channels), and the two branches are selected to balance the calculation efficiency and feature diversity. Channel division can reduce the computational complexity of each branch while maintaining sufficient representation ability to capture complex features. After processing, the features generated by the two branches are fused through adaptive gating and channel rearrangement to enhance the feature complementarity and model robustness.
[0011] 3. Constructing a time domain perspective network The one-dimensional signal obtained after preprocessing is input into the time domain perspective network, and the following operations are sequentially performed: First, for the original vibration sequence , a multi-scale set is used to construct a segmentation operator to divide the signal into subsequences with a length of . Then, a wide kernel convolution is sampled for each branch sequence to capture the cross-cycle correlation, and a point-by-point projection and a sigmoid function Sigmoid processing are performed to obtain the segmented features. The feature sequence is generated through gating modulation, and then the channel dimension maximum and average pooling of the feature sequence is performed to generate weights and multiply the original sequence point by point to perform bit field re-labeling. The reset features are calculated to strengthen the fault-sensitive components in the channel dimension. Finally, the maximum and average pooling in the spatial dimension are performed on the reset features, and the two are connected to form the final feature vector, realizing efficient aggregation of global and local features.
[0012] 4. Multi-level self-decoupling fusion mechanism for fault diagnosis The feature vectors in the time-frequency domain and the time domain are spliced in the channel dimension to construct a multi-perspective joint representation. On this basis, a channel-level self-decoupling strategy is used to reconstruct and reverse map the features in the high-level semantic space, generate spatial features that match the semantics, and pass them into the next level of encoder and decoder, so as to achieve adaptive complementation and progressive optimization of multi-perspective semantics.
[0013] The effects of the present application are further verified by the following experiments: On the HIT aero-engine bearing and CWRU bearing data sets, the average accuracy of the present application method reaches 99.98% and 100.0%, respectively, which is significantly better than the comparative models (ResNet18, FasterNet and Transformer, etc.); In the anti-noise experiment, under the condition of-5dB strong noise, the present application method still maintains an accuracy of 94.75% on the HIT data set, and the accuracy can be restored to 99.49% at a signal-to-noise ratio of 5dB; The ablation experiment proves the necessity of the cooperation of each module, and the accuracy of the single time domain perspective network on the HIT data set is 96.53%, and the accuracy is increased to 99.98% after adding the time-frequency domain perspective network, which verifies the absolute advantage of multi-perspective signal feature and physical prior perception network in complex fault feature extraction. Through the above experiments, the fault diagnosis effect of the present application is further verified.
Claims
1. A method for diagnosing aero-engine faults based on multi-view signals and physical prior perception, characterized in that, The method includes the following steps: (1) Construct a time-frequency domain perspective network, and suppress noise components and refine sensitive features by using a wavelet domain physical prior perception module and threshold-guided spatiotemporal cross-domain attention gating modulation; (2) Construct a time-domain perspective network, design a parallel multi-scale bit-domain modulation architecture, perform multi-scale frequency band parallel processing on the signal in the time domain, and use a wide-kernel convolution coupled bit-domain recalibration module to reconstruct the multi-scale features and provide cross-cycle representation for the time-frequency domain perspective. (3) Construct a multi-level self-decoupling feature fusion mechanism, splice the feature vectors generated in the time-frequency domain and the time domain in the channel dimension to construct a multi-view joint representation; on this basis, use the channel-level self-decoupling strategy to perform feature reconstruction and reverse mapping in the high-level semantic space, generate spatial features for semantic matching and pass them into the next level encoder and decoder, thereby achieving adaptive complementarity and progressive optimization of multi-view semantics.
2. The aero-engine fault diagnosis method based on multi-view signals and physical prior perception according to claim 1, characterized in that... (1) The construction of the time-frequency domain perspective network, through the wavelet domain physical prior perception module and threshold-guided spatiotemporal cross-domain attention gating modulation, suppresses noise components and refines sensitive features. The steps are as follows: (2a) Wavelet domain physical prior information perception Let the features of the input wavelet time-frequency image be... ,in For batch quantity, These represent the height of the image's time dimension and the width of its frequency dimension, respectively. The number of channels after image decomposition corresponds to time-frequency components at different scales; the feature dimension is expanded to [number] through linear projection. ,in , This is the channel expansion coefficient, set to 2 to balance feature capacity and computational efficiency; and it is split into feature branches. With gated branches : Next Convert to channel priority format Furthermore, depthwise convolution is used to enhance local feature edges, avoiding confusion of faulty texture features caused by cross-channel interference. The expression is shown below: Where, groups= , The kernel size is [size]. Zero-padding values are used to ensure that the feature scale after convolution is consistent with the input. Activation functions enhance the nonlinear representation of fault characteristics through nonlinear variations; Then, to Perform wavelet decomposition to obtain frequency scale information, assuming For two-dimensional wavelet decomposition functions, the following is adopted: wavelet basis Layer decomposition yields low-frequency approximation coefficients. With high-frequency detail coefficient ,in These are the high-frequency coefficients in the horizontal, vertical, and diagonal directions, respectively. This decomposition result will serve as the physical constraint basis for the selective state space, directly guiding the processing strategy of state parameters for different frequency components; further... Combined with the frequency scaling information obtained from wavelet decomposition, it is reconstructed into a feature flow in four directions: in Indicates the dimension of spatial flattening. This is the low-frequency enhancement factor, whose value is dynamically adjusted according to the intensity of high-frequency noise. When the noise increases, Increase the weights of low- and mid-frequency features to enhance them, and conversely, maintain feature balance. Based on the above multi-directional-frequency joint feature flow The core of the selective state-space model's dynamic parameter update process lies in embedding the wavelet domain's high-frequency noise suppression mechanism into the calculation of the state transition matrix and dynamic time step, enabling the state learning process to actively follow the physical laws of suppressing high-frequency noise and preserving mid-to-low-frequency fault characteristics. The continuous-domain state equation of the selective state space is first modeled based on the dynamic characteristics of the vibration signal, assuming the state variables are... The state update formula for the continuous domain is: Where the state transition matrix Incorporating prior wavelet physical parameters, the traditional selective state space... Matrix stability is ensured only by negative exponents, while this module is based on... Sparsity dynamic adjustment The attenuation intensity is mathematically expressed as: in These are the learnable basic state matrix parameters. This is the high-frequency attenuation coefficient. The Frobenius norm is used; to adapt to discrete image data, the above continuous domain model needs to be discretized by introducing a dynamic time step. This module further incorporates wavelet domain physical information constraints; sparsity adjustment The specific calculations are as follows: in For time step bias, This is the noise step size adjustment factor. for Flattened vector; Discretized state transition matrix With input projection matrix for: (2b) Dynamic threshold channel allocation For the input wavelet time-frequency features are ,in Indicates batch size, Indicates the number of channels. and These represent the height and width of the time-frequency plot, respectively. First, a global spatial mean is calculated along the channel dimension of the input to characterize the overall response intensity of each channel. The specific steps are as follows: in For channel-level response vectors, Indicates the first One sample in the channel Spatial location The eigenvalues; then through Depth-Wise convolutions are used to capture local temporal-frequency textures within channels, followed by... Point convolution performs channel projection and linear combination to obtain projection features. : Then, the global mean operation is repeated on the projected features to obtain the channel statistics after projection. And define the channel response differential: in Indicates in The time-to-channel linear mapping of the original statistics, difference It directly reflects the increasing or decreasing trend of the response of each channel after the noise channels interact. This indicates that the channel is amplified in the aggregated representation, making it more discriminative, and then a learnable threshold parameter is used. And thus construct a dynamic binarization mask: Here, 1 indicates that the channel with significant difference is selected, and 0 indicates that it is suppressed; then, the difference values after mask filtering are probabilistically normalized along the channel dimension: in And for each sample The sum over the channel dimension is 1; Subsequently, the original projection response, probabilistic gain, and differential nonlinear mapping are fused to construct the final channel attention score: in Attention weights for the channel dimension. for , Used to capture the nonlinear relationship of the difference; channel attention obtains channel-weighted features by scaling the principal elements of the projected features after broadcasting in the spatial dimension: After obtaining the channel enhancement features, the channel response based on the channel mean and its corresponding enhanced spatial response are further calculated in parallel: in These are the channel average spatial mappings before and after fusion; constructing spatial differences. And apply to the flattened spatial position vector To obtain the spatial location weights: in For each sample, a probability distribution is formed at its spatial location, and then reshaped into... Shapes are obtained by broadcasting according to channels, resulting in spatially weighted features. Channels and spatial branches are controlled by learnable gating coefficients. Blend proportionally: in The weighted features are the result of fusion; to preserve the details and numerical stability of the original wavelet image, the fusion result is finally fed back to the input via residual connections: in This is a learnable scaling factor used to control the strength of the compensation of the fused features to the original input; (2c) Spatiotemporal cross-domain attention computation Calculate the channel mean at each spatial location as a global spatial statistic; for the input feature map Global spatial statistics are defined as follows: in This represents a global spatial statistics graph, indicating the average channel response intensity at each spatial location. Next, adopt Point convolution The input feature map is transformed; simultaneously, the global spatial statistics of the transformed feature map are calculated. (The calculation method is the same as above, and the input is...) This serves as a reference for spatial enhancement; to highlight the spatial region of the fault, a spatial differential signal is defined: in This reflects the degree of change in spatial position after point convolution processing; through The function normalizes the spatial dimension of the differences to obtain the spatial augmentation weights: in Along spatial dimensions ( The operation-enhancing weights satisfy a probability distribution in space; For the transformed feature map Spatial max pooling and average pooling are performed separately, and then the pooling results are concatenated along the channel dimension. Convolution and Function generates spatial attention graph: in Capture spatial contextual semantic information by combining local receptive fields; Next to Perform global average pooling to compress the spatial dimension, then... Convolution and After activation, via The function generates channel attention vectors; the specific calculation process is shown below: in It is a feature map after compressing spatial dimensions. It is the channel attention vector; Finally, the feature maps are weighted by attention scores using assigned weights to obtain the output. The specific calculation process is as follows: in Attention guided by spatial difference enhancement (Indicates broadcast to channel dimension). For the additive fusion of channel-space attention, Normalization of the spatial enhancement part, This represents the final weighted output of the feature maps.
3. The aero-engine fault diagnosis method based on multi-view signals and physical prior perception according to claim 1, characterized in that... (2) The construction of the time-domain perspective network, the design of the parallel multi-scale bit-domain modulation architecture, the multi-scale frequency band parallel processing of the signal in the time domain, and the use of a wide-kernel convolution coupled bit-domain recalibration module to reconstruct the multi-scale features, providing cross-cycle representation for the time-frequency domain perspective, are carried out as follows: (3a) Parallel multi-scale frequency band processing The original vibration sequence is (batch ,aisle Sampling length First, in the time scale set A multi-scale segmentation operator is constructed to segment the signal sequence in parallel. A length of The subsequence is processed as follows: in Indicates the first Scale No. Segment truncation operator; Then for each scale branch First, a wide-kernel convolution operator is used to capture long-term correlation and broadband coupling effects. Then, through point-by-point projection and nonlinear activation, segmented features are obtained. The calculation process is as follows: in For the first The length of the wide core of the branch ( ), This represents the number of channels output by this branch. For point-by-point projection, For activation functions; Next, the outputs of each segment are fused using gated modulation. The calculation process is as follows: in These are the characteristics after gated modulation; (3b) Bit-field gated recalibration The fused features are then subjected to max pooling and average pooling along the channel dimension, and the two are fused into a sequence before bit-domain interaction; the specific calculation process is as follows: Then, bit-domain recalibration is performed to obtain the reset 3D features. The specific calculation formula is as follows: in The reset 3D features; Subsequently, the location recalibration features were defined. Perform max pooling in the spatial dimension. With average pooling The two are processed and concatenated to form the final feature vector. in These are the eigenvectors.
4. The aero-engine fault diagnosis method based on multi-view signals and physical prior perception according to claim 1, characterized in that... Step (3) describes the construction of a multi-level self-decoupling feature fusion mechanism, which concatenates the feature vectors generated in the time-frequency domain and the time domain in the channel dimension to construct a multi-view joint representation. Based on this, a channel-level self-decoupling strategy is used to perform feature reconstruction and reverse mapping in the high-level semantic space, generating spatial features for semantic matching and feeding them into the next level encoder / decoder, thereby achieving adaptive complementarity and progressive optimization of multi-view semantics. The steps are as follows: (4a) Channel-level splicing and joint characterization construction First, feature concatenation along the channel dimension is used to fuse the initial feature vectors generated in the time-frequency domain and the time domain; specifically, for the first... Layer (of which) = , (Total number of layers), features extracted from signal branches and The data are cascaded along the channel dimension to generate a multi-view joint representation. : This ensures the initial integration of complementary information from the time-frequency domain and the time-domain perspective in the high-level semantic space, forming a unified multi-perspective representation and laying the foundation for subsequent optimization. (4b) Feature reconstruction and inverse mapping under channel-level self-decoupling Based on the above joint characterization Through parameterized fusion operator right Reconstruct the high-level semantic space to generate decoupled fused features. : Then Divide into two subsets evenly along the channel dimension and Each through a reverse mapping transformer and Spatial feature generation for semantic matching between viewpoints: (4c) Progressive optimization The semantic space features generated above and It is passed to the next level of codec as the first The input of each layer ensures progressive link optimization. This mechanism enables adaptive complementarity of multi-perspective semantics and progressive optimization layer by layer, which significantly improves the robustness and generalization ability of the model.