Artificial intelligence-based compressor fault diagnosis method

By employing a fault diagnosis method based on a multi-scale coupled spatiotemporal feature extraction network and metric learning, the problems of weak anti-interference capability and small sample fault identification in compressor fault diagnosis are solved. This method achieves high-precision fault type classification and root cause localization, meeting the intelligent operation and maintenance needs of industrial sites.

CN122365008APending Publication Date: 2026-07-10JINAN BAIFU REFRIGERATION EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINAN BAIFU REFRIGERATION EQUIP CO LTD
Filing Date
2026-04-16
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing compressor fault diagnosis technologies have weak anti-interference capabilities, making it difficult to separate complex noise from effective fault features, unable to uncover fault coupling correlations between multi-source signals, difficult to capture long-term time-series evolution characteristics of faults, and have poor generalization ability for rare faults with small samples. They also lack the ability to locate the root cause of faults, making it difficult to meet the needs of intelligent and precise operation and maintenance in industrial sites.

Method used

A fault diagnosis method based on multi-scale coupled spatiotemporal feature extraction network and metric learning is adopted. By preprocessing and enhancing various compressor operating signals, local time series, cross-source coupling and long-range time series features are extracted, and a fault diagnosis model based on metric learning is constructed to support accurate classification and root cause localization of small sample fault types.

Benefits of technology

It improves the accuracy and generalization of fault diagnosis, adapts to complex and changing industrial conditions, reduces false alarms and missed alarms, and provides reliable support for the safe and stable operation of compressors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365008A_ABST
    Figure CN122365008A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of fault diagnosis technology, and particularly relates to an artificial intelligence-based method for compressor fault diagnosis. The method collects multi-source operating signals of compressor vibration, current, and temperature, and enhances weak fault features through preprocessing such as adaptive hybrid noise separation, frequency band optimization, and periodic synchronous averaging. A multi-scale coupled spatiotemporal feature extraction network is constructed to extract multi-scale local temporal, cross-source fault coupling, and long-range temporal dependency features, and then adaptively weights and fuses them. A prototype network is built based on metric learning, and a general feature embedding space is constructed using large-sample base class data for pre-training. Rare fault prototypes are then dynamically updated using a small-sample support set to achieve accurate fault classification. This invention has strong anti-interference capabilities, can mine multi-source signal coupling correlations and long-term evolution features, overcomes the challenge of identifying rare faults with small samples, has high diagnostic accuracy and strong generalization, and is adaptable to complex industrial operating conditions, providing reliable support for intelligent operation and maintenance of compressors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of fault diagnosis technology, and in particular relates to a compressor fault diagnosis method based on artificial intelligence. Background Technology

[0002] Compressors are core equipment in industrial refrigeration, petrochemical, and energy power industries. Operating under continuous high loads and complex variable conditions for extended periods, they are prone to problems such as abnormal vibration, bearing wear, and motor failure. Minor faults reduce equipment efficiency and increase energy consumption, while severe faults can lead to unplanned downtime, production interruptions, and economic losses. Efficient and accurate fault diagnosis is crucial for ensuring the safe and stable operation of compressors. Current traditional fault diagnosis methods rely heavily on single-signal analysis and manual experience, exhibiting weak anti-interference capabilities and difficulty in separating complex noise from effective fault features. Existing AI-based diagnostic solutions have significant shortcomings: they often employ single-scale feature extraction, failing to uncover fault coupling correlations between multi-source signals and struggling to capture long-term temporal evolution characteristics of faults; they have poor generalization ability for rare faults with small sample sizes, and the models do not support dynamic incremental updates; furthermore, most methods can only classify faults, lacking the ability to pinpoint root causes, failing to accurately locate sensitive fault signal segments and causative components, and thus failing to meet the actual needs of intelligent and precise operation and maintenance in industrial settings. Summary of the Invention

[0003] In view of the technical problems existing in the above-mentioned background art, the present invention proposes a compressor fault diagnosis method based on artificial intelligence.

[0004] To achieve the above objectives, the technical solution adopted by the present invention includes the following steps:

[0005] S1. Collect various compressor operating signals to construct a signal set, and perform preprocessing and enhancement operations on each time-series signal in the signal set to obtain a noise-reduced and enhanced signal set;

[0006] S2. Construct a multi-scale coupled spatiotemporal feature extraction network, input the noise-reduced and enhanced signal set into the network, extract local temporal features of different fault-sensitive scales through a multi-scale temporal convolution module, mine fault coupling correlation features between different signal sources through a cross-source attention module, extract long-range temporal dependency features through a bidirectional gated recurrent unit module, and obtain a fault discrimination fusion feature set after adaptive weighted fusion of the three types of features through convolutional layer feature mapping.

[0007] S3. Construct a fault diagnosis model based on metric learning. Input the fault discrimination fusion feature set into the model. First, obtain the general feature embedding space through base class pre-training. Then, update the category prototype through the support set for small sample fault categories to complete the accurate classification of fault types and obtain accurate fault diagnosis results.

[0008] Preferably, the compressor operating signals in step S1 include vibration signals, current signals, and temperature signals.

[0009] Preferably, the specific implementation of preprocessing and enhancement operations performed on each time-series signal in the signal set in step S1 to obtain the noise-reduced and enhanced signal set is as follows:

[0010] S11. Adaptive hybrid noise separation and modal component filtering, targeting each raw time-series signal of the acquired compressor. n=1,2,...,N, where N is the number of signal sampling points. An improved adaptive noise complete set empirical mode decomposition is performed, and the decomposition process satisfies: Where I is the total number of intrinsic mode function components obtained from the decomposition. Let i be the i-th order intrinsic mode function component. To decompose the residual components;

[0011] S12. Calculate the components of each natural mode function. permutation entropy With spectral kurtosis Preset permutation entropy threshold With spectral kurtosis threshold All components are classified and differentiated. If it is determined to be a dominant noise component, it is retained after performing singular value thresholding noise reduction; if and If it is determined to be a fault-sensitive component, it will be preserved completely without distortion; if and It is determined to be a stationary background component, and a moving average smoothing process is performed on it.

[0012] S13. Reconstruct all processed components and residual components to obtain the preliminary denoised signal. ;

[0013] S14. Initial noise reduction signal Perform fast spectral kurtosis calculation, construct a frequency band division system based on a 1 / 3 binary tree filter bank, calculate the spectral kurtosis value of each divided frequency band, and generate a three-dimensional distribution map of spectral kurtosis, center frequency, and bandwidth. Traverse the spectral kurtosis values ​​of all frequency bands in the map to locate the optimal sensitive frequency band corresponding to the global maximum spectral kurtosis value. ,in, This is the cutoff frequency in the frequency band. This is the cutoff frequency in the frequency band;

[0014] S15, Initial noise reduction signal The input is a Chebyshev bandpass filter, which filters out irrelevant background noise and harmonic interference outside the optimal sensitive frequency band, and outputs a frequency band optimized signal. ;

[0015] S16, Optimize the frequency band signal Demodulation using the discrete Teager energy operator yields an energy sequence, calculated as follows: ,in, For the Teager energy operator;

[0016] S17. Perform cyclic stationary autocorrelation analysis on the energy sequence and calculate the cyclic stationary autocorrelation function: ,in, The cycle frequency, For time delay, This is a conjugate operation. For a cyclically stationary autocorrelation function, For the complex exponential modulation term, T is the observation duration of the signal; by peak search, the cyclic frequency corresponding to the significant peak in the cyclic stationary autocorrelation function is located, and the reference period of the fault characteristics is adaptively identified;

[0017] S18. Using the reference period as the period length, perform periodic synchronous averaging on the energy sequence to finally obtain the noise-reduced and enhanced signal of a single signal; after traversing the time-series signals of all acquisition channels of the compressor and completing the full-channel preprocessing, construct the noise-reduced and enhanced signal set.

[0018] Preferably, if the component is determined to be noise-dominant in step S12, the specific implementation of singular value threshold denoising processing is as follows:

[0019] S121. For the dominant noise component to be processed, construct a Hankel matrix with row dimension matching the embedding dimension and column dimension matching the component sampling length, with fixed embedding dimension and unit delay time, and perform zero-mean normalization on the matrix.

[0020] S122. Perform singular value decomposition on the standardized Hankel matrix to obtain a sequence of singular values ​​arranged in descending order of magnitude; calculate the cumulative energy contribution rate of the singular value sequence, and use the position where the cumulative energy contribution rate first reaches the preset threshold as the dividing point to adaptively separate the first column of effective singular values ​​from the last column of pure noise singular values.

[0021] S123. For valid singular values ​​before the boundary point, perform soft thresholding based on the absolute deviation of the median; for pure noise singular values ​​after the boundary point, filter them directly.

[0022] S124. Based on the singular value sequence after thresholding, the denoised Hankel matrix is ​​reconstructed by combining the left and right unitary matrices corresponding to the singular value decomposition. The reconstructed matrix is ​​then inversely transformed into a time-domain sequence using the diagonal averaging method to complete the denoising process.

[0023] Preferably, the multi-scale coupled spatiotemporal feature extraction network in step S2 specifically includes:

[0024] S21, Multi-scale temporal convolution module: Set up 3 groups of one-dimensional causal convolution layers with different dilation coefficients, namely 1, 2 and 4, to extract local temporal features of short period, medium period and long period in the signal respectively, and output multi-scale local features;

[0025] S22, Cross-source attention module: Using multi-scale local features of different signals as query, key, and value vectors, calculate the attention weight matrix across signals, obtain the fault coupling correlation features between different signals, and output the cross-source coupling features;

[0026] S23, Bidirectional Gated Cyclic Unit Module: Input the cross-source coupling characteristics into the bidirectional gated cyclic unit, extract the forward and backward long-range time-series dependency features of the signal, capture the time-series evolution law of the fault features, and output the long-range time-series features.

[0027] S24. Adaptive Feature Fusion Module: Dynamic weights are assigned to multi-scale local features, cross-source coupling features, and long-range temporal features through an attention mechanism. After weighted fusion, a fault discrimination fusion feature set is obtained through convolutional layer feature mapping. The weights are positively correlated with the contribution of the features to fault discrimination.

[0028] Preferably, the fault diagnosis model based on metric learning in step S3 specifically includes:

[0029] S311. Collect historical datasets, using the large number of normal states and fault types in the datasets as base class datasets. Use prototype networks as backbone networks and pre-train on the base class datasets to learn to map the input fault discrimination fusion features to a highly separable feature embedding space. Calculate the feature mean of all samples in the base class in this embedding space to obtain the pre-trained prototypes of each category of the base class. Build an initial prototype library and fix the embedding function parameters of the prototype network to complete the pre-training.

[0030] S312. For rare fault types with a small sample size, construct a support set containing a small number of labeled samples. Input the support set samples into a pre-trained embedding function to obtain the feature embedding vectors of the support set samples. Calculate the mean of the support set embedding vectors to obtain a new prototype for the rare fault category. Add the new prototype to the initial prototype library to complete the dynamic update of the prototype library.

[0031] S313. Input the fault discrimination fusion features of the sample to be diagnosed into the pre-trained embedding function and map them into the feature embedding space to obtain the feature embedding vector of the sample to be diagnosed; calculate the Euclidean distance between the embedding vector and each category prototype in the prototype library, and take the category with the smallest distance as the fault classification result of the sample to be diagnosed.

[0032] Compared with existing technologies, the advantages and positive effects of this invention are as follows: It employs refined preprocessing of multi-source signals, including noise reduction, frequency band optimization, and period enhancement, to accurately highlight weak fault characteristics and significantly improve anti-interference capabilities; a multi-scale coupled spatiotemporal feature network can simultaneously extract local temporal, cross-source coupling, and long-range evolution features, comprehensively capturing fault information; combined with metric learning, it overcomes the challenge of identifying rare faults with small samples, supports dynamic updates of the prototype library without degrading base class diagnostic capabilities; and boasts high overall diagnostic accuracy and strong generalization, adapting to complex and variable industrial conditions, effectively reducing misjudgments and missed judgments, and providing reliable support for the safe and stable operation and intelligent maintenance of compressors. Attached Figure Description

[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a flowchart illustrating an AI-based compressor fault diagnosis method. Detailed Implementation

[0035] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0036] Numerous specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways than those described herein, and therefore the invention is not limited to the specific embodiments disclosed in the following specification.

[0037] In this embodiment, compressors operate under continuous high loads, complex conditions, and high noise levels for extended periods, making them prone to problems such as bearing wear and motor failure. Early fault characteristics are often subtle and difficult to detect. Traditional fault diagnosis relies on human experience and single-signal analysis, resulting in poor anti-interference capabilities, difficulty in extracting effective fault features, and a high risk of misdiagnosis and missed diagnosis. Existing AI-based diagnostic methods often employ single-scale feature extraction, failing to uncover fault coupling correlations between multi-source signals, lacking sufficient capture of long-term evolution features, exhibiting low accuracy in identifying rare faults with small samples, and being unable to incrementally update models. Most solutions can only classify faults, not pinpoint root causes, making it difficult to meet the precise operation and maintenance needs of industrial sites. This invention provides an AI-based compressor fault diagnosis method. Specific implementation details are as follows... Figure 1 As shown.

[0038] First, various compressor operating signals are collected to construct a signal set. Then, preprocessing and enhancement operations are performed on each time-series signal in the signal set to obtain a noise-reduced and enhanced signal set.

[0039] Furthermore, the compressor operating signals include vibration signals, current signals, and temperature signals.

[0040] Furthermore, preprocessing and enhancement operations are performed on each time-series signal in the signal set to obtain a noise-reduced and enhanced signal set. The specific implementation involves adaptive hybrid noise separation and modal component filtering, applied to each original time-series signal of the acquired compressor. n=1,2,...,N, where N is the number of signal sampling points. An improved adaptive noise complete set empirical mode decomposition is performed, and the decomposition process satisfies: Where I is the total number of intrinsic mode function components obtained from the decomposition. Let i be the i-th order intrinsic mode function component. To decompose the residual components, calculate the components of each natural mode function. permutation entropy With spectral kurtosis Wherein, the permutation entropy is calculated as Where J is the total number of sequence distributions in the reconstructed phase space. The probability of the j-th sequence distribution occurring; permutation entropy is used to characterize the randomness of the components, the larger the value, the stronger the random noise characteristics of the corresponding components; the spectral kurtosis The calculation method is as follows: ,in, For components The short-time Fourier transform result, where E{} is the mathematical expectation operator. The spectral kurtosis of the component at frequency f is used to characterize the impact characteristics of the component; the larger the value, the more significant the fault impact characteristics contained in the component; a preset permutation entropy threshold is used. With spectral kurtosis threshold All components are classified and differentiated. If it is determined to be a dominant noise component, it is retained after performing singular value thresholding noise reduction; if and If it is determined to be a fault-sensitive component, it will be preserved completely without distortion; if and It was determined to be a stationary background component, and a moving average smoothing process was performed on it.

[0041] The processed components are reconstructed with the residual components to obtain the preliminary denoised signal. For the initial noise reduction signal Perform fast spectral kurtosis calculation, construct a frequency band division system based on a 1 / 3 binary tree filter bank, calculate the spectral kurtosis value of each divided frequency band, and generate a three-dimensional distribution map of spectral kurtosis, center frequency, and bandwidth. Traverse the spectral kurtosis values ​​of all frequency bands in the map to locate the optimal sensitive frequency band corresponding to the global maximum spectral kurtosis value. ,in, This is the cutoff frequency in the frequency band. The cutoff frequency in the frequency band; the initial noise reduction signal The input is a Chebyshev bandpass filter, which filters out irrelevant background noise and harmonic interference outside the optimal sensitive frequency band, and outputs a frequency band optimized signal. .

[0042] Finally, the frequency band optimization signal was analyzed. Demodulation using the discrete Teager energy operator yields an energy sequence, calculated as follows: ,in, The Teager energy operator is used; cyclic stationary autocorrelation analysis is performed on the energy sequence to calculate the cyclic stationary autocorrelation function: ,in, The cycle frequency, For time delay, This is a conjugate operation. For a cyclically stationary autocorrelation function, For the complex exponential modulation term, T is the observation duration of the signal; by peak search, the cyclic frequency corresponding to the significant peak in the cyclic stationary autocorrelation function is located, and the reference period of the fault characteristics is adaptively identified; with the reference period as the period length, the energy sequence is subjected to periodic synchronous averaging processing, and finally the noise-reduced and enhanced signal of a single signal is obtained; after traversing the time-series signals of all acquisition channels of the compressor and completing the full-channel preprocessing, the noise-reduced and enhanced signal set is constructed.

[0043] Furthermore, the specific implementation of performing singular value threshold denoising processing on a component determined to be noise-dominant is as follows:

[0044] For the dominant noise component to be processed, a Hankel matrix is ​​constructed with a fixed embedding dimension and unit delay time, where the row dimension matches the embedding dimension and the column dimension matches the component sampling length. Zero-mean normalization is then performed on the matrix. Specifically, the dominant noise component to be processed is a single-channel one-dimensional temporal intrinsic mode function component obtained through improved adaptive noise complete set empirical mode decomposition, with a sequence sampling length of N. First, the fixed embedding dimension and unit delay time are determined. The embedding dimension is a fixed integer value between 10 and 50, based on the number of sampling points corresponding to the minimum period of the compressor fault characteristics. The unit delay time is fixed at 1 to ensure the integrity of the temporal correlation of the sequence. Based on these parameters, a Hankel matrix is ​​constructed. The row dimension of the matrix is ​​consistent with the embedding dimension, and the column dimension is N minus the embedding dimension plus 1. The element in the i-th row and j-th column of the matrix corresponds to the (j+i-1)-th sampling point in the original component sequence, fully preserving the temporal evolution correlation of the sequence. The constructed Hankel matrix is ​​then subjected to zero-mean standardization. First, the arithmetic mean and standard deviation of all elements in the matrix are calculated. Then, a standardization transformation is performed on each element in the matrix. The transformation method is to subtract the arithmetic mean from the element and then divide by the standard deviation, so that the mean of the matrix elements is 0 and the standard deviation is 1, thus eliminating the interference of magnitude differences on subsequent singular value decomposition.

[0045] Singular value decomposition (SVD) is performed on the standardized Hankel matrix, resulting in a sequence of singular values ​​arranged in descending order of magnitude. The cumulative energy contribution rate of the singular value sequence is calculated, and the position where the cumulative energy contribution rate first reaches a preset threshold is used as the dividing point to adaptively separate the first column of valid singular values ​​from the subsequent column of pure noise singular values. Specifically, the SVD operation is performed on the Hankel matrix that has undergone zero-mean standardization, and the decomposition formula is as follows: Where H is the normalized Hankel matrix, U is a left unitary matrix with the same order as the embedding dimension, and V is a right unitary matrix with the same order as the column dimension. Let V be the conjugate transpose of V. A diagonal matrix matching the dimension of the H matrix is ​​used, with the diagonal elements being singular values ​​arranged in descending order of amplitude, thus obtaining an ordered singular value sequence. The cumulative energy contribution rate of the singular value sequence is then calculated. First, the energy value corresponding to each singular value is calculated as the square of its amplitude. The total energy of the sequence is the sum of the squares of all singular values. The cumulative energy contribution rate of the first k singular values ​​is the ratio of the sum of the squares of the first k singular values ​​to the total energy of the sequence. A preset cumulative energy contribution rate threshold of 90% is set. The ordered singular value sequence is traversed to find the minimum value k that makes the cumulative energy contribution rate reach the preset threshold for the first time. Using this k value as a dividing point, the first k singular values ​​are classified as valid singular values, corresponding to the fault-related valid information contained in the components. The k+1 and subsequent singular values ​​are classified as pure noise singular values, corresponding to the random noise components in the components, achieving an adaptive and precise division between valid information and noise.

[0046] For valid singular values ​​before the cutoff point, soft thresholding based on the median absolute deviation is performed; for pure noise singular values ​​after the cutoff point, they are directly filtered out. Specifically, for valid singular values ​​before the cutoff point, soft thresholding based on the median absolute deviation is performed. First, the median absolute deviation (MAD) of the valid singular value sequence is calculated. Then, the median of the valid singular value sequence is calculated, and the absolute deviation of each singular value from this median is calculated, resulting in an absolute deviation sequence. The median of this absolute deviation sequence is taken and multiplied by a correction factor of 1.4826 to adapt to the statistical characteristics of normally distributed noise, yielding the final MAD value. A soft threshold is then calculated based on the MAD value. The calculation formula is: ,in This is the threshold adjustment coefficient, ranging from 1.0 to 1.5. It can be adaptively adjusted according to the ambient noise intensity at the compressor site. The stronger the ambient noise, the better. The value increases accordingly. Soft thresholding is performed on each valid singular value; when the absolute value of the singular value is greater than the soft threshold... When, the output value is When the absolute value of a singular value is less than or equal to When the output value is 0, this processing can suppress the small noise component remaining in the effective singular values ​​while preserving the effective fault impulse characteristics, thus avoiding the signal oscillation problem caused by hard thresholding. Pure noise singular values ​​after the boundary point are directly set to zero and filtered out, completely eliminating random noise components.

[0047] Finally, based on the thresholded singular value sequence, the denoised Hankel matrix is ​​reconstructed using the left and right unitary matrices corresponding to the singular value decomposition. The reconstructed matrix is ​​then inversely transformed into a time-domain sequence using the diagonal averaging method, completing the denoising process. Specifically, based on the thresholded singular value sequence, a new diagonal matrix is ​​constructed. This matrix has the same dimension as the original diagonal matrix obtained from the singular value decomposition. The first k positions on the diagonal are the effective singular values ​​after soft thresholding shrinkage, and the remaining positions are all 0. Combining the left unitary matrix U and the conjugate transpose of the right unitary matrix obtained from the singular value decomposition, a matrix reconstruction operation is performed to obtain the denoised Hankel matrix. Subsequently, the reconstructed Hankel matrix is ​​inversely transformed into a one-dimensional time-domain sequence using the diagonal averaging method. Specifically, the arithmetic mean of all elements in the denoised Hankel matrix whose row and column numbers sum to a fixed value is taken. This mean is used as the sample value at the corresponding position in the inversely transformed sequence. The average is calculated by traversing all diagonal elements of the matrix, resulting in a denoised time-domain sequence with the same sampling length as the original dominant noise component. This inverse transform method can completely restore the temporal structure of the sequence, avoid temporal misalignment or length distortion during reconstruction, and finally complete the full-process denoising of the noise-dominant component, providing high-quality temporal data for subsequent component reconstruction.

[0048] Then, a multi-scale coupled spatiotemporal feature extraction network is constructed. The noise-reduced and enhanced signal set is input into the network. Local temporal features at different fault-sensitive scales are extracted through a multi-scale temporal convolution module. Fault coupling correlation features between different signal sources are mined through a cross-source attention module. Long-range temporal dependency features are extracted through a bidirectional gated recurrent unit module. After adaptive weighted fusion of the three types of features, a fault discrimination fusion feature set is obtained through convolutional layer feature mapping.

[0049] Furthermore, the multi-scale coupled spatiotemporal feature extraction network includes: a multi-scale temporal convolution module: setting three groups of one-dimensional causal convolution layers with different dilation coefficients, namely 1, 2 and 4, to extract local temporal features of short period, medium period and long period in the signal respectively, and output multi-scale local features.

[0050] Specifically, the multi-branch causal dilation convolution branch construction involves setting up three parallel one-dimensional causal dilation convolution branches, with dilation coefficients set to [values ​​to be filled in]. The kernel size of each branch is uniformly k, the stride is 1, and the zero-padding size matches the dilation coefficient to ensure that the output temporal length is consistent with the input; the formula for calculating a single-branch one-dimensional causal dilated convolution is: ,in, Let d be the convolution output of the branch at time t. Let be the weight value of the i-th convolution kernel. Causal convolution constrains the output at the current time step to depend only on the inputs at the current and historical time steps. The convolution outputs of the three branches are then subjected to non-linear activation and feature extraction via gated linear units (GLUs): ,in, The branch features after activation, For gating learnable weights and biases, The activation function is Sigmoid. Next, the refined features from the three branches are processed using a scale attention mechanism to calculate adaptive weights. The weight values ​​are positively correlated with the feature's contribution to fault detection. Weighted fusion yields multi-scale local features, calculated as follows: ,in, For the adaptive weights of the d-th branch, ,in, This is a global average pooling operation. The weight mapping layer provides learnable parameters; the above processing is performed by traversing all types of signals to obtain multi-scale local feature sets corresponding to various signals. , where m is the m-th signal type.

[0051] Cross-source attention module: Using multi-scale local features of different signals as query, key, and value vectors, calculate the attention weight matrix across signals, obtain the fault coupling correlation features between different signals, and output the cross-source coupling features.

[0052] Specifically, cross-source querying and key-value vector mapping are performed. The multi-scale local features of all central signals are mapped to query Q, key K, and value V vectors respectively through a linear transformation layer. For the m-th signal, its own features are used as the query vector, and the concatenated features of the remaining M-1 signals are used as the key-value vectors, where M is the total number of signal types. The calculation formula is: , , ,in, All are learnable linear mapping weight matrices. An H-head parallel attention mechanism is employed, with each head... Calculate attention weights The outputs of the H-head attention are concatenated along the channel dimension, aggregated through a linear transformation layer to obtain the cross-source coupling features of the m-th signal, and then the feature distribution is optimized through layer normalization and residual connections. The calculation formula is as follows: ,in, To output the mapping weight matrix, Let m be the cross-source coupling feature of the m-th signal; traverse all signal sources to obtain the cross-source coupling feature set.

[0053] The bidirectional gated recurrent unit (GRU) module: Inputting cross-source coupling features into the bidirectional gated recurrent unit extracts the forward and backward long-range temporal dependence features of the signal, captures the temporal evolution of fault features, and outputs long-range temporal features. Specifically, it constructs forward and backward GRU branches. The forward branch inputs cross-source coupling features in forward temporal order to capture the forward evolution dependence of the fault from its early stages to degradation; the backward branch inputs cross-source coupling features in reverse temporal order to capture the retrospective correlation and precursor features after the fault occurs. The core calculation method for a single branch is as follows: , , , ,in, Let be the cross-source coupling feature vector of the m-th signal at time t. To reset the door, To update the door, In the candidate hidden state, Output the hidden state at time t. For learnable weight matrix, For bias terms, The product is element-wise. Irrelevant historical information is filtered out by resetting the gates, and the gates are updated adaptively to preserve long-range temporal dependencies related to faults. The forward-order hidden state output from the forward GRU branch and the reverse-order hidden state output from the backward GRU branch are concatenated along the channel dimension to obtain bidirectional temporal features. Then, the bidirectional features are filtered through gated linear units to suppress stationary background features that are unrelated to fault evolution. The calculation formula is as follows: ,in, These are the learnable weights and biases of the gating layer, respectively. The filtered bidirectional temporal features are used. For the filtered full-temporal bidirectional features, an adaptive weight is calculated for each temporal position using a temporal attention mechanism. The weight value is positively correlated with the contribution of the feature at that position to fault detection. The final long-range temporal features are obtained after weighted summation, calculated as follows: ,in, Let be the weight of the temporal position at time t. ,in, The parameters are learnable for the time-series weight mapping layer; by traversing all signal sources, a long-range time-series feature set is obtained.

[0054] The adaptive feature fusion module assigns dynamic weights to multi-scale local features, cross-source coupled features, and long-range temporal features through an attention mechanism. After weighted fusion, a fault-detection fusion feature set is obtained through convolutional layer feature mapping. The weight magnitude is positively correlated with the feature's contribution to fault detection. Specifically, firstly, dimension alignment processing is performed on the input multi-scale local features, cross-source coupled features, and long-range temporal features. A linear transformation layer maps the temporal length and channel dimension of the three types of features to a completely consistent dimensional space, ensuring the dimensionality matching of feature fusion and eliminating dimensional differences in the outputs of different feature extraction branches. Subsequently, an adaptive weight generation branch is constructed, performing global average pooling on the three aligned features to compress the temporal dimension and obtain the global context representation vector of the corresponding feature. The three representation vectors are input into a shared two-layer fully connected mapping network. After nonlinear transformation through the ReLU activation function, they are normalized by Softmax to generate three dynamic weight coefficients that sum to 1. The magnitude of the weight coefficient is positively correlated with the corresponding feature's contribution to fault detection. The three types of features are summed after performing element-wise weighted operations with their corresponding weight coefficients to obtain the initial fused features. The initial fused features are then input into a one-dimensional convolutional layer to perform feature mapping. Local feature aggregation and redundant information filtering are completed by using convolutional kernels adapted to the time sequence length. The final fault discrimination fused features are then output after passing through the ReLU activation function. After fusion processing of the signal features from all acquisition channels of the compressor, a complete fault discrimination fused feature set is constructed.

[0055] A fault diagnosis model based on metric learning is constructed. The fault discrimination fusion feature set is input into the model. First, a general feature embedding space is obtained through pre-training with base classes. Then, for small sample fault categories, the category prototype is updated through the support set to complete the accurate classification of fault types. At the same time, based on gradient weighted class activation mapping, the signal sensitive segment and feature dimension corresponding to the fault are located to realize the root cause of the fault.

[0056] Furthermore, the fault diagnosis model based on metric learning specifically includes: collecting historical datasets, using a large number of normal states and fault types in the datasets as base class datasets, employing a prototype network as the backbone network, pre-training on the base class datasets, learning to map the input fault discrimination fusion features to a highly separable feature embedding space, calculating the feature mean of all samples in the base class in this embedding space, obtaining pre-trained prototypes for each category of the base class, constructing an initial prototype library, and simultaneously fixing the embedding function parameters of the prototype network to complete pre-training. Specifically, historical datasets of the compressor's entire lifecycle operation are collected. The datasets cover the compressor's normal operating state and fault discrimination fusion features corresponding to various common fault types. All samples have been labeled through fault simulation experiments and cross-validation of on-site operation and maintenance records to ensure labeling accuracy and data reliability. Normal states and common fault types with a single class sample size of no less than 500 groups are selected as base class datasets, and divided into training and validation sets in an 8:2 ratio to ensure the consistency of the distribution of training and validation data. A prototype network is used as the backbone network, constructing an embedding function consisting of three fully connected layers with batch normalization and ReLU activation functions added between layers. The output dimension of the embedding function is fixed at 256 dimensions to ensure class discrimination in the feature embedding space. A joint loss function combining cross-entropy loss and triplet loss is used, and pre-training is performed on the base class dataset using the Adam optimizer. The training epochs are set to 100. Training is stopped early when the classification accuracy on the validation set does not improve by a set threshold for 10 consecutive epochs. After pre-training, all weight parameters of the embedding function are fixed. The arithmetic mean of the feature vectors of all training samples of each class in the base class dataset after being mapped by the embedding function is calculated to obtain the pre-trained prototype for that class. A unique class code is assigned to each prototype, and all base class prototypes are aggregated to build an initial prototype library, completing the entire pre-training process.

[0057] Next, for rare fault types with small sample sizes, a support set containing a small number of labeled samples is constructed. The support set samples are input into a pre-trained embedding function to obtain feature embedding vectors for the support set samples. The mean of these embedding vectors is calculated to obtain a new prototype for that rare fault category. The new prototype is then added to the initial prototype library to complete the dynamic updating of the prototype library. Specifically, this step targets rare fault types in compressors with small sample sizes (no more than 50 samples per category), including crankshaft microcracks, early inter-turn short circuits in stator windings, and minor valve plate fractures—fault types for which sufficient samples are difficult to collect during routine maintenance. A labeled support set for each category is constructed. All support set samples have undergone disassembly verification and fault reproduction tests for accurate labeling. Each fault type's support set contains 5 to 20 labeled samples, covering the fault characteristics under different load and speed conditions of the compressor, eliminating the interference of operating condition fluctuations on prototype accuracy. The fault discrimination fusion features of all samples in the support set are input into a pre-trained embedding function with fixed parameters, mapping them to a unified feature embedding space constructed through pre-training, resulting in a 256-dimensional feature embedding vector corresponding to each support set sample. Calculate the arithmetic mean of all feature embedding vectors in the support set for this category to obtain a new prototype for this rare fault category. Assign a unique category code to the new prototype that does not overlap with the base class prototype, and label the corresponding fault type name and fault feature description. Add the new prototype directly to the initial prototype library. During the update process, do not modify the values ​​of the base class pre-trained prototype or adjust the parameters of the embedding function to avoid the degradation of the base class fault recognition capability caused by incremental updates, thus completing the dynamic incremental update of the prototype library.

[0058] Finally, the fault discrimination fusion features of the sample to be diagnosed are input into a pre-trained embedding function and mapped to the feature embedding space to obtain the feature embedding vector of the sample to be diagnosed. The Euclidean distance between this embedding vector and each category prototype in the prototype library is calculated, and the category with the smallest distance is taken as the fault classification result of the sample to be diagnosed. Specifically, for the compressor operation sample to be diagnosed, the fault discrimination fusion features extracted in the above steps are input into a pre-trained embedding function with fixed parameters and mapped to a unified feature embedding space constructed in the pre-training stage to obtain a 256-dimensional feature embedding vector of the sample to be diagnosed that is completely consistent with the prototype dimension of the prototype library. The prototypes corresponding to all categories in the prototype library are traversed, and the Euclidean distance between the feature embedding vector of the sample to be diagnosed and each category prototype is calculated. The Euclidean distance is the square root of the sum of the squares of the differences between the corresponding dimensions of the two vectors. The smaller the distance value, the higher the feature similarity between the sample to be diagnosed and the corresponding category. Pre-set fault classification and unknown fault identification thresholds. If the calculated minimum Euclidean distance is less than or equal to the classification threshold, the category of the prototype corresponding to the minimum distance is used as the fault classification result for the sample to be diagnosed. If the minimum Euclidean distance is greater than the unknown fault identification threshold, the sample is determined to be an unknown fault type, and an unknown fault prompt is output simultaneously. Finally, the category name, unique code, and classification confidence score calculated based on the inverse of the distance are output simultaneously to complete the accurate classification of the fault type of the sample to be diagnosed and obtain the fault diagnosis result.

[0059] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments for application in other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A compressor fault diagnosis method based on artificial intelligence, characterized in that, Includes the following steps: S1. Collect various compressor operating signals to construct a signal set, and perform preprocessing and enhancement operations on each time-series signal in the signal set to obtain a noise-reduced and enhanced signal set; S2. Construct a multi-scale coupled spatiotemporal feature extraction network, input the noise-reduced and enhanced signal set into the network, extract local temporal features of different fault-sensitive scales through a multi-scale temporal convolution module, mine fault coupling correlation features between different signal sources through a cross-source attention module, extract long-range temporal dependency features through a bidirectional gated recurrent unit module, and obtain a fault discrimination fusion feature set after adaptive weighted fusion of the three types of features through convolutional layer feature mapping. S3. Construct a fault diagnosis model based on metric learning. Input the fault discrimination fusion feature set into the model. First, obtain the general feature embedding space through base class pre-training. Then, update the category prototype through the support set for small sample fault categories to complete the accurate classification of fault types and obtain accurate fault diagnosis results.

2. The compressor fault diagnosis method based on artificial intelligence according to claim 1, characterized in that, The compressor operating signals in step S1 include vibration signals, current signals, and temperature signals.

3. The compressor fault diagnosis method based on artificial intelligence according to claim 1, characterized in that, The specific implementation of preprocessing and enhancement operations performed on each time-series signal in the signal set in step S1 to obtain the noise-reduced and enhanced signal set is as follows: S11. Adaptive hybrid noise separation and modal component filtering, targeting each raw time-series signal of the acquired compressor. n=1,2,...,N, where N is the number of signal sampling points. An improved adaptive noise complete set empirical mode decomposition is performed, and the decomposition process satisfies: Where I is the total number of intrinsic mode function components obtained from the decomposition. Let i be the i-th order intrinsic mode function component. To decompose the residual components; S12. Calculate the components of each natural mode function. permutation entropy With spectral kurtosis Preset permutation entropy threshold With spectral kurtosis threshold All components are classified and differentiated. If it is determined to be a dominant noise component, it is retained after performing singular value thresholding noise reduction; if and If it is determined to be a fault-sensitive component, it will be preserved completely without distortion; if and It is determined to be a stationary background component, and a moving average smoothing process is performed on it. S13. Reconstruct all processed components and residual components to obtain the preliminary denoised signal. ; S14. Initial noise reduction signal Perform fast spectral kurtosis calculation, construct a frequency band division system based on a 1 / 3 binary tree filter bank, calculate the spectral kurtosis value of each divided frequency band, and generate a three-dimensional distribution map of spectral kurtosis, center frequency, and bandwidth. Traverse the spectral kurtosis values ​​of all frequency bands in the map to locate the optimal sensitive frequency band corresponding to the global maximum spectral kurtosis value. ,in, This is the cutoff frequency in the frequency band. This is the cutoff frequency in the frequency band; S15, Initial noise reduction signal The input is a Chebyshev bandpass filter, which filters out irrelevant background noise and harmonic interference outside the optimal sensitive frequency band, and outputs a frequency band optimized signal. ; S16, Optimize the frequency band signal Demodulation using the discrete Teager energy operator yields an energy sequence, calculated as follows: ,in, For the Teager energy operator; S17. Perform cyclic stationary autocorrelation analysis on the energy sequence and calculate the cyclic stationary autocorrelation function: ,in, The cycle frequency, For time delay, This is a conjugate operation. For a cyclically stationary autocorrelation function, For the complex exponential modulation term, T is the observation duration of the signal; by peak search, the cyclic frequency corresponding to the significant peak in the cyclic stationary autocorrelation function is located, and the reference period of the fault characteristics is adaptively identified; S18. Using the reference period as the period length, perform periodic synchronous averaging on the energy sequence to finally obtain the noise-reduced and enhanced signal of a single signal; after traversing the time-series signals of all acquisition channels of the compressor and completing the full-channel preprocessing, construct the noise-reduced and enhanced signal set.

4. The compressor fault diagnosis method based on artificial intelligence according to claim 3, characterized in that, If the component is determined to be noise-dominant in step S12, the specific implementation of singular value threshold denoising is as follows: S121. For the dominant noise component to be processed, construct a Hankel matrix with row dimension matching the embedding dimension and column dimension matching the component sampling length, with fixed embedding dimension and unit delay time, and perform zero-mean normalization on the matrix. S122. Perform singular value decomposition on the standardized Hankel matrix to obtain a sequence of singular values ​​arranged in descending order of magnitude; calculate the cumulative energy contribution rate of the singular value sequence, and use the position where the cumulative energy contribution rate first reaches the preset threshold as the dividing point to adaptively separate the first column of effective singular values ​​from the last column of pure noise singular values. S123. For valid singular values ​​before the boundary point, perform soft thresholding based on the absolute deviation of the median; for pure noise singular values ​​after the boundary point, filter them directly. S124. Based on the singular value sequence after thresholding, the denoised Hankel matrix is ​​reconstructed by combining the left and right unitary matrices corresponding to the singular value decomposition. The reconstructed matrix is ​​then inversely transformed into a time-domain sequence using the diagonal averaging method to complete the denoising process.

5. The compressor fault diagnosis method based on artificial intelligence according to claim 1, characterized in that, The multi-scale coupled spatiotemporal feature extraction network in step S2 specifically includes: S21, Multi-scale temporal convolution module: Set up 3 groups of one-dimensional causal convolution layers with different dilation coefficients, namely 1, 2 and 4, to extract local temporal features of short period, medium period and long period in the signal respectively, and output multi-scale local features; S22, Cross-source attention module: Using multi-scale local features of different signals as query, key, and value vectors, calculate the attention weight matrix across signals, obtain the fault coupling correlation features between different signals, and output the cross-source coupling features; S23, Bidirectional Gated Cyclic Unit Module: Input the cross-source coupling characteristics into the bidirectional gated cyclic unit, extract the forward and backward long-range time-series dependency features of the signal, capture the time-series evolution law of the fault features, and output the long-range time-series features. S24. Adaptive Feature Fusion Module: Dynamic weights are assigned to multi-scale local features, cross-source coupling features, and long-range temporal features through an attention mechanism. After weighted fusion, a fault discrimination fusion feature set is obtained through convolutional layer feature mapping. The weights are positively correlated with the contribution of the features to fault discrimination.

6. The compressor fault diagnosis method based on artificial intelligence according to claim 1, characterized in that, The fault diagnosis model based on metric learning in step S3 specifically includes: S311. Collect historical datasets, using the large number of normal states and fault types in the datasets as base class datasets. Use prototype networks as backbone networks and pre-train on the base class datasets to learn to map the input fault discrimination fusion features to a highly separable feature embedding space. Calculate the feature mean of all samples in the base class in this embedding space to obtain the pre-trained prototypes of each category of the base class. Build an initial prototype library and fix the embedding function parameters of the prototype network to complete the pre-training. S312. For rare fault types with a small sample size, construct a support set containing a small number of labeled samples. Input the support set samples into a pre-trained embedding function to obtain the feature embedding vectors of the support set samples. Calculate the mean of the support set embedding vectors to obtain a new prototype for the rare fault category. Add the new prototype to the initial prototype library to complete the dynamic update of the prototype library. S313. Input the fault discrimination fusion features of the sample to be diagnosed into the pre-trained embedding function and map them into the feature embedding space to obtain the feature embedding vector of the sample to be diagnosed; calculate the Euclidean distance between the embedding vector and each category prototype in the prototype library, and take the category with the smallest distance as the fault classification result of the sample to be diagnosed.