An electromechanical fault rapid diagnosis method based on adaptive multi-scale fusion

CN122548419APending Publication Date: 2026-08-11CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

传统信号处理方法在处理非平稳振动信号时,难以精准捕捉故障相关的关键信息,频带选择往往缺乏针对性,容易受到无关噪声的干扰导致特征提取效果不佳

Benefits of technology

1、本发明在关键频带筛选阶段,创新性地设计并应用了几何分位概率权重峭度特征指数(GQ-PWKI),传统方法大多仅依赖能量峰值或简单的包络谱幅值,在复杂工业现场的强环境噪声及随机大冲击干扰下,容易出现频带误选,本发明通过引入频率加权差分算子与自相关序列,有效突出了被噪声掩盖的循环平稳冲击特征;同时,结合几何分位映射和概率权重统计矩的无偏估计,克服了传统峭度对偶发性离群点敏感的缺陷,将能量分布值与该特征指数进行加权融合,实现了能量集中度与故障冲击显著度的双重优化,极大提升了模型在背景噪声重灾区精准锁定最佳解调频带的能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548419A_ABST
    Figure CN122548419A_ABST
Patent Text Reader

Abstract

This invention relates to the field of industrial automation equipment monitoring and fault diagnosis technology, and particularly to a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion. The method involves collecting raw vibration signals from electromechanical equipment, generating a standardized two-dimensional time-frequency map through quality inspection and adaptive time-frequency preprocessing, and then augmenting the data using time warping and spectral masking. Multi-scale features are then fused through a deep feature extraction network composed of dual convolutional paths and multi-granularity attention modules. Finally, a comprehensive reliability score is calculated based on the confidence score output by the classifier and historical diagnostic records, outputting the final fault diagnosis category. This invention effectively improves the accuracy of fault feature extraction, model generalization ability, and reliability of diagnostic results, and is applicable to fault diagnosis scenarios for various industrial electromechanical equipment, providing technical support for stable equipment operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial automation equipment monitoring and fault diagnosis technology, and in particular to a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion. Background Technology

[0002] With the continuous improvement of industrial automation, electromechanical equipment is increasingly widely used in production and daily life. The stability of its operating status directly affects production efficiency and safety, and electromechanical fault diagnosis technology has also gradually developed accordingly. Early fault diagnosis relied heavily on the practical experience of maintenance personnel, judging equipment abnormalities through sensory perception, which was highly subjective and inefficient. Subsequently, signal processing-based diagnostic methods gradually emerged, using techniques such as Fourier transform to analyze equipment vibration signals and extract fault features for diagnosis. However, these methods have poor adaptability to signals under complex operating conditions. In recent years, the integration of machine learning and deep learning technologies has driven the development of diagnostic technology towards intelligence. By constructing neural network models to automatically learn fault features, diagnostic efficiency has been improved to some extent, but overall, it is still in a stage of continuous optimization and improvement.

[0003] Despite significant advancements in electromechanical fault diagnosis technology, numerous challenges remain in practical applications. Traditional signal processing methods struggle to accurately capture critical fault-related information when handling non-stationary vibration signals. Frequency band selection often lacks specificity, making them susceptible to interference from irrelevant noise and resulting in poor feature extraction. Furthermore, fault sample data is often scarce in industrial settings, and existing data augmentation methods are insufficiently effective, limiting the model's generalization ability across different operating conditions and hindering its adaptation to complex and changing real-world environments. Moreover, most diagnostic methods rely solely on a single confidence level from the model output to determine reliability, failing to adequately integrate historical equipment operation and diagnostic data. This can lead to misjudgments and fail to provide sufficient reliable support for equipment maintenance. Summary of the Invention

[0004] The purpose of this invention is to overcome the above-mentioned problems and provide a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion. To achieve the above objective, this invention adopts the following technical solution: A rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion includes the following steps: Step S1: Obtain the original vibration signal of the electromechanical equipment, perform quality inspection and adaptive time-frequency preprocessing on the original vibration signal, and generate a standardized two-dimensional time-frequency diagram; Step S2: Perform data augmentation on the standardized two-dimensional time-frequency graph to generate an enhanced time-frequency graph; Step S3: The enhanced time-frequency image is input into the deep feature extraction network, which extracts and fuses multi-scale features to generate a fused feature vector. Step S4: Input the fused feature vector into the classifier to obtain the initial fault diagnosis result; Step S5: Calculate the comprehensive confidence score based on the confidence score of the initial fault diagnosis result and the historical diagnosis records, and output the final fault diagnosis category based on the comprehensive confidence score.

[0005] Furthermore, in step S1, the quality inspection and adaptive time-frequency preprocessing are implemented in the following way: Step S11, perform short-time Fourier transform on the original vibration signal sequence to obtain the initial time spectrum; ; in, The initial frequency spectrum, The original vibration signal sequence varies over time. For time, For window functions, For frequency, The center point of the time window; Step S12: Calculate the energy distribution of the initial spectrum in the frequency dimension; ; in, In frequency Total energy at the location; Step S13, based on energy distribution Select key frequency bands with concentrated energy; Step S14: Construct a set of bandpass filters based on the boundary frequencies of the key frequency bands; Step S15: The original vibration signal sequence is filtered using a bandpass filter bank, and the filtered signal is subjected to short-time Fourier transform and size standardization to generate a standardized two-dimensional time-frequency diagram.

[0006] Furthermore, in step S13, based on the energy distribution... Selecting key frequency bands includes the following steps: Step S131, calculate energy distribution The global average value is used to mark frequency bands whose energy values ​​are continuously higher than the global average value as candidate frequency bands; Step S132: Obtain the physical parameters and operating speed of the electromechanical equipment, and calculate the theoretical fault characteristic frequency related to the potential fault type; ; in, The theoretical characteristic frequency corresponding to a specific failure mode. The fundamental frequency corresponding to the operating speed of the equipment. These are constant coefficients related to the equipment's physical parameters and fault types; Step S133: Using the theoretical fault characteristic frequency as a reference, and combining the fault impact characteristics of the candidate frequency bands, the key frequency bands are selected from the candidate frequency bands.

[0007] Furthermore, in step S133, the fault impact characteristics are obtained by constructing a geometric quantile probability weighted kurtosis feature index, specifically including the following steps: Step S1331: Calculate the frequency-weighted differential envelope vector for the signal sequence within each candidate frequency band. And based on this vector, calculate its corresponding autocorrelation normalized energy sequence. ; Step S1332: Based on the inverse function of the cumulative probability distribution, perform a geometric quantile mapping operation with preset coefficients on the above vectors respectively to obtain the mapped denoised envelope sequence and the mapped autocorrelation sequence. Step S1333: Using the unbiased estimation method of probability weight statistical moments, calculate the linear skewness kurtosis values ​​of the mapped denoised envelope sequence and the mapped autocorrelation sequence respectively. Step S1334: Construct the geometric quantile probability weighted kurtosis feature index based on the two obtained linear skewed kurtosis values. This is used as an indicator of fault impact energy. Step S1335, the energy distribution values ​​of the candidate frequency bands Geometric quantile probability weighted kurtosis feature index Weighted fusion is performed to obtain a comprehensive evaluation value. : ; in, , For the preset weighting coefficients, and Select comprehensive evaluation value The highest candidate frequency band is selected as the key frequency band.

[0008] Further, in step S2, time warping is achieved by performing a nonlinear time-scale transformation on the original vibration signal sequence. The nonlinear time-scale transformation is controlled by a randomly generated scaling factor, which is uniformly and randomly selected within a preset range. After the transformation, the signal length is restored to the original length by linear interpolation. Spectral masking is achieved by randomly masking a continuous rectangular region on a standardized two-dimensional time-frequency map. The size of the rectangular region does not exceed the preset ratio of the corresponding dimension of the time-frequency map. The coordinates of the upper left corner of the rectangular region are randomly determined within the range of the time-frequency map. All pixel values ​​within the masked region are set to zero.

[0009] Further, in step S3, the deep feature extraction network includes a first convolutional path, a second convolutional path, a feature fusion module, and a multi-granularity attention module; the first convolutional path includes a first convolutional layer and a first pooling layer, the second convolutional path includes a second convolutional layer and a second pooling layer, and the kernel size of the first convolutional layer is larger than the kernel size of the second convolutional layer; the feature fusion module concatenates the first feature map output by the first convolutional path and the second feature map output by the second convolutional path in the channel dimension, and processes them through a convolutional layer with a kernel size of 1×1 to generate an initial fused feature; the multi-granularity attention module sequentially performs channel attention weighting and spatial attention weighting on the initial fused feature to output a fused feature vector.

[0010] Furthermore, channel attention weighting performs global average pooling and global max pooling on each channel of the initial fused feature map, concatenates the two feature vectors obtained, inputs them into a multilayer perceptron, and outputs a channel weight vector; spatial attention weighting performs average pooling and max pooling on the channel-weighted feature map in the channel dimension, concatenates the two feature maps, processes them through a convolutional layer, and outputs a spatial weight matrix.

[0011] Furthermore, the classifier employs an adaptive learning rate adjustment strategy during the model training phase. This strategy dynamically adjusts the learning rate based on the relative rate of change of the loss value between adjacent training cycles. When the relative rate of change is lower than a preset threshold, the learning rate is reduced by a preset decay coefficient.

[0012] Further, in step S5, the comprehensive confidence score is calculated and the final fault diagnosis category is output, including the following steps: Step S51, obtain the confidence score of the initial fault diagnosis result output by the classifier; Step S52: Retrieve several historical diagnostic records corresponding to the current electromechanical equipment from the historical diagnostic database, and calculate the consistency index between the current diagnostic category and the category in the historical diagnostic records. Step S53: Calculate the overall credibility score based on the confidence score and the consistency index; ; in, To determine the overall credibility score, The maximum probability value in the initial fault diagnosis results output by the classifier. As a consistency indicator, These are preset weighting coefficients; Step S54, when the overall credibility score is... If the result exceeds the preset confidence threshold, the initial fault diagnosis result is output as the final fault diagnosis category.

[0013] The advantages of this invention are: 1. In the critical frequency band selection stage, this invention innovatively designs and applies the Geometric Quantile Probability Weighted Kurtosis Feature Index (GQ-PWKI). Traditional methods mostly rely on energy peak values ​​or simple envelope spectrum amplitudes, which are prone to frequency band misselection under strong environmental noise and random large-impact interference in complex industrial sites. This invention effectively highlights the cyclic stationary impact characteristics masked by noise by introducing a frequency-weighted difference operator and autocorrelation sequence. At the same time, by combining the geometric quantile mapping and unbiased estimation of probability weighted statistical moments, it overcomes the defect of traditional kurtosis being sensitive to occasional outliers. By weightedly fusing the energy distribution value with this feature index, it achieves dual optimization of energy concentration and fault impact significance, greatly improving the model's ability to accurately lock the optimal demodulation frequency band in areas with heavy background noise.

[0014] 2. This invention performs nonlinear time-scale transformation on the original vibration signal through time warping to simulate speed fluctuations during equipment operation, and recovers the signal length through linear interpolation; combined with spectral masking, it randomly masks continuous rectangular areas on a two-dimensional time-frequency graph to simulate local information loss or distortion during signal acquisition; the two data augmentation operations work together to expand data diversity in the time and frequency domains respectively, enabling the model to learn the invariance of fault features under speed fluctuations and to identify fault features from incomplete time-frequency information, effectively solving the problem of insufficient model generalization ability caused by the scarcity of fault samples in industrial scenarios.

[0015] 3. This invention extracts features at different scales through multi-scale convolutional paths, generates initial fused features through a feature fusion module, and then sequentially performs channel attention weighting and spatial attention weighting through a multi-granularity attention module. Channel attention weighting evaluates the contribution of each channel through global average pooling and global max pooling, while spatial attention weighting focuses on the fault region in the feature map, forming a hierarchical focusing mechanism from coarse-grained to fine-grained, thereby achieving refined enhancement of fault features and significantly improving the ability of deep feature extraction networks to identify weak faults and complex interferences.

[0016] 4. This invention obtains the initial fault diagnosis confidence score output by the classifier, retrieves the historical diagnosis database to calculate the consistency index between the current diagnosis category and the historical records, and weights and fuses the two to obtain a comprehensive confidence score. The final diagnosis result is output only when the score is higher than a preset threshold, realizing the collaborative verification of a single diagnosis result and the historical status of the equipment, effectively reducing the risk of misjudgment caused by occasional noise or model errors. Attached Figure Description

[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application.

[0018] In the attached diagram: Figure 1 This is a flowchart of a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion in Example 1.

[0019] Figure 2 This is a flowchart of the quality inspection and adaptive time-frequency preprocessing process of a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion in Example 1.

[0020] Figure 3 This is a flowchart of the key frequency band screening based on the geometric quantile probability weighted kurtosis feature index in an electromechanical fault rapid diagnosis method based on adaptive multi-scale fusion in Example 1.

[0021] Figure 4 This is a flowchart of the data augmentation operation of a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion in Example 1.

[0022] Figure 5 This is a detailed flowchart of the deep feature extraction process for a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion in Example 1.

[0023] Figure 6 This is a detailed flowchart of the initial classification of a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion in Example 1.

[0024] Figure 7 This is a detailed flowchart of the credibility integration and decision-making process for a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion in Example 1. Detailed Implementation

[0025] The present invention will now be described in detail and specifically through specific embodiments to enable a better understanding of the invention. However, the following embodiments do not limit the scope of protection of the present invention.

[0026] Example 1 like Figure 1 As shown, a rapid electromechanical fault diagnosis method based on adaptive multi-scale fusion includes the following steps: Step S1: Obtain the original vibration signal of the electromechanical equipment, perform quality inspection and adaptive time-frequency preprocessing on the original vibration signal, and generate a standardized two-dimensional time-frequency diagram; Step S2: Perform data augmentation on the standardized two-dimensional time-frequency graph to generate an enhanced time-frequency graph; Step S3: The enhanced time-frequency image is input into the deep feature extraction network, which extracts and fuses multi-scale features to generate a fused feature vector. Step S4: Input the fused feature vector into the classifier to obtain the initial fault diagnosis result; Step S5: Calculate the comprehensive confidence score based on the confidence score of the initial fault diagnosis result and the historical diagnosis records, and output the final fault diagnosis category based on the comprehensive confidence score.

[0027] In this specific embodiment, a commercially available piezoelectric accelerometer is used to collect the raw vibration signal. The sensor range is set to ±5g, and the sampling frequency is fixed at 10kHz. This parameter combination can cover the vibration measurement needs of most industrial electromechanical equipment. The sensor is fixed to the vibration-sensitive area of ​​the equipment bearing housing or shell with bolts. A twisted-pair shielded cable is used, with one end connected to the sensor signal output and the other end connected to an industrial-grade data acquisition card. The acquisition card communicates with an industrial computer via a PCIe interface, reducing electromagnetic interference and attenuation during signal transmission. The quality inspection process adopts the mature three-criteria: first, the signal mean μ and standard deviation σ are calculated; then, excessive spikes are smoothed using a moving average method; and segments with more than 5 consecutive missing sampling points are completed using linear interpolation to ensure signal integrity. During the adaptive time-frequency preprocessing, the process is dynamically adjusted by calculating the signal-to-noise ratio. When the signal-to-noise ratio is below 20dB, wavelet threshold noise reduction preprocessing is used before further operation.

[0028] The standardized two-dimensional time-frequency graphs were uniformly adjusted to a size of 224×224 pixels using a bilinear interpolation algorithm, a size that fits the input requirements of mainstream deep learning networks. Data augmentation was implemented using the open-source Albumentations library in Python, the deep feature extraction network was built on the PyTorch framework, and feature processing was completed through forward propagation. The classifier was optimized using the cross-entropy loss function, and historical diagnostic records were stored in a MySQL database. The entire process can be fully implemented using existing hardware and software.

[0029] Furthermore, such as Figure 2 As shown, in step S1, the quality inspection and adaptive time-frequency preprocessing are implemented in the following way: Step S11, perform short-time Fourier transform on the original vibration signal sequence to obtain the initial time spectrum; ; in, The initial frequency spectrum, The original vibration signal sequence varies over time. For time, For window functions, For frequency, The center point of the time window; Step S12: Calculate the energy distribution of the initial spectrum in the frequency dimension; ; in, In frequency Total energy at the location; Step S13, based on energy distribution Select key frequency bands with concentrated energy; Step S14: Construct a set of bandpass filters based on the boundary frequencies of the key frequency bands; Step S15: The original vibration signal sequence is filtered using a bandpass filter bank, and the filtered signal is subjected to short-time Fourier transform and size standardization to generate a standardized two-dimensional time-frequency diagram.

[0030] In a specific embodiment, the Hanning window, commonly used in industrial signal processing, is selected for the short-time Fourier transform. The window length is set to 256 sampling points. This window function effectively suppresses spectral leakage and achieves a balance between time resolution and frequency resolution. The continuous time-domain vibration signal collected by the sensor slides along the time axis with a step size of 128 sampling points to ensure coverage of the entire signal sequence and reasonable overlap between adjacent time windows. The frequency range is limited to 0-5kHz, matching the sensor sampling frequency. The initial time spectrum is calculated using a numerical integration algorithm, transforming the one-dimensional time-domain signal into a two-dimensional time-frequency matrix.

[0031] The calculation is performed by iterating through each frequency point and accumulating the corresponding values ​​for all time windows at that frequency. The squared value of the modulus visually represents the energy proportion of each frequency component. The bandpass filter adopts a mature Chebyshev Type I filter with an 8th order. The passband ripple is controlled within 1dB, and the stopband attenuation is greater than 40dB. The passband range of each filter is determined according to the upper and lower boundary frequencies of the key frequency band. The number of filters is consistent with the number of key frequency bands, which can accurately filter out high-frequency noise and low-frequency interference.

[0032] The filtered signal is subjected to a short-time Fourier transform again with the same parameters. The size is standardized using the Min-Max normalization algorithm, which maps the pixel values ​​to the 0-1 range to ensure a uniform data format. All processing algorithms can be implemented using Python's NumPy and SciPy libraries.

[0033] Furthermore, in step S13, based on the energy distribution... Selecting key frequency bands includes the following steps: Step S131, calculate energy distribution The global average value is used to mark frequency bands whose energy values ​​are continuously higher than the global average value as candidate frequency bands; Step S132: Obtain the physical parameters and operating speed of the electromechanical equipment, and calculate the theoretical fault characteristic frequency related to the potential fault type; ; in, The theoretical characteristic frequency corresponding to a specific failure mode. Let n be the fundamental frequency corresponding to the operating speed of the device. These are constant coefficients related to the equipment's physical parameters and fault types; Step S133: Using the theoretical fault characteristic frequency as a reference, and combining the fault impact characteristics of the candidate frequency bands, the key frequency bands are selected from the candidate frequency bands.

[0034] This embodiment provides a specific implementation method for selecting key frequency bands and calculating theoretical fault characteristic frequencies based on energy distribution E(f), as follows: In the process of selecting key frequency bands, the arithmetic mean of the energy of all frequency points in the 0-5kHz frequency range is first calculated to obtain the global average energy value, which is used as the screening benchmark. Frequency bands where the energy values ​​of three or more consecutive frequency points are higher than the global average value are marked as candidate frequency bands. This setting can effectively avoid misjudgment caused by a single high energy point and ensure that the candidate frequency bands can truly reflect the energy concentration area of ​​the equipment, providing a reliable basis for fault feature extraction.

[0035] The calculation of theoretical fault characteristic frequencies depends on the equipment's fundamental frequency and fault-related parameters, where the equipment's fundamental frequency... The speed is acquired using an existing, mature photoelectric speed sensor. The sensor's sampling frequency is set to 1Hz, and the acquired speed values ​​are expressed in r / min. This is then calculated using the formula... This is converted to the fundamental frequency in Hertz (Hz) units to ensure the accuracy of the fundamental frequency calculation.

[0036] The equipment's physical parameters, including the number of gear teeth, the number of bearing rolling elements, the pitch circle diameter, and the rolling element diameter, can be obtained directly from the equipment's manufacturer's manual and do not require additional measurement; coefficients The fault type is determined by consulting the mechanical fault diagnosis manual; different fault types correspond to different... Values, specific examples are as follows: Cracks at the root of the gear tooth correspond to... Values ​​of harmonic orders 1, 2, 3, etc., can cover the main characteristic frequencies of gear crack faults; the spalling of the bearing inner ring corresponds to... The formula for calculating the value is: number of rolling elements The outer ring of the bearing peels off. The formula for calculating the value is: number of rolling elements The theoretical fault characteristic frequency was calculated. .

[0037] In step S133, using the calculated theoretical fault characteristic frequency as a reference, and combining it with the fault impact characteristics of the candidate frequency bands, key frequency bands are selected from the candidate frequency bands. The selection process sets the following criteria: the frequency band center and the theoretical fault characteristic frequency. The deviation is no more than 5%, and candidate frequency bands that meet this condition are retained as critical frequency bands. This deviation range has been verified by engineering to ensure accurate capture of equipment fault characteristics while taking into account frequency fluctuations in actual industrial conditions, thereby improving the reliability and adaptability of the screening results.

[0038] Furthermore, such as Figure 3 As shown, in step S133, the fault impact characteristics are obtained by constructing a geometric quantile probability weighted kurtosis feature index, specifically including the following steps: Step S1331, for the filtered signal within each candidate frequency band Traditional envelope methods are susceptible to strong random shocks and background noise interference. This embodiment introduces a frequency-weighted differential operator to improve transient energy sensing capability. Its frequency-weighted differential envelope vector The calculation is as follows: ; in, Represents the Hilbert transform. To further highlight the periodic impact characteristics of being submerged in strong noise, its autocorrelation-normalized energy sequence was calculated. : ; in, represent The mean, This is the signal length.

[0039] Step S1332: To mitigate the impact of large-amplitude random noise on statistical characteristics, this embodiment introduces the concept of geometric quantile mapping, assuming the cumulative probability distribution function of the sequence is... The geometric quantile mapping operator is defined by its inverse function (quantile function). : ; in, For the preset geometric partitioning coefficients (preferred range 1000 to 2000), respectively, and Input the mapping operator to obtain the noise-reduced and interference-resistant mapping sequence. and .

[0040] Step S1333: To address the drawback of traditional kurtosis being susceptible to outliers, this embodiment employs a linear skewed kurtosis evaluation value based on probability-weighted statistical moments (similar to the L-Kurtosis mechanism). For ordered samples of the input sequence, its... Unbiased Estimation Parameters The calculation is as follows: ; Based on this, the linear skewness kurtosis value of the sequence The calculation is as follows: ; The mapping sequences are obtained using this formula. kurtosis value as well as kurtosis value .

[0041] Step S1334: The obtained kurtosis values ​​are nonlinearly combined to construct a geometric quantile probability weighted kurtosis characteristic index that is extremely sensitive to periodic fault impacts and has strong robustness to Gaussian white noise and random large impacts. : ; Step S1335, the energy distribution values ​​of the candidate frequency bands With this characteristic index Perform weighted fusion: Select the comprehensive evaluation value The highest candidate frequency band is used as the key frequency band for input to the deep network, and the weight coefficients are... Take 0.4, Take 0.6; in this step, the focus is more on... The ability to accurately locate the transient impact characteristics of early-stage minor faults.

[0042] Further, in step S2, time warping is achieved by performing a nonlinear time-scale transformation on the original vibration signal sequence. The nonlinear time-scale transformation is controlled by a randomly generated scaling factor, which is uniformly and randomly selected within a preset range. After the transformation, the signal length is restored to the original length by linear interpolation. Spectral masking is achieved by randomly masking a continuous rectangular region on a standardized two-dimensional time-frequency map. The size of the rectangular region does not exceed the preset ratio of the corresponding dimension of the time-frequency map. The coordinates of the upper left corner of the rectangular region are randomly determined within the range of the time-frequency map. All pixel values ​​within the masked region are set to zero.

[0043] In a specific embodiment, the data augmentation operation includes two independently designed operations: time warp and spectral masking. Each time data augmentation is performed, one operation can be randomly selected or the two operations can be used in combination. By expanding the diversity of training data, the deep feature extraction network is exposed to more diverse signal patterns, which improves the model's generalization ability and anti-interference ability. All operations can be quickly implemented through Python's Albumentations open-source library without the need to develop additional dedicated tools.

[0044] Time warp operation: This is achieved by performing a nonlinear time-scale transformation on the original vibration signal sequence. An elastic deformation algorithm is employed, and the specific calculations are performed using the interpolation module of Python's SciPy library. The amplitude of the nonlinear time-scale transformation is controlled by a randomly generated scaling factor, which is randomly generated based on a uniform distribution between 0.8 and 1.2. This range references the slight speed fluctuations in actual industrial equipment operation, ensuring that the signal is not excessively warped, thus avoiding distortion of fault characteristics, while realistically simulating signal changes caused by slight fluctuations in equipment speed in an industrial setting. After the time warp transformation, the signal length is restored to its original length through linear interpolation, ensuring that the duration of the transformed signal remains consistent with the original signal and does not affect time-frequency transformation and feature extraction.

[0045] Spectral masking operation: A continuous rectangular region is randomly masked on the standardized 2D time-frequency plot. The size of the rectangular region does not exceed one-quarter of the corresponding dimension of the time-frequency plot. The coordinates of the top-left corner of the rectangular region are randomly determined within the time-frequency plot range. All pixel values ​​within the masked region are set to zero. This size limitation has been verified in engineering to effectively avoid the loss of key fault features due to an excessively large masking region. It also simulates scenarios such as local electromagnetic interference or transient sensor contact problems that may occur during signal transmission, improving the model's adaptability to signals under complex operating conditions.

[0046] Further, in step S3, the deep feature extraction network includes a first convolutional path, a second convolutional path, a feature fusion module, and a multi-granularity attention module; the first convolutional path includes a first convolutional layer and a first pooling layer, the second convolutional path includes a second convolutional layer and a second pooling layer, and the kernel size of the first convolutional layer is larger than the kernel size of the second convolutional layer; the feature fusion module concatenates the first feature map output by the first convolutional path and the second feature map output by the second convolutional path in the channel dimension, and processes them through a convolutional layer with a kernel size of 1×1 to generate an initial fused feature; the multi-granularity attention module sequentially performs channel attention weighting and spatial attention weighting on the initial fused feature to output a fused feature vector.

[0047] In a specific embodiment, the deep feature extraction network is built on the PyTorch framework. The overall structure includes a first convolutional path, a second convolutional path, a feature fusion module, and a multi-granularity attention module. These modules work together to achieve deep feature extraction from the time-frequency map of key frequency bands. The specific structural parameters are as follows: The first convolutional path is used to quickly capture global large-scale features of the signal while controlling the computational load of the network. The first convolutional layer has 64 convolutional layers. The convolutional kernel is of size [size missing], with a stride of 2, and the padding method is "same". The activation function is ReLU. This configuration can quickly extract global features while avoiding feature loss. The first pooling layer uses max pooling, with a kernel size of [size missing]. With a step size of 2, the feature map dimension is effectively reduced while retaining key features, thus reducing the amount of subsequent computation.

[0048] The second convolutional path is used to extract local fine features of the signal, complementing the first convolutional path and improving the comprehensiveness of feature extraction. The second convolutional layer consists of 128 layers. The first convolutional layer uses a kernel of a certain size, a stride of 2, and the same padding method. The activation function is also ReLU, focusing on capturing subtle local fault features. The second pooling layer uses the same parameter configuration as the first pooling layer: max pooling and pooling kernel. A step size of 2 is used to ensure the consistency of the feature map dimensions.

[0049] The first feature map output by the first convolutional path has a size of 56×56×64, and the second feature map output by the second convolutional path has a size of 56×56×128. The feature fusion module first concatenates the two feature maps along the channel dimension to obtain a concatenated feature map with a size of 56×56×192. Then, a 1×1 convolutional layer is used to process the concatenated feature map. This 1×1 convolutional layer has 128 convolutional kernels and the activation function is ReLU to achieve the fusion of channel information and dimensionality compression, and finally generates an initial fused feature with a size of 56×56×128.

[0050] Multi-granularity attention module: This module sequentially performs channel attention weighting and spatial attention weighting operations on the initial fused features, outputting the final fused feature vector. This module's design allows the network to simultaneously focus on important feature channels and important spatial regions, strengthening the representation of fault features, suppressing interference from irrelevant noise, and improving the targeting and effectiveness of feature extraction.

[0051] Furthermore, channel attention weighting performs global average pooling and global max pooling on each channel of the initial fused feature map, concatenates the two feature vectors obtained, inputs them into a multilayer perceptron, and outputs a channel weight vector; spatial attention weighting performs average pooling and max pooling on the channel-weighted feature map in the channel dimension, concatenates the two feature maps, processes them through a convolutional layer, and outputs a spatial weight matrix.

[0052] In a specific embodiment, channel attention weighting is implemented: for each channel of the initial fused feature map output by the feature fusion module, global average pooling and global max pooling operations are performed respectively, resulting in two feature vectors of size . Global average pooling captures the overall response of each channel, while global max pooling captures the strongest response of each channel. The combination of these two methods comprehensively evaluates the importance of each channel, avoiding feature omissions caused by a single pooling method. The two feature vectors are concatenated along the channel dimension and then input into two fully connected layers for processing: the first fully connected layer has 32 neurons, which is 1 / 4 of the initial fused feature channel count of 128, and uses the ReLU activation function; the second fully connected layer has 128 neurons and uses the Sigmoid activation function, ultimately outputting a channel weight vector of size . This channel weight vector is multiplied channel-by-channel with the original initial fused feature map to achieve adaptive enhancement of the channel dimension, strengthening the response of important feature channels and suppressing interference from irrelevant channels.

[0053] Spatial attention weighting: For the channel-weighted feature map, average pooling and max pooling operations are performed along the channel dimension to obtain two features of different sizes. The feature maps are concatenated to form a feature map of size 1. The feature map is constructed using average pooling to aggregate spatial context information and max pooling to highlight salient regions in the space. The combination of these two methods comprehensively evaluates the importance of each spatial location in the feature map. Then, through a... The convolutional layer processes the concatenated feature maps. This convolutional layer uses the same padding method, the sigmoid activation function, and has an output size of [size missing]. The spatial weight matrix is ​​then multiplied pixel-by-pixel with the channel-weighted feature map to enhance the feature response at the location of the fault impact, suppress interference from background noise regions, and further optimize the feature representation.

[0054] All computations in the attention module are implemented using tensor operations in the PyTorch framework. The order of channel attention weighting and spatial attention weighting has been verified through multiple engineering studies, which can effectively optimize feature representation and improve the accuracy of fault classification.

[0055] Furthermore, the classifier employs an adaptive learning rate adjustment strategy during the model training phase. This strategy dynamically adjusts the learning rate based on the relative rate of change of the loss value between adjacent training cycles. When the relative rate of change is lower than a preset threshold, the learning rate is reduced by a preset decay coefficient.

[0056] In a specific embodiment, the classifier structure consists of two fully connected layers, used to transform the fused feature vector output by the multi-granularity attention module into a fault category probability distribution. The first fully connected layer has an input dimension of 1024 and an output dimension of 256, with ReLU activation. The second fully connected layer has an input dimension of 256, an output dimension equal to the number of fault categories in the industrial equipment, and Softmax activation. The final output is the probability distribution of each fault category, with the category having the highest probability value being the initial fault diagnosis result.

[0057] Model training parameters: The model is trained using the Adam optimizer, and the optimizer parameters are set as follows: Set to 0.9, Set to 0.999. Set as The cross-entropy loss function is selected to measure the difference between the probability distribution of the model output and the real fault category, guiding model optimization; the initial learning rate is set to 0.001 to ensure the convergence speed of the initial training of the model.

[0058] Adaptive learning rate adjustment strategy: A training cycle is defined as a complete traversal of the training dataset. The formula for calculating the relative rate of change of the loss value between two adjacent training cycles is: ,in Let be the loss value for the t-th training period. This represents the loss value in the (t-1)th training epoch. The preset threshold for the relative rate of change of loss is 0.01. When the calculated relative rate of change is lower than this threshold, it indicates that the model has entered a convergence plateau. At this point, the current learning rate is multiplied by a decay coefficient of 0.9 to reduce the learning rate and achieve fine-tuning of the model. Simultaneously, a lower limit for the learning rate is set. This avoids training stagnation caused by excessively low learning rates. This adaptive adjustment strategy ensures stable training while maintaining model convergence speed, thereby improving the model's generalization ability and diagnostic accuracy.

[0059] Furthermore, such as Figure 7 As shown, in step S5, the comprehensive confidence score is calculated and the final fault diagnosis category is output, including the following steps: Step S51, obtain the confidence score of the initial fault diagnosis result output by the classifier; Step S52: Retrieve several historical diagnostic records corresponding to the current electromechanical equipment from the historical diagnostic database, and calculate the consistency index between the current diagnostic category and the category in the historical diagnostic records. Step S53: Calculate the overall credibility score based on the confidence score and the consistency index; ; in, To determine the overall credibility score, The maximum probability value in the initial fault diagnosis results output by the classifier. As a consistency indicator, These are preset weighting coefficients; Step S54, when the overall credibility score is... If the result exceeds the preset confidence threshold, the initial fault diagnosis result is output as the final fault diagnosis category.

[0060] In a specific embodiment, the confidence score is the maximum value of the Softmax probability output by the classifier, directly reflecting the model's confidence in the initial diagnostic results. The historical diagnostic database is built using MySQL, with data tables indexed by fields such as equipment number, speed, load, diagnostic timestamp, and diagnostic category for easy and fast retrieval. During retrieval, the current equipment number is used as the core keyword to filter the 10 most recent historical records matching the current operating parameters. The operating condition matching criteria are speed deviation ±5% and load deviation ±10%, which conform to the conventional requirements for judging the operating conditions of industrial equipment.

[0061] The consistency index C is calculated by dividing the number of identical categories in the current diagnostic category and the historical records by the total number of historical records retrieved. Weighting coefficient. The confidence score was optimized to 0.6 through validation set experiments, emphasizing both the core role of the model output and the reference value of historical data. The confidence threshold was set to 0.85, a value determined based on the actual requirements for false positive rates in industrial scenarios. When the overall confidence score... If the value is higher than 0.85, the initial diagnostic result is output directly; if it is lower than the threshold, the industrial control interface will prompt that further signal acquisition or manual verification is required. The construction, retrieval, and calculation of the database are all based on existing database technology, and the entire process conforms to the actual operational logic of industrial applications.

[0062] The specific embodiments of the present invention have been described in detail above, but they are merely examples, and the present invention is not equivalent to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.

Claims

1. A method for rapid diagnosis of mechanical and electrical faults based on adaptive multi-scale fusion, characterized in that, Includes the following steps: Step S1: Obtain the original vibration signal of the electromechanical equipment, perform quality inspection and adaptive time-frequency preprocessing on the original vibration signal, and generate a standardized two-dimensional time-frequency diagram; Step S2: Perform data augmentation on the standardized two-dimensional time-frequency graph to generate an enhanced time-frequency graph; Step S3: The enhanced time-frequency image is input into a deep feature extraction network, which extracts and fuses multi-scale features to generate a fused feature vector. Step S4: Input the fused feature vector into the classifier to obtain the initial fault diagnosis result; Step S5: Calculate the comprehensive confidence score based on the confidence score of the initial fault diagnosis result and the historical diagnosis records, and output the final fault diagnosis category based on the comprehensive confidence score.

2. The method according to claim 1, wherein, In step S1, the quality inspection and adaptive time-frequency preprocessing are implemented in the following ways: Step S11: Perform a short-time Fourier transform on the original vibration signal sequence to obtain the initial time spectrum; ; wherein, is the initial time-frequency spectrum, is the original vibration signal sequence over time, is time, is a window function, is frequency, is the center point of the time window. Step S12: Calculate the energy distribution of the initial spectrum in the frequency dimension; ; wherein is the total energy at frequency ; Step S13, depending on the energy distribution selecting a key band of energy concentration; Step S14: Construct a set of bandpass filters based on the boundary frequencies of the key frequency bands; Step S15: The original vibration signal sequence is filtered using a bandpass filter bank, and the filtered signal is subjected to short-time Fourier transform and size standardization to generate a standardized two-dimensional time-frequency diagram.

3. The method according to claim 2, wherein, In step S13, based on the energy distribution Selecting key frequency bands includes the following steps: Step S131, calculate energy distribution The global average value is used to mark frequency bands whose energy values ​​are continuously higher than the global average value as candidate frequency bands; Step S132: Obtain the physical parameters and operating speed of the electromechanical equipment, and calculate the theoretical fault characteristic frequency related to the potential fault type; ; wherein, is a theoretical characteristic frequency corresponding to a specific failure mode, is a base frequency corresponding to a running speed of the device, is a constant coefficient related to physical parameters of the device and the failure type; Step S133: Using the theoretical fault characteristic frequency as a reference, and combining the fault impact characteristics of the candidate frequency bands, the key frequency bands are selected from the candidate frequency bands.

4. The method according to claim 3, wherein, In step S133, the fault impact characteristics are obtained by constructing a geometric quantile probability weighted kurtosis feature index, specifically including the following steps: Step S1331, compute a frequency-weighted differential envelope vector for the signal sequence within each candidate band and compute its corresponding autocorrelation-normalized energy sequence based on the vector ; Step S1332: Based on the inverse function of the cumulative probability distribution, perform a geometric quantile mapping operation with preset coefficients on the above vectors respectively to obtain the mapped denoised envelope sequence and the mapped autocorrelation sequence. Step S1333: Using the unbiased estimation method of probability weight statistical moments, calculate the linear skewness kurtosis values ​​of the mapped denoised envelope sequence and the mapped autocorrelation sequence respectively. Step S1334: Construct the geometric quantile probability weighted kurtosis feature index based on the two obtained linear skewed kurtosis values. This is used as an indicator of fault impact energy. Step S1335, the energy distribution value of the candidate frequency band with the geometric quantile probability weight kurtosis feature index weighted fusion is performed to obtain a comprehensive evaluation value : ; in, , For the preset weighting coefficients, and Select comprehensive evaluation value The highest candidate frequency band is selected as the key frequency band.

5. The method according to claim 4, wherein, In step S2, the time warp is achieved by performing a nonlinear time-scale transformation on the original vibration signal sequence. The nonlinear time-scale transformation is controlled by a randomly generated scaling factor, which is uniformly and randomly selected within a preset range. After the transformation, the signal length is restored to the original length by linear interpolation. The spectral masking is achieved by randomly masking a continuous rectangular region on a standardized two-dimensional time-frequency map. The size of the rectangular region does not exceed a preset ratio of the corresponding dimension of the time-frequency map. The coordinates of the upper left corner of the rectangular region are randomly determined within the range of the time-frequency map. All pixel values ​​within the masked region are set to zero.

6. The method according to claim 5, wherein, In step S3, the deep feature extraction network includes a first convolutional path, a second convolutional path, a feature fusion module, and a multi-granularity attention module; the first convolutional path includes a first convolutional layer and a first pooling layer, the second convolutional path includes a second convolutional layer and a second pooling layer, and the kernel size of the first convolutional layer is larger than the kernel size of the second convolutional layer; the feature fusion module concatenates the first feature map output by the first convolutional path and the second feature map output by the second convolutional path in the channel dimension, and processes them through a convolutional layer with a kernel size of 1×1 to generate initial fused features; The multi-granularity attention module sequentially performs channel attention weighting and spatial attention weighting on the initial fused features, and outputs a fused feature vector.

7. The method according to claim 6, wherein, The channel attention weighting method performs global average pooling and global max pooling on each channel of the initial fused feature map, concatenates the two feature vectors, inputs them into a multilayer perceptron, and outputs a channel weight vector. The spatial attention weighting method performs average pooling and max pooling on the channel-weighted feature map in the channel dimension, concatenates the two feature maps, processes them through a convolutional layer, and outputs a spatial weight matrix.

8. The method according to claim 7, wherein, The classifier employs an adaptive learning rate adjustment strategy during the model training phase. This strategy dynamically adjusts the learning rate based on the relative rate of change of the loss value in adjacent training cycles. When the relative rate of change is lower than a preset threshold, the learning rate is reduced by a preset decay coefficient.

9. The method according to claim 8, wherein, In step S5, the comprehensive confidence score is calculated and the final fault diagnosis category is output, including the following steps: Step S51: Obtain the confidence score of the initial fault diagnosis result output by the classifier; Step S52: Retrieve several historical diagnostic records corresponding to the current electromechanical equipment from the historical diagnostic database, and calculate the consistency index between the current diagnostic category and the category in the historical diagnostic records. Step S53: Calculate the overall credibility score based on the confidence score and the consistency index; ; wherein, is a comprehensive credibility score, is a maximum probability value in an initial fault diagnosis result output by the classifier, is a consistency index, is a preset weight coefficient; Step S54, when the overall credibility score is... If the result exceeds the preset confidence threshold, the initial fault diagnosis result is output as the final fault diagnosis category.