Gearbox Fault Detection Method, Device, and Equipment Based on Neural Network Model

By combining wavelet transform and diffusion model data enhancement with residual network feature extraction, the problems of data imbalance and feature extraction difficulties in gearbox fault diagnosis are solved, thereby improving the accuracy and applicability of fault identification.

CN120632642BActive Publication Date: 2025-10-31ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511115807.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-31
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing gearbox fault diagnosis methods struggle to achieve high-precision fault identification and classification when faced with weak fault signals, complex fault modes, or complex noise interference. Furthermore, traditional methods are prone to overfitting or loss of key information.

Method used

This paper adopts a technical approach that combines wavelet transform, diffusion model data augmentation, and residual network feature extraction. The vibration signal is converted into a time-frequency image through wavelet transform and then subjected to grayscale processing and normalization. High-quality synthetic fault samples are generated using the diffusion model, and a residual network model is constructed for training to extract multi-scale features.

Benefits of technology

It significantly improves the accuracy and engineering applicability of gearbox fault diagnosis, enhances the ability to identify weak fault characteristics and complex fault modes, and improves the model's generalization ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632642B_ABST
    Figure CN120632642B_ABST
Patent Text Reader

Abstract

This disclosure relates to the technical field of fault detection based on neural networks, and discloses a method, apparatus, and device for gearbox fault detection based on a neural network model. The method includes: performing wavelet transform on sample vibration signals to generate time-frequency images; obtaining normal and original fault samples through grayscale conversion and normalization; enhancing the original fault samples using a diffusion model to generate synthetic samples; calculating distribution differences based on a multi-Gaussian kernel function and extracting deep features; evaluating the quality of the synthetic samples; if qualified, using them for training; otherwise, regenerating them; constructing a training dataset containing normal, original, and synthetic samples; training an initial residual network to obtain the target model; and inputting the target gearbox vibration signal into the model after performing the same preprocessing to achieve fault type detection. This method solves the problems of data imbalance, difficulty in feature extraction, and insufficient diagnostic accuracy in gearbox fault diagnosis, improving the fault identification accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of fault detection technology based on neural networks, and for example to a method, apparatus and equipment for gearbox fault detection based on a neural network model. Background Technology

[0002] As a core component of mechanical transmission systems, gearboxes are widely used in critical fields such as wind power generation, rail transportation, and industrial manufacturing. Their operating status directly affects the reliability and safety of the entire system. However, due to prolonged exposure to complex operating environments such as high speed, heavy load, and variable working conditions, gearboxes are prone to failure due to gear wear, shaft misalignment, and bearing damage. Failure to detect and effectively diagnose these issues in a timely manner can lead to a chain reaction of equipment damage and even serious safety accidents.

[0003] Traditional gearbox fault diagnosis methods primarily rely on shallow signal analysis techniques such as threshold judgment, empirical mode decomposition, and short-time Fourier transform. While these methods are simple to implement, they have significant limitations when dealing with weak fault signals, complex fault modes, or complex noise interference, making it difficult to achieve high-precision fault identification and classification. In practical engineering applications, the difficulty in collecting fault samples leads to a significant imbalance in the fault diagnosis dataset, with normal samples far outnumbering fault samples. Against this backdrop, traditional diagnostic models are prone to overfitting, resulting in a decreased ability to identify fault categories, leading to missed or false diagnoses and severely impacting the reliability of the diagnostic system.

[0004] To alleviate the problem of data imbalance, existing technologies often employ oversampling or undersampling methods for data augmentation. Oversampling generates new fault samples by interpolating in the feature space, but the generated samples are of low quality, easily introducing noise or unreasonable sample distributions, exacerbating model overfitting. Undersampling achieves data balance by reducing the number of majority class samples, but may lead to the loss of key fault information, affecting the model's generalization ability. Furthermore, with the development of deep learning technology, fault diagnosis methods based on convolutional neural networks have shown certain advantages in automatic feature extraction. However, traditional network structures are sensitive to noise interference in vibration signals and struggle to effectively extract multi-scale temporal features, resulting in the submergence of weak fault features and insufficient ability to distinguish complex fault modes. Therefore, there is an urgent need to propose a gearbox fault diagnosis method that balances data balance optimization with robust feature extraction capabilities to improve the diagnostic accuracy and generalizability of the model under complex operating conditions. Summary of the Invention

[0005] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0006] This disclosure provides a gearbox fault detection method, apparatus, and device based on a neural network model. By combining wavelet transform, diffusion model data enhancement, and residual network feature extraction, it effectively solves the problems of data imbalance, difficulty in feature extraction, and insufficient diagnostic accuracy in gearbox fault diagnosis, and significantly improves the fault identification accuracy and engineering applicability of the model.

[0007] According to a first aspect of this disclosure, a gearbox fault detection method based on a neural network model is provided, comprising:

[0008] The vibration signal of the sample gearbox during operation is preprocessed. The vibration signal is converted into a time-frequency image by wavelet transform, and then grayscale processing and normalization are performed to obtain normal samples under normal conditions and original fault samples under fault conditions.

[0009] A diffusion model is used to augment a preset number of original fault samples, and synthetic fault samples are generated through forward noise addition and reverse noise reduction mechanisms.

[0010] By calculating the distribution difference between the original fault samples and the corresponding synthetic fault samples, and extracting the deep features of each original fault sample and synthetic fault sample, the overall mean difference value and the overall silhouette coefficient are obtained respectively, and the quality of the synthetic fault samples is determined based on the overall mean difference value and the overall silhouette coefficient.

[0011] Synthetic fault samples are permitted to be used when their quality is deemed acceptable.

[0012] When it is determined that the quality of the synthesized faulty sample is unqualified, the random sampling strategy of the diffusion model is adjusted, and the synthesized faulty sample is regenerated using the diffusion model.

[0013] A training dataset is constructed based on normal samples, original fault samples and synthetic fault samples. An initial residual network model is trained based on the training dataset to obtain a well-trained target residual network model.

[0014] The vibration signal during the operation of the target gearbox is preprocessed. The vibration signal is converted into a time-frequency image by wavelet transform, and after grayscale processing and normalization, it is input into the target residual network model. The fault type of the target gearbox is detected based on the target residual network model.

[0015] In some embodiments, the vibration signal of the sample gearbox during operation is preprocessed by converting the vibration signal into a time-frequency image using wavelet transform, followed by grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions, including:

[0016] During the operation of the sample gearbox, vibration signals under normal conditions and vibration signals under various fault conditions were collected.

[0017] The vibration signal is segmented into segments with a fixed window length, and a continuous wavelet transform is applied to each segment to convert it into a time-frequency image.

[0018] Each time-frequency image is processed into a single-channel grayscale image after grayscale processing. Each single-channel grayscale image is then normalized to eliminate dimensional differences.

[0019] The single-channel grayscale image corresponding to the vibration signal under normal conditions is used as the normal sample, and the single-channel grayscale image corresponding to the vibration signal under fault conditions is used as the original fault sample.

[0020] In some embodiments, a diffusion model is used to augment a preset number of original fault samples, and synthetic fault samples are generated through a forward noise addition and backward noise reduction mechanism, including:

[0021] The original input fault samples are subjected to forward noise addition using a diffusion model, and Gaussian noise is iteratively introduced to obtain noisy fault samples.

[0022] The diffusion model is used to perform reverse denoising on the noisy fault samples, iteratively removing noise from the fault samples until a synthetic fault sample with the same distribution as the original fault sample is obtained.

[0023] In some embodiments, the process of introducing Gaussian noise in a single iteration can be expressed by the following formula: , This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the cumulative noise proportionality factor. Indicates the original fault sample. Indicates Gaussian noise distribution;

[0024] It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed.

[0025] In some embodiments, the process of removing noise from faulty samples in a single iteration can be expressed by the following formula: ;

[0026] Indicates the ()th step in the reverse denoising operation Synthetic fault samples obtained through iterative steps This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the noise in the prediction. This represents the cumulative noise proportionality factor. Indicates standard Gaussian noise. To represent the first step in the reverse denoising operation The predefined variance of each iteration;

[0027] It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed;

[0028] It is calculated using the following formula: , It is a hyperparameter. .

[0029] In some embodiments, by calculating the distribution difference between the original fault samples and the corresponding synthetic fault samples, and extracting the deep features of each original fault sample and synthetic fault sample, the overall mean difference value and the overall silhouette coefficient are obtained respectively, and the quality of the synthetic fault samples is determined based on the overall mean difference value and the overall silhouette coefficient, including:

[0030] The distribution difference between the original fault sample and the corresponding synthetic fault sample is calculated based on multiple different Gaussian kernel functions to obtain multiple maximum mean difference values.

[0031] The average of the multiple largest mean differences is taken as the overall mean difference.

[0032] The deep features of each original fault sample and the synthetic fault sample are extracted using a classifier network, and the extracted features are merged into fusion features.

[0033] Cluster analysis is performed on the feature space after dimensionality reduction of the fused features to calculate the contour coefficient of each synthetic fault sample;

[0034] The average profile coefficient of the synthesized fault samples is used as the overall profile coefficient.

[0035] The quality of the synthesized fault samples is determined based on the overall mean difference and the overall profile coefficient.

[0036] In some embodiments, the maximum mean difference value based on a kernel function can be obtained by the following formula:

[0037] ,

[0038] This represents the square of the largest mean difference. Indicates the first One original fault sample, Indicates the use of the first Pair one original fault sample with another original fault sample to compute the kernel function. Indicates the first A synthetic fault sample, Indicates the use of the first For each synthetic fault sample, a kernel function is computed on a corresponding synthetic fault sample. Indicates the number of original fault samples. Indicates the number of synthesized fault samples. Represents the kernel function;

[0039] The contour coefficients of the synthesized fault samples can be obtained using the following formula: , Indicates the first The profile coefficients of a synthetic fault sample. Indicates the first The average distance between a synthetic faulty sample and other synthetic faulty samples in the same cluster Indicates the first The average distance between a synthetic faulty sample and all synthetic faulty samples in the nearest other clusters.

[0040] In some embodiments, the vibration signal during the operation of the target gearbox is preprocessed, converted into a time-frequency image by wavelet transform, and then input into the target residual network model after grayscale processing and normalization. The fault type of the target gearbox is detected based on the target residual network model, including:

[0041] Vibration signals were collected during the operation of the target gearbox;

[0042] The vibration signal is segmented into segments with a fixed window length, and a continuous wavelet transform is applied to each segment to convert it into a time-frequency image.

[0043] Each time-frequency image is processed into a single-channel grayscale image after grayscale processing. Each single-channel grayscale image is then normalized to eliminate dimensional differences.

[0044] Each single-channel grayscale image is input into the target residual network model, and the fault type of the target gearbox is detected based on the target residual network model.

[0045] According to a second aspect of this disclosure, a gearbox fault detection device based on a neural network model is provided, comprising:

[0046] The sample acquisition module is configured to: preprocess the vibration signal of the sample gearbox during operation, convert the vibration signal into a time-frequency image through wavelet transform, and perform grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions.

[0047] The sample generation module is configured to: augment a preset number of original fault samples using a diffusion model; generate synthetic fault samples through forward noise addition and reverse denoising mechanisms; obtain the overall mean difference value and overall profile coefficient by calculating the distribution difference between the original fault samples and the corresponding synthetic fault samples, and extracting the deep features of each original fault sample and synthetic fault sample; determine the quality of the synthetic fault samples based on the overall mean difference value and overall profile coefficient; allow the use of synthetic fault samples when the quality of the synthetic fault samples is deemed acceptable; and adjust the random sampling strategy of the diffusion model and regenerate synthetic fault samples using the diffusion model when the quality of the synthetic fault samples is deemed unacceptable.

[0048] The model training module is configured to: construct a training dataset based on normal samples, original fault samples and synthetic fault samples; train an initial residual network model based on the training dataset; and obtain a trained target residual network model.

[0049] The fault detection module is configured to: preprocess the vibration signal during the operation of the target gearbox, convert the vibration signal into a time-frequency image through wavelet transform, and input it into the target residual network model after grayscale processing and normalization, and detect the fault type of the target gearbox based on the target residual network model.

[0050] According to a third aspect of this disclosure, an electronic device is provided, including a processor and a memory storing program instructions, the processor being configured to execute, when running the program instructions, a gearbox fault detection method based on a neural network model provided according to a first aspect of this disclosure.

[0051] The gearbox fault detection method, apparatus, and device based on a neural network model provided in this disclosure can achieve the following technical effects:

[0052] The gearbox fault detection method based on a neural network model provided in this disclosure employs a diffusion model to augment the original fault samples. Through forward noise addition and reverse denoising mechanisms, high-quality, highly diverse synthetic fault samples are generated. These synthetic fault samples are more closely distributed than the real data, effectively avoiding the introduction of noise samples while retaining key features of the original fault samples (such as abnormal patterns in time-frequency images), significantly improving sample quality. Furthermore, the number of synthetic samples is controllable, avoiding the loss of key information that may occur with undersampling methods, thus effectively alleviating the shortage of fault samples in the training dataset and improving the model's generalization ability and diagnostic stability. In addition, wavelet transform is used to convert the vibration signal into a time-frequency image, thereby capturing the local time-frequency features of the signal and effectively extracting non-stationary components in the vibration signal, such as weak fault signals and sudden impacts. Subsequently, the image undergoes grayscale processing and normalization to unify the input format and enhance feature representation capabilities. Further, a residual network model is constructed and trained, utilizing its deep structure to extract multi-scale features. The residual network alleviates the vanishing gradient problem through residual connections, supports efficient extraction and fusion of deep features, improves the model's robustness to noise interference, and enhances its ability to identify weak fault features and complex fault modes. Furthermore, by constructing a balanced training dataset containing normal samples, original fault samples, and synthetic fault samples, the model gains more learning opportunities during training, avoiding overfitting to the majority class. The trained target residual network model exhibits good generalization ability, accurately identifying the fault type of the target gearbox vibration signal after the same preprocessing in actual detection. This significantly improves the model's adaptability to different fault modes, enhances its ability to model complex features, and increases the accuracy of fault identification. The combined approach of wavelet transform, diffusion model data augmentation, and residual network feature extraction effectively solves the problems of data imbalance, difficult feature extraction, and insufficient diagnostic accuracy in gearbox fault diagnosis, significantly improving the model's fault identification accuracy and engineering applicability.

[0053] The above general description and the description below are exemplary and illustrative only and are not intended to limit this disclosure. Attached Figure Description

[0054] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:

[0055] Figure 1 This is a schematic flowchart of a gearbox fault detection method based on a neural network model provided in an embodiment of this disclosure;

[0056] Figure 2 This is a schematic flowchart of another gearbox fault detection method based on a neural network model provided in this disclosure embodiment;

[0057] Figure 3 This is a schematic flowchart of another gearbox fault detection method based on a neural network model provided in this disclosure embodiment;

[0058] Figure 4 This is a schematic diagram illustrating the process of generating synthetic fault samples using the diffusion model provided in this embodiment of the disclosure;

[0059] Figure 5 This is a schematic flowchart of another gearbox fault detection method based on a neural network model provided in this disclosure embodiment;

[0060] Figure 6 This is a schematic diagram of the structure of the residual network model provided in the embodiments of this disclosure;

[0061] Figure 7 This is a schematic diagram of the structure of a gearbox fault detection device based on a neural network model provided in an embodiment of this disclosure;

[0062] Figure 8 This is a schematic diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0063] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0064] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0065] Unless otherwise stated, the term "multiple" means two or more.

[0066] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0067] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0068] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0069] This disclosure provides a gearbox fault detection method based on a neural network model, such as... Figure 1 As shown, the gearbox fault detection method based on a neural network model includes:

[0070] S101 preprocesses the vibration signal during the operation of the sample gearbox, converts the vibration signal into a time-frequency image through wavelet transform, and performs grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions.

[0071] S102, use a diffusion model to perform data augmentation on a preset number of original fault samples, and generate synthetic fault samples through forward noise addition and reverse noise reduction mechanisms.

[0072] S103, by calculating the distribution difference between the original fault sample and the corresponding synthetic fault sample, and extracting the deep features of each original fault sample and synthetic fault sample, the overall mean difference value and the overall profile coefficient are obtained respectively, and the quality of the synthetic fault sample is determined based on the overall mean difference value and the overall profile coefficient.

[0073] S104. Synthetic fault samples are permitted to be used when the quality of the synthesized fault samples is determined to be acceptable.

[0074] S105, when it is determined that the quality of the synthesized fault sample is unqualified, adjust the random sampling strategy of the diffusion model and regenerate the synthesized fault sample using the diffusion model.

[0075] S106. A training dataset is constructed based on normal samples, original fault samples, and synthetic fault samples. An initial residual network model is trained based on the training dataset to obtain a well-trained target residual network model.

[0076] S107 preprocesses the vibration signal during the operation of the target gearbox, converts the vibration signal into a time-frequency image through wavelet transform, and inputs it into the target residual network model after grayscale processing and normalization. The fault type of the target gearbox is detected based on the target residual network model.

[0077] The gearbox fault detection method based on a neural network model provided in this disclosure employs a diffusion model to augment the original fault samples. Through forward noise addition and reverse denoising mechanisms, high-quality, highly diverse synthetic fault samples are generated. These synthetic fault samples are more closely distributed than the real data, effectively avoiding the introduction of noise samples while retaining key features of the original fault samples (such as abnormal patterns in time-frequency images), significantly improving sample quality. Furthermore, the number of synthetic samples is controllable, avoiding the loss of key information that may occur with undersampling methods, thus effectively alleviating the shortage of fault samples in the training dataset and improving the model's generalization ability and diagnostic stability. In addition, wavelet transform is used to convert the vibration signal into a time-frequency image, thereby capturing the local time-frequency features of the signal and effectively extracting non-stationary components in the vibration signal, such as weak fault signals and sudden impacts. Subsequently, the image undergoes grayscale processing and normalization to unify the input format and enhance feature representation capabilities. Further, a residual network model is constructed and trained, utilizing its deep structure to extract multi-scale features. The residual network alleviates the vanishing gradient problem through residual connections, supports efficient extraction and fusion of deep features, improves the model's robustness to noise interference, and enhances its ability to identify weak fault features and complex fault modes. Furthermore, by constructing a balanced training dataset containing normal samples, original fault samples, and synthetic fault samples, the model gains more learning opportunities during training, avoiding overfitting to the majority class. The trained target residual network model exhibits good generalization ability, accurately identifying the fault type of the target gearbox vibration signal after the same preprocessing in actual detection. This significantly improves the model's adaptability to different fault modes, enhances its ability to model complex features, and increases the accuracy of fault identification. The combined approach of wavelet transform, diffusion model data augmentation, and residual network feature extraction effectively solves the problems of data imbalance, difficult feature extraction, and insufficient diagnostic accuracy in gearbox fault diagnosis, significantly improving the model's fault identification accuracy and engineering applicability.

[0078] In some embodiments, the vibration signals during the operation of the sample gearbox are preprocessed by converting the vibration signals into time-frequency images using wavelet transform, followed by grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions. This includes: collecting vibration signals under normal conditions and vibration signals under various fault conditions during the operation of the sample gearbox; segmenting the vibration signals into segments with a fixed window length, and applying continuous wavelet transform to each segment to convert it into a time-frequency image; converting each time-frequency image into a single-channel grayscale image after grayscale processing, and normalizing each single-channel grayscale image to eliminate dimensional differences; using the single-channel grayscale image corresponding to the vibration signal under normal conditions as the normal sample, and the single-channel grayscale image corresponding to the vibration signal under fault conditions as the original fault sample.

[0079] This disclosure provides another gearbox fault detection method based on a neural network model, such as... Figure 2 As shown, the gearbox fault detection method based on a neural network model includes:

[0080] S201 collects vibration signals under normal conditions and vibration signals under various fault conditions during the operation of the sample gearbox.

[0081] Here, the fault conditions include inner ring wear, outer ring wear, rolling element wear, etc.

[0082] S202, the vibration signal is segmented into segments with a fixed window length, and a continuous wavelet transform is applied to each segment of the vibration signal to convert it into a time-frequency image.

[0083] S203 converts each time-frequency image into a single-channel grayscale image after grayscale processing, and performs normalization processing on each single-channel grayscale image to eliminate dimensional differences.

[0084] S204, take the single-channel grayscale image corresponding to the vibration signal under normal conditions as the normal sample, and take the single-channel grayscale image corresponding to the vibration signal under fault conditions as the original fault sample.

[0085] S205 uses a diffusion model to augment the data of a preset number of original fault samples, and generates synthetic fault samples through forward noise addition and reverse noise reduction mechanisms.

[0086] S206. By calculating the distribution difference between the original fault sample and the corresponding synthetic fault sample, and extracting the deep features of each original fault sample and synthetic fault sample, the overall mean difference value and the overall profile coefficient are obtained respectively, and the quality of the synthetic fault sample is determined based on the overall mean difference value and the overall profile coefficient.

[0087] S207, Synthetic fault samples are permitted to be used when the quality of the synthesized fault samples is determined to be acceptable.

[0088] S208, when it is determined that the quality of the synthesized fault sample is unqualified, adjust the random sampling strategy of the diffusion model and regenerate the synthesized fault sample using the diffusion model.

[0089] S209: Construct a training dataset based on normal samples, original fault samples, and synthetic fault samples; train an initial residual network model based on the training dataset; and obtain a well-trained target residual network model.

[0090] S210 preprocesses the vibration signal during the operation of the target gearbox, converts the vibration signal into a time-frequency image through wavelet transform, and inputs it into the target residual network model after grayscale processing and normalization. The fault type of the target gearbox is detected based on the target residual network model.

[0091] In some embodiments, a diffusion model is used to perform data augmentation on a preset number of original fault samples, and a synthetic fault sample is generated through a forward noise addition and reverse noise reduction mechanism. This includes: using a diffusion model to perform a forward noise addition operation on the input original fault samples, iteratively introducing Gaussian noise to obtain noisy fault samples; using a diffusion model to perform a reverse noise reduction operation on the noisy fault samples, iteratively removing noise from the fault samples until a synthetic fault sample with the same distribution as the original fault sample is obtained.

[0092] This disclosure provides another gearbox fault detection method based on a neural network model, such as... Figure 3 and Figure 4 As shown, the gearbox fault detection method based on a neural network model includes:

[0093] S301 preprocesses the vibration signal during the operation of the sample gearbox, converts the vibration signal into a time-frequency image through wavelet transform, and performs grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions.

[0094] S302 uses a diffusion model to perform forward noise addition on the input original fault samples, iteratively introducing Gaussian noise to obtain noisy fault samples.

[0095] S303 uses a diffusion model to perform reverse denoising on the noisy fault samples, iteratively removing noise from the fault samples until a synthetic fault sample with the same distribution as the original fault sample is obtained.

[0096] S304. By calculating the distribution difference between the original fault sample and the corresponding synthetic fault sample, and extracting the deep features of each original fault sample and the synthetic fault sample, the overall mean difference value and the overall profile coefficient are obtained respectively, and the quality of the synthetic fault sample is determined based on the overall mean difference value and the overall profile coefficient.

[0097] S305, when it is determined that the quality of the synthesized fault sample is acceptable, the synthesized fault sample may be used.

[0098] S306, when it is determined that the quality of the synthesized fault sample is unqualified, adjust the random sampling strategy of the diffusion model and regenerate the synthesized fault sample using the diffusion model.

[0099] S307: Construct a training dataset based on normal samples, original fault samples, and synthetic fault samples. Train an initial residual network model based on the training dataset to obtain a well-trained target residual network model.

[0100] S308 preprocesses the vibration signal during the operation of the target gearbox, converts the vibration signal into a time-frequency image through wavelet transform, and inputs it into the target residual network model after grayscale processing and normalization. The fault type of the target gearbox is detected based on the target residual network model.

[0101] In this embodiment of the disclosure, the process of introducing Gaussian noise in one iteration can be expressed by the following formula: , This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the cumulative noise proportionality factor. Indicates the original fault sample. This represents Gaussian noise. Figure 4 In This represents the final noise fault sample.

[0102] In this embodiment of the disclosure, It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed.

[0103] In this embodiment of the disclosure, the process of removing noise from faulty samples in one iteration can be expressed by the following formula: . Indicates the ()th step in the reverse denoising operation Synthetic fault samples obtained through iterative steps This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the noise in the prediction. This represents the cumulative noise proportionality factor. Indicates standard Gaussian noise. To represent the first step in the reverse denoising operation The predefined variance of each iteration;

[0104] In this embodiment of the disclosure, It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed;

[0105] In this embodiment of the disclosure, It is calculated using the following formula: , It is a hyperparameter. .

[0106] In this embodiment, the core of the diffusion model is a noise prediction network based on U-Net. U-Net, a classic convolutional neural network architecture, comprises three core parts: an encoder (contraction path) responsible for feature extraction, intermediate layers processing high-level semantics, and a decoder (expansion path) for feature reconstruction. This architecture introduces a cross-layer connection mechanism to fuse the feature maps extracted at each stage of the encoder with the corresponding layers of the decoder. This unique structural design effectively integrates feature information at different scales, thereby capturing both local details and global semantic information during image generation. Its loss function is:

[0107] ,

[0108] in, This indicates the noise that the diffusion model needs to estimate and train. This represents the noise in the prediction. In the forward process, in step The image obtained after adding noise. The mathematical expression of the diffusion model clearly reveals its core mechanism: the model achieves image reconstruction by progressively learning its ability to predict the distribution of noise in the image. Specifically, its training process focuses on accurately estimating the noise superimposed in each denoising step. By gradually optimizing the difference between the predicted noise and the actual noise, the system ultimately possesses the ability to progressively restore a structured image from a random noise distribution. This optimization paradigm based on error backpropagation essentially constructs a multi-scale feature learning system from coarse to fine.

[0109] In some embodiments, by calculating the distribution difference between the original fault sample and the corresponding synthetic fault sample, and extracting the deep features of each original fault sample and synthetic fault sample, an overall mean difference value and an overall silhouette coefficient are obtained, and the quality of the synthetic fault sample is determined based on the overall mean difference value and the overall silhouette coefficient. This includes: calculating the distribution difference between the original fault sample and the corresponding synthetic fault sample based on multiple different Gaussian kernel functions to obtain multiple maximum mean difference values; using the average of the multiple maximum mean difference values ​​as the overall mean difference value; extracting the deep features of each original fault sample and synthetic fault sample using a classifier network, and merging the extracted features into a fusion feature; performing cluster analysis on the feature space after dimensionality reduction of the fusion feature to calculate the silhouette coefficient of each synthetic fault sample; using the average of the silhouette coefficients of the synthetic fault samples as the overall silhouette coefficient; and determining the quality of the synthetic fault sample based on the overall mean difference value and the overall silhouette coefficient.

[0110] Optionally, if the overall mean difference is low and the overall silhouette coefficient is high, it indicates that the synthesized fault samples simultaneously satisfy both realism and diversity, thus determining that the quality of the synthesized fault samples is acceptable. If the overall mean difference is low and the overall silhouette coefficient is low, it may indicate that the synthesized fault samples overfit the true distribution and lack diversity, thus determining that the quality of the synthesized fault samples is unacceptable. In this case, the random sampling strategy of the diffusion model should be adjusted, and the synthesized fault samples should be regenerated using the diffusion model. Adjusting the random sampling strategy of the diffusion model can involve increasing the number of diffusion iterations.

[0111] In some embodiments, the maximum mean difference value based on a kernel function can be obtained by the following formula:

[0112] ,

[0113] This represents the square of the largest mean difference. Indicates the first One original fault sample, Indicates the use of the first Pair one original fault sample with another original fault sample to compute the kernel function. Indicates the first A synthetic fault sample, Indicates the use of the first For each synthetic fault sample, a kernel function is computed on a corresponding synthetic fault sample. Indicates the number of original fault samples. Indicates the number of synthesized fault samples. Represents the kernel function;

[0114] The contour coefficients of the synthesized fault samples can be obtained using the following formula: , Indicates the first The profile coefficients of a synthetic fault sample. Indicates the first The average distance between a synthetic faulty sample and other synthetic faulty samples in the same cluster Indicates the first The average distance between a synthetic faulty sample and all synthetic faulty samples in the nearest other clusters.

[0115] In some embodiments, the vibration signal during the operation of the target gearbox is preprocessed, and the vibration signal is converted into a time-frequency image by wavelet transform. After grayscale processing and normalization, the image is input into the target residual network model. The fault type of the target gearbox is detected based on the target residual network model. This includes: acquiring vibration signals during the operation of the target gearbox; segmenting the vibration signal into segments according to a fixed window length, and applying continuous wavelet transform to each segment to convert it into a time-frequency image; converting each time-frequency image into a single-channel grayscale image after grayscale processing, and normalizing each single-channel grayscale image to eliminate dimensional differences; inputting each single-channel grayscale image into the target residual network model, and detecting the fault type of the target gearbox based on the target residual network model.

[0116] This disclosure provides another gearbox fault detection method based on a neural network model, such as... Figure 5 As shown, the gearbox fault detection method based on a neural network model includes:

[0117] S501 preprocesses the vibration signal during the operation of the sample gearbox, converts the vibration signal into a time-frequency image through wavelet transform, and performs grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions.

[0118] S502 uses a diffusion model to augment the data of a preset number of original fault samples and generates synthetic fault samples through forward noise addition and reverse noise reduction mechanisms.

[0119] S503, by calculating the distribution difference between the original fault sample and the corresponding synthetic fault sample, and extracting the deep features of each original fault sample and synthetic fault sample, the overall mean difference value and the overall profile coefficient are obtained respectively, and the quality of the synthetic fault sample is determined based on the overall mean difference value and the overall profile coefficient.

[0120] S504, when it is determined that the quality of the synthesized fault sample is acceptable, the synthesized fault sample may be used.

[0121] S505, when it is determined that the quality of the synthesized fault sample is unqualified, adjust the random sampling strategy of the diffusion model and regenerate the synthesized fault sample using the diffusion model.

[0122] S506: Construct a training dataset based on normal samples, original fault samples, and synthetic fault samples; train an initial residual network model based on the training dataset; and obtain a well-trained target residual network model.

[0123] S507 collects vibration signals during the operation of the target gearbox.

[0124] S508 divides the vibration signal into segments of a fixed window length and applies a continuous wavelet transform to each segment to convert it into a time-frequency image.

[0125] S509 converts each time-frequency image into a single-channel grayscale image after grayscale processing, and performs normalization processing on each single-channel grayscale image to eliminate dimensional differences.

[0126] S510 inputs each single-channel grayscale image into the target residual network model, and detects the fault type of the target gearbox based on the target residual network model.

[0127] In this embodiment of the disclosure, combined with Figure 6 As shown, the residual network model includes convolutional kernels, a vision transformer (ViT), a batch normalization (BN) layer, a max pooling layer, an attention mechanism layer, and a global average pooling (GAP) layer. The global average pooling layer outputs categories C1, C2, ..., Cn. The attention mechanism layer includes a channel attention module that receives input features and a spatial attention module that outputs refined features.

[0128] When training the initial residual network model, samples from the training dataset are input into the model. The input tensor is first cross-correlationd using convolutional kernels to generate the output tensor. Specifically, in a two-dimensional convolution operation, the convolution window slides from the top left corner of the input tensor. Each time the window slides, the tensor within the window is multiplied by the corresponding element of the convolution kernel, and the sum is obtained to obtain a single scalar value. The formula for calculating the output size is:

[0129] ,

[0130] in, For output size, The kernel size is [size]. The sliding step size, For filling parameters, The height and width of the input feature map are given.

[0131] The output tensor is fed into the visual transformer. During the feature extraction stage, the image patch encoding module of the visual transformer is used to divide the time-frequency image into blocks. Cross-region feature associations are captured through a multi-head attention layer, which improves the ability to perceive the global structure in the time-frequency image.

[0132] Feature extraction is then performed using batch normalization and max pooling layers. The processed information then enters the channel attention module, enabling the model to more intelligently focus on key features and suppress irrelevant information. This is achieved through channel-level attention in both channel and spatial dimensions.

[0133] The channel attention module determines the feature value of different channels and assigns high weights to key channels (such as frequency channels containing core fault information in a fault frequency map), allowing the model to focus on these channels. Specifically, it uses global average pooling and max pooling layers to compress the spatial dimension of the feature map to one dimension, resulting in two "channel descriptors" (one based on average and one based on max), comprehensively capturing the global statistical information of channel features. The two descriptors are concatenated, and a convolutional layer extracts spatial features to generate a spatial attention map. The map is then scaled to 0-1 using the sigmoid function, representing the importance weight of each spatial location.

[0134] The spatial attention module's role is to locate key spatial regions in the feature map (e.g., the specific time-frequency location of a fault in a fault time-frequency map), assign high weights to these regions, and allow the model to focus on important locations. Specifically, it performs average pooling and max pooling along the channel axes to obtain two "spatial descriptors" (one based on channel average and one based on channel max), capturing channel information about spatial location. These two descriptors are then input into a shared multilayer perceptron (MLP) to learn the dependencies between channels, and finally fused to output a channel attention map. The sigmoid function is used to shrink the values ​​of the attention map to between 0 and 1, representing the importance weight of each channel.

[0135] Finally, the system enters the residual block, which skips the convolutional layers, mitigating gradient vanishing and supporting deep network training. Through the synergy of these modules, high-precision diagnostics can be achieved in imbalanced data scenarios. Its loss function is set to Adaptive Focal Loss, and the loss function is defined as follows: for:

[0136] ,

[0137] in, To balance the weights for each category, To focus parameters, For the model to class The predicted probability, This represents the total number of categories.

[0138] By utilizing the forward noise addition and backward denoising mechanism of a diffusion model, synthetic samples are generated that highly match the distribution of real fault samples. The maximum mean difference (MMD) metric verifies that the distributional difference is significantly lower than traditional oversampling techniques (such as SMOTE). Furthermore, t-SNE visualization and silhouette coefficient evaluation methods ensure that the synthetic samples possess both realism (closeness to the original distribution) and sufficient diversity (preventing pattern collapse). Compared to the overfitting or undersampling information loss problems caused by low-quality samples generated using methods such as SMOTE, this invention retains key fault features (e.g., time-frequency domain anomaly patterns) while balancing the dataset, providing high-quality training data for diagnostic models.

[0139] The improved residual network employs residual connections to support deep network training and combines Patch Embedding and multi-head attention mechanisms from VisionTransformer to capture cross-regional feature associations between local details (such as specific frequency components) and global structures (such as fault mode distribution) in the time-frequency image of vibration signals. The dual-branch mechanism, consisting of a Channel Attention Module (CAM) and a Spatial Attention Module (SAM), dynamically suppresses noisy channels and irrelevant regions, thereby enhancing the representation weights of weak fault features (such as the high-frequency components of early wear), improving the accuracy of composite fault mode recognition by more than 20%. This design effectively solves the problem of "overwhelming" when traditional networks process complex features.

[0140] The end-to-end training process, from data preprocessing (including wavelet transform and time-frequency image generation), diffusion model-based data augmentation, quality assessment to feature extraction and classification decision-making, significantly reduces the cost of manual parameter tuning and meets the needs of real-time industrial diagnostics. A multi-dimensional evaluation system based on MMD (for assessing distribution consistency) and silhouette coefficient (for assessing sample diversity) provides objective quantitative indicators for the quality of generated samples (e.g., MMD < 0.1, silhouette coefficient > 0.8), avoiding the problem of traditional methods relying on experience-based tuning and ensuring the stability and reliability of the model in applications such as wind power and rail transit.

[0141] Combination Figure 7As shown, this disclosure provides a gearbox fault detection device 1000 based on a neural network model. The gearbox fault detection device 1000 based on a neural network model includes a sample acquisition module 1001, a sample generation module 1002, a model training module 1003, and a fault detection module 1004.

[0142] The sample acquisition module 1001 is configured to: preprocess the vibration signal during the operation of the sample gearbox, convert the vibration signal into a time-frequency image through wavelet transform, and perform grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions.

[0143] The sample generation module 1002 is configured to: perform data augmentation on a preset number of original fault samples using a diffusion model; generate synthetic fault samples through forward noise addition and reverse noise reduction mechanisms; obtain the overall mean difference value and the overall contour coefficient by calculating the distribution difference between the original fault samples and the corresponding synthetic fault samples, and extracting the deep features of each original fault sample and the synthetic fault sample; determine the quality of the synthetic fault samples based on the overall mean difference value and the overall contour coefficient; allow the use of synthetic fault samples when the quality of the synthetic fault samples is determined to be qualified; and adjust the random sampling strategy of the diffusion model and regenerate synthetic fault samples using the diffusion model when the quality of the synthetic fault samples is determined to be unqualified.

[0144] The model training module 1003 is configured to: construct a training dataset based on normal samples, original fault samples and synthetic fault samples, train an initial residual network model based on the training dataset, and obtain a trained target residual network model.

[0145] The fault detection module 1004 is configured to: preprocess the vibration signal during the operation of the target gearbox, convert the vibration signal into a time-frequency image through wavelet transform, and input it into the target residual network model after grayscale processing and normalization, and detect the fault type of the target gearbox based on the target residual network model.

[0146] The gearbox fault detection device based on a neural network model provided in this disclosure employs a diffusion model to augment the original fault samples. Through forward noise addition and reverse denoising mechanisms, it generates high-quality, highly diverse synthetic fault samples. These synthetic samples are more closely distributed than the real data, effectively avoiding the introduction of noise samples while retaining key features of the original fault samples (such as abnormal patterns in time-frequency images), significantly improving sample quality. Furthermore, the number of synthetic samples is controllable, avoiding the loss of key information that may occur with undersampling methods. This effectively alleviates the shortage of fault samples in the training dataset, improving the model's generalization ability and diagnostic stability. In addition, wavelet transform converts the vibration signal into a time-frequency image, capturing the local time-frequency features of the signal and effectively extracting non-stationary components of the vibration signal, such as weak fault signals and sudden impacts. Subsequently, the image undergoes grayscale processing and normalization to unify the input format and enhance feature representation. Further, a residual network model is constructed and trained, utilizing its deep structure to extract multi-scale features. The residual network alleviates the vanishing gradient problem through residual connections, supports efficient extraction and fusion of deep features, improves the model's robustness to noise interference, and enhances its ability to identify weak fault features and complex fault modes. Furthermore, by constructing a balanced training dataset containing normal samples, original fault samples, and synthetic fault samples, the model gains more learning opportunities during training, avoiding overfitting to the majority class. The trained target residual network model exhibits good generalization ability, accurately identifying the fault type of the target gearbox vibration signal after the same preprocessing in actual detection. This significantly improves the model's adaptability to different fault modes, enhances its ability to model complex features, and increases the accuracy of fault identification. The combined approach of wavelet transform, diffusion model data augmentation, and residual network feature extraction effectively solves the problems of data imbalance, difficult feature extraction, and insufficient diagnostic accuracy in gearbox fault diagnosis, significantly improving the model's fault identification accuracy and engineering applicability.

[0147] In some embodiments, the sample acquisition module 1001 is configured to:

[0148] During the operation of the sample gearbox, vibration signals under normal conditions and vibration signals under various fault conditions were collected.

[0149] The vibration signal is segmented into segments with a fixed window length, and a continuous wavelet transform is applied to each segment to convert it into a time-frequency image.

[0150] Each time-frequency image is processed into a single-channel grayscale image after grayscale processing. Each single-channel grayscale image is then normalized to eliminate dimensional differences.

[0151] The single-channel grayscale image corresponding to the vibration signal under normal conditions is used as the normal sample, and the single-channel grayscale image corresponding to the vibration signal under fault conditions is used as the original fault sample.

[0152] In some embodiments, the sample generation module 1002 is configured to:

[0153] The original input fault samples are subjected to forward noise addition using a diffusion model, and Gaussian noise is iteratively introduced to obtain noisy fault samples.

[0154] The diffusion model is used to perform reverse denoising on the noisy fault samples, iteratively removing noise from the fault samples until a synthetic fault sample with the same distribution as the original fault sample is obtained.

[0155] In some embodiments, the process of introducing Gaussian noise in a single iteration can be expressed by the following formula: , This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the cumulative noise proportionality factor. Indicates the original fault sample. Indicates Gaussian noise distribution;

[0156] It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed.

[0157] In some embodiments, the process of removing noise from faulty samples in a single iteration can be expressed by the following formula:

[0158] ,

[0159] Indicates the ()th step in the reverse denoising operation Synthetic fault samples obtained through iterative steps This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the noise in the prediction. This represents the cumulative noise proportionality factor. Indicates standard Gaussian noise. To represent the first step in the reverse denoising operation The predefined variance of each iteration;

[0160] It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed;

[0161] It is calculated using the following formula: , It is a hyperparameter. .

[0162] In some embodiments, the sample generation module 1002 is configured to: calculate the distribution difference between the original fault sample and the corresponding synthetic fault sample based on multiple different Gaussian kernel functions to obtain multiple maximum mean difference values; take the average of the multiple maximum mean difference values ​​as the overall mean difference value; extract the deep features of each original fault sample and synthetic fault sample using a classifier network, and merge the extracted features into fusion features; perform cluster analysis on the feature space after dimensionality reduction of the fusion features to calculate the silhouette coefficient of each synthetic fault sample; take the average of the silhouette coefficients of the synthetic fault samples as the overall silhouette coefficient; and determine the quality of the synthetic fault sample based on the overall mean difference value and the overall silhouette coefficient.

[0163] In some embodiments, the maximum mean difference value based on a kernel function can be obtained by the following formula:

[0164] ,

[0165] This represents the square of the largest mean difference. Indicates the first One original fault sample, Indicates the use of the first Pair one original fault sample with another original fault sample to compute the kernel function. Indicates the first A synthetic fault sample, Indicates the use of the first For each synthetic fault sample, a kernel function is computed on a corresponding synthetic fault sample. Indicates the number of original fault samples. Indicates the number of synthesized fault samples. Represents the kernel function;

[0166] The contour coefficients of the synthesized fault samples can be obtained using the following formula: , Indicates the first The profile coefficients of a synthetic fault sample. Indicates the first The average distance between a synthetic faulty sample and other synthetic faulty samples in the same cluster Indicates the first The average distance between a synthetic faulty sample and all synthetic faulty samples in the nearest other clusters.

[0167] In some embodiments, the fault detection module 1004 is configured to:

[0168] Vibration signals were collected during the operation of the target gearbox;

[0169] The vibration signal is segmented into segments with a fixed window length, and a continuous wavelet transform is applied to each segment to convert it into a time-frequency image.

[0170] Each time-frequency image is processed into a single-channel grayscale image after grayscale processing. Each single-channel grayscale image is then normalized to eliminate dimensional differences.

[0171] Each single-channel grayscale image is input into the target residual network model, and the fault type of the target gearbox is detected based on the target residual network model.

[0172] Combination Figure 8 As shown, this disclosure provides an electronic device 2000, which includes a processor 2001 and a memory 2002. Optionally, the device 2000 may further include a communication interface 2003 and a bus 2004. The processor 2001, communication interface 2003, and memory 2002 can communicate with each other via the bus 2004. The communication interface 2003 can be used for information transmission. The processor 2001 can call logical instructions in the memory 2002 to execute the gearbox fault detection method based on a neural network model described in the above embodiment. Furthermore, the logical instructions in the memory 2002 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0173] The memory 2002, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 2001 executes functional applications and data processing by running the program instructions / modules stored in the memory 2002, that is, it implements the gearbox fault detection method based on the neural network model in the above embodiments.

[0174] The memory 2002 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 2002 may include high-speed random access memory and may also include non-volatile memory.

[0175] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, then the relevant parts can be referred to the description of the method section.

[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A gearbox fault detection method based on a neural network model, characterized in that, include: The vibration signal of the sample gearbox during operation is preprocessed. The vibration signal is converted into a time-frequency image by wavelet transform, and then grayscale processing and normalization are performed to obtain normal samples under normal conditions and original fault samples under fault conditions. A diffusion model is used to augment a preset number of original fault samples, and synthetic fault samples are generated through forward noise addition and reverse noise reduction mechanisms. By calculating the distribution difference between the original fault samples and the corresponding synthetic fault samples, and extracting the deep features of each original fault sample and synthetic fault sample, the overall mean difference value and the overall silhouette coefficient are obtained respectively, and the quality of the synthetic fault samples is determined based on the overall mean difference value and the overall silhouette coefficient. Synthetic fault samples are permitted to be used when their quality is deemed acceptable. When it is determined that the quality of the synthesized faulty sample is unqualified, the random sampling strategy of the diffusion model is adjusted, and the synthesized faulty sample is regenerated using the diffusion model. A training dataset is constructed based on normal samples, original fault samples and synthetic fault samples. An initial residual network model is trained based on the training dataset to obtain a well-trained target residual network model. The vibration signal of the target gearbox during operation is preprocessed. Wavelet transform is used to convert the vibration signal into a time-frequency image, which is then processed for grayscale and normalized before being input into the target residual network model. The fault type of the target gearbox is detected based on the target residual network model. The target residual network model includes convolutional kernels, a visual transformer, a batch normalization layer, a max pooling layer, an attention mechanism layer, and a global average pooling layer. The attention mechanism layer includes a channel attention module that receives input features and a spatial attention module that outputs refined features.

2. The gearbox fault detection method according to claim 1, characterized in that, The vibration signals of the sample gearbox during operation are preprocessed. Wavelet transform is used to convert the vibration signals into time-frequency images, followed by grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions, including: During the operation of the sample gearbox, vibration signals under normal conditions and vibration signals under various fault conditions were collected. The vibration signal is segmented into segments with a fixed window length, and a continuous wavelet transform is applied to each segment to convert it into a time-frequency image. Each time-frequency image is processed into a single-channel grayscale image after grayscale processing. Each single-channel grayscale image is then normalized to eliminate dimensional differences. The single-channel grayscale image corresponding to the vibration signal under normal conditions is used as the normal sample, and the single-channel grayscale image corresponding to the vibration signal under fault conditions is used as the original fault sample.

3. The gearbox fault detection method according to claim 1, characterized in that, Data augmentation is performed on a predetermined number of original fault samples using a diffusion model. Synthetic fault samples are generated through forward noise addition and backward noise reduction mechanisms, including: The original input fault samples are subjected to forward noise addition using a diffusion model, and Gaussian noise is iteratively introduced to obtain noisy fault samples. The diffusion model is used to perform reverse denoising on the noisy fault samples, iteratively removing noise from the fault samples until a synthetic fault sample with the same distribution as the original fault sample is obtained.

4. The gearbox fault detection method according to claim 3, characterized in that, The process of introducing Gaussian noise in a single iteration can be expressed by the following formula: , This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the cumulative noise proportionality factor. Indicates the original fault sample. Indicates Gaussian noise distribution; It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed.

5. The gearbox fault detection method according to claim 3, characterized in that, The process of removing noise from faulty samples in a single iteration can be expressed by the following formula: ; Indicates the ()th step in the reverse denoising operation Synthetic fault samples obtained through iterative steps This represents the noise fault sample obtained in the t-th iteration during the forward noise addition operation. This represents the noise in the prediction. This represents the cumulative noise proportionality factor. Indicates standard Gaussian noise. To represent the first step in the reverse denoising operation The predefined variance of each iteration; It is calculated using the following formula: , This represents the noise scaling factor for one iteration. Indicates the number of iterations that have been performed; It is calculated using the following formula: , It is a hyperparameter. .

6. The gearbox fault detection method according to claim 1, characterized in that, By calculating the distribution difference between the original fault samples and the corresponding synthetic fault samples, and extracting the deep features of each original fault sample and synthetic fault sample, the overall mean difference value and the overall silhouette coefficient are obtained respectively. Based on the overall mean difference value and the overall silhouette coefficient, the quality of the synthetic fault samples is determined, including: The distribution difference between the original fault sample and the corresponding synthetic fault sample is calculated based on multiple different Gaussian kernel functions to obtain multiple maximum mean difference values. The average of the multiple largest mean differences is taken as the overall mean difference. The deep features of each original fault sample and the synthetic fault sample are extracted using a classifier network, and the extracted features are merged into fusion features. Cluster analysis is performed on the feature space after dimensionality reduction of the fused features to calculate the contour coefficient of each synthetic fault sample; The average profile coefficient of the synthesized fault samples is used as the overall profile coefficient. The quality of the synthesized fault samples is determined based on the overall mean difference and the overall profile coefficient.

7. The gearbox fault detection method according to claim 6, characterized in that, The maximum mean difference value based on a kernel function can be obtained using the following formula: ; This represents the square of the largest mean difference. Indicates the first One original fault sample, Indicates the use of the first Pair one original fault sample with another original fault sample to compute the kernel function. Indicates the first A synthetic fault sample, Indicates the use of the first For each synthetic fault sample, a kernel function is computed on a corresponding synthetic fault sample. Indicates the number of original fault samples. Indicates the number of synthesized fault samples. Represents the kernel function; The contour coefficients of the synthesized fault samples can be obtained using the following formula: , Indicates the first The profile coefficients of a synthetic fault sample. Indicates the first The average distance between a synthetic faulty sample and other synthetic faulty samples in the same cluster Indicates the first The average distance between a synthetic faulty sample and all synthetic faulty samples in the nearest other clusters.

8. The gearbox fault detection method based on a neural network model according to claim 1, characterized in that, The vibration signal during the operation of the target gearbox is preprocessed. Wavelet transform is used to convert the vibration signal into a time-frequency image, which is then processed for grayscale and normalized before being input into the target residual network model. Based on the target residual network model, the fault types of the target gearbox are detected, including: Vibration signals were collected during the operation of the target gearbox; The vibration signal is segmented into segments with a fixed window length, and a continuous wavelet transform is applied to each segment to convert it into a time-frequency image. Each time-frequency image is processed into a single-channel grayscale image after grayscale processing. Each single-channel grayscale image is then normalized to eliminate dimensional differences. Each single-channel grayscale image is input into the target residual network model, and the fault type of the target gearbox is detected based on the target residual network model.

9. A gearbox fault detection device based on a neural network model, characterized in that, include: The sample acquisition module is configured to: preprocess the vibration signal of the sample gearbox during operation, convert the vibration signal into a time-frequency image through wavelet transform, and perform grayscale processing and normalization to obtain normal samples under normal conditions and original fault samples under fault conditions. The sample generation module is configured to: augment a preset number of original fault samples using a diffusion model; generate synthetic fault samples through forward noise addition and reverse denoising mechanisms; obtain the overall mean difference value and overall profile coefficient by calculating the distribution difference between the original fault samples and the corresponding synthetic fault samples, and extracting the deep features of each original fault sample and synthetic fault sample; determine the quality of the synthetic fault samples based on the overall mean difference value and overall profile coefficient; allow the use of synthetic fault samples when the quality of the synthetic fault samples is deemed acceptable; and adjust the random sampling strategy of the diffusion model and regenerate synthetic fault samples using the diffusion model when the quality of the synthetic fault samples is deemed unacceptable. The model training module is configured to: construct a training dataset based on normal samples, original fault samples and synthetic fault samples; train an initial residual network model based on the training dataset; and obtain a trained target residual network model. The fault detection module is configured to: preprocess the vibration signal during the operation of the target gearbox, convert the vibration signal into a time-frequency image through wavelet transform, perform grayscale processing and normalization, and then input it into the target residual network model; and detect the fault type of the target gearbox based on the target residual network model. The target residual network model includes convolutional kernels, a visual transformer, a batch normalization layer, a max pooling layer, an attention mechanism layer, and a global average pooling layer. The attention mechanism layer includes a channel attention module that receives input features and a spatial attention module that outputs refined features.

10. An electronic device comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the gearbox fault detection method based on a neural network model as described in any one of claims 1 to 8 when running program instructions.

Citation Information

Patent Citations

  • Tunnel disease monitoring data enhancement method and system based on CTGAN

    CN119598301A

  • Hydraulic system intelligent fault diagnosis method based on diffusion model data enhancement framework

    CN120332289A