Equipment fault diagnosis model training method and device based on semi-supervised learning

Through the equipment fault diagnosis model training method based on semi-supervised learning, the problems of difficult data labeling and model instability are solved, high-precision fault diagnosis in complex industrial environments is achieved, and the robustness and adaptability of the model are enhanced.

CN120611191AActive Publication Date: 2025-09-09SHANDONG ENERGY DIGITAL CLOUD TECH CO LTD

Patent Information

Application Number
CN202511121340.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-09
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing technologies in industrial equipment fault diagnosis have problems such as difficult data labeling, unstable models, and poor anti-interference capabilities. In particular, in vibration signal processing, strong noise, non-stationarity, and local mutation characteristics lead to feature distortion and model misjudgment, and the accuracy of pseudo-labels is low, making it difficult to adapt to complex and changing industrial environments.

Method used

A semi-supervised learning-based equipment fault diagnosis model training method is adopted. By obtaining equipment vibration data, data preprocessing and feature extraction are performed, pseudo labels are generated, and the final training sample set is constructed. The neural network model is trained using adaptive activation function and consistency loss function to realize the diagnosis of equipment faults.

Benefits of technology

It improves the accuracy and stability of equipment fault diagnosis, can adapt to sample differences under different working conditions, sensor channels and sampling frequencies, and enhances the robustness and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611191A_ABST
    Figure CN120611191A_ABST
Patent Text Reader

Abstract

The invention provides an equipment fault diagnosis model training method and device based on semi-supervised learning, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the collection of equipment vibration data in a preset scene, and carrying out the training of an equipment fault diagnosis model based on the feature similarity between a target sample in a corresponding target feature vector and a preset training sample set; and generating a pseudo tag corresponding to the target sample. Limited high-quality labeled samples (supervised learning) and a large number of unlabeled samples (unsupervised learning) are organically combined, the method does not depend on manually set working condition rules, and corresponding pseudo labels are generated after sample similarity analysis is performed from feature level analysis. According to the method, knowledge migration with points as surfaces can be achieved, and on the basis, after an equipment fault diagnosis model is constructed, the model can effectively cope with sample differences under different working conditions, sensor channels and sampling frequencies, adapt to complex and changeable industrial environment requirements and improve the equipment fault diagnosis precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and device for training an equipment fault diagnosis model based on semi-supervised learning. Background Art

[0002] With the continuous advancement of industrial automation, various rotating and reciprocating mechanical equipment are widely used in key sectors such as manufacturing, energy, and transportation. The reliability of equipment operation is directly related to production efficiency and safety. Vibration signals, one of the most commonly used monitoring methods in fault diagnosis, contain rich health status information in both time and frequency dimensions. They are particularly valuable in early fault detection and early warning of critical component failure.

[0003] However, in real industrial environments, vibration data often exhibits complex characteristics such as strong noise, high nonstationarity, and localized mutations. Conventional signal processing and feature extraction methods are insufficiently robust to these non-ideal data, which can easily lead to feature distortion and model misjudgment. Furthermore, equipment failure samples are often difficult to obtain in large quantities, and manual labeling is costly and time-consuming, resulting in a large number of unlabeled samples in the training data. This limits the application of traditional fully supervised learning methods in industrial diagnostic scenarios.

[0004] Although some studies have attempted to introduce semi-supervised learning methods to alleviate the problem of label scarcity, existing methods generally have problems such as low pseudo-label accuracy, unstable training process, and model sensitivity to noise, making them difficult to adapt to the complex and changing needs of industrial environments. Summary of the Invention

[0005] In view of this, the present invention provides a method and device for training an equipment fault diagnosis model based on semi-supervised learning, which effectively solves key technical problems in industrial equipment fault diagnosis, such as difficult data labeling, unstable model, and poor anti-interference ability.

[0006] In a first aspect, an embodiment of the present invention provides a method for training an equipment fault diagnosis model based on semi-supervised learning, wherein the method includes: obtaining equipment vibration data of a preset mechanical equipment in a preset scenario, performing data preprocessing on the equipment vibration data, and determining a target feature vector corresponding to the equipment vibration data; wherein the equipment vibration data includes dynamic operation signals pre-collected from key parts of the preset mechanical equipment; extracting target samples from the target feature vector according to a pre-constructed training sample set; generating pseudo labels corresponding to the target samples based on feature similarity between the target samples and the training sample set; constructing a final training sample set corresponding to the preset mechanical equipment based on the pseudo labels; inputting the final training sample set into a preset neural network model to train the neural network model; constructing an equipment fault diagnosis model based on the trained neural network model, so as to perform equipment fault diagnosis on the target mechanical equipment using the equipment fault diagnosis model.

[0007] In combination with the first aspect, an embodiment of the present invention provides a first implementation method of the first aspect, wherein the steps of performing data preprocessing on the device vibration data and determining the target feature vector corresponding to the device vibration data include: denoising the device vibration data based on the scale change corresponding to the device vibration data, and determining the denoised samples corresponding to the device vibration data; performing sample-by-sample extreme value calculation on the denoised samples, and dynamically normalizing the denoised samples based on the calculation results to determine the initial preprocessed data; performing multi-scale time-frequency feature extraction on the initial preprocessed data to determine the target feature vector corresponding to the device vibration data.

[0008] In combination with the first aspect, an embodiment of the present invention provides a second implementation method of the first aspect, wherein, based on the scale changes corresponding to the device vibration data, the device vibration data is denoised, and the steps of determining the denoised samples corresponding to the device vibration data include: using wavelet decomposition to decompose the device vibration data according to different scales, and determining multiple sub-band data corresponding to the device vibration data; calculating the adaptive threshold corresponding to each sub-band data based on the overall standard deviation of the samples corresponding to the device vibration data; extracting local frequency domain features from the sub-band data based on the adaptive threshold and a preset noise suppression coefficient, and determining the denoised samples corresponding to the device vibration data.

[0009] In combination with the first aspect, an embodiment of the present invention provides a third implementation method of the first aspect, wherein the steps of performing multi-scale time-frequency feature extraction on the initial preprocessed data and determining the target feature vector corresponding to the equipment vibration data include: performing wavelet packet transform on the initial preprocessed data to extract the energy distribution of the initial preprocessed data in different sub-bands; performing time domain feature statistics on the energy distribution corresponding to each sub-band to determine the data characteristics of the initial preprocessed data in the current sub-band; and combining the data characteristics and energy distribution of each sub-band to determine the target feature vector corresponding to the equipment vibration data.

[0010] In combination with the first aspect, an embodiment of the present invention provides a fourth implementation of the first aspect, wherein the step of extracting a target sample from a target feature vector based on a pre-constructed training sample set includes: identifying an unlabeled sample from the target feature vector based on the pre-constructed training sample set, and determining the unlabeled sample as the target sample; based on the feature similarity between the target sample and the training sample set, generating a pseudo label corresponding to the target sample includes: obtaining the target sample and its neighboring samples in the feature space, and calculating the similarity weight between the neighboring samples and the target sample; determining the sample label corresponding to the neighboring sample from the training sample set; determining the predicted label output corresponding to the target sample based on the sample label and the similarity weight; using a preset confidence threshold, performing a threshold comparison on the similarity weight, retaining the predicted label output corresponding to the similarity weight that meets the confidence threshold, and determining the pseudo label corresponding to the target sample.

[0011] In combination with the first aspect, an embodiment of the present invention provides a fifth implementation of the first aspect, wherein the final training sample set is input into a preset neural network model, and the step of training the neural network model includes: calculating the label feature mean corresponding to the sample distribution of each category of the final training sample set; determining the initialization weight matrix corresponding to the neural network model based on the label feature mean and the clustering correction term corresponding to the final training sample set; adding Gaussian noise to the final training sample set, and using a preset adaptive activation function to perform forward propagation on the final training sample set and the final training sample set with Gaussian noise added, respectively; and determining the first model output corresponding to the final training sample set and the second model output corresponding to the final training sample set with Gaussian noise added; based on the first model output, calculating the supervision loss corresponding to the final training sample set; based on the second model output, calculating the consistency loss corresponding to the final training sample set; calculating the total loss function of the neural network model based on the supervision loss and the consistency loss; based on the total loss function, iteratively training the neural network model, and, based on the momentum vector of the neural network model during the iterative training process, updating the parameters of the neural network model.

[0012] In combination with the first aspect, an embodiment of the present invention provides a sixth implementation of the first aspect, wherein the adaptive activation function includes a first activation function and a second activation function, and the preset adaptive activation function is used to perform forward propagation steps on the final training sample set and the final training sample set with Gaussian noise added, respectively, including: judging the positive and negative values ​​of the input samples; when the input samples are greater than zero, performing nonlinear activation processing on the input samples based on a preset sinusoidal term; wherein the sinusoidal term includes a preset sinusoidal modulation intensity and a preset sinusoidal frequency factor parameter; otherwise, performing nonlinear activation processing on the input samples based on a preset exponential term; wherein the exponential term includes an exponential attenuation coefficient.

[0013] In combination with the first aspect, an embodiment of the present invention provides a seventh implementation of the first aspect, wherein the step of calculating the consistency loss corresponding to the final training sample set based on the second model output includes: calculating the Euclidean distance corresponding to the first model output and the second model output; and calculating the consistency loss corresponding to the final training sample set based on the Euclidean distance.

[0014] In combination with the first aspect, an embodiment of the present invention provides an eighth implementation of the first aspect, wherein the steps of determining the first model output corresponding to the final training sample set and the second model output corresponding to the final training sample set with Gaussian noise added include: obtaining the original output value corresponding to the neural network model; determining the calibration probability output corresponding to the original output value based on preset calibration hyperparameters; and determining the calibration probability output as the final model output.

[0015] In a second aspect, an embodiment of the present invention provides an equipment fault diagnosis model training device based on semi-supervised learning, wherein the device includes: a data processing module for obtaining equipment vibration data of a preset mechanical equipment in a preset scenario, performing data preprocessing on the equipment vibration data, and determining a target feature vector corresponding to the equipment vibration data; wherein the equipment vibration data includes dynamic operation signals pre-collected from key parts of the preset mechanical equipment; a label generation module for extracting target samples from the target feature vector based on a pre-constructed training sample set; generating pseudo labels corresponding to the target samples based on feature similarity between the target samples and the training sample set; an execution module for constructing a final training sample set corresponding to the preset mechanical equipment based on the pseudo labels; a training module for inputting the final training sample set into a preset neural network model to train the neural network model; a construction module for constructing an equipment fault diagnosis model based on the trained neural network model, so as to perform equipment fault diagnosis on the target mechanical equipment using the equipment fault diagnosis model.

[0016] The embodiments of the present invention provide a method and apparatus for training an equipment fault diagnosis model based on semi-supervised learning. After collecting equipment vibration data under a preset scenario, the method generates a pseudo-label corresponding to the target sample based on the feature similarity between the target sample in the corresponding target feature vector and the preset training sample set. It organically combines a limited number of high-quality labeled samples (supervised learning) with a large number of unlabeled samples (unsupervised learning). It does not rely on manually set working condition rules, and generates corresponding pseudo-labels after performing sample similarity analysis at the feature level. It can achieve "point-to-surface" knowledge transfer. Based on this, after constructing an equipment fault diagnosis model, the model can effectively cope with sample differences under different working conditions, sensor channels, and sampling frequencies, adapt to the needs of complex and changing industrial environments, and improve the accuracy of equipment fault diagnosis.

[0017] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0018] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 A flowchart of a method for training a device fault diagnosis model based on semi-supervised learning provided by an embodiment of the present invention; Figure 2 A flowchart of another method for training a device fault diagnosis model based on semi-supervised learning provided by an embodiment of the present invention; Figure 3 A schematic diagram of a comparison of diagnostic accuracy results under different annotation ratios provided by an embodiment of the present invention; Figure 4 A schematic diagram of the comparison results of F1 scores under different noise levels provided by an embodiment of the present invention; Figure 5 A schematic diagram of a comparison result of recall rates corresponding to different fault categories provided by an embodiment of the present invention; Figure 6 A schematic diagram of the relationship between training time and diagnostic accuracy provided by an embodiment of the present invention; Figure 7 A schematic diagram of a denoising effect provided by an embodiment of the present invention; Figure 8 A schematic diagram of the structure of a device fault diagnosis model training device based on semi-supervised learning provided by an embodiment of the present invention; Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0022] To solve the above technical problems, an embodiment of the present invention provides a method and device for training an equipment fault diagnosis model based on semi-supervised learning, which effectively solves key technical problems in industrial equipment fault diagnosis, such as difficult data labeling, unstable model, and poor anti-interference ability.

[0023] For ease of understanding, an embodiment of the present invention first provides a method for training a device fault diagnosis model based on semi-supervised learning to illustrate. Figure 1 The flowchart corresponding to the embodiment of the present invention is shown. Figure 1 , including the following steps: Step S102 : obtaining equipment vibration data of a preset mechanical equipment in a preset scenario, performing data preprocessing on the equipment vibration data, and determining a target feature vector corresponding to the equipment vibration data.

[0024] Step S104 : extracting a target sample from the target feature vector according to the pre-constructed training sample set; and generating a pseudo label corresponding to the target sample based on the feature similarity between the target sample and the training sample set.

[0025] Step S106: constructing a final training sample set corresponding to the preset mechanical equipment based on the pseudo labels.

[0026] In some scenarios, there are few labeled samples in the vibration data, and only some fault types are labeled. For example, in a semi-supervised scenario, only key fault samples are manually labeled, such as fault samples of sudden severe damage. For the remaining samples, conventional means automatically generate weak labels through working condition association, such as "normal-abnormal" binary labels. However, this labeling method lacks specific fault type information, making it difficult to support fine classification tasks, and the label quality is low; in addition, the "abnormal" label may be triggered by environmental noise or atypical disturbances, and the misjudgment rate is high. Pseudo-labels constructed based on weak labels are prone to introduce errors, polluting model training, and causing continuous decline in model performance. Moreover, the model has poor generalization ability, making it difficult to identify unseen fault types, and the diagnostic accuracy is limited.

[0027] To this end, embodiments of the present invention acquire fault samples of pre-set mechanical equipment under specific scenarios. These fault samples include manually labeled critical fault samples, such as sudden severe damage, fracture, and eccentricity, as well as a large number of unlabeled samples, which may come from different operating conditions, operating states, or sensor channels. Furthermore, embodiments of the present invention utilize a pre-set training sample set to distinguish unlabeled samples from the collected fault samples, and use the training sample set as a "reference sample library" for subsequent feature matching and pseudo-label generation.

[0028] In specific implementations, the equipment vibration data collected by embodiments of the present invention includes dynamic operating signals pre-collected from key components of pre-defined mechanical equipment. Dynamic operating signals can be obtained from key components of various rotating or reciprocating machinery (such as motors, pumps, fans, and gearboxes) on industrial sites. These signals are typically collected from areas where vibration energy is concentrated, such as bearing seats, housings, and spindle connections. In specific implementations, a rigidly mounted triaxial accelerometer can be placed in direct contact with the equipment surface to capture the equipment's vibration acceleration signals in the X, Y, and Z directions.

[0029] Data acquisition can be performed by continuously recording time-domain waveforms at a fixed sampling frequency, typically at least 10 times the device's fundamental frequency, such as 10kHz–20kHz. Operating parameters such as device speed and load conditions are also recorded simultaneously. All collected data is stored in a time-series database and strictly aligned with device startup, shutdown, and maintenance records using timestamps.

[0030] The data source for the training sample set is similar to that for the equipment vibration data, so I won't elaborate on this here. The training sample set consists of labeled samples, which can be manually labeled based on the equipment failure mechanism. The labeled categories can be divided into five categories: normal state, rotor imbalance, bearing damage, mechanical looseness, and gear failure.

[0031] This approach involves preprocessing the equipment vibration data to obtain the target feature vector, which is then matched against samples in the training set for similarity. Alternatively, sample clustering can be employed, assigning each cluster a reasonable pseudo-label to identify the target sample. Furthermore, based on the feature similarity between the labeled and unlabeled samples, a pseudo-label is generated for the unlabeled sample.

[0032] Specifically, conventional pseudo-label generation methods rely on the understanding of the equipment's operating environment, such as changes in operating parameters such as temperature, pressure, and speed, and use this as a basis for inferring the type of fault. This method is based on rules or simple statistical analysis. However, the embodiment of the present invention generates pseudo-labels based on the feature similarity between unlabeled samples and labeled samples, does not rely on manually set operating rules, and analyzes the similarity between unlabeled samples and labeled samples from a feature level. Among them, the embodiment of the present invention organically combines limited high-quality labeled samples (supervised learning) with a large number of unlabeled samples (unsupervised learning) to achieve "point-to-surface" knowledge transfer, which can effectively deal with sample differences under different operating conditions, sensor channels, and sampling frequencies.

[0033] Step S108: input the final training sample set into the preset neural network model to train the neural network model.

[0034] Step S110 : constructing an equipment fault diagnosis model based on the trained neural network model, so as to perform equipment fault diagnosis on the target mechanical equipment using the equipment fault diagnosis model.

[0035] Furthermore, after generating pseudo labels for the unlabeled data, a final training sample set is constructed and the model is trained to use the model for equipment fault diagnosis.

[0036] Furthermore, the existing technology has the following technical problems: (1) Conventional self-training methods generate pseudo labels based on model prediction results, lack semantic similarity measurement between samples, have a high risk of error propagation, and are prone to pseudo label contamination problems when labels are scarce; (2) Commonly used z-score or min-max normalization ignores time-varying features and local mutations, and the wavelet denoising process does not adopt multi-scale threshold dynamic control, making it difficult to handle high-frequency strong noise scenes; (3) Fixed activation functions such as ReLU / Sigmoid have expression bottlenecks when processing inputs with small fault activation intervals, and are insensitive to negative gradients, resulting in weak feature learning capabilities; (4) Most current methods lack a consistency constraint mechanism or use input layer perturbations, resulting in low robustness; the fixed parameter design of the momentum optimization process is difficult to deal with the oscillation problem in the early stages of training, and the training time is long and unstable. To solve the above problems, based on the above embodiments, the embodiments of the present invention also provide another equipment fault diagnosis model training method based on semi-supervised learning, Figure 2 The flowchart corresponding to the embodiment of the present invention is shown. Figure 2 , including the following steps: Step S202 : obtaining equipment vibration data of a preset mechanical equipment in a preset scenario, performing data preprocessing on the equipment vibration data, and determining a target feature vector corresponding to the equipment vibration data.

[0037] In a specific implementation, the embodiment of the present invention performs the following preprocessing operations on the collected data: 1) Based on the scale changes corresponding to the equipment vibration data, the equipment vibration data is denoised and the denoised samples corresponding to the equipment vibration data are determined.

[0038] The monitoring data of the vibration sensor has high noise, non-stationary and local mutation characteristics. Due to the complex equipment operating environment and strong noise interference, conventional normalization methods such as z-score will amplify the noise and ignore the time-varying characteristics, resulting in unstable model training and overfitting. The embodiment of the present invention performs wavelet threshold denoising on the original vibration sensor data. Specifically, the embodiment of the present invention uses wavelet decomposition to decompose the equipment vibration data according to different scales, and determines multiple sub-band data corresponding to the equipment vibration data; according to the overall standard deviation of the samples corresponding to the equipment vibration data, the adaptive threshold corresponding to each sub-band data is calculated; based on the adaptive threshold and the preset noise suppression coefficient, the local frequency domain features are extracted from the sub-band data to determine the denoised samples corresponding to the equipment vibration data.

[0039] In the specific implementation, the vibration data of each sample at each moment is used as the basis, and wavelet decomposition is used to decompose it into sub-bands of different scales and positions. Then, for each sub-band, the wavelet coefficients are processed according to the adaptive threshold, and the noise components are removed through operations such as noise suppression coefficient and sign function to obtain the denoised data, which is expressed as:

[0040] Where y i (t) is the vibration data of the ith sample at the tth moment after denoising, which serves as the normalized input and represents the clean signal after noise removal; x i (t) is the vibration sensor data of the original i-th sample at the t-th moment, which contains noise and mutation; |x i (t)| represents the absolute value of the vibration sensor data of the original i-th sample at the t-th moment; J is the number of wavelet decomposition layers, which is set adaptively based on data characteristics and represents the depth of multi-scale decomposition. For non-stationarity, multi-layer decomposition can capture noise in different frequency bands. For example, if it is set to 5 layers, the frequency band of rotating machinery faults is dispersed, and too shallow a layer cannot separate noise and fault harmonics. j is a positive integer. is the wavelet coefficient scaling factor of the jth layer; It is a wavelet basis function that characterizes the extraction of local frequency domain features and is used to separate noise and useful signals. For example, Daubechies wavelet is selected. The compact support of Daubechies wavelet is suitable for capturing transient events, such as local mutations of gear tooth breakage. is the translation parameter of the wavelet of the jth layer; is the scale parameter of the wavelet of the jth layer; is the noise suppression coefficient, which controls the denoising strength. For example, setting ; is a sign function to ensure the stability of threshold processing; is the maximum value function; is the adaptive threshold of the jth layer. The threshold is set according to the overall standard deviation of the sample to adapt to different equipment working conditions, such as vibration amplitude fluctuations caused by load changes, while avoiding excessive filtering of the fixed threshold under weak fault signals. The calculation method is expressed as ; is the standard deviation function; N is the total number of samples. It should be noted that The term is used to retain the fault impulse component, ensuring that pulses with amplitudes greater than the threshold are not filtered, while suppressing Gaussian noise and providing high signal-to-noise ratio input for subsequent feature extraction.

[0041] 2) Calculate the extreme value of each denoised sample, and perform dynamic normalization on the denoised samples based on the calculation results to determine the initial preprocessed data.

[0042] Conventional techniques typically perform direct normalization, which can lead to feature distortion due to uneven scales in the denoised data. This invention dynamically normalizes the denoised data, calculating the difference between the value of each sample at each moment and the minimum value of that sample at all moments, and then dividing it by the difference between the maximum and minimum values ​​of that sample at all moments, normalizing the data to the range [0, 1], expressed as:

[0043] Where, is the normalized data of the i-th sample at the t-th moment; is the time series minimum function, It is a time series maximum function, which is calculated sample by sample to preserve local dynamics for non-stationarity.

[0044] 3) Perform multi-scale time-frequency feature extraction on the initial preprocessed data to determine the target feature vector corresponding to the equipment vibration data.

[0045] Vibration data contains fault-related frequency domain patterns, such as harmonic components, but conventional fast Fourier transform ignores local correlations in the time domain, resulting in redundant features and inability to capture transient events; The embodiment of the present invention performs feature extraction through the following steps: a-Perform wavelet packet transform on the initial pre-processed data to extract the energy distribution of the initial pre-processed data in different sub-bands; The present invention performs wavelet packet transform on the normalized data, maps it to different sub-bands, and calculates the energy characteristics of each sub-band. That is, the transformation results in each sub-band are squared and summed, which is expressed as:

[0046] Where, For the The first sample Sub-band energy characteristics; is a positive integer; is a positive integer; T is the time series length; is the wavelet packet transform function, which maps the input to the For example, the db4 wavelet basis is selected as the wavelet packet transform function.

[0047] b- Perform time domain feature statistics on the energy distribution corresponding to each sub-band to determine the data characteristics of the initial pre-processed data in the current sub-band.

[0048] c- Combine the data characteristics and energy distribution of each sub-band to determine the target feature vector corresponding to the equipment vibration data.

[0049] Furthermore, the time domain statistics of the normalized data, including skewness and kurtosis, are calculated. Then, the sub-band energy features are combined with the time domain statistics to form a combined feature vector for each sample, which is expressed as:

[0050] Where, is the combined feature vector of the i-th sample; For the The first sample Sub-band energy characteristics; For the The first sample Sub-band energy characteristics; For the The first sample Sub-band energy characteristics; is the skewness calculation function; is the kurtosis calculation function; After normalization, Sample data.

[0051] It should be noted that the combined feature vector represents the multimodal fusion of time-frequency features and time-domain statistics, overcoming the limitations of single features, such as the loss of time-domain information by fast Fourier transform and the neglect of frequency-domain patterns by time-domain statistics.

[0052] Step S204 : identifying unlabeled samples from the target feature vector according to the pre-constructed training sample set, and determining the unlabeled samples as target samples.

[0053] Step S206: Generate a pseudo label corresponding to the target sample based on the feature similarity between the target sample and the training sample set.

[0054] In specific implementation, the present invention dynamically generates pseudo labels based on feature similarity. For unlabeled samples, the similarity weight between them and labeled samples is calculated, combined with the true labels of the labeled samples, to calculate the pseudo labels of the unlabeled samples. Specifically, the following steps are included: 1) Obtain the target sample and its neighboring samples in the feature space, and calculate the similarity weight between the neighboring samples and the target sample; 2) Determine the sample labels corresponding to the nearest neighbor samples from the training sample set; 3) Determine the predicted label output corresponding to the target sample based on the sample label and similarity weight; Specifically, among the k nearest neighbors of the unlabeled sample, the similarity weight of each neighbor sample is multiplied by the corresponding true label and then summed, and then divided by the sum of the similarity weights of all neighbors, which is expressed as:

[0055] Where, is the pseudo label of the u-th unlabeled sample, used to expand the training set; is the true label of the vth labeled sample; u is a positive integer, defining the uth sample as an unlabeled sample; v is a positive integer, defining the vth sample as a labeled sample; is the k-nearest neighbor set of the u-th sample, for example, the value of k in the k-nearest neighbor is set to 5, and is calculated by the Euclidean distance of the combined feature vector; is the similarity weight between the u-th sample and the v-th sample, and the calculation method is expressed as , calculated based on the distance of the combined feature vector, which is more consistent with the similarity of fault modes than the Euclidean distance; is an exponential function with a natural constant as its base; is the similarity attenuation factor, which controls the similarity attenuation speed. For example, it is set to 0.5; is the combined feature vector of the u-th sample; is the combined feature vector of the vth sample; is the L2 norm.

[0056] 4) Using the preset confidence threshold, perform threshold comparison on the similarity weights, retain the predicted label output corresponding to the similarity weight that meets the confidence threshold, and determine the pseudo label corresponding to the target sample.

[0057] Furthermore, confidence threshold filtering is added to perform confidence threshold filtering on pseudo labels. Only when the maximum value of the similarity weight in the neighbor set is greater than the confidence threshold, the pseudo label is retained, otherwise it is discarded, which is expressed as:

[0058] Where, is the final pseudo label; is the confidence threshold, ensuring that pseudo labels are only used for high-confidence samples to avoid low-quality data contamination of training, such as setting ; Represents the maximum similarity weight in the neighbor set of the u-th sample.

[0059] Step S208: constructing a final training sample set corresponding to the preset mechanical equipment based on the pseudo labels.

[0060] In summary, a final training sample set is constructed and used to train a model to perform fault diagnosis on the device. In specific implementation, refer to the following steps S210-S220.

[0061] Step S210, calculating the label feature mean corresponding to the distribution of each category sample of the final training sample set; based on the label feature mean and the clustering correction item corresponding to the final training sample set, determining the initialization weight matrix corresponding to the neural network model.

[0062] The neural network model structure used for equipment fault diagnosis in the present invention is a variant of a multi-layer perceptron, comprising an input layer, multiple hidden layers and an output layer. The number of the multiple hidden layers can be flexibly set, for example, to 5 layers.

[0063] The input layer is used to receive the combined feature vector. The input vector of each sample is composed of multi-scale frequency domain energy features and time domain statistical features. The dimension of the input layer is consistent with the length of the combined feature and is fully connected to the first hidden layer. The number of neurons in the hidden layer is dynamically set according to the input dimension and the complexity of the fault category, shrinking layer by layer and having the ability of dimensionality reduction and compression.

[0064] Furthermore, within each hidden layer, the present invention employs adaptive activation functions, enabling different nonlinear mapping strategies for different input intervals. Oscillatory modulation is employed for positive inputs to enhance expressiveness, while exponential decay is employed for negative inputs to control the gradient, addressing the expressive bottlenecks of traditional activation functions in scenarios with sparse fault features. Furthermore, batch normalization and dropout regularization can be employed between hidden layers in this embodiment of the present invention to enhance model stability and generalization capabilities. The output layer adopts a classification output structure with a number of neurons equal to the number of fault categories. The output is the unnormalized raw score, which is then calculated using a temperature calibration function to calibrate the category probabilities.

[0065] The samples input to the neural network are samples corresponding to the combined feature vectors. Although effective preprocessing has been performed at the feature level, there may be an imbalance in the distribution of categories, such as more normal samples and fewer faulty samples. In this case, conventional random initialization can easily bias the model towards the majority class. The present invention calculates the label feature mean based on the combined feature vector of the label samples and multiplies it with the identity matrix to obtain the basic weight. Then, the basic weight is added to the clustering correction term to obtain the initialization weight matrix of the neural network, which is expressed as:

[0066] Where, is the initialization weight matrix of the neural network; is the combined feature vector of the i-th sample; is the number of labeled samples, Represent the mean of label features and offset random deviations; is the identity matrix, ensuring that the weight dimensions match; This is a clustering correction term that addresses distribution imbalance. Conventional global averaging ignores inter-class differences, but this formula introduces a clustering term to enhance the representation of the minority class.

[0067] Furthermore, the sample set corresponding to the combined feature vector is clustered, and the cluster center and number of samples of each cluster are calculated. The number of samples corresponding to the cluster center offset is weighted to obtain the cluster correction term, which is expressed as:

[0068] Where, For clustering learning rate, for example, set to ; C is the number of clusters. By default, C is defined to be the same as the number of fault categories. It is obtained by clustering the sample set corresponding to the combined feature vector through k-means; is the cluster center of the cth class; is the number of samples in the cth category; c is a positive integer.

[0069] It should be noted that Characterize the number of samples in the cth class weighted by the class center offset , so that the weights are biased towards sparse fault classes.

[0070] Step S212: Add Gaussian noise to the final training sample set, and use a preset adaptive activation function to perform forward propagation on the final training sample set and the final training sample set with Gaussian noise added, respectively.

[0071] Vibration characteristics have complex nonlinear patterns. Conventional neural networks using Sigmoid or ReLU activation functions are prone to neuronal inactivation under sparse fault characteristics. The present invention designs a piecewise threshold adaptive activation function for forward propagation. In this embodiment, the present invention determines whether the input sample is positive or negative. When the input sample is greater than zero, the input sample is nonlinearly activated based on a preset sinusoidal term. Otherwise, the input sample is nonlinearly activated based on a preset exponential term. The sinusoidal term includes a preset sinusoidal modulation intensity and a preset sinusoidal frequency factor parameter, and the exponential term includes an exponential attenuation coefficient.

[0072] Vibration signals usually contain asymmetric impact features, such as the impact response of the positive half-cycle is more significant. The embodiment of the present invention uses an adaptive activation function to force differential processing of positive and negative regions, so that the network automatically separates bidirectional vibration information, focuses on extracting impact fault features in the positive half-cycle, and filters redundant information in the negative half-cycle. In the embodiment of the present invention, the adaptive activation function adopts a "linear + sinusoidal modulation" structure for the positive input, and enhances the sensitivity to periodic fault features through the sinusoidal term. At the same time, combined with the sinusoidal modulation intensity and the sinusoidal frequency factor, it can enable the network to adaptively match the fault feature scale; the negative input adopts an exponential attenuation form, which can effectively extract the negative half-cycle of the vibration signal, which often contains environmental noise or secondary vibration components. At the same time, the exponential attenuation coefficient is combined to control the attenuation intensity, effectively suppressing irrelevant noise interference.

[0073] In the specific implementation, for the neuron input value of the neural network, if it is a positive input, the sine term combined with the sine modulation intensity and sine frequency factor parameters is multiplied by the input value; if it is a negative input, it is multiplied by the exponential term combined with the exponential attenuation coefficient to obtain the corresponding output value, which is expressed as:

[0074] Where, is the adaptive activation function; is the neuron input value of the neural network. For the first layer of the neural network, its input is the combined feature vector of the sample; is the sine modulation intensity, for example, set to ; is the sine frequency factor, for example, set to ; is the exponential decay coefficient, such as, set to ; is a sine function, Characterizing the capture of oscillatory patterns that enhance positive input; is an exponential function with a natural constant as its base, Characterizes impulse noise for smooth negative inputs.

[0075] It should be noted that the adaptive activation function overcomes the neuron inactivation problem under the sparse fault characteristics of the ReLU activation function, while enhancing the ability to express the oscillation and impact patterns of the vibration signal.

[0076] Step S214: Determine the first model output corresponding to the final training sample set and the second model output corresponding to the final training sample set with Gaussian noise added.

[0077] The final output of this embodiment of the present invention needs to distinguish multiple fault types, but the pseudo-labels in semi-supervised data may contain noise. To address this, this embodiment of the present invention also calibrates the output of the neural network. Specifically, this embodiment obtains the raw output values ​​corresponding to the neural network model; determines the calibrated probability output corresponding to the raw output values ​​based on preset calibration hyperparameters; and determines the calibrated probability output as the final model output.

[0078] In a specific implementation, the embodiment of the present invention obtains the original output value of the last layer of the neural network for each fault category, and performs calculations based on the original output values ​​of all fault categories and calibration hyperparameters to obtain the calibrated probability output for each fault category, which is expressed as:

[0079] Where, is the calibrated probability output of the kth type of fault; is the original output value of the last layer of the neural network; K is the number of fault categories; To calibrate hyperparameters, avoid the model's overconfidence in noise labels, and improve generalization, for example, set it to .

[0080] Furthermore, the fault classification category of this iteration is obtained based on the calibration probability output. That is, for each input sample, the fault diagnosis category is determined to be the category with the largest probability value in the calibration probability output, that is, the fault category with the highest probability is selected as the classification result.

[0081] Step S216: Calculate the supervision loss corresponding to the final training sample set based on the first model output; and calculate the consistency loss corresponding to the final training sample set based on the second model output.

[0082] Step S218, calculating the total loss function of the neural network model based on the supervision loss and consistency loss.

[0083] Conventional neural networks only use supervised loss during training and cannot utilize unlabeled data. Conventional consistency loss is sensitive to vibration data disturbances. The total loss function of the neural network of the present invention is calculated based on supervised loss and consistency loss, and is expressed as:

[0084] Where, is the total loss of the neural network; To monitor losses; is the unsupervised weight to prevent low-confidence pseudo-labels from dominating the optimization direction, such as setting it to ; It is a consistency loss based on unlabeled data. Among them, the cross entropy loss can be used to calculate the error between the label sample output and the true label to determine the corresponding supervision loss. Furthermore, the embodiment of the present invention also calculates the Euclidean distance corresponding to the first model output and the second model output; based on the Euclidean distance, calculates the consistency loss corresponding to the final training sample set. Specifically, the present invention sets unsupervised weights, adds Gaussian noise to the combined feature vector of each unlabeled sample, inputs the combined feature vector of the unlabeled sample into the neural network, obtains its hidden layer output, and then inputs the combined feature vector of the unlabeled sample with Gaussian noise added into the neural network to obtain its hidden layer output, and then calculates the square of the Euclidean distance between the two hidden layer outputs, and averages the calculated results of all unlabeled samples to obtain the feature-level consistency loss, which is expressed as:

[0085] Where, Represents the hidden layer output of the neural network for the combined feature vector of the u-th sample; represents the hidden layer output of the combined feature vector of the u-th sample after the network adds noise, Representation reduces noise sensitivity by perturbing features rather than inputs; is Gaussian noise, i.e., sampled from a normal distribution with mean 0 and variance 1; is the number of unlabeled samples.

[0086] In step S220 , the neural network model is iteratively trained based on the total loss function, and parameters of the neural network model are updated based on the momentum vector of the neural network model during the iterative training process.

[0087] When training neural networks using vibration data, the momentum term in conventional stochastic gradient descent is a fixed hyperparameter (such as 0.9) that maintains the influence of historical gradient directions. When the input data distribution undergoes sudden changes (e.g., when the device operating state switches or sensor channels are replaced), traditional fixed momentum can easily lead to delayed model updates, poor adaptability to dynamic feature changes, and prone to local optima. Embodiments of the present invention use a momentum vector to update neural network parameters. By modeling momentum as a vector, parameter-level momentum control is implemented. This allows for adaptive adjustment of historical gradient retention based on different parameter dimensions, enabling parameter updates to more quickly adapt to feature changes in new environments and improving adaptability to complex feature spaces.

[0088] In its specific implementation, the present invention calculates a dynamic momentum coefficient for the second-order derivative of the neural network weights based on the total loss of the neural network. The momentum vector is then updated based on the gradient of the total loss and the dynamic momentum coefficient. This dynamic momentum coefficient adjusts the momentum strength in real time using the second-order derivative (curvature information) of the loss function with respect to the weights, thereby balancing convergence speed and stability during training. When the curvature of the loss surface is large (such as in a steep canyon or near a local minimum), the absolute value of the second-order derivative increases, dynamically increasing the momentum coefficient and enhancing the inertia in the direction of the historical gradient, suppressing oscillations and accelerating traversal of flat areas. When the curvature is small, the momentum strength is reduced to avoid missing fine optimization points due to excessive inertia. This approach improves the robustness of non-convex optimization for equipment vibration fault diagnosis tasks during model parameter updates while reducing the need for hyperparameter tuning.

[0089] Specifically, the new momentum vector is expressed by the following formula:

[0090] Where, is the new momentum vector; is the old momentum vector; is the total loss gradient of the neural network; Update the learning rate for momentum, for example, set it to ; is the dynamic momentum coefficient, which is calculated as ; As the basic momentum, set it as ; is the curvature sensitivity factor, which controls the momentum adjustment amplitude to avoid oscillation. For example, it is set to ; Characterizes the second-order derivative of the total loss of the neural network with respect to the weights of the neural network.

[0091] Furthermore, the trainable parameters of the neural network are updated based on the new momentum vector, which is expressed as:

[0092] Where, is the updated neural network weight matrix; is the neural network weight matrix before updating.

[0093] Step S222: constructing an equipment fault diagnosis model based on the trained neural network model, so as to perform equipment fault diagnosis on the target mechanical equipment using the equipment fault diagnosis model.

[0094] Correspondingly, the forward propagation, loss calculation and weight update steps are repeated until the training termination condition is met, such as reaching the preset maximum number of iterations, in which case the training is terminated.

[0095] Furthermore, the trained model is used to diagnose equipment faults. Combining the above steps, new vibration sensor samples can be preprocessed using the same process as the training phase, sequentially performing wavelet threshold denoising, dynamic normalization, and multi-scale time-frequency feature extraction to generate a corresponding combined feature vector. This feature vector is then fed into the trained neural network model, which then undergoes forward propagation calculations, extracting nonlinear expressions in each hidden layer using adaptive activation functions. The model's output layer converts the raw scores for each fault category into a calibrated probability distribution using a temperature calibration function. The category with the highest probability value is then used as the final fault diagnosis result for the sample, completing the classification reasoning for the new sample.

[0096] In summary, another equipment fault diagnosis model training method based on semi-supervised learning provided by an embodiment of the present invention has the following characteristics: (1) It is proposed to calculate the similarity by combining the Euclidean distance of the feature vectors, and to generate pseudo labels by fusing the labels of high-confidence neighboring samples. Then, the confidence threshold is combined to filter the low-reliability samples, thus breaking through the error accumulation problem of the traditional self-training method.

[0097] (2) For vibration signals with non-stationary and local mutation characteristics, a joint processing strategy of wavelet multi-layer denoising and sample-level dynamic normalization is designed to enhance the adaptability to low signal-to-noise ratio conditions and improve the stability and effectiveness of the input data.

[0098] (3) In view of the nonlinear and sparse characteristics of vibration data, an adaptive activation function that performs sinusoidal modulation on positive inputs and exponential decay on negative inputs is combined with a cluster center correction strategy to initialize the network weights, which significantly enhances the recognition ability of minority class faults.

[0099] (4) Constructing a joint loss function of supervision loss and feature perturbation consistency loss, combined with a momentum optimization method based on dynamic adjustment of second-order derivatives, effectively alleviates the instability and local optimal dilemma in the early stages of training and improves training efficiency.

[0100] In summary, in order to evaluate the performance difference between the technology of the present invention and the conventional method in the semi-supervised scenario where labeled data is scarce, the embodiment of the present invention also provides a schematic diagram of the comparison results of diagnostic accuracy under different labeling ratios, referring to Figure 3 The embodiment of the present invention compares the changes in the accuracy of each method when the proportion of labeled data changes from low to high. The horizontal axis represents the proportion of labeled data, and the vertical axis represents the test accuracy. The four broken lines in the figure represent the traditional fast Fourier transform fully supervised method, the traditional wavelet feature self-training method, the traditional wavelet feature fully supervised method and the technical method of the present invention. The experimental results show that when the labeling ratio is low (<20%), the accuracy of the technology of the present invention is significantly higher than that of other methods, and its broken line is always at the top. As the labeling ratio increases, the curve of the technology of the present invention rises slowly but still maintains its leading advantage, indicating that the technology of the present invention dynamically generates pseudo-labels and confidence filtering mechanisms through feature similarity, can effectively utilize unlabeled data, reduce dependence on manual labeling, and has outstanding advantages in scenarios where the cost of industrial field labeling is high.

[0101] In order to test the robustness of each method when the vibration sensor data is interfered by noise, the embodiment of the present invention also provides a schematic diagram of the comparison results of F1 scores under different noise levels. Figure 4 The horizontal axis represents the standard deviation of the added Gaussian noise, and the vertical axis represents the F1 score of fault diagnosis. The four groups of columns correspond to the traditional fast Fourier transform method, the traditional wavelet feature method, the traditional wavelet self-training method and the technology of the present invention. The experimental results show that when the noise level is low, the performance of all methods is similar, but as the noise increases, the F1 score of the traditional method drops sharply, especially the traditional fast Fourier transform method, while the columnar drop of the technology of the present invention is the smallest, proving that the wavelet threshold denoising preprocessing and adaptive activation function design of the technology of the present invention can effectively suppress noise interference and maintain stable diagnostic performance in industrial environments with poor sensor signal quality, solving the core defect of the traditional method being sensitive to noise.

[0102] In order to analyze the detection capabilities of each method for the five types of fault conditions, especially the fault types with sparse samples, the embodiment of the present invention also provides a schematic diagram of the comparison results of recall rates corresponding to different fault categories. Figure 5The horizontal axis shows five types of faults, namely normal state, rotor imbalance, bearing damage, mechanical looseness, and gear failure. The vertical axis is the recall rate index. The three groups of columns represent the traditional fast Fourier transform method, the traditional wavelet feature method and the technology of the present invention. The experimental results show that for common faults, such as rotor imbalance, the performance differences of each method are small; but for bearing damage and gear failure faults with sparse samples, the dark columns of the technology of the present invention are significantly higher than other methods, which verifies that the clustering correction weight initialization strategy of the technology of the present invention can effectively solve the category imbalance problem, enhance the feature learning ability of rare faults, and avoid the missed detection of minority faults by traditional methods.

[0103] In order to analyze the training efficiency and performance balance of each method, the embodiment of the present invention also provides a diagram showing the relationship between training time and diagnostic accuracy. Figure 6 The horizontal axis represents the training time, the vertical axis represents the test accuracy, and the four groups of scatter points and their trend lines correspond to the four methods respectively. The experimental results show that the traditional method requires a long training time to approach performance saturation, and the right end of the curve tends to be flat, while the orange scatter points of the technology of the present invention are concentrated in the upper left area, and the initial slope of its trend line is steeper and reaches the plateau faster, indicating that the present invention can quickly improve the performance in the early stage of training through dynamic momentum coefficient optimization and feature-level consistency loss design, and only requires about 60% of the training time of the traditional method to achieve the same accuracy. In industrial fault diagnosis scenarios that require rapid model iteration, the technology of the present invention has significant efficiency advantages.

[0104] In order to verify the advantages of the wavelet threshold denoising method proposed in the present invention in processing industrial vibration signals, the embodiment of the present invention also provides a denoising effect schematic diagram. Figure 7 The embodiment of the present invention uses heat maps to respectively display the time-frequency characteristics of the original noisy signal, the denoising result of the traditional method, and the denoising result of the technology of the present invention. Obvious broadband noise interference can be seen in the original signal diagram, and the key fault characteristics are submerged by the noise. Although the noise of the signal after processing by the traditional method is reduced, it also blurs the important transient event characteristics. The processing result of the technology of the present invention clearly retains the fault characteristic frequency and its harmonic components, and effectively suppresses the background noise. This shows the superiority of the adaptive wavelet threshold algorithm of the technology of the present invention in separating noise and useful signals. The multi-scale decomposition and adaptive threshold design can effectively suppress the noise characteristics of different frequency bands.

[0105] Furthermore, based on the above embodiment, the embodiment of the present invention also provides a device fault diagnosis model training device based on semi-supervised learning. Figure 8 FIG1 shows a schematic diagram of the structure corresponding to an embodiment of the present invention. Figure 8The device includes: a data processing module 100, which is used to obtain equipment vibration data of a preset mechanical equipment in a preset scenario, perform data preprocessing on the equipment vibration data, and determine a target feature vector corresponding to the equipment vibration data; wherein the equipment vibration data includes dynamic operation signals collected in advance from key parts of the preset mechanical equipment; a label generation module 200, which is used to extract target samples from the target feature vector based on a pre-constructed training sample set; based on the feature similarity between the target sample and the training sample set, generate a pseudo label corresponding to the target sample; an execution module 300, which is used to construct a final training sample set corresponding to the preset mechanical equipment based on the pseudo label; a training module 400, which is used to input the final training sample set into a preset neural network model to train the neural network model; a construction module 500, which is used to construct an equipment fault diagnosis model based on the trained neural network model, so as to perform equipment fault diagnosis on the target mechanical equipment using the equipment fault diagnosis model.

[0106] An embodiment of the present invention provides an equipment fault diagnosis model training device based on semi-supervised learning, the implementation principle and technical effects of which are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.

[0107] The data processing module 100 is further configured to: denoise the device vibration data based on the scale changes corresponding to the device vibration data, and determine the denoised samples corresponding to the device vibration data; calculate the extreme values ​​of the denoised samples sample by sample, and dynamically normalize the denoised samples based on the calculation results to determine initial preprocessed data; extract multi-scale time-frequency features from the initial preprocessed data to determine the target feature vector corresponding to the device vibration data. The data processing module 100 is further configured to: decompose the device vibration data according to different scales using wavelet decomposition to determine multiple sub-band data corresponding to the device vibration data; calculate the adaptive threshold corresponding to each sub-band data based on the overall standard deviation of the samples corresponding to the device vibration data; extract local frequency domain features from the sub-band data based on the adaptive threshold and a preset noise suppression coefficient to determine the denoised samples corresponding to the device vibration data. The above-mentioned data processing module 100 is also used to: perform wavelet packet transform on the initial preprocessed data to extract the energy distribution of the initial preprocessed data in different sub-bands; perform time domain feature statistics on the energy distribution corresponding to each sub-band to determine the data characteristics of the initial preprocessed data in the current sub-band; and combine the data characteristics and energy distribution of each sub-band to determine the target feature vector corresponding to the equipment vibration data.

[0108] The above-mentioned label generation module 200 is also used to: identify unlabeled samples from the target feature vector based on a pre-constructed training sample set, and determine the unlabeled samples as target samples; the above-mentioned label generation module 200 is also used to: obtain the target sample and its neighboring samples in the feature space, and calculate the similarity weight between the neighboring samples and the target sample; determine the sample labels corresponding to the neighboring samples from the training sample set; determine the predicted label output corresponding to the target sample based on the sample labels and the similarity weights; use a preset confidence threshold to perform threshold comparison on the similarity weights, retain the predicted label output corresponding to the similarity weights that meet the confidence threshold, and determine the pseudo label corresponding to the target sample.

[0109] The above-mentioned training module 400 is also used to: calculate the label feature mean corresponding to the sample distribution of each category of the final training sample set; determine the initialization weight matrix corresponding to the neural network model based on the label feature mean and the clustering correction term corresponding to the final training sample set; add Gaussian noise to the final training sample set, and use the preset adaptive activation function to perform forward propagation on the final training sample set and the final training sample set with Gaussian noise added, respectively; and determine the first model output corresponding to the final training sample set and the second model output corresponding to the final training sample set with Gaussian noise added; based on the first model output, calculate the supervision loss corresponding to the final training sample set; based on the second model output, calculate the consistency loss corresponding to the final training sample set; calculate the total loss function of the neural network model based on the supervision loss and the consistency loss; based on the total loss function, iteratively train the neural network model, and, based on the momentum vector of the neural network model during the iterative training process, update the parameters of the neural network model.

[0110] The above-mentioned training module 400 is also used to: judge the positive and negative values ​​of the input samples; when the input samples are greater than zero, perform nonlinear activation processing on the input samples based on the preset sinusoidal term; wherein the sinusoidal term includes a preset sinusoidal modulation intensity and a preset sinusoidal frequency factor parameter; otherwise, perform nonlinear activation processing on the input samples based on the preset exponential term; wherein the exponential term includes an exponential attenuation coefficient. The above-mentioned training module 400 is also used to: calculate the Euclidean distance corresponding to the first model output and the second model output; based on the Euclidean distance, calculate the consistency loss corresponding to the final training sample set. The above-mentioned training module 400 is also used to: obtain the original output value corresponding to the neural network model; determine the calibration probability output corresponding to the original output value based on the preset calibration hyperparameters; and determine the calibration probability output as the final model output.

[0111] Furthermore, an embodiment of the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to realize the above-mentioned Figures 1 to 2The embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to execute the above Figures 1 to 2 The embodiment of the present invention also provides a structural diagram of an electronic device, such as Figure 9 FIG. 1 is a schematic diagram of the structure of the electronic device, wherein the electronic device includes a processor 91 and a memory 90, the memory 90 stores computer executable instructions that can be executed by the processor 91, and the processor 91 executes the computer executable instructions to implement the above Figures 1 to 2 Either of the methods shown. Figure 9 In the illustrated embodiment, the electronic device further includes a bus 92 and a communication interface 93 , wherein the processor 91 , the communication interface 93 and the memory 90 are connected via the bus 92 .

[0112] Among them, the memory 90 may include high-speed random access memory (RAM), and may also include non-volatile memory (non-volatile memory), such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 93 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 92 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc., and can also be an AMBA (Advanced Microcontroller Bus Architecture, on-chip bus standard) bus, wherein AMBA defines three types of buses, including APB (Advanced Peripheral Bus) bus, AHB (Advanced High-performance Bus) bus and AXI (Advanced eXtensible Interface) bus. The bus 92 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0113] The processor 91 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 91 or by software instructions. The processor 91 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor 91 reads the information in the memory and combines its hardware to complete the above Figures 1 to 2 Any of the methods shown.

[0114] The computer program product of a method and apparatus for training a device fault diagnosis model based on semi-supervised learning, provided in an embodiment of the present invention, includes a computer-readable storage medium storing program code. The program code includes instructions that can be used to execute the methods described in the aforementioned method embodiments. For specific implementations, please refer to the method embodiments and will not be described in detail here. Those skilled in the art will clearly understand that, for ease of description and brevity, the specific operating process of the system described above can refer to the corresponding process in the aforementioned method embodiments and will not be described in detail here. If the functions described are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for causing a computer device (such as a personal computer, server, or network device) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0115] In the description of the present invention, it should be noted that the terms "first", "second" and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. Finally, it should be noted that the above embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that any person skilled in the art who is familiar with this technical field can still modify the technical solutions described in the aforementioned embodiments within the technical scope disclosed by the present invention, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for training a device fault diagnosis model based on semi-supervised learning, characterized in that: The method comprises: Acquiring equipment vibration data of a preset mechanical device in a preset scenario, performing data preprocessing on the equipment vibration data, and determining a target feature vector corresponding to the equipment vibration data; wherein the equipment vibration data includes dynamic operation signals pre-collected from key parts of the preset mechanical device; Extracting a target sample from the target feature vector according to a pre-constructed training sample set; generating a pseudo label corresponding to the target sample based on feature similarity between the target sample and the training sample set; Based on the pseudo labels, construct a final training sample set corresponding to the preset mechanical equipment; Inputting the final training sample set into a preset neural network model to train the neural network model; An equipment fault diagnosis model is constructed based on the trained neural network model, so as to perform equipment fault diagnosis on the target mechanical equipment using the equipment fault diagnosis model.

2. The method according to claim 1, characterized in that The step of performing data preprocessing on the device vibration data to determine a target feature vector corresponding to the device vibration data includes: performing denoising processing on the device vibration data based on scale changes corresponding to the device vibration data, and determining denoised samples corresponding to the device vibration data; Performing sample-by-sample extreme value calculation on the denoised samples, and performing dynamic normalization processing on the denoised samples according to the calculation results to determine initial preprocessed data; Multi-scale time-frequency feature extraction is performed on the initial pre-processed data to determine a target feature vector corresponding to the equipment vibration data.

3. The method according to claim 2, characterized in that The step of performing denoising processing on the device vibration data based on the scale change corresponding to the device vibration data and determining denoised samples corresponding to the device vibration data includes: Decomposing the device vibration data according to different scales using wavelet decomposition to determine a plurality of sub-band data corresponding to the device vibration data; Calculating an adaptive threshold corresponding to each sub-band data according to the overall standard deviation of the samples corresponding to the device vibration data; Based on the adaptive threshold and a preset noise suppression coefficient, local frequency domain features are extracted from the sub-band data to determine denoised samples corresponding to the device vibration data.

4. The method according to claim 2, characterized in that The step of performing multi-scale time-frequency feature extraction on the initial pre-processed data to determine a target feature vector corresponding to the device vibration data includes: Performing wavelet packet transform on the initial preprocessed data to extract energy distribution of the initial preprocessed data in different sub-bands; Performing time domain feature statistics on the energy distribution corresponding to each of the sub-frequency bands to determine data characteristics of the initial pre-processed data in the current sub-frequency band; The data characteristics and the energy distribution of each sub-frequency band are combined to determine a target feature vector corresponding to the device vibration data.

5. The method according to claim 1, wherein The step of extracting target samples from the target feature vector according to the pre-constructed training sample set includes: Identifying unlabeled samples from the target feature vector according to a pre-constructed training sample set, and determining the unlabeled samples as target samples; The step of generating a pseudo label corresponding to the target sample based on feature similarity between the target sample and the training sample set includes: Obtaining the target sample and its neighboring samples in the feature space, and calculating the similarity weights between the neighboring samples and the target sample; Determine the sample labels corresponding to the neighboring samples from the training sample set; Determining a predicted label output corresponding to the target sample based on the sample label and the similarity weight; Using a preset confidence threshold, the similarity weights are threshold-matched, the predicted labels corresponding to the similarity weights that meet the confidence threshold are output and retained, and the pseudo labels corresponding to the target samples are determined.

6. The method according to claim 1, characterized in that Inputting the final training sample set into a preset neural network model, and training the neural network model comprises: Calculating the label feature mean corresponding to the distribution of each category sample of the final training sample set; determining the initialization weight matrix corresponding to the neural network model based on the label feature mean and the cluster correction term corresponding to the final training sample set; Adding Gaussian noise to the final training sample set, and performing forward propagation on the final training sample set and the final training sample set with the Gaussian noise added using a preset adaptive activation function; and determining a first model output corresponding to the final training sample set and a second model output corresponding to the final training sample set with the Gaussian noise added; Based on the output of the first model, calculating the supervision loss corresponding to the final training sample set; based on the output of the second model, calculating the consistency loss corresponding to the final training sample set; Calculating a total loss function of the neural network model based on the supervision loss and the consistency loss; Based on the total loss function, the neural network model is iteratively trained, and based on the momentum vector of the neural network model during the iterative training process, the parameters of the neural network model are updated.

7. The method according to claim 6, characterized in that The steps of performing forward propagation on the final training sample set and the final training sample set with Gaussian noise added by using a preset adaptive activation function include: Make positive and negative judgments on the input samples; When the input sample is greater than zero, nonlinear activation processing is performed on the input sample based on a preset sinusoidal term; wherein the sinusoidal term includes a preset sinusoidal modulation intensity and a preset sinusoidal frequency factor parameter; Otherwise, nonlinear activation processing is performed on the input sample based on a preset exponential term; wherein the exponential term includes an exponential decay coefficient.

8. The method according to claim 6, characterized in that The step of calculating the consistency loss corresponding to the final training sample set based on the second model output includes: Calculating the Euclidean distance between the first model output and the second model output; Based on the Euclidean distance, the consistency loss corresponding to the final training sample set is calculated.

9. The method according to claim 6, characterized in that The step of determining a first model output corresponding to the final training sample set and a second model output corresponding to the final training sample set with Gaussian noise added comprises: Obtaining the original output value corresponding to the neural network model; Determining a calibrated probability output corresponding to the original output value based on preset calibration hyperparameters; The calibrated probability output is determined as the final model output.

10. A device for training a fault diagnosis model based on semi-supervised learning, characterized in that: The device comprises: a data processing module configured to obtain vibration data of a preset mechanical device under a preset scenario, perform data preprocessing on the vibration data, and determine a target feature vector corresponding to the vibration data; wherein the vibration data includes dynamic operation signals pre-collected from key parts of the preset mechanical device; A label generation module is used to extract a target sample from the target feature vector according to a pre-constructed training sample set; and generate a pseudo label corresponding to the target sample based on the feature similarity between the target sample and the training sample set; An execution module, configured to construct a final training sample set corresponding to the preset mechanical equipment based on the pseudo labels; A training module, configured to input the final training sample set into a preset neural network model to train the neural network model; The construction module is used to construct an equipment fault diagnosis model based on the trained neural network model, so as to use the equipment fault diagnosis model to perform equipment fault diagnosis on the target mechanical equipment.

Citation Information

Patent Citations

  • Semi-supervised identification method based on clustering

    CN111695612A

  • Semi-supervised identification method based on relational network label sample expansion

    CN112257862A

  • Equipment fault classification method and device, equipment and medium

    CN114548254A

  • Bearing diagnosis method and system based on multi-modal and multi-scale fusion network

    CN117763494A

  • Method and system for predicting health of battery pack

    CN119959781A

Cited By

  • Machine tool cutting stability prediction method based on similarity measurement and semi-supervised learning

    CN121167321A

  • Data processing method and electronic equipment

    CN121213800A